Repository files navigation

VBVR-Pro-Bench

Project PagearXivCodeDatasetDatasetDatasetBench DataLeaderboardCode LicenseData License

The evaluation kit for VBVR-Pro-Bench — 100 visual-reasoning tasks, each with its own specific evaluator.

Every score is computed with classical computer vision and combinatorial algorithms (OpenCV + NumPy, Hungarian matching, multi-object tracking). There is no neural judge and no API call anywhere in this repository — the same input always produces the same score, and each score comes with a details dict recording how it was derived.

The benchmark ships in two settings, scored by two entry points:

SettingEntry pointModel output
Video (TI2V)run_evaluation_video.pyone .mp4 per instance
Image (interleaved)run_evaluation_image.pyone directory of .png steps per instance

1. Quick Start

1.1 Install

git clone https://github.com/Video-Reason/VBVR-Pro-Bench.git
cd VBVR-Pro-Bench
pip install -r requirements.txt

Python 3.9+. requirements.txt covers everything the evaluators need.

Some tasks read on-screen text or digits with EasyOCR, which downloads its weights to ~/.EasyOCR on first use. To point it at a copy you already have:

export VBVR_EASYOCR_MODELS=/path/to/easyocr_models

1.2 Download Ground Truth Data

Dataset:https://huggingface.co/datasets/Video-Reason/VBVR-Pro-Bench

# Install huggingface_hub if needed
pip install huggingface_hub
# Download the dataset
huggingface-cli download Video-Reason/VBVR-Pro-Bench --repo-type dataset --local-dir /path/to/VBVR-Pro-Bench
cd /path/to/VBVR-Pro-Bench
tar xzf VBVR-Pro-Bench-Video.tar.gz # video setting
tar xzf VBVR-Pro-Bench-Image.tar.gz # image setting

100 tasks (In-Domain_50 50 + Out-of-Domain_50 50), 5 instances each, so 500 instances per setting. Resolution 1024 × 1024; reference videos are 16 fps.

After extracting, the video setting looks like this:

/path/to/VBVR-Pro-Bench-Video/
├── In-Domain_50/
│ ├── G-131_select_next_figure_.../
│ │ ├── 00000/
│ │ │ ├── first_frame.png # Input image (condition frame)
│ │ │ ├── prompt.txt # Text prompt
│ │ │ ├── ground_truth.mp4 # Reference video
│ │ │ ├── final_frame.png # Last frame of the reference video
│ │ │ └── metadata.json # Task parameters and symbolic ground truth
│ │ ├── 00001/
│ │ └── ... # 5 instances per task
│ └── ... # 50 tasks
└── Out-of-Domain_50/
└── ... # 50 tasks

The image setting has the same shape, with the reference output as an image sequence instead of a video:

/path/to/VBVR-Pro-Bench-Image/
├── In-Domain_50/
│ ├── G-131_select_next_figure_.../
│ │ ├── 00000/
│ │ │ ├── first_frame.png # Input image (condition frame)
│ │ │ ├── prompt.txt # Text prompt
│ │ │ ├── frame_1.png # Reference output, step 1
│ │ │ ├── frame_2.png # ... step 2, when the task has more steps
│ │ │ ├── ... # up to frame_N.png
│ │ │ └── metadata.json # Task parameters and symbolic ground truth
│ │ └── ...
│ └── ...
└── Out-of-Domain_50/
└── ...

Note:metadata.json is required. It carries the symbolic ground truth (grid layouts, object positions, correct answers) that the evaluators check against — scoring will not work without it.

1.3 Generate Model Output (Inference)

For each instance, condition on first_frame.png and prompt.txt and produce one video (or one image sequence).

Video setting — one .mp4 per instance:

/path/to/model_outputs/
├── In-Domain_50/
│ ├── G-131_select_next_figure_.../
│ │ ├── 00000.mp4 # Generated video for instance 0
│ │ ├── 00001.mp4
│ │ ├── 00002.mp4
│ │ ├── 00003.mp4
│ │ └── 00004.mp4 # 5 videos per task
│ └── ... # Same task folders as GT
└── Out-of-Domain_50/
└── ... # Same task folders as GT

Image setting — one directory per instance, holding its output steps:

/path/to/model_outputs/
├── In-Domain_50/
│ ├── G-131_select_next_figure_.../
│ │ ├── 00000/
│ │ │ ├── 001.png # Step 1
│ │ │ ├── 002.png # Step 2, if the task has more steps
│ │ │ └── ...
│ │ ├── 00001/
│ │ └── ... # 5 instances per task
│ └── ...
└── Out-of-Domain_50/
└── ...

Note: The folder names (In-Domain_50/, Out-of-Domain_50/, task names) and instance ids (0000000004) must match the ground truth structure exactly. Only .mp4 is recognised in the video setting.

Every .png inside an instance directory is one output step, ordered by the number in its filename — so 1.png, 2.png, 10.png and 001.png, 002.png, 010.png both order correctly. Single-step tasks just have one file.

1.4 Run Evaluation

Video

python run_evaluation_video.py \
--model_path /path/to/model_outputs \
--gt_base /path/to/VBVR-Pro-Bench-Video \
--output_dir ./video_eval_results

Image

python run_evaluation_image.py \
--model_path /path/to/model_outputs \
--gt_image_base /path/to/VBVR-Pro-Bench-Image \
--output_dir ./image_eval_results

Batch evaluation (multiple models). Point at the directory that contains them, and optionally restrict the list:

python run_evaluation_video.py \
--models_base /path/to/all_model_outputs \
--gt_base /path/to/VBVR-Pro-Bench-Video
python run_evaluation_video.py \
--models_base /path/to/all_model_outputs \
--gt_base /path/to/VBVR-Pro-Bench-Video \
--models model_A model_B

Arguments

ArgumentApplies toDescription
--model_pathbothPath to a single model's output directory
--models_basebothBase directory containing multiple model folders
--modelsbothSpecific model names to evaluate (with --models_base)
--gt_basevideo(Required) Path to VBVR-Pro-Bench-Video
--gt_image_baseimage(Required) Path to VBVR-Pro-Bench-Image
--output_dirbothOutput directory for results
--devicebothcuda or cpu (default: cuda); affects OCR only

--model_path and --models_base are mutually exclusive, and one is required.

Use a GPU if you have one. The geometric and graph algorithms are CPU-only, but the OCR step runs on the GPU when --device cuda (the default) is set, and is noticeably slower without it. --device cpu still works if no GPU is available.


2. Output Format

Each model produces {output_dir}/{model_name}_vbvr_results.json. With --models_base, the video entry point additionally writes all_models_summary.json alongside it.

{
"model_name": "my_model",
"summary": {
"In_Domain": { "mean_score": 0.48, "num_samples": 250, "by_task": {}, "by_category": {} },
"Out_of_Domain": { "mean_score": 0.65, "num_samples": 250, "by_task": {}, "by_category": {} },
"overall": { "mean_score": 0.56, "num_samples": 500, "by_task": {}, "by_category": {} }
},
"samples": [
{
"task_name": "G-131_select_next_...",
"video_file": "00000.mp4",
"folder": "In-Domain_50",
"split": "In_Domain",
"category": "Abstraction",
"score": 0.85,
"dimensions": { "task_specific": 0.85 },
"error": null
}
]
}

Aggregates are reported per split, per task, and per cognitive category. The 100 tasks break down as:

CategoryTasks
Perception28
Abstraction25
Knowledge19
Spatiality16
Transformation12

Report Out_of_Domain for generalization: those task families do not appear in the VBVR-Pro training splits.


3. Repository Structure

VBVR-Pro-Bench/
├── run_evaluation_video.py # Entry point: video (I2V) setting
├── run_evaluation_image.py # Entry point: interleaved image setting
├── requirements.txt # Python dependencies
└── vbvr_bench/
├── __init__.py # VBVRBench class
├── utils.py # Shared CV primitives (color/shape/frame ops)
└── evaluators/
├── __init__.py # Evaluator registry, task categories, split definitions
├── base_evaluator.py # BaseEvaluator: dimensions, weights, frame handling
├── In_Domain_50_part1..5.py # 50 in-domain task evaluators
└── Out_of_Domain_50_part1..5.py # 50 out-of-domain task evaluators

TASK_EVALUATOR_MAP in vbvr_bench/evaluators/__init__.py maps each of the 100 task names to its evaluator class. To score a single instance directly:

fromvbvr_bench.evaluatorsimportget_evaluatortask="G-45_key_door_matching_data-generator"gt_dir=f"/path/to/VBVR-Pro-Bench-Video/In-Domain_50/{task}/00000"evaluator=get_evaluator(task, device="cpu")
result=evaluator.evaluate({
"video_path": "pred/00000.mp4",
"task_name": task,
"gt_path": gt_dir,
"metafile_path": [f"{gt_dir}/metadata.json"],
}, task_specific_only=True)
print(result["score"]) # float in [0, 1]print(result["details"]) # per-evaluator diagnostics

License

VBVR-Pro source code, scripts, configuration files and task-specific scoring software — including everything in this repository — are licensed under the Apache License 2.0. VBVR-Pro data and benchmark materials are separately licensed under CC BY-NC 4.0. Model weights and third-party materials remain subject to their applicable model-card and upstream terms. See LICENSE.md for details.

Citation

@misc{xu2026vbvrproscalableverifiablesuite,
title={VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning},
author={Junxiang Xu and Ruisi Wang and Fanyi Pu and Maijunxian Wang and Ran Ji and Tongxi Zhou and Chenyang Gu and Jing Zuo and Hongcan Xiao and Yimeng Geng and Wanqi Yin and Wei Chen and Oscar Qian and Zhengan Yan and Ziqi Huang and Haiwen Diao and Liang Pan and Bo Li and Xiangyu Fan and Dezhi Luo and Fengyuan Yu and Zehong Zhao and Qingying Gao and Tinghui Zhu and Yilan Zhang and Jingqi Tong and Pinyuan Feng and Zhengze Jiang and Letian Wang and Ziyu Guo and Renrui Zhang and Jieneng Chen and Sonia Joseph and Constantin Venhoff and Saman Motamed and Mengyue Yang and Chandra Sripada and Alan Yuille and Philip Torr and Lvmin Zhang and Vikash Kumar and Daniel Khashabi and Nikolaus Kriegeskorte and Rapha\"{e}l Milli\`{e}re and Vincent C. M\"{u}ller and Anyi Rao and Quan Wang and Ziwei Liu and Dahua Lin and Lei Yang and Hokin Deng and Zhongang Cai},
year={2026},
eprint={2608.26105},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2608.26105},
}

About

Rule-based evaluation kit for VBVR-Pro-Bench — 100 visual-reasoning tasks, one hand-written evaluator each

Topics

Resources

Stars

12 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

VBVR-Pro-Bench

Project PagearXivCodeDatasetDatasetDatasetBench DataLeaderboardCode LicenseData License

The evaluation kit for VBVR-Pro-Bench — 100 visual-reasoning tasks, each with its own specific evaluator.

Every score is computed with classical computer vision and combinatorial algorithms (OpenCV + NumPy, Hungarian matching, multi-object tracking). There is no neural judge and no API call anywhere in this repository — the same input always produces the same score, and each score comes with a details dict recording how it was derived.

The benchmark ships in two settings, scored by two entry points:

SettingEntry pointModel output
Video (TI2V)run_evaluation_video.pyone .mp4 per instance
Image (interleaved)run_evaluation_image.pyone directory of .png steps per instance

1. Quick Start

1.1 Install

git clone https://github.com/Video-Reason/VBVR-Pro-Bench.git
cd VBVR-Pro-Bench
pip install -r requirements.txt

Python 3.9+. requirements.txt covers everything the evaluators need.

Some tasks read on-screen text or digits with EasyOCR, which downloads its weights to ~/.EasyOCR on first use. To point it at a copy you already have:

export VBVR_EASYOCR_MODELS=/path/to/easyocr_models

1.2 Download Ground Truth Data

Dataset:https://huggingface.co/datasets/Video-Reason/VBVR-Pro-Bench

# Install huggingface_hub if needed
pip install huggingface_hub
# Download the dataset
huggingface-cli download Video-Reason/VBVR-Pro-Bench --repo-type dataset --local-dir /path/to/VBVR-Pro-Bench
cd /path/to/VBVR-Pro-Bench
tar xzf VBVR-Pro-Bench-Video.tar.gz # video setting
tar xzf VBVR-Pro-Bench-Image.tar.gz # image setting

100 tasks (In-Domain_50 50 + Out-of-Domain_50 50), 5 instances each, so 500 instances per setting. Resolution 1024 × 1024; reference videos are 16 fps.

After extracting, the video setting looks like this:

/path/to/VBVR-Pro-Bench-Video/
├── In-Domain_50/
│ ├── G-131_select_next_figure_.../
│ │ ├── 00000/
│ │ │ ├── first_frame.png # Input image (condition frame)
│ │ │ ├── prompt.txt # Text prompt
│ │ │ ├── ground_truth.mp4 # Reference video
│ │ │ ├── final_frame.png # Last frame of the reference video
│ │ │ └── metadata.json # Task parameters and symbolic ground truth
│ │ ├── 00001/
│ │ └── ... # 5 instances per task
│ └── ... # 50 tasks
└── Out-of-Domain_50/
└── ... # 50 tasks

The image setting has the same shape, with the reference output as an image sequence instead of a video:

/path/to/VBVR-Pro-Bench-Image/
├── In-Domain_50/
│ ├── G-131_select_next_figure_.../
│ │ ├── 00000/
│ │ │ ├── first_frame.png # Input image (condition frame)
│ │ │ ├── prompt.txt # Text prompt
│ │ │ ├── frame_1.png # Reference output, step 1
│ │ │ ├── frame_2.png # ... step 2, when the task has more steps
│ │ │ ├── ... # up to frame_N.png
│ │ │ └── metadata.json # Task parameters and symbolic ground truth
│ │ └── ...
│ └── ...
└── Out-of-Domain_50/
└── ...

Note:metadata.json is required. It carries the symbolic ground truth (grid layouts, object positions, correct answers) that the evaluators check against — scoring will not work without it.

1.3 Generate Model Output (Inference)

For each instance, condition on first_frame.png and prompt.txt and produce one video (or one image sequence).

Video setting — one .mp4 per instance:

/path/to/model_outputs/
├── In-Domain_50/
│ ├── G-131_select_next_figure_.../
│ │ ├── 00000.mp4 # Generated video for instance 0
│ │ ├── 00001.mp4
│ │ ├── 00002.mp4
│ │ ├── 00003.mp4
│ │ └── 00004.mp4 # 5 videos per task
│ └── ... # Same task folders as GT
└── Out-of-Domain_50/
└── ... # Same task folders as GT

Image setting — one directory per instance, holding its output steps:

/path/to/model_outputs/
├── In-Domain_50/
│ ├── G-131_select_next_figure_.../
│ │ ├── 00000/
│ │ │ ├── 001.png # Step 1
│ │ │ ├── 002.png # Step 2, if the task has more steps
│ │ │ └── ...
│ │ ├── 00001/
│ │ └── ... # 5 instances per task
│ └── ...
└── Out-of-Domain_50/
└── ...

Note: The folder names (In-Domain_50/, Out-of-Domain_50/, task names) and instance ids (0000000004) must match the ground truth structure exactly. Only .mp4 is recognised in the video setting.

Every .png inside an instance directory is one output step, ordered by the number in its filename — so 1.png, 2.png, 10.png and 001.png, 002.png, 010.png both order correctly. Single-step tasks just have one file.

1.4 Run Evaluation

Video

python run_evaluation_video.py \
--model_path /path/to/model_outputs \
--gt_base /path/to/VBVR-Pro-Bench-Video \
--output_dir ./video_eval_results

Image

python run_evaluation_image.py \
--model_path /path/to/model_outputs \
--gt_image_base /path/to/VBVR-Pro-Bench-Image \
--output_dir ./image_eval_results

Batch evaluation (multiple models). Point at the directory that contains them, and optionally restrict the list:

python run_evaluation_video.py \
--models_base /path/to/all_model_outputs \
--gt_base /path/to/VBVR-Pro-Bench-Video
python run_evaluation_video.py \
--models_base /path/to/all_model_outputs \
--gt_base /path/to/VBVR-Pro-Bench-Video \
--models model_A model_B

Arguments

ArgumentApplies toDescription
--model_pathbothPath to a single model's output directory
--models_basebothBase directory containing multiple model folders
--modelsbothSpecific model names to evaluate (with --models_base)
--gt_basevideo(Required) Path to VBVR-Pro-Bench-Video
--gt_image_baseimage(Required) Path to VBVR-Pro-Bench-Image
--output_dirbothOutput directory for results
--devicebothcuda or cpu (default: cuda); affects OCR only

--model_path and --models_base are mutually exclusive, and one is required.

Use a GPU if you have one. The geometric and graph algorithms are CPU-only, but the OCR step runs on the GPU when --device cuda (the default) is set, and is noticeably slower without it. --device cpu still works if no GPU is available.


2. Output Format

Each model produces {output_dir}/{model_name}_vbvr_results.json. With --models_base, the video entry point additionally writes all_models_summary.json alongside it.

{
"model_name": "my_model",
"summary": {
"In_Domain": { "mean_score": 0.48, "num_samples": 250, "by_task": {}, "by_category": {} },
"Out_of_Domain": { "mean_score": 0.65, "num_samples": 250, "by_task": {}, "by_category": {} },
"overall": { "mean_score": 0.56, "num_samples": 500, "by_task": {}, "by_category": {} }
},
"samples": [
{
"task_name": "G-131_select_next_...",
"video_file": "00000.mp4",
"folder": "In-Domain_50",
"split": "In_Domain",
"category": "Abstraction",
"score": 0.85,
"dimensions": { "task_specific": 0.85 },
"error": null
}
]
}

Aggregates are reported per split, per task, and per cognitive category. The 100 tasks break down as:

CategoryTasks
Perception28
Abstraction25
Knowledge19
Spatiality16
Transformation12

Report Out_of_Domain for generalization: those task families do not appear in the VBVR-Pro training splits.


3. Repository Structure

VBVR-Pro-Bench/
├── run_evaluation_video.py # Entry point: video (I2V) setting
├── run_evaluation_image.py # Entry point: interleaved image setting
├── requirements.txt # Python dependencies
└── vbvr_bench/
├── __init__.py # VBVRBench class
├── utils.py # Shared CV primitives (color/shape/frame ops)
└── evaluators/
├── __init__.py # Evaluator registry, task categories, split definitions
├── base_evaluator.py # BaseEvaluator: dimensions, weights, frame handling
├── In_Domain_50_part1..5.py # 50 in-domain task evaluators
└── Out_of_Domain_50_part1..5.py # 50 out-of-domain task evaluators

TASK_EVALUATOR_MAP in vbvr_bench/evaluators/__init__.py maps each of the 100 task names to its evaluator class. To score a single instance directly:

fromvbvr_bench.evaluatorsimportget_evaluatortask="G-45_key_door_matching_data-generator"gt_dir=f"/path/to/VBVR-Pro-Bench-Video/In-Domain_50/{task}/00000"evaluator=get_evaluator(task, device="cpu")
result=evaluator.evaluate({
"video_path": "pred/00000.mp4",
"task_name": task,
"gt_path": gt_dir,
"metafile_path": [f"{gt_dir}/metadata.json"],
}, task_specific_only=True)
print(result["score"]) # float in [0, 1]print(result["details"]) # per-evaluator diagnostics

License

VBVR-Pro source code, scripts, configuration files and task-specific scoring software — including everything in this repository — are licensed under the Apache License 2.0. VBVR-Pro data and benchmark materials are separately licensed under CC BY-NC 4.0. Model weights and third-party materials remain subject to their applicable model-card and upstream terms. See LICENSE.md for details.

Citation

@misc{xu2026vbvrproscalableverifiablesuite,
title={VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning},
author={Junxiang Xu and Ruisi Wang and Fanyi Pu and Maijunxian Wang and Ran Ji and Tongxi Zhou and Chenyang Gu and Jing Zuo and Hongcan Xiao and Yimeng Geng and Wanqi Yin and Wei Chen and Oscar Qian and Zhengan Yan and Ziqi Huang and Haiwen Diao and Liang Pan and Bo Li and Xiangyu Fan and Dezhi Luo and Fengyuan Yu and Zehong Zhao and Qingying Gao and Tinghui Zhu and Yilan Zhang and Jingqi Tong and Pinyuan Feng and Zhengze Jiang and Letian Wang and Ziyu Guo and Renrui Zhang and Jieneng Chen and Sonia Joseph and Constantin Venhoff and Saman Motamed and Mengyue Yang and Chandra Sripada and Alan Yuille and Philip Torr and Lvmin Zhang and Vikash Kumar and Daniel Khashabi and Nikolaus Kriegeskorte and Rapha\"{e}l Milli\`{e}re and Vincent C. M\"{u}ller and Anyi Rao and Quan Wang and Ziwei Liu and Dahua Lin and Lei Yang and Hokin Deng and Zhongang Cai},
year={2026},
eprint={2608.26105},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2608.26105},
}

About

Rule-based evaluation kit for VBVR-Pro-Bench — 100 visual-reasoning tasks, one hand-written evaluator each

Topics

Resources

Stars

12 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

VBVR-Pro-Bench

Project PagearXivCodeDatasetDatasetDatasetBench DataLeaderboardCode LicenseData License

The evaluation kit for VBVR-Pro-Bench — 100 visual-reasoning tasks, each with its own specific evaluator.

Every score is computed with classical computer vision and combinatorial algorithms (OpenCV + NumPy, Hungarian matching, multi-object tracking). There is no neural judge and no API call anywhere in this repository — the same input always produces the same score, and each score comes with a details dict recording how it was derived.

The benchmark ships in two settings, scored by two entry points:

SettingEntry pointModel output
Video (TI2V)run_evaluation_video.pyone .mp4 per instance
Image (interleaved)run_evaluation_image.pyone directory of .png steps per instance

1. Quick Start

1.1 Install

git clone https://github.com/Video-Reason/VBVR-Pro-Bench.git
cd VBVR-Pro-Bench
pip install -r requirements.txt

Python 3.9+. requirements.txt covers everything the evaluators need.

Some tasks read on-screen text or digits with EasyOCR, which downloads its weights to ~/.EasyOCR on first use. To point it at a copy you already have:

export VBVR_EASYOCR_MODELS=/path/to/easyocr_models

1.2 Download Ground Truth Data

Dataset:https://huggingface.co/datasets/Video-Reason/VBVR-Pro-Bench

# Install huggingface_hub if needed
pip install huggingface_hub
# Download the dataset
huggingface-cli download Video-Reason/VBVR-Pro-Bench --repo-type dataset --local-dir /path/to/VBVR-Pro-Bench
cd /path/to/VBVR-Pro-Bench
tar xzf VBVR-Pro-Bench-Video.tar.gz # video setting
tar xzf VBVR-Pro-Bench-Image.tar.gz # image setting

100 tasks (In-Domain_50 50 + Out-of-Domain_50 50), 5 instances each, so 500 instances per setting. Resolution 1024 × 1024; reference videos are 16 fps.

After extracting, the video setting looks like this:

/path/to/VBVR-Pro-Bench-Video/
├── In-Domain_50/
│ ├── G-131_select_next_figure_.../
│ │ ├── 00000/
│ │ │ ├── first_frame.png # Input image (condition frame)
│ │ │ ├── prompt.txt # Text prompt
│ │ │ ├── ground_truth.mp4 # Reference video
│ │ │ ├── final_frame.png # Last frame of the reference video
│ │ │ └── metadata.json # Task parameters and symbolic ground truth
│ │ ├── 00001/
│ │ └── ... # 5 instances per task
│ └── ... # 50 tasks
└── Out-of-Domain_50/
└── ... # 50 tasks

The image setting has the same shape, with the reference output as an image sequence instead of a video:

/path/to/VBVR-Pro-Bench-Image/
├── In-Domain_50/
│ ├── G-131_select_next_figure_.../
│ │ ├── 00000/
│ │ │ ├── first_frame.png # Input image (condition frame)
│ │ │ ├── prompt.txt # Text prompt
│ │ │ ├── frame_1.png # Reference output, step 1
│ │ │ ├── frame_2.png # ... step 2, when the task has more steps
│ │ │ ├── ... # up to frame_N.png
│ │ │ └── metadata.json # Task parameters and symbolic ground truth
│ │ └── ...
│ └── ...
└── Out-of-Domain_50/
└── ...

Note:metadata.json is required. It carries the symbolic ground truth (grid layouts, object positions, correct answers) that the evaluators check against — scoring will not work without it.

1.3 Generate Model Output (Inference)

For each instance, condition on first_frame.png and prompt.txt and produce one video (or one image sequence).

Video setting — one .mp4 per instance:

/path/to/model_outputs/
├── In-Domain_50/
│ ├── G-131_select_next_figure_.../
│ │ ├── 00000.mp4 # Generated video for instance 0
│ │ ├── 00001.mp4
│ │ ├── 00002.mp4
│ │ ├── 00003.mp4
│ │ └── 00004.mp4 # 5 videos per task
│ └── ... # Same task folders as GT
└── Out-of-Domain_50/
└── ... # Same task folders as GT

Image setting — one directory per instance, holding its output steps:

/path/to/model_outputs/
├── In-Domain_50/
│ ├── G-131_select_next_figure_.../
│ │ ├── 00000/
│ │ │ ├── 001.png # Step 1
│ │ │ ├── 002.png # Step 2, if the task has more steps
│ │ │ └── ...
│ │ ├── 00001/
│ │ └── ... # 5 instances per task
│ └── ...
└── Out-of-Domain_50/
└── ...

Note: The folder names (In-Domain_50/, Out-of-Domain_50/, task names) and instance ids (0000000004) must match the ground truth structure exactly. Only .mp4 is recognised in the video setting.

Every .png inside an instance directory is one output step, ordered by the number in its filename — so 1.png, 2.png, 10.png and 001.png, 002.png, 010.png both order correctly. Single-step tasks just have one file.

1.4 Run Evaluation

Video

python run_evaluation_video.py \
--model_path /path/to/model_outputs \
--gt_base /path/to/VBVR-Pro-Bench-Video \
--output_dir ./video_eval_results

Image

python run_evaluation_image.py \
--model_path /path/to/model_outputs \
--gt_image_base /path/to/VBVR-Pro-Bench-Image \
--output_dir ./image_eval_results

Batch evaluation (multiple models). Point at the directory that contains them, and optionally restrict the list:

python run_evaluation_video.py \
--models_base /path/to/all_model_outputs \
--gt_base /path/to/VBVR-Pro-Bench-Video
python run_evaluation_video.py \
--models_base /path/to/all_model_outputs \
--gt_base /path/to/VBVR-Pro-Bench-Video \
--models model_A model_B

Arguments

ArgumentApplies toDescription
--model_pathbothPath to a single model's output directory
--models_basebothBase directory containing multiple model folders
--modelsbothSpecific model names to evaluate (with --models_base)
--gt_basevideo(Required) Path to VBVR-Pro-Bench-Video
--gt_image_baseimage(Required) Path to VBVR-Pro-Bench-Image
--output_dirbothOutput directory for results
--devicebothcuda or cpu (default: cuda); affects OCR only

--model_path and --models_base are mutually exclusive, and one is required.

Use a GPU if you have one. The geometric and graph algorithms are CPU-only, but the OCR step runs on the GPU when --device cuda (the default) is set, and is noticeably slower without it. --device cpu still works if no GPU is available.


2. Output Format

Each model produces {output_dir}/{model_name}_vbvr_results.json. With --models_base, the video entry point additionally writes all_models_summary.json alongside it.

{
"model_name": "my_model",
"summary": {
"In_Domain": { "mean_score": 0.48, "num_samples": 250, "by_task": {}, "by_category": {} },
"Out_of_Domain": { "mean_score": 0.65, "num_samples": 250, "by_task": {}, "by_category": {} },
"overall": { "mean_score": 0.56, "num_samples": 500, "by_task": {}, "by_category": {} }
},
"samples": [
{
"task_name": "G-131_select_next_...",
"video_file": "00000.mp4",
"folder": "In-Domain_50",
"split": "In_Domain",
"category": "Abstraction",
"score": 0.85,
"dimensions": { "task_specific": 0.85 },
"error": null
}
]
}

Aggregates are reported per split, per task, and per cognitive category. The 100 tasks break down as:

CategoryTasks
Perception28
Abstraction25
Knowledge19
Spatiality16
Transformation12

Report Out_of_Domain for generalization: those task families do not appear in the VBVR-Pro training splits.


3. Repository Structure

VBVR-Pro-Bench/
├── run_evaluation_video.py # Entry point: video (I2V) setting
├── run_evaluation_image.py # Entry point: interleaved image setting
├── requirements.txt # Python dependencies
└── vbvr_bench/
├── __init__.py # VBVRBench class
├── utils.py # Shared CV primitives (color/shape/frame ops)
└── evaluators/
├── __init__.py # Evaluator registry, task categories, split definitions
├── base_evaluator.py # BaseEvaluator: dimensions, weights, frame handling
├── In_Domain_50_part1..5.py # 50 in-domain task evaluators
└── Out_of_Domain_50_part1..5.py # 50 out-of-domain task evaluators

TASK_EVALUATOR_MAP in vbvr_bench/evaluators/__init__.py maps each of the 100 task names to its evaluator class. To score a single instance directly:

fromvbvr_bench.evaluatorsimportget_evaluatortask="G-45_key_door_matching_data-generator"gt_dir=f"/path/to/VBVR-Pro-Bench-Video/In-Domain_50/{task}/00000"evaluator=get_evaluator(task, device="cpu")
result=evaluator.evaluate({
"video_path": "pred/00000.mp4",
"task_name": task,
"gt_path": gt_dir,
"metafile_path": [f"{gt_dir}/metadata.json"],
}, task_specific_only=True)
print(result["score"]) # float in [0, 1]print(result["details"]) # per-evaluator diagnostics

License

VBVR-Pro source code, scripts, configuration files and task-specific scoring software — including everything in this repository — are licensed under the Apache License 2.0. VBVR-Pro data and benchmark materials are separately licensed under CC BY-NC 4.0. Model weights and third-party materials remain subject to their applicable model-card and upstream terms. See LICENSE.md for details.

Citation

@misc{xu2026vbvrproscalableverifiablesuite,
title={VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning},
author={Junxiang Xu and Ruisi Wang and Fanyi Pu and Maijunxian Wang and Ran Ji and Tongxi Zhou and Chenyang Gu and Jing Zuo and Hongcan Xiao and Yimeng Geng and Wanqi Yin and Wei Chen and Oscar Qian and Zhengan Yan and Ziqi Huang and Haiwen Diao and Liang Pan and Bo Li and Xiangyu Fan and Dezhi Luo and Fengyuan Yu and Zehong Zhao and Qingying Gao and Tinghui Zhu and Yilan Zhang and Jingqi Tong and Pinyuan Feng and Zhengze Jiang and Letian Wang and Ziyu Guo and Renrui Zhang and Jieneng Chen and Sonia Joseph and Constantin Venhoff and Saman Motamed and Mengyue Yang and Chandra Sripada and Alan Yuille and Philip Torr and Lvmin Zhang and Vikash Kumar and Daniel Khashabi and Nikolaus Kriegeskorte and Rapha\"{e}l Milli\`{e}re and Vincent C. M\"{u}ller and Anyi Rao and Quan Wang and Ziwei Liu and Dahua Lin and Lei Yang and Hokin Deng and Zhongang Cai},
year={2026},
eprint={2608.26105},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2608.26105},
}

About

Rule-based evaluation kit for VBVR-Pro-Bench — 100 visual-reasoning tasks, one hand-written evaluator each

Topics

Resources

Stars

12 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

VBVR-Pro-Bench

Project PagearXivCodeDatasetDatasetDatasetBench DataLeaderboardCode LicenseData License

The evaluation kit for VBVR-Pro-Bench — 100 visual-reasoning tasks, each with its own specific evaluator.

Every score is computed with classical computer vision and combinatorial algorithms (OpenCV + NumPy, Hungarian matching, multi-object tracking). There is no neural judge and no API call anywhere in this repository — the same input always produces the same score, and each score comes with a details dict recording how it was derived.

The benchmark ships in two settings, scored by two entry points:

SettingEntry pointModel output
Video (TI2V)run_evaluation_video.pyone .mp4 per instance
Image (interleaved)run_evaluation_image.pyone directory of .png steps per instance

1. Quick Start

1.1 Install

git clone https://github.com/Video-Reason/VBVR-Pro-Bench.git
cd VBVR-Pro-Bench
pip install -r requirements.txt

Python 3.9+. requirements.txt covers everything the evaluators need.

Some tasks read on-screen text or digits with EasyOCR, which downloads its weights to ~/.EasyOCR on first use. To point it at a copy you already have:

export VBVR_EASYOCR_MODELS=/path/to/easyocr_models

1.2 Download Ground Truth Data

Dataset:https://huggingface.co/datasets/Video-Reason/VBVR-Pro-Bench

# Install huggingface_hub if needed
pip install huggingface_hub
# Download the dataset
huggingface-cli download Video-Reason/VBVR-Pro-Bench --repo-type dataset --local-dir /path/to/VBVR-Pro-Bench
cd /path/to/VBVR-Pro-Bench
tar xzf VBVR-Pro-Bench-Video.tar.gz # video setting
tar xzf VBVR-Pro-Bench-Image.tar.gz # image setting

100 tasks (In-Domain_50 50 + Out-of-Domain_50 50), 5 instances each, so 500 instances per setting. Resolution 1024 × 1024; reference videos are 16 fps.

After extracting, the video setting looks like this:

/path/to/VBVR-Pro-Bench-Video/
├── In-Domain_50/
│ ├── G-131_select_next_figure_.../
│ │ ├── 00000/
│ │ │ ├── first_frame.png # Input image (condition frame)
│ │ │ ├── prompt.txt # Text prompt
│ │ │ ├── ground_truth.mp4 # Reference video
│ │ │ ├── final_frame.png # Last frame of the reference video
│ │ │ └── metadata.json # Task parameters and symbolic ground truth
│ │ ├── 00001/
│ │ └── ... # 5 instances per task
│ └── ... # 50 tasks
└── Out-of-Domain_50/
└── ... # 50 tasks

The image setting has the same shape, with the reference output as an image sequence instead of a video:

/path/to/VBVR-Pro-Bench-Image/
├── In-Domain_50/
│ ├── G-131_select_next_figure_.../
│ │ ├── 00000/
│ │ │ ├── first_frame.png # Input image (condition frame)
│ │ │ ├── prompt.txt # Text prompt
│ │ │ ├── frame_1.png # Reference output, step 1
│ │ │ ├── frame_2.png # ... step 2, when the task has more steps
│ │ │ ├── ... # up to frame_N.png
│ │ │ └── metadata.json # Task parameters and symbolic ground truth
│ │ └── ...
│ └── ...
└── Out-of-Domain_50/
└── ...

Note:metadata.json is required. It carries the symbolic ground truth (grid layouts, object positions, correct answers) that the evaluators check against — scoring will not work without it.

1.3 Generate Model Output (Inference)

For each instance, condition on first_frame.png and prompt.txt and produce one video (or one image sequence).

Video setting — one .mp4 per instance:

/path/to/model_outputs/
├── In-Domain_50/
│ ├── G-131_select_next_figure_.../
│ │ ├── 00000.mp4 # Generated video for instance 0
│ │ ├── 00001.mp4
│ │ ├── 00002.mp4
│ │ ├── 00003.mp4
│ │ └── 00004.mp4 # 5 videos per task
│ └── ... # Same task folders as GT
└── Out-of-Domain_50/
└── ... # Same task folders as GT

Image setting — one directory per instance, holding its output steps:

/path/to/model_outputs/
├── In-Domain_50/
│ ├── G-131_select_next_figure_.../
│ │ ├── 00000/
│ │ │ ├── 001.png # Step 1
│ │ │ ├── 002.png # Step 2, if the task has more steps
│ │ │ └── ...
│ │ ├── 00001/
│ │ └── ... # 5 instances per task
│ └── ...
└── Out-of-Domain_50/
└── ...

Note: The folder names (In-Domain_50/, Out-of-Domain_50/, task names) and instance ids (0000000004) must match the ground truth structure exactly. Only .mp4 is recognised in the video setting.

Every .png inside an instance directory is one output step, ordered by the number in its filename — so 1.png, 2.png, 10.png and 001.png, 002.png, 010.png both order correctly. Single-step tasks just have one file.

1.4 Run Evaluation

Video

python run_evaluation_video.py \
--model_path /path/to/model_outputs \
--gt_base /path/to/VBVR-Pro-Bench-Video \
--output_dir ./video_eval_results

Image

python run_evaluation_image.py \
--model_path /path/to/model_outputs \
--gt_image_base /path/to/VBVR-Pro-Bench-Image \
--output_dir ./image_eval_results

Batch evaluation (multiple models). Point at the directory that contains them, and optionally restrict the list:

python run_evaluation_video.py \
--models_base /path/to/all_model_outputs \
--gt_base /path/to/VBVR-Pro-Bench-Video
python run_evaluation_video.py \
--models_base /path/to/all_model_outputs \
--gt_base /path/to/VBVR-Pro-Bench-Video \
--models model_A model_B

Arguments

ArgumentApplies toDescription
--model_pathbothPath to a single model's output directory
--models_basebothBase directory containing multiple model folders
--modelsbothSpecific model names to evaluate (with --models_base)
--gt_basevideo(Required) Path to VBVR-Pro-Bench-Video
--gt_image_baseimage(Required) Path to VBVR-Pro-Bench-Image
--output_dirbothOutput directory for results
--devicebothcuda or cpu (default: cuda); affects OCR only

--model_path and --models_base are mutually exclusive, and one is required.

Use a GPU if you have one. The geometric and graph algorithms are CPU-only, but the OCR step runs on the GPU when --device cuda (the default) is set, and is noticeably slower without it. --device cpu still works if no GPU is available.


2. Output Format

Each model produces {output_dir}/{model_name}_vbvr_results.json. With --models_base, the video entry point additionally writes all_models_summary.json alongside it.

{
"model_name": "my_model",
"summary": {
"In_Domain": { "mean_score": 0.48, "num_samples": 250, "by_task": {}, "by_category": {} },
"Out_of_Domain": { "mean_score": 0.65, "num_samples": 250, "by_task": {}, "by_category": {} },
"overall": { "mean_score": 0.56, "num_samples": 500, "by_task": {}, "by_category": {} }
},
"samples": [
{
"task_name": "G-131_select_next_...",
"video_file": "00000.mp4",
"folder": "In-Domain_50",
"split": "In_Domain",
"category": "Abstraction",
"score": 0.85,
"dimensions": { "task_specific": 0.85 },
"error": null
}
]
}

Aggregates are reported per split, per task, and per cognitive category. The 100 tasks break down as:

CategoryTasks
Perception28
Abstraction25
Knowledge19
Spatiality16
Transformation12

Report Out_of_Domain for generalization: those task families do not appear in the VBVR-Pro training splits.


3. Repository Structure

VBVR-Pro-Bench/
├── run_evaluation_video.py # Entry point: video (I2V) setting
├── run_evaluation_image.py # Entry point: interleaved image setting
├── requirements.txt # Python dependencies
└── vbvr_bench/
├── __init__.py # VBVRBench class
├── utils.py # Shared CV primitives (color/shape/frame ops)
└── evaluators/
├── __init__.py # Evaluator registry, task categories, split definitions
├── base_evaluator.py # BaseEvaluator: dimensions, weights, frame handling
├── In_Domain_50_part1..5.py # 50 in-domain task evaluators
└── Out_of_Domain_50_part1..5.py # 50 out-of-domain task evaluators

TASK_EVALUATOR_MAP in vbvr_bench/evaluators/__init__.py maps each of the 100 task names to its evaluator class. To score a single instance directly:

fromvbvr_bench.evaluatorsimportget_evaluatortask="G-45_key_door_matching_data-generator"gt_dir=f"/path/to/VBVR-Pro-Bench-Video/In-Domain_50/{task}/00000"evaluator=get_evaluator(task, device="cpu")
result=evaluator.evaluate({
"video_path": "pred/00000.mp4",
"task_name": task,
"gt_path": gt_dir,
"metafile_path": [f"{gt_dir}/metadata.json"],
}, task_specific_only=True)
print(result["score"]) # float in [0, 1]print(result["details"]) # per-evaluator diagnostics

License

VBVR-Pro source code, scripts, configuration files and task-specific scoring software — including everything in this repository — are licensed under the Apache License 2.0. VBVR-Pro data and benchmark materials are separately licensed under CC BY-NC 4.0. Model weights and third-party materials remain subject to their applicable model-card and upstream terms. See LICENSE.md for details.

Citation

@misc{xu2026vbvrproscalableverifiablesuite,
title={VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning},
author={Junxiang Xu and Ruisi Wang and Fanyi Pu and Maijunxian Wang and Ran Ji and Tongxi Zhou and Chenyang Gu and Jing Zuo and Hongcan Xiao and Yimeng Geng and Wanqi Yin and Wei Chen and Oscar Qian and Zhengan Yan and Ziqi Huang and Haiwen Diao and Liang Pan and Bo Li and Xiangyu Fan and Dezhi Luo and Fengyuan Yu and Zehong Zhao and Qingying Gao and Tinghui Zhu and Yilan Zhang and Jingqi Tong and Pinyuan Feng and Zhengze Jiang and Letian Wang and Ziyu Guo and Renrui Zhang and Jieneng Chen and Sonia Joseph and Constantin Venhoff and Saman Motamed and Mengyue Yang and Chandra Sripada and Alan Yuille and Philip Torr and Lvmin Zhang and Vikash Kumar and Daniel Khashabi and Nikolaus Kriegeskorte and Rapha\"{e}l Milli\`{e}re and Vincent C. M\"{u}ller and Anyi Rao and Quan Wang and Ziwei Liu and Dahua Lin and Lei Yang and Hokin Deng and Zhongang Cai},
year={2026},
eprint={2608.26105},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2608.26105},
}

About

Rule-based evaluation kit for VBVR-Pro-Bench — 100 visual-reasoning tasks, one hand-written evaluator each

Topics

Resources

Stars

12 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

VBVR-Pro-Bench

Project PagearXivCodeDatasetDatasetDatasetBench DataLeaderboardCode LicenseData License

The evaluation kit for VBVR-Pro-Bench — 100 visual-reasoning tasks, each with its own specific evaluator.

Every score is computed with classical computer vision and combinatorial algorithms (OpenCV + NumPy, Hungarian matching, multi-object tracking). There is no neural judge and no API call anywhere in this repository — the same input always produces the same score, and each score comes with a details dict recording how it was derived.

The benchmark ships in two settings, scored by two entry points:

SettingEntry pointModel output
Video (TI2V)run_evaluation_video.pyone .mp4 per instance
Image (interleaved)run_evaluation_image.pyone directory of .png steps per instance

1. Quick Start

1.1 Install

git clone https://github.com/Video-Reason/VBVR-Pro-Bench.git
cd VBVR-Pro-Bench
pip install -r requirements.txt

Python 3.9+. requirements.txt covers everything the evaluators need.

Some tasks read on-screen text or digits with EasyOCR, which downloads its weights to ~/.EasyOCR on first use. To point it at a copy you already have:

export VBVR_EASYOCR_MODELS=/path/to/easyocr_models

1.2 Download Ground Truth Data

Dataset:https://huggingface.co/datasets/Video-Reason/VBVR-Pro-Bench

# Install huggingface_hub if needed
pip install huggingface_hub
# Download the dataset
huggingface-cli download Video-Reason/VBVR-Pro-Bench --repo-type dataset --local-dir /path/to/VBVR-Pro-Bench
cd /path/to/VBVR-Pro-Bench
tar xzf VBVR-Pro-Bench-Video.tar.gz # video setting
tar xzf VBVR-Pro-Bench-Image.tar.gz # image setting

100 tasks (In-Domain_50 50 + Out-of-Domain_50 50), 5 instances each, so 500 instances per setting. Resolution 1024 × 1024; reference videos are 16 fps.

After extracting, the video setting looks like this:

/path/to/VBVR-Pro-Bench-Video/
├── In-Domain_50/
│ ├── G-131_select_next_figure_.../
│ │ ├── 00000/
│ │ │ ├── first_frame.png # Input image (condition frame)
│ │ │ ├── prompt.txt # Text prompt
│ │ │ ├── ground_truth.mp4 # Reference video
│ │ │ ├── final_frame.png # Last frame of the reference video
│ │ │ └── metadata.json # Task parameters and symbolic ground truth
│ │ ├── 00001/
│ │ └── ... # 5 instances per task
│ └── ... # 50 tasks
└── Out-of-Domain_50/
└── ... # 50 tasks

The image setting has the same shape, with the reference output as an image sequence instead of a video:

/path/to/VBVR-Pro-Bench-Image/
├── In-Domain_50/
│ ├── G-131_select_next_figure_.../
│ │ ├── 00000/
│ │ │ ├── first_frame.png # Input image (condition frame)
│ │ │ ├── prompt.txt # Text prompt
│ │ │ ├── frame_1.png # Reference output, step 1
│ │ │ ├── frame_2.png # ... step 2, when the task has more steps
│ │ │ ├── ... # up to frame_N.png
│ │ │ └── metadata.json # Task parameters and symbolic ground truth
│ │ └── ...
│ └── ...
└── Out-of-Domain_50/
└── ...

Note:metadata.json is required. It carries the symbolic ground truth (grid layouts, object positions, correct answers) that the evaluators check against — scoring will not work without it.

1.3 Generate Model Output (Inference)

For each instance, condition on first_frame.png and prompt.txt and produce one video (or one image sequence).

Video setting — one .mp4 per instance:

/path/to/model_outputs/
├── In-Domain_50/
│ ├── G-131_select_next_figure_.../
│ │ ├── 00000.mp4 # Generated video for instance 0
│ │ ├── 00001.mp4
│ │ ├── 00002.mp4
│ │ ├── 00003.mp4
│ │ └── 00004.mp4 # 5 videos per task
│ └── ... # Same task folders as GT
└── Out-of-Domain_50/
└── ... # Same task folders as GT

Image setting — one directory per instance, holding its output steps:

/path/to/model_outputs/
├── In-Domain_50/
│ ├── G-131_select_next_figure_.../
│ │ ├── 00000/
│ │ │ ├── 001.png # Step 1
│ │ │ ├── 002.png # Step 2, if the task has more steps
│ │ │ └── ...
│ │ ├── 00001/
│ │ └── ... # 5 instances per task
│ └── ...
└── Out-of-Domain_50/
└── ...

Note: The folder names (In-Domain_50/, Out-of-Domain_50/, task names) and instance ids (0000000004) must match the ground truth structure exactly. Only .mp4 is recognised in the video setting.

Every .png inside an instance directory is one output step, ordered by the number in its filename — so 1.png, 2.png, 10.png and 001.png, 002.png, 010.png both order correctly. Single-step tasks just have one file.

1.4 Run Evaluation

Video

python run_evaluation_video.py \
--model_path /path/to/model_outputs \
--gt_base /path/to/VBVR-Pro-Bench-Video \
--output_dir ./video_eval_results

Image

python run_evaluation_image.py \
--model_path /path/to/model_outputs \
--gt_image_base /path/to/VBVR-Pro-Bench-Image \
--output_dir ./image_eval_results

Batch evaluation (multiple models). Point at the directory that contains them, and optionally restrict the list:

python run_evaluation_video.py \
--models_base /path/to/all_model_outputs \
--gt_base /path/to/VBVR-Pro-Bench-Video
python run_evaluation_video.py \
--models_base /path/to/all_model_outputs \
--gt_base /path/to/VBVR-Pro-Bench-Video \
--models model_A model_B

Arguments

ArgumentApplies toDescription
--model_pathbothPath to a single model's output directory
--models_basebothBase directory containing multiple model folders
--modelsbothSpecific model names to evaluate (with --models_base)
--gt_basevideo(Required) Path to VBVR-Pro-Bench-Video
--gt_image_baseimage(Required) Path to VBVR-Pro-Bench-Image
--output_dirbothOutput directory for results
--devicebothcuda or cpu (default: cuda); affects OCR only

--model_path and --models_base are mutually exclusive, and one is required.

Use a GPU if you have one. The geometric and graph algorithms are CPU-only, but the OCR step runs on the GPU when --device cuda (the default) is set, and is noticeably slower without it. --device cpu still works if no GPU is available.


2. Output Format

Each model produces {output_dir}/{model_name}_vbvr_results.json. With --models_base, the video entry point additionally writes all_models_summary.json alongside it.

{
"model_name": "my_model",
"summary": {
"In_Domain": { "mean_score": 0.48, "num_samples": 250, "by_task": {}, "by_category": {} },
"Out_of_Domain": { "mean_score": 0.65, "num_samples": 250, "by_task": {}, "by_category": {} },
"overall": { "mean_score": 0.56, "num_samples": 500, "by_task": {}, "by_category": {} }
},
"samples": [
{
"task_name": "G-131_select_next_...",
"video_file": "00000.mp4",
"folder": "In-Domain_50",
"split": "In_Domain",
"category": "Abstraction",
"score": 0.85,
"dimensions": { "task_specific": 0.85 },
"error": null
}
]
}

Aggregates are reported per split, per task, and per cognitive category. The 100 tasks break down as:

CategoryTasks
Perception28
Abstraction25
Knowledge19
Spatiality16
Transformation12

Report Out_of_Domain for generalization: those task families do not appear in the VBVR-Pro training splits.


3. Repository Structure

VBVR-Pro-Bench/
├── run_evaluation_video.py # Entry point: video (I2V) setting
├── run_evaluation_image.py # Entry point: interleaved image setting
├── requirements.txt # Python dependencies
└── vbvr_bench/
├── __init__.py # VBVRBench class
├── utils.py # Shared CV primitives (color/shape/frame ops)
└── evaluators/
├── __init__.py # Evaluator registry, task categories, split definitions
├── base_evaluator.py # BaseEvaluator: dimensions, weights, frame handling
├── In_Domain_50_part1..5.py # 50 in-domain task evaluators
└── Out_of_Domain_50_part1..5.py # 50 out-of-domain task evaluators

TASK_EVALUATOR_MAP in vbvr_bench/evaluators/__init__.py maps each of the 100 task names to its evaluator class. To score a single instance directly:

fromvbvr_bench.evaluatorsimportget_evaluatortask="G-45_key_door_matching_data-generator"gt_dir=f"/path/to/VBVR-Pro-Bench-Video/In-Domain_50/{task}/00000"evaluator=get_evaluator(task, device="cpu")
result=evaluator.evaluate({
"video_path": "pred/00000.mp4",
"task_name": task,
"gt_path": gt_dir,
"metafile_path": [f"{gt_dir}/metadata.json"],
}, task_specific_only=True)
print(result["score"]) # float in [0, 1]print(result["details"]) # per-evaluator diagnostics

License

VBVR-Pro source code, scripts, configuration files and task-specific scoring software — including everything in this repository — are licensed under the Apache License 2.0. VBVR-Pro data and benchmark materials are separately licensed under CC BY-NC 4.0. Model weights and third-party materials remain subject to their applicable model-card and upstream terms. See LICENSE.md for details.

Citation

@misc{xu2026vbvrproscalableverifiablesuite,
title={VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning},
author={Junxiang Xu and Ruisi Wang and Fanyi Pu and Maijunxian Wang and Ran Ji and Tongxi Zhou and Chenyang Gu and Jing Zuo and Hongcan Xiao and Yimeng Geng and Wanqi Yin and Wei Chen and Oscar Qian and Zhengan Yan and Ziqi Huang and Haiwen Diao and Liang Pan and Bo Li and Xiangyu Fan and Dezhi Luo and Fengyuan Yu and Zehong Zhao and Qingying Gao and Tinghui Zhu and Yilan Zhang and Jingqi Tong and Pinyuan Feng and Zhengze Jiang and Letian Wang and Ziyu Guo and Renrui Zhang and Jieneng Chen and Sonia Joseph and Constantin Venhoff and Saman Motamed and Mengyue Yang and Chandra Sripada and Alan Yuille and Philip Torr and Lvmin Zhang and Vikash Kumar and Daniel Khashabi and Nikolaus Kriegeskorte and Rapha\"{e}l Milli\`{e}re and Vincent C. M\"{u}ller and Anyi Rao and Quan Wang and Ziwei Liu and Dahua Lin and Lei Yang and Hokin Deng and Zhongang Cai},
year={2026},
eprint={2608.26105},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2608.26105},
}

About

Rule-based evaluation kit for VBVR-Pro-Bench — 100 visual-reasoning tasks, one hand-written evaluator each

Topics

Resources

Stars

12 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

VBVR-Pro-Bench

Project PagearXivCodeDatasetDatasetDatasetBench DataLeaderboardCode LicenseData License

The evaluation kit for VBVR-Pro-Bench — 100 visual-reasoning tasks, each with its own specific evaluator.

Every score is computed with classical computer vision and combinatorial algorithms (OpenCV + NumPy, Hungarian matching, multi-object tracking). There is no neural judge and no API call anywhere in this repository — the same input always produces the same score, and each score comes with a details dict recording how it was derived.

The benchmark ships in two settings, scored by two entry points:

SettingEntry pointModel output
Video (TI2V)run_evaluation_video.pyone .mp4 per instance
Image (interleaved)run_evaluation_image.pyone directory of .png steps per instance

1. Quick Start

1.1 Install

git clone https://github.com/Video-Reason/VBVR-Pro-Bench.git
cd VBVR-Pro-Bench
pip install -r requirements.txt

Python 3.9+. requirements.txt covers everything the evaluators need.

Some tasks read on-screen text or digits with EasyOCR, which downloads its weights to ~/.EasyOCR on first use. To point it at a copy you already have:

export VBVR_EASYOCR_MODELS=/path/to/easyocr_models

1.2 Download Ground Truth Data

Dataset:https://huggingface.co/datasets/Video-Reason/VBVR-Pro-Bench

# Install huggingface_hub if needed
pip install huggingface_hub
# Download the dataset
huggingface-cli download Video-Reason/VBVR-Pro-Bench --repo-type dataset --local-dir /path/to/VBVR-Pro-Bench
cd /path/to/VBVR-Pro-Bench
tar xzf VBVR-Pro-Bench-Video.tar.gz # video setting
tar xzf VBVR-Pro-Bench-Image.tar.gz # image setting

100 tasks (In-Domain_50 50 + Out-of-Domain_50 50), 5 instances each, so 500 instances per setting. Resolution 1024 × 1024; reference videos are 16 fps.

After extracting, the video setting looks like this:

/path/to/VBVR-Pro-Bench-Video/
├── In-Domain_50/
│ ├── G-131_select_next_figure_.../
│ │ ├── 00000/
│ │ │ ├── first_frame.png # Input image (condition frame)
│ │ │ ├── prompt.txt # Text prompt
│ │ │ ├── ground_truth.mp4 # Reference video
│ │ │ ├── final_frame.png # Last frame of the reference video
│ │ │ └── metadata.json # Task parameters and symbolic ground truth
│ │ ├── 00001/
│ │ └── ... # 5 instances per task
│ └── ... # 50 tasks
└── Out-of-Domain_50/
└── ... # 50 tasks

The image setting has the same shape, with the reference output as an image sequence instead of a video:

/path/to/VBVR-Pro-Bench-Image/
├── In-Domain_50/
│ ├── G-131_select_next_figure_.../
│ │ ├── 00000/
│ │ │ ├── first_frame.png # Input image (condition frame)
│ │ │ ├── prompt.txt # Text prompt
│ │ │ ├── frame_1.png # Reference output, step 1
│ │ │ ├── frame_2.png # ... step 2, when the task has more steps
│ │ │ ├── ... # up to frame_N.png
│ │ │ └── metadata.json # Task parameters and symbolic ground truth
│ │ └── ...
│ └── ...
└── Out-of-Domain_50/
└── ...

Note:metadata.json is required. It carries the symbolic ground truth (grid layouts, object positions, correct answers) that the evaluators check against — scoring will not work without it.

1.3 Generate Model Output (Inference)

For each instance, condition on first_frame.png and prompt.txt and produce one video (or one image sequence).

Video setting — one .mp4 per instance:

/path/to/model_outputs/
├── In-Domain_50/
│ ├── G-131_select_next_figure_.../
│ │ ├── 00000.mp4 # Generated video for instance 0
│ │ ├── 00001.mp4
│ │ ├── 00002.mp4
│ │ ├── 00003.mp4
│ │ └── 00004.mp4 # 5 videos per task
│ └── ... # Same task folders as GT
└── Out-of-Domain_50/
└── ... # Same task folders as GT

Image setting — one directory per instance, holding its output steps:

/path/to/model_outputs/
├── In-Domain_50/
│ ├── G-131_select_next_figure_.../
│ │ ├── 00000/
│ │ │ ├── 001.png # Step 1
│ │ │ ├── 002.png # Step 2, if the task has more steps
│ │ │ └── ...
│ │ ├── 00001/
│ │ └── ... # 5 instances per task
│ └── ...
└── Out-of-Domain_50/
└── ...

Note: The folder names (In-Domain_50/, Out-of-Domain_50/, task names) and instance ids (0000000004) must match the ground truth structure exactly. Only .mp4 is recognised in the video setting.

Every .png inside an instance directory is one output step, ordered by the number in its filename — so 1.png, 2.png, 10.png and 001.png, 002.png, 010.png both order correctly. Single-step tasks just have one file.

1.4 Run Evaluation

Video

python run_evaluation_video.py \
--model_path /path/to/model_outputs \
--gt_base /path/to/VBVR-Pro-Bench-Video \
--output_dir ./video_eval_results

Image

python run_evaluation_image.py \
--model_path /path/to/model_outputs \
--gt_image_base /path/to/VBVR-Pro-Bench-Image \
--output_dir ./image_eval_results

Batch evaluation (multiple models). Point at the directory that contains them, and optionally restrict the list:

python run_evaluation_video.py \
--models_base /path/to/all_model_outputs \
--gt_base /path/to/VBVR-Pro-Bench-Video
python run_evaluation_video.py \
--models_base /path/to/all_model_outputs \
--gt_base /path/to/VBVR-Pro-Bench-Video \
--models model_A model_B

Arguments

ArgumentApplies toDescription
--model_pathbothPath to a single model's output directory
--models_basebothBase directory containing multiple model folders
--modelsbothSpecific model names to evaluate (with --models_base)
--gt_basevideo(Required) Path to VBVR-Pro-Bench-Video
--gt_image_baseimage(Required) Path to VBVR-Pro-Bench-Image
--output_dirbothOutput directory for results
--devicebothcuda or cpu (default: cuda); affects OCR only

--model_path and --models_base are mutually exclusive, and one is required.

Use a GPU if you have one. The geometric and graph algorithms are CPU-only, but the OCR step runs on the GPU when --device cuda (the default) is set, and is noticeably slower without it. --device cpu still works if no GPU is available.


2. Output Format

Each model produces {output_dir}/{model_name}_vbvr_results.json. With --models_base, the video entry point additionally writes all_models_summary.json alongside it.

{
"model_name": "my_model",
"summary": {
"In_Domain": { "mean_score": 0.48, "num_samples": 250, "by_task": {}, "by_category": {} },
"Out_of_Domain": { "mean_score": 0.65, "num_samples": 250, "by_task": {}, "by_category": {} },
"overall": { "mean_score": 0.56, "num_samples": 500, "by_task": {}, "by_category": {} }
},
"samples": [
{
"task_name": "G-131_select_next_...",
"video_file": "00000.mp4",
"folder": "In-Domain_50",
"split": "In_Domain",
"category": "Abstraction",
"score": 0.85,
"dimensions": { "task_specific": 0.85 },
"error": null
}
]
}

Aggregates are reported per split, per task, and per cognitive category. The 100 tasks break down as:

CategoryTasks
Perception28
Abstraction25
Knowledge19
Spatiality16
Transformation12

Report Out_of_Domain for generalization: those task families do not appear in the VBVR-Pro training splits.


3. Repository Structure

VBVR-Pro-Bench/
├── run_evaluation_video.py # Entry point: video (I2V) setting
├── run_evaluation_image.py # Entry point: interleaved image setting
├── requirements.txt # Python dependencies
└── vbvr_bench/
├── __init__.py # VBVRBench class
├── utils.py # Shared CV primitives (color/shape/frame ops)
└── evaluators/
├── __init__.py # Evaluator registry, task categories, split definitions
├── base_evaluator.py # BaseEvaluator: dimensions, weights, frame handling
├── In_Domain_50_part1..5.py # 50 in-domain task evaluators
└── Out_of_Domain_50_part1..5.py # 50 out-of-domain task evaluators

TASK_EVALUATOR_MAP in vbvr_bench/evaluators/__init__.py maps each of the 100 task names to its evaluator class. To score a single instance directly:

fromvbvr_bench.evaluatorsimportget_evaluatortask="G-45_key_door_matching_data-generator"gt_dir=f"/path/to/VBVR-Pro-Bench-Video/In-Domain_50/{task}/00000"evaluator=get_evaluator(task, device="cpu")
result=evaluator.evaluate({
"video_path": "pred/00000.mp4",
"task_name": task,
"gt_path": gt_dir,
"metafile_path": [f"{gt_dir}/metadata.json"],
}, task_specific_only=True)
print(result["score"]) # float in [0, 1]print(result["details"]) # per-evaluator diagnostics

License

VBVR-Pro source code, scripts, configuration files and task-specific scoring software — including everything in this repository — are licensed under the Apache License 2.0. VBVR-Pro data and benchmark materials are separately licensed under CC BY-NC 4.0. Model weights and third-party materials remain subject to their applicable model-card and upstream terms. See LICENSE.md for details.

Citation

@misc{xu2026vbvrproscalableverifiablesuite,
title={VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning},
author={Junxiang Xu and Ruisi Wang and Fanyi Pu and Maijunxian Wang and Ran Ji and Tongxi Zhou and Chenyang Gu and Jing Zuo and Hongcan Xiao and Yimeng Geng and Wanqi Yin and Wei Chen and Oscar Qian and Zhengan Yan and Ziqi Huang and Haiwen Diao and Liang Pan and Bo Li and Xiangyu Fan and Dezhi Luo and Fengyuan Yu and Zehong Zhao and Qingying Gao and Tinghui Zhu and Yilan Zhang and Jingqi Tong and Pinyuan Feng and Zhengze Jiang and Letian Wang and Ziyu Guo and Renrui Zhang and Jieneng Chen and Sonia Joseph and Constantin Venhoff and Saman Motamed and Mengyue Yang and Chandra Sripada and Alan Yuille and Philip Torr and Lvmin Zhang and Vikash Kumar and Daniel Khashabi and Nikolaus Kriegeskorte and Rapha\"{e}l Milli\`{e}re and Vincent C. M\"{u}ller and Anyi Rao and Quan Wang and Ziwei Liu and Dahua Lin and Lei Yang and Hokin Deng and Zhongang Cai},
year={2026},
eprint={2608.26105},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2608.26105},
}

About

Rule-based evaluation kit for VBVR-Pro-Bench — 100 visual-reasoning tasks, one hand-written evaluator each

Topics

Resources

Stars

12 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

VBVR-Pro-Bench

Project PagearXivCodeDatasetDatasetDatasetBench DataLeaderboardCode LicenseData License

The evaluation kit for VBVR-Pro-Bench — 100 visual-reasoning tasks, each with its own specific evaluator.

Every score is computed with classical computer vision and combinatorial algorithms (OpenCV + NumPy, Hungarian matching, multi-object tracking). There is no neural judge and no API call anywhere in this repository — the same input always produces the same score, and each score comes with a details dict recording how it was derived.

The benchmark ships in two settings, scored by two entry points:

SettingEntry pointModel output
Video (TI2V)run_evaluation_video.pyone .mp4 per instance
Image (interleaved)run_evaluation_image.pyone directory of .png steps per instance

1. Quick Start

1.1 Install

git clone https://github.com/Video-Reason/VBVR-Pro-Bench.git
cd VBVR-Pro-Bench
pip install -r requirements.txt

Python 3.9+. requirements.txt covers everything the evaluators need.

Some tasks read on-screen text or digits with EasyOCR, which downloads its weights to ~/.EasyOCR on first use. To point it at a copy you already have:

export VBVR_EASYOCR_MODELS=/path/to/easyocr_models

1.2 Download Ground Truth Data

Dataset:https://huggingface.co/datasets/Video-Reason/VBVR-Pro-Bench

# Install huggingface_hub if needed
pip install huggingface_hub
# Download the dataset
huggingface-cli download Video-Reason/VBVR-Pro-Bench --repo-type dataset --local-dir /path/to/VBVR-Pro-Bench
cd /path/to/VBVR-Pro-Bench
tar xzf VBVR-Pro-Bench-Video.tar.gz # video setting
tar xzf VBVR-Pro-Bench-Image.tar.gz # image setting

100 tasks (In-Domain_50 50 + Out-of-Domain_50 50), 5 instances each, so 500 instances per setting. Resolution 1024 × 1024; reference videos are 16 fps.

After extracting, the video setting looks like this:

/path/to/VBVR-Pro-Bench-Video/
├── In-Domain_50/
│ ├── G-131_select_next_figure_.../
│ │ ├── 00000/
│ │ │ ├── first_frame.png # Input image (condition frame)
│ │ │ ├── prompt.txt # Text prompt
│ │ │ ├── ground_truth.mp4 # Reference video
│ │ │ ├── final_frame.png # Last frame of the reference video
│ │ │ └── metadata.json # Task parameters and symbolic ground truth
│ │ ├── 00001/
│ │ └── ... # 5 instances per task
│ └── ... # 50 tasks
└── Out-of-Domain_50/
└── ... # 50 tasks

The image setting has the same shape, with the reference output as an image sequence instead of a video:

/path/to/VBVR-Pro-Bench-Image/
├── In-Domain_50/
│ ├── G-131_select_next_figure_.../
│ │ ├── 00000/
│ │ │ ├── first_frame.png # Input image (condition frame)
│ │ │ ├── prompt.txt # Text prompt
│ │ │ ├── frame_1.png # Reference output, step 1
│ │ │ ├── frame_2.png # ... step 2, when the task has more steps
│ │ │ ├── ... # up to frame_N.png
│ │ │ └── metadata.json # Task parameters and symbolic ground truth
│ │ └── ...
│ └── ...
└── Out-of-Domain_50/
└── ...

Note:metadata.json is required. It carries the symbolic ground truth (grid layouts, object positions, correct answers) that the evaluators check against — scoring will not work without it.

1.3 Generate Model Output (Inference)

For each instance, condition on first_frame.png and prompt.txt and produce one video (or one image sequence).

Video setting — one .mp4 per instance:

/path/to/model_outputs/
├── In-Domain_50/
│ ├── G-131_select_next_figure_.../
│ │ ├── 00000.mp4 # Generated video for instance 0
│ │ ├── 00001.mp4
│ │ ├── 00002.mp4
│ │ ├── 00003.mp4
│ │ └── 00004.mp4 # 5 videos per task
│ └── ... # Same task folders as GT
└── Out-of-Domain_50/
└── ... # Same task folders as GT

Image setting — one directory per instance, holding its output steps:

/path/to/model_outputs/
├── In-Domain_50/
│ ├── G-131_select_next_figure_.../
│ │ ├── 00000/
│ │ │ ├── 001.png # Step 1
│ │ │ ├── 002.png # Step 2, if the task has more steps
│ │ │ └── ...
│ │ ├── 00001/
│ │ └── ... # 5 instances per task
│ └── ...
└── Out-of-Domain_50/
└── ...

Note: The folder names (In-Domain_50/, Out-of-Domain_50/, task names) and instance ids (0000000004) must match the ground truth structure exactly. Only .mp4 is recognised in the video setting.

Every .png inside an instance directory is one output step, ordered by the number in its filename — so 1.png, 2.png, 10.png and 001.png, 002.png, 010.png both order correctly. Single-step tasks just have one file.

1.4 Run Evaluation

Video

python run_evaluation_video.py \
--model_path /path/to/model_outputs \
--gt_base /path/to/VBVR-Pro-Bench-Video \
--output_dir ./video_eval_results

Image

python run_evaluation_image.py \
--model_path /path/to/model_outputs \
--gt_image_base /path/to/VBVR-Pro-Bench-Image \
--output_dir ./image_eval_results

Batch evaluation (multiple models). Point at the directory that contains them, and optionally restrict the list:

python run_evaluation_video.py \
--models_base /path/to/all_model_outputs \
--gt_base /path/to/VBVR-Pro-Bench-Video
python run_evaluation_video.py \
--models_base /path/to/all_model_outputs \
--gt_base /path/to/VBVR-Pro-Bench-Video \
--models model_A model_B

Arguments

ArgumentApplies toDescription
--model_pathbothPath to a single model's output directory
--models_basebothBase directory containing multiple model folders
--modelsbothSpecific model names to evaluate (with --models_base)
--gt_basevideo(Required) Path to VBVR-Pro-Bench-Video
--gt_image_baseimage(Required) Path to VBVR-Pro-Bench-Image
--output_dirbothOutput directory for results
--devicebothcuda or cpu (default: cuda); affects OCR only

--model_path and --models_base are mutually exclusive, and one is required.

Use a GPU if you have one. The geometric and graph algorithms are CPU-only, but the OCR step runs on the GPU when --device cuda (the default) is set, and is noticeably slower without it. --device cpu still works if no GPU is available.


2. Output Format

Each model produces {output_dir}/{model_name}_vbvr_results.json. With --models_base, the video entry point additionally writes all_models_summary.json alongside it.

{
"model_name": "my_model",
"summary": {
"In_Domain": { "mean_score": 0.48, "num_samples": 250, "by_task": {}, "by_category": {} },
"Out_of_Domain": { "mean_score": 0.65, "num_samples": 250, "by_task": {}, "by_category": {} },
"overall": { "mean_score": 0.56, "num_samples": 500, "by_task": {}, "by_category": {} }
},
"samples": [
{
"task_name": "G-131_select_next_...",
"video_file": "00000.mp4",
"folder": "In-Domain_50",
"split": "In_Domain",
"category": "Abstraction",
"score": 0.85,
"dimensions": { "task_specific": 0.85 },
"error": null
}
]
}

Aggregates are reported per split, per task, and per cognitive category. The 100 tasks break down as:

CategoryTasks
Perception28
Abstraction25
Knowledge19
Spatiality16
Transformation12

Report Out_of_Domain for generalization: those task families do not appear in the VBVR-Pro training splits.


3. Repository Structure

VBVR-Pro-Bench/
├── run_evaluation_video.py # Entry point: video (I2V) setting
├── run_evaluation_image.py # Entry point: interleaved image setting
├── requirements.txt # Python dependencies
└── vbvr_bench/
├── __init__.py # VBVRBench class
├── utils.py # Shared CV primitives (color/shape/frame ops)
└── evaluators/
├── __init__.py # Evaluator registry, task categories, split definitions
├── base_evaluator.py # BaseEvaluator: dimensions, weights, frame handling
├── In_Domain_50_part1..5.py # 50 in-domain task evaluators
└── Out_of_Domain_50_part1..5.py # 50 out-of-domain task evaluators

TASK_EVALUATOR_MAP in vbvr_bench/evaluators/__init__.py maps each of the 100 task names to its evaluator class. To score a single instance directly:

fromvbvr_bench.evaluatorsimportget_evaluatortask="G-45_key_door_matching_data-generator"gt_dir=f"/path/to/VBVR-Pro-Bench-Video/In-Domain_50/{task}/00000"evaluator=get_evaluator(task, device="cpu")
result=evaluator.evaluate({
"video_path": "pred/00000.mp4",
"task_name": task,
"gt_path": gt_dir,
"metafile_path": [f"{gt_dir}/metadata.json"],
}, task_specific_only=True)
print(result["score"]) # float in [0, 1]print(result["details"]) # per-evaluator diagnostics

License

VBVR-Pro source code, scripts, configuration files and task-specific scoring software — including everything in this repository — are licensed under the Apache License 2.0. VBVR-Pro data and benchmark materials are separately licensed under CC BY-NC 4.0. Model weights and third-party materials remain subject to their applicable model-card and upstream terms. See LICENSE.md for details.

Citation

@misc{xu2026vbvrproscalableverifiablesuite,
title={VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning},
author={Junxiang Xu and Ruisi Wang and Fanyi Pu and Maijunxian Wang and Ran Ji and Tongxi Zhou and Chenyang Gu and Jing Zuo and Hongcan Xiao and Yimeng Geng and Wanqi Yin and Wei Chen and Oscar Qian and Zhengan Yan and Ziqi Huang and Haiwen Diao and Liang Pan and Bo Li and Xiangyu Fan and Dezhi Luo and Fengyuan Yu and Zehong Zhao and Qingying Gao and Tinghui Zhu and Yilan Zhang and Jingqi Tong and Pinyuan Feng and Zhengze Jiang and Letian Wang and Ziyu Guo and Renrui Zhang and Jieneng Chen and Sonia Joseph and Constantin Venhoff and Saman Motamed and Mengyue Yang and Chandra Sripada and Alan Yuille and Philip Torr and Lvmin Zhang and Vikash Kumar and Daniel Khashabi and Nikolaus Kriegeskorte and Rapha\"{e}l Milli\`{e}re and Vincent C. M\"{u}ller and Anyi Rao and Quan Wang and Ziwei Liu and Dahua Lin and Lei Yang and Hokin Deng and Zhongang Cai},
year={2026},
eprint={2608.26105},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2608.26105},
}

About

Rule-based evaluation kit for VBVR-Pro-Bench — 100 visual-reasoning tasks, one hand-written evaluator each

Topics

Resources

Stars

12 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

VBVR-Pro-Bench

Project PagearXivCodeDatasetDatasetDatasetBench DataLeaderboardCode LicenseData License

The evaluation kit for VBVR-Pro-Bench — 100 visual-reasoning tasks, each with its own specific evaluator.

Every score is computed with classical computer vision and combinatorial algorithms (OpenCV + NumPy, Hungarian matching, multi-object tracking). There is no neural judge and no API call anywhere in this repository — the same input always produces the same score, and each score comes with a details dict recording how it was derived.

The benchmark ships in two settings, scored by two entry points:

SettingEntry pointModel output
Video (TI2V)run_evaluation_video.pyone .mp4 per instance
Image (interleaved)run_evaluation_image.pyone directory of .png steps per instance

1. Quick Start

1.1 Install

git clone https://github.com/Video-Reason/VBVR-Pro-Bench.git
cd VBVR-Pro-Bench
pip install -r requirements.txt

Python 3.9+. requirements.txt covers everything the evaluators need.

Some tasks read on-screen text or digits with EasyOCR, which downloads its weights to ~/.EasyOCR on first use. To point it at a copy you already have:

export VBVR_EASYOCR_MODELS=/path/to/easyocr_models

1.2 Download Ground Truth Data

Dataset:https://huggingface.co/datasets/Video-Reason/VBVR-Pro-Bench

# Install huggingface_hub if needed
pip install huggingface_hub
# Download the dataset
huggingface-cli download Video-Reason/VBVR-Pro-Bench --repo-type dataset --local-dir /path/to/VBVR-Pro-Bench
cd /path/to/VBVR-Pro-Bench
tar xzf VBVR-Pro-Bench-Video.tar.gz # video setting
tar xzf VBVR-Pro-Bench-Image.tar.gz # image setting

100 tasks (In-Domain_50 50 + Out-of-Domain_50 50), 5 instances each, so 500 instances per setting. Resolution 1024 × 1024; reference videos are 16 fps.

After extracting, the video setting looks like this:

/path/to/VBVR-Pro-Bench-Video/
├── In-Domain_50/
│ ├── G-131_select_next_figure_.../
│ │ ├── 00000/
│ │ │ ├── first_frame.png # Input image (condition frame)
│ │ │ ├── prompt.txt # Text prompt
│ │ │ ├── ground_truth.mp4 # Reference video
│ │ │ ├── final_frame.png # Last frame of the reference video
│ │ │ └── metadata.json # Task parameters and symbolic ground truth
│ │ ├── 00001/
│ │ └── ... # 5 instances per task
│ └── ... # 50 tasks
└── Out-of-Domain_50/
└── ... # 50 tasks

The image setting has the same shape, with the reference output as an image sequence instead of a video:

/path/to/VBVR-Pro-Bench-Image/
├── In-Domain_50/
│ ├── G-131_select_next_figure_.../
│ │ ├── 00000/
│ │ │ ├── first_frame.png # Input image (condition frame)
│ │ │ ├── prompt.txt # Text prompt
│ │ │ ├── frame_1.png # Reference output, step 1
│ │ │ ├── frame_2.png # ... step 2, when the task has more steps
│ │ │ ├── ... # up to frame_N.png
│ │ │ └── metadata.json # Task parameters and symbolic ground truth
│ │ └── ...
│ └── ...
└── Out-of-Domain_50/
└── ...

Note:metadata.json is required. It carries the symbolic ground truth (grid layouts, object positions, correct answers) that the evaluators check against — scoring will not work without it.

1.3 Generate Model Output (Inference)

For each instance, condition on first_frame.png and prompt.txt and produce one video (or one image sequence).

Video setting — one .mp4 per instance:

/path/to/model_outputs/
├── In-Domain_50/
│ ├── G-131_select_next_figure_.../
│ │ ├── 00000.mp4 # Generated video for instance 0
│ │ ├── 00001.mp4
│ │ ├── 00002.mp4
│ │ ├── 00003.mp4
│ │ └── 00004.mp4 # 5 videos per task
│ └── ... # Same task folders as GT
└── Out-of-Domain_50/
└── ... # Same task folders as GT

Image setting — one directory per instance, holding its output steps:

/path/to/model_outputs/
├── In-Domain_50/
│ ├── G-131_select_next_figure_.../
│ │ ├── 00000/
│ │ │ ├── 001.png # Step 1
│ │ │ ├── 002.png # Step 2, if the task has more steps
│ │ │ └── ...
│ │ ├── 00001/
│ │ └── ... # 5 instances per task
│ └── ...
└── Out-of-Domain_50/
└── ...

Note: The folder names (In-Domain_50/, Out-of-Domain_50/, task names) and instance ids (0000000004) must match the ground truth structure exactly. Only .mp4 is recognised in the video setting.

Every .png inside an instance directory is one output step, ordered by the number in its filename — so 1.png, 2.png, 10.png and 001.png, 002.png, 010.png both order correctly. Single-step tasks just have one file.

1.4 Run Evaluation

Video

python run_evaluation_video.py \
--model_path /path/to/model_outputs \
--gt_base /path/to/VBVR-Pro-Bench-Video \
--output_dir ./video_eval_results

Image

python run_evaluation_image.py \
--model_path /path/to/model_outputs \
--gt_image_base /path/to/VBVR-Pro-Bench-Image \
--output_dir ./image_eval_results

Batch evaluation (multiple models). Point at the directory that contains them, and optionally restrict the list:

python run_evaluation_video.py \
--models_base /path/to/all_model_outputs \
--gt_base /path/to/VBVR-Pro-Bench-Video
python run_evaluation_video.py \
--models_base /path/to/all_model_outputs \
--gt_base /path/to/VBVR-Pro-Bench-Video \
--models model_A model_B

Arguments

ArgumentApplies toDescription
--model_pathbothPath to a single model's output directory
--models_basebothBase directory containing multiple model folders
--modelsbothSpecific model names to evaluate (with --models_base)
--gt_basevideo(Required) Path to VBVR-Pro-Bench-Video
--gt_image_baseimage(Required) Path to VBVR-Pro-Bench-Image
--output_dirbothOutput directory for results
--devicebothcuda or cpu (default: cuda); affects OCR only

--model_path and --models_base are mutually exclusive, and one is required.

Use a GPU if you have one. The geometric and graph algorithms are CPU-only, but the OCR step runs on the GPU when --device cuda (the default) is set, and is noticeably slower without it. --device cpu still works if no GPU is available.


2. Output Format

Each model produces {output_dir}/{model_name}_vbvr_results.json. With --models_base, the video entry point additionally writes all_models_summary.json alongside it.

{
"model_name": "my_model",
"summary": {
"In_Domain": { "mean_score": 0.48, "num_samples": 250, "by_task": {}, "by_category": {} },
"Out_of_Domain": { "mean_score": 0.65, "num_samples": 250, "by_task": {}, "by_category": {} },
"overall": { "mean_score": 0.56, "num_samples": 500, "by_task": {}, "by_category": {} }
},
"samples": [
{
"task_name": "G-131_select_next_...",
"video_file": "00000.mp4",
"folder": "In-Domain_50",
"split": "In_Domain",
"category": "Abstraction",
"score": 0.85,
"dimensions": { "task_specific": 0.85 },
"error": null
}
]
}

Aggregates are reported per split, per task, and per cognitive category. The 100 tasks break down as:

CategoryTasks
Perception28
Abstraction25
Knowledge19
Spatiality16
Transformation12

Report Out_of_Domain for generalization: those task families do not appear in the VBVR-Pro training splits.


3. Repository Structure

VBVR-Pro-Bench/
├── run_evaluation_video.py # Entry point: video (I2V) setting
├── run_evaluation_image.py # Entry point: interleaved image setting
├── requirements.txt # Python dependencies
└── vbvr_bench/
├── __init__.py # VBVRBench class
├── utils.py # Shared CV primitives (color/shape/frame ops)
└── evaluators/
├── __init__.py # Evaluator registry, task categories, split definitions
├── base_evaluator.py # BaseEvaluator: dimensions, weights, frame handling
├── In_Domain_50_part1..5.py # 50 in-domain task evaluators
└── Out_of_Domain_50_part1..5.py # 50 out-of-domain task evaluators

TASK_EVALUATOR_MAP in vbvr_bench/evaluators/__init__.py maps each of the 100 task names to its evaluator class. To score a single instance directly:

fromvbvr_bench.evaluatorsimportget_evaluatortask="G-45_key_door_matching_data-generator"gt_dir=f"/path/to/VBVR-Pro-Bench-Video/In-Domain_50/{task}/00000"evaluator=get_evaluator(task, device="cpu")
result=evaluator.evaluate({
"video_path": "pred/00000.mp4",
"task_name": task,
"gt_path": gt_dir,
"metafile_path": [f"{gt_dir}/metadata.json"],
}, task_specific_only=True)
print(result["score"]) # float in [0, 1]print(result["details"]) # per-evaluator diagnostics

License

VBVR-Pro source code, scripts, configuration files and task-specific scoring software — including everything in this repository — are licensed under the Apache License 2.0. VBVR-Pro data and benchmark materials are separately licensed under CC BY-NC 4.0. Model weights and third-party materials remain subject to their applicable model-card and upstream terms. See LICENSE.md for details.

Citation

@misc{xu2026vbvrproscalableverifiablesuite,
title={VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning},
author={Junxiang Xu and Ruisi Wang and Fanyi Pu and Maijunxian Wang and Ran Ji and Tongxi Zhou and Chenyang Gu and Jing Zuo and Hongcan Xiao and Yimeng Geng and Wanqi Yin and Wei Chen and Oscar Qian and Zhengan Yan and Ziqi Huang and Haiwen Diao and Liang Pan and Bo Li and Xiangyu Fan and Dezhi Luo and Fengyuan Yu and Zehong Zhao and Qingying Gao and Tinghui Zhu and Yilan Zhang and Jingqi Tong and Pinyuan Feng and Zhengze Jiang and Letian Wang and Ziyu Guo and Renrui Zhang and Jieneng Chen and Sonia Joseph and Constantin Venhoff and Saman Motamed and Mengyue Yang and Chandra Sripada and Alan Yuille and Philip Torr and Lvmin Zhang and Vikash Kumar and Daniel Khashabi and Nikolaus Kriegeskorte and Rapha\"{e}l Milli\`{e}re and Vincent C. M\"{u}ller and Anyi Rao and Quan Wang and Ziwei Liu and Dahua Lin and Lei Yang and Hokin Deng and Zhongang Cai},
year={2026},
eprint={2608.26105},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2608.26105},
}

About

Rule-based evaluation kit for VBVR-Pro-Bench — 100 visual-reasoning tasks, one hand-written evaluator each

Topics

Resources

Stars

12 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages