Skip to content

Repository files navigation

📐 GAGE: General AI evaluation and Gauge Engine

PythonCode StyleLicenseStatus

English · 中文

📧 Contact:zhangrongjunchen@myhexin.com

Overview · Sample Schema · Smart Defaults · Run Reports · Game Arena · Arena Visual Control · AgentKitV2 · External Harness · Benchmark · Contributing · Standards


GAGE is a unified, extensible evaluation framework for large language models, multimodal models, audio models, diffusion models, agents, and game environments. It provides one evaluation engine for datasets, model backends, metrics, arena runtimes, structured outputs, and replayable artifacts.

Game Arena Showcase

Gomoku GameArena demoDoudizhu GameArena demoMahjong GameArena demo

Space Invaders demoMario demoVizDoom demo

Why GAGE?

  • Fast evaluation engine: Run local smoke tests, model-backed jobs, and larger benchmark batches through the same pipeline shape.
  • Unified evaluation surface: Datasets, backends, role adapters, metrics, and output contracts are configured instead of hand-wired per benchmark.
  • Game and agent sandboxing: Game Arena, AgentKitV2, AppWorld, SWE-bench-style agent tasks, GUI interaction, and tool-augmented workflows share the same run/output model.
  • External harness integration: Delegate task-batch benchmarks to Harbor, then import trial evidence back into standard GAGE samples, metrics, reports, and raw artifacts.
  • Replayable GameKit runtime: Gomoku, Tic-Tac-Toe, Doudizhu, Mahjong, PettingZoo Space Invaders, Retro Mario, and ViZDoom now emit structured arena traces plus arena_visual sessions.
  • Operational visibility: Runs write summary.json, sample outputs, logs, visual artifacts, and a static report pack so failures can be inspected after the fact.

Design Overview

Core design philosophy: everything is a step, everything is configurable.

Architecture Design

End-to-end flow

Orchestration Design

Step view

Game Arena Design

GameArena runtime core design

AgentKitV2 Design

AgentKit v2 pipeline design

External Harness Design

External Harness (Harbor) pipeline design

Quick Start

1. Installation

# If you are in the mono-repo root:cd gage-eval-main
# Python 3.10+ recommended
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

For Game Arena LLM configs, use the *_openai_gamekit.yaml variants and export OPENAI_API_KEY. The model defaults to gpt-5.4; set GAGE_GAME_ARENA_LLM_MODEL to override it, or set OPENAI_API_BASE for an OpenAI-compatible endpoint.

2. Run a Basic Demo

python run.py \
--config config/run_configs/demo_echo_run_1.yaml \
--output-dir runs \
--run-id demo_echo

3. View Reports

Default output structure:

runs/<run_id>/
events.jsonl
samples.jsonl
summary.json
samples/
<namespace>/
<sample_id>.json
report_pack/
report.html
report_context.json
report_context.md
prompt.txt
diagnostics.json
assets_manifest.json

Open runs/<run_id>/report_pack/report.html for the execution-aware report: primary metrics, key findings, scenario profiles, evidence links, media previews, diagnostics, and reason-code explanations. See Run Reports.

Advanced Configurations

ScenarioConfig ExampleDescription
GameArena Human-vs-AIconfig/custom/doudizhu/doudizhu_human_visual_gamekit.yamlBrowser-controlled Doudizhu match against LLM players
GameArena Pure Human Controlconfig/custom/retro_mario/retro_mario_human_visual_gamekit.yamlBrowser-controlled real-time Retro Mario session
AgentKitV2 Tau2config/custom/manual_e2e/agentkit_v2_tau2_local_lmstudio.yamlNative per-sample local-process Tau2 1-case smoke run
AgentKitV2 SWE-bench Proconfig/custom/manual_e2e/agentkit_v2_swebench_pro_docker_lmstudio_smoke1_qutebrowser.yamlNative Docker-backed SWE-bench Pro smoke run
External Harness Harborconfig/custom/external_harness_kits/harbor_terminal_bench2_lmstudio_1case.yamlDelegates a Terminal-Bench 2.0 task to Harbor and imports results
AgentKitV2 AppWorldconfig/custom/appworld/appworld_official_jsonl.yamlAppWorld sandbox evaluation through the native AgentKitV2 path
Textconfig/custom/aime24/aime2024_chat.yamlAIME, GPQA, Math500, and related text benchmarks
Multimodalconfig/custom/mathvista/chat.yamlMathVista and related multimodal benchmarks
LLM Judgeconfig/custom/examples/single_task_local_judge_qwen.yamlLocal LLM judge example

Roadmap

  • Agent evaluation: Continue hardening AgentKitV2 and External Harness trace import, failure diagnostics, and reproducible live smoke configs.
  • Game Arena expansion: Grow the GameKit catalog and keep browser control, replay, and output contracts consistent.
  • Gage-Client: Add a client tool for configuration management, failure diagnostics, and benchmark onboarding.
  • Distributed inference: Support multi-node task sharding and load balancing for large runs.
  • Benchmark expansion: Continue adding benchmark configs, metrics, and troubleshooting guidance.

Status

This project is in internal validation; APIs, configs, and docs may change rapidly.

About

General AI evaluation and Gauge Engine. A unified evaluation engine for LLMs, MLLMs, audio, and diffusion models.

Topics

Resources

Contributing

Stars

52 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - HiThink-Research/GAGE: General AI evaluation and Gauge Engine. A unified evaluation engine for LLMs, MLLMs, audio, and diffusion models. · GitHub
Skip to content

Repository files navigation

📐 GAGE: General AI evaluation and Gauge Engine

PythonCode StyleLicenseStatus

English · 中文

📧 Contact:zhangrongjunchen@myhexin.com

Overview · Sample Schema · Smart Defaults · Run Reports · Game Arena · Arena Visual Control · AgentKitV2 · External Harness · Benchmark · Contributing · Standards


GAGE is a unified, extensible evaluation framework for large language models, multimodal models, audio models, diffusion models, agents, and game environments. It provides one evaluation engine for datasets, model backends, metrics, arena runtimes, structured outputs, and replayable artifacts.

Game Arena Showcase

Gomoku GameArena demoDoudizhu GameArena demoMahjong GameArena demo

Space Invaders demoMario demoVizDoom demo

Why GAGE?

  • Fast evaluation engine: Run local smoke tests, model-backed jobs, and larger benchmark batches through the same pipeline shape.
  • Unified evaluation surface: Datasets, backends, role adapters, metrics, and output contracts are configured instead of hand-wired per benchmark.
  • Game and agent sandboxing: Game Arena, AgentKitV2, AppWorld, SWE-bench-style agent tasks, GUI interaction, and tool-augmented workflows share the same run/output model.
  • External harness integration: Delegate task-batch benchmarks to Harbor, then import trial evidence back into standard GAGE samples, metrics, reports, and raw artifacts.
  • Replayable GameKit runtime: Gomoku, Tic-Tac-Toe, Doudizhu, Mahjong, PettingZoo Space Invaders, Retro Mario, and ViZDoom now emit structured arena traces plus arena_visual sessions.
  • Operational visibility: Runs write summary.json, sample outputs, logs, visual artifacts, and a static report pack so failures can be inspected after the fact.

Design Overview

Core design philosophy: everything is a step, everything is configurable.

Architecture Design

End-to-end flow

Orchestration Design

Step view

Game Arena Design

GameArena runtime core design

AgentKitV2 Design

AgentKit v2 pipeline design

External Harness Design

External Harness (Harbor) pipeline design

Quick Start

1. Installation

# If you are in the mono-repo root:cd gage-eval-main
# Python 3.10+ recommended
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

For Game Arena LLM configs, use the *_openai_gamekit.yaml variants and export OPENAI_API_KEY. The model defaults to gpt-5.4; set GAGE_GAME_ARENA_LLM_MODEL to override it, or set OPENAI_API_BASE for an OpenAI-compatible endpoint.

2. Run a Basic Demo

python run.py \
--config config/run_configs/demo_echo_run_1.yaml \
--output-dir runs \
--run-id demo_echo

3. View Reports

Default output structure:

runs/<run_id>/
events.jsonl
samples.jsonl
summary.json
samples/
<namespace>/
<sample_id>.json
report_pack/
report.html
report_context.json
report_context.md
prompt.txt
diagnostics.json
assets_manifest.json

Open runs/<run_id>/report_pack/report.html for the execution-aware report: primary metrics, key findings, scenario profiles, evidence links, media previews, diagnostics, and reason-code explanations. See Run Reports.

Advanced Configurations

ScenarioConfig ExampleDescription
GameArena Human-vs-AIconfig/custom/doudizhu/doudizhu_human_visual_gamekit.yamlBrowser-controlled Doudizhu match against LLM players
GameArena Pure Human Controlconfig/custom/retro_mario/retro_mario_human_visual_gamekit.yamlBrowser-controlled real-time Retro Mario session
AgentKitV2 Tau2config/custom/manual_e2e/agentkit_v2_tau2_local_lmstudio.yamlNative per-sample local-process Tau2 1-case smoke run
AgentKitV2 SWE-bench Proconfig/custom/manual_e2e/agentkit_v2_swebench_pro_docker_lmstudio_smoke1_qutebrowser.yamlNative Docker-backed SWE-bench Pro smoke run
External Harness Harborconfig/custom/external_harness_kits/harbor_terminal_bench2_lmstudio_1case.yamlDelegates a Terminal-Bench 2.0 task to Harbor and imports results
AgentKitV2 AppWorldconfig/custom/appworld/appworld_official_jsonl.yamlAppWorld sandbox evaluation through the native AgentKitV2 path
Textconfig/custom/aime24/aime2024_chat.yamlAIME, GPQA, Math500, and related text benchmarks
Multimodalconfig/custom/mathvista/chat.yamlMathVista and related multimodal benchmarks
LLM Judgeconfig/custom/examples/single_task_local_judge_qwen.yamlLocal LLM judge example

Roadmap

  • Agent evaluation: Continue hardening AgentKitV2 and External Harness trace import, failure diagnostics, and reproducible live smoke configs.
  • Game Arena expansion: Grow the GameKit catalog and keep browser control, replay, and output contracts consistent.
  • Gage-Client: Add a client tool for configuration management, failure diagnostics, and benchmark onboarding.
  • Distributed inference: Support multi-node task sharding and load balancing for large runs.
  • Benchmark expansion: Continue adding benchmark configs, metrics, and troubleshooting guidance.

Status

This project is in internal validation; APIs, configs, and docs may change rapidly.

About

General AI evaluation and Gauge Engine. A unified evaluation engine for LLMs, MLLMs, audio, and diffusion models.

Topics

Resources

Contributing

Stars

52 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - HiThink-Research/GAGE: General AI evaluation and Gauge Engine. A unified evaluation engine for LLMs, MLLMs, audio, and diffusion models. · GitHub
Skip to content

Repository files navigation

📐 GAGE: General AI evaluation and Gauge Engine

PythonCode StyleLicenseStatus

English · 中文

📧 Contact:zhangrongjunchen@myhexin.com

Overview · Sample Schema · Smart Defaults · Run Reports · Game Arena · Arena Visual Control · AgentKitV2 · External Harness · Benchmark · Contributing · Standards


GAGE is a unified, extensible evaluation framework for large language models, multimodal models, audio models, diffusion models, agents, and game environments. It provides one evaluation engine for datasets, model backends, metrics, arena runtimes, structured outputs, and replayable artifacts.

Game Arena Showcase

Gomoku GameArena demoDoudizhu GameArena demoMahjong GameArena demo

Space Invaders demoMario demoVizDoom demo

Why GAGE?

  • Fast evaluation engine: Run local smoke tests, model-backed jobs, and larger benchmark batches through the same pipeline shape.
  • Unified evaluation surface: Datasets, backends, role adapters, metrics, and output contracts are configured instead of hand-wired per benchmark.
  • Game and agent sandboxing: Game Arena, AgentKitV2, AppWorld, SWE-bench-style agent tasks, GUI interaction, and tool-augmented workflows share the same run/output model.
  • External harness integration: Delegate task-batch benchmarks to Harbor, then import trial evidence back into standard GAGE samples, metrics, reports, and raw artifacts.
  • Replayable GameKit runtime: Gomoku, Tic-Tac-Toe, Doudizhu, Mahjong, PettingZoo Space Invaders, Retro Mario, and ViZDoom now emit structured arena traces plus arena_visual sessions.
  • Operational visibility: Runs write summary.json, sample outputs, logs, visual artifacts, and a static report pack so failures can be inspected after the fact.

Design Overview

Core design philosophy: everything is a step, everything is configurable.

Architecture Design

End-to-end flow

Orchestration Design

Step view

Game Arena Design

GameArena runtime core design

AgentKitV2 Design

AgentKit v2 pipeline design

External Harness Design

External Harness (Harbor) pipeline design

Quick Start

1. Installation

# If you are in the mono-repo root:cd gage-eval-main
# Python 3.10+ recommended
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

For Game Arena LLM configs, use the *_openai_gamekit.yaml variants and export OPENAI_API_KEY. The model defaults to gpt-5.4; set GAGE_GAME_ARENA_LLM_MODEL to override it, or set OPENAI_API_BASE for an OpenAI-compatible endpoint.

2. Run a Basic Demo

python run.py \
--config config/run_configs/demo_echo_run_1.yaml \
--output-dir runs \
--run-id demo_echo

3. View Reports

Default output structure:

runs/<run_id>/
events.jsonl
samples.jsonl
summary.json
samples/
<namespace>/
<sample_id>.json
report_pack/
report.html
report_context.json
report_context.md
prompt.txt
diagnostics.json
assets_manifest.json

Open runs/<run_id>/report_pack/report.html for the execution-aware report: primary metrics, key findings, scenario profiles, evidence links, media previews, diagnostics, and reason-code explanations. See Run Reports.

Advanced Configurations

ScenarioConfig ExampleDescription
GameArena Human-vs-AIconfig/custom/doudizhu/doudizhu_human_visual_gamekit.yamlBrowser-controlled Doudizhu match against LLM players
GameArena Pure Human Controlconfig/custom/retro_mario/retro_mario_human_visual_gamekit.yamlBrowser-controlled real-time Retro Mario session
AgentKitV2 Tau2config/custom/manual_e2e/agentkit_v2_tau2_local_lmstudio.yamlNative per-sample local-process Tau2 1-case smoke run
AgentKitV2 SWE-bench Proconfig/custom/manual_e2e/agentkit_v2_swebench_pro_docker_lmstudio_smoke1_qutebrowser.yamlNative Docker-backed SWE-bench Pro smoke run
External Harness Harborconfig/custom/external_harness_kits/harbor_terminal_bench2_lmstudio_1case.yamlDelegates a Terminal-Bench 2.0 task to Harbor and imports results
AgentKitV2 AppWorldconfig/custom/appworld/appworld_official_jsonl.yamlAppWorld sandbox evaluation through the native AgentKitV2 path
Textconfig/custom/aime24/aime2024_chat.yamlAIME, GPQA, Math500, and related text benchmarks
Multimodalconfig/custom/mathvista/chat.yamlMathVista and related multimodal benchmarks
LLM Judgeconfig/custom/examples/single_task_local_judge_qwen.yamlLocal LLM judge example

Roadmap

  • Agent evaluation: Continue hardening AgentKitV2 and External Harness trace import, failure diagnostics, and reproducible live smoke configs.
  • Game Arena expansion: Grow the GameKit catalog and keep browser control, replay, and output contracts consistent.
  • Gage-Client: Add a client tool for configuration management, failure diagnostics, and benchmark onboarding.
  • Distributed inference: Support multi-node task sharding and load balancing for large runs.
  • Benchmark expansion: Continue adding benchmark configs, metrics, and troubleshooting guidance.

Status

This project is in internal validation; APIs, configs, and docs may change rapidly.

About

General AI evaluation and Gauge Engine. A unified evaluation engine for LLMs, MLLMs, audio, and diffusion models.

Topics

Resources

Contributing

Stars

52 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - HiThink-Research/GAGE: General AI evaluation and Gauge Engine. A unified evaluation engine for LLMs, MLLMs, audio, and diffusion models. · GitHub
Skip to content

Repository files navigation

📐 GAGE: General AI evaluation and Gauge Engine

PythonCode StyleLicenseStatus

English · 中文

📧 Contact:zhangrongjunchen@myhexin.com

Overview · Sample Schema · Smart Defaults · Run Reports · Game Arena · Arena Visual Control · AgentKitV2 · External Harness · Benchmark · Contributing · Standards


GAGE is a unified, extensible evaluation framework for large language models, multimodal models, audio models, diffusion models, agents, and game environments. It provides one evaluation engine for datasets, model backends, metrics, arena runtimes, structured outputs, and replayable artifacts.

Game Arena Showcase

Gomoku GameArena demoDoudizhu GameArena demoMahjong GameArena demo

Space Invaders demoMario demoVizDoom demo

Why GAGE?

  • Fast evaluation engine: Run local smoke tests, model-backed jobs, and larger benchmark batches through the same pipeline shape.
  • Unified evaluation surface: Datasets, backends, role adapters, metrics, and output contracts are configured instead of hand-wired per benchmark.
  • Game and agent sandboxing: Game Arena, AgentKitV2, AppWorld, SWE-bench-style agent tasks, GUI interaction, and tool-augmented workflows share the same run/output model.
  • External harness integration: Delegate task-batch benchmarks to Harbor, then import trial evidence back into standard GAGE samples, metrics, reports, and raw artifacts.
  • Replayable GameKit runtime: Gomoku, Tic-Tac-Toe, Doudizhu, Mahjong, PettingZoo Space Invaders, Retro Mario, and ViZDoom now emit structured arena traces plus arena_visual sessions.
  • Operational visibility: Runs write summary.json, sample outputs, logs, visual artifacts, and a static report pack so failures can be inspected after the fact.

Design Overview

Core design philosophy: everything is a step, everything is configurable.

Architecture Design

End-to-end flow

Orchestration Design

Step view

Game Arena Design

GameArena runtime core design

AgentKitV2 Design

AgentKit v2 pipeline design

External Harness Design

External Harness (Harbor) pipeline design

Quick Start

1. Installation

# If you are in the mono-repo root:cd gage-eval-main
# Python 3.10+ recommended
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

For Game Arena LLM configs, use the *_openai_gamekit.yaml variants and export OPENAI_API_KEY. The model defaults to gpt-5.4; set GAGE_GAME_ARENA_LLM_MODEL to override it, or set OPENAI_API_BASE for an OpenAI-compatible endpoint.

2. Run a Basic Demo

python run.py \
--config config/run_configs/demo_echo_run_1.yaml \
--output-dir runs \
--run-id demo_echo

3. View Reports

Default output structure:

runs/<run_id>/
events.jsonl
samples.jsonl
summary.json
samples/
<namespace>/
<sample_id>.json
report_pack/
report.html
report_context.json
report_context.md
prompt.txt
diagnostics.json
assets_manifest.json

Open runs/<run_id>/report_pack/report.html for the execution-aware report: primary metrics, key findings, scenario profiles, evidence links, media previews, diagnostics, and reason-code explanations. See Run Reports.

Advanced Configurations

ScenarioConfig ExampleDescription
GameArena Human-vs-AIconfig/custom/doudizhu/doudizhu_human_visual_gamekit.yamlBrowser-controlled Doudizhu match against LLM players
GameArena Pure Human Controlconfig/custom/retro_mario/retro_mario_human_visual_gamekit.yamlBrowser-controlled real-time Retro Mario session
AgentKitV2 Tau2config/custom/manual_e2e/agentkit_v2_tau2_local_lmstudio.yamlNative per-sample local-process Tau2 1-case smoke run
AgentKitV2 SWE-bench Proconfig/custom/manual_e2e/agentkit_v2_swebench_pro_docker_lmstudio_smoke1_qutebrowser.yamlNative Docker-backed SWE-bench Pro smoke run
External Harness Harborconfig/custom/external_harness_kits/harbor_terminal_bench2_lmstudio_1case.yamlDelegates a Terminal-Bench 2.0 task to Harbor and imports results
AgentKitV2 AppWorldconfig/custom/appworld/appworld_official_jsonl.yamlAppWorld sandbox evaluation through the native AgentKitV2 path
Textconfig/custom/aime24/aime2024_chat.yamlAIME, GPQA, Math500, and related text benchmarks
Multimodalconfig/custom/mathvista/chat.yamlMathVista and related multimodal benchmarks
LLM Judgeconfig/custom/examples/single_task_local_judge_qwen.yamlLocal LLM judge example

Roadmap

  • Agent evaluation: Continue hardening AgentKitV2 and External Harness trace import, failure diagnostics, and reproducible live smoke configs.
  • Game Arena expansion: Grow the GameKit catalog and keep browser control, replay, and output contracts consistent.
  • Gage-Client: Add a client tool for configuration management, failure diagnostics, and benchmark onboarding.
  • Distributed inference: Support multi-node task sharding and load balancing for large runs.
  • Benchmark expansion: Continue adding benchmark configs, metrics, and troubleshooting guidance.

Status

This project is in internal validation; APIs, configs, and docs may change rapidly.

About

General AI evaluation and Gauge Engine. A unified evaluation engine for LLMs, MLLMs, audio, and diffusion models.

Topics

Resources

Contributing

Stars

52 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - HiThink-Research/GAGE: General AI evaluation and Gauge Engine. A unified evaluation engine for LLMs, MLLMs, audio, and diffusion models. · GitHub
Skip to content

Repository files navigation

📐 GAGE: General AI evaluation and Gauge Engine

PythonCode StyleLicenseStatus

English · 中文

📧 Contact:zhangrongjunchen@myhexin.com

Overview · Sample Schema · Smart Defaults · Run Reports · Game Arena · Arena Visual Control · AgentKitV2 · External Harness · Benchmark · Contributing · Standards


GAGE is a unified, extensible evaluation framework for large language models, multimodal models, audio models, diffusion models, agents, and game environments. It provides one evaluation engine for datasets, model backends, metrics, arena runtimes, structured outputs, and replayable artifacts.

Game Arena Showcase

Gomoku GameArena demoDoudizhu GameArena demoMahjong GameArena demo

Space Invaders demoMario demoVizDoom demo

Why GAGE?

  • Fast evaluation engine: Run local smoke tests, model-backed jobs, and larger benchmark batches through the same pipeline shape.
  • Unified evaluation surface: Datasets, backends, role adapters, metrics, and output contracts are configured instead of hand-wired per benchmark.
  • Game and agent sandboxing: Game Arena, AgentKitV2, AppWorld, SWE-bench-style agent tasks, GUI interaction, and tool-augmented workflows share the same run/output model.
  • External harness integration: Delegate task-batch benchmarks to Harbor, then import trial evidence back into standard GAGE samples, metrics, reports, and raw artifacts.
  • Replayable GameKit runtime: Gomoku, Tic-Tac-Toe, Doudizhu, Mahjong, PettingZoo Space Invaders, Retro Mario, and ViZDoom now emit structured arena traces plus arena_visual sessions.
  • Operational visibility: Runs write summary.json, sample outputs, logs, visual artifacts, and a static report pack so failures can be inspected after the fact.

Design Overview

Core design philosophy: everything is a step, everything is configurable.

Architecture Design

End-to-end flow

Orchestration Design

Step view

Game Arena Design

GameArena runtime core design

AgentKitV2 Design

AgentKit v2 pipeline design

External Harness Design

External Harness (Harbor) pipeline design

Quick Start

1. Installation

# If you are in the mono-repo root:cd gage-eval-main
# Python 3.10+ recommended
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

For Game Arena LLM configs, use the *_openai_gamekit.yaml variants and export OPENAI_API_KEY. The model defaults to gpt-5.4; set GAGE_GAME_ARENA_LLM_MODEL to override it, or set OPENAI_API_BASE for an OpenAI-compatible endpoint.

2. Run a Basic Demo

python run.py \
--config config/run_configs/demo_echo_run_1.yaml \
--output-dir runs \
--run-id demo_echo

3. View Reports

Default output structure:

runs/<run_id>/
events.jsonl
samples.jsonl
summary.json
samples/
<namespace>/
<sample_id>.json
report_pack/
report.html
report_context.json
report_context.md
prompt.txt
diagnostics.json
assets_manifest.json

Open runs/<run_id>/report_pack/report.html for the execution-aware report: primary metrics, key findings, scenario profiles, evidence links, media previews, diagnostics, and reason-code explanations. See Run Reports.

Advanced Configurations

ScenarioConfig ExampleDescription
GameArena Human-vs-AIconfig/custom/doudizhu/doudizhu_human_visual_gamekit.yamlBrowser-controlled Doudizhu match against LLM players
GameArena Pure Human Controlconfig/custom/retro_mario/retro_mario_human_visual_gamekit.yamlBrowser-controlled real-time Retro Mario session
AgentKitV2 Tau2config/custom/manual_e2e/agentkit_v2_tau2_local_lmstudio.yamlNative per-sample local-process Tau2 1-case smoke run
AgentKitV2 SWE-bench Proconfig/custom/manual_e2e/agentkit_v2_swebench_pro_docker_lmstudio_smoke1_qutebrowser.yamlNative Docker-backed SWE-bench Pro smoke run
External Harness Harborconfig/custom/external_harness_kits/harbor_terminal_bench2_lmstudio_1case.yamlDelegates a Terminal-Bench 2.0 task to Harbor and imports results
AgentKitV2 AppWorldconfig/custom/appworld/appworld_official_jsonl.yamlAppWorld sandbox evaluation through the native AgentKitV2 path
Textconfig/custom/aime24/aime2024_chat.yamlAIME, GPQA, Math500, and related text benchmarks
Multimodalconfig/custom/mathvista/chat.yamlMathVista and related multimodal benchmarks
LLM Judgeconfig/custom/examples/single_task_local_judge_qwen.yamlLocal LLM judge example

Roadmap

  • Agent evaluation: Continue hardening AgentKitV2 and External Harness trace import, failure diagnostics, and reproducible live smoke configs.
  • Game Arena expansion: Grow the GameKit catalog and keep browser control, replay, and output contracts consistent.
  • Gage-Client: Add a client tool for configuration management, failure diagnostics, and benchmark onboarding.
  • Distributed inference: Support multi-node task sharding and load balancing for large runs.
  • Benchmark expansion: Continue adding benchmark configs, metrics, and troubleshooting guidance.

Status

This project is in internal validation; APIs, configs, and docs may change rapidly.

About

General AI evaluation and Gauge Engine. A unified evaluation engine for LLMs, MLLMs, audio, and diffusion models.

Topics

Resources

Contributing

Stars

52 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - HiThink-Research/GAGE: General AI evaluation and Gauge Engine. A unified evaluation engine for LLMs, MLLMs, audio, and diffusion models. · GitHub
Skip to content

Repository files navigation

📐 GAGE: General AI evaluation and Gauge Engine

PythonCode StyleLicenseStatus

English · 中文

📧 Contact:zhangrongjunchen@myhexin.com

Overview · Sample Schema · Smart Defaults · Run Reports · Game Arena · Arena Visual Control · AgentKitV2 · External Harness · Benchmark · Contributing · Standards


GAGE is a unified, extensible evaluation framework for large language models, multimodal models, audio models, diffusion models, agents, and game environments. It provides one evaluation engine for datasets, model backends, metrics, arena runtimes, structured outputs, and replayable artifacts.

Game Arena Showcase

Gomoku GameArena demoDoudizhu GameArena demoMahjong GameArena demo

Space Invaders demoMario demoVizDoom demo

Why GAGE?

  • Fast evaluation engine: Run local smoke tests, model-backed jobs, and larger benchmark batches through the same pipeline shape.
  • Unified evaluation surface: Datasets, backends, role adapters, metrics, and output contracts are configured instead of hand-wired per benchmark.
  • Game and agent sandboxing: Game Arena, AgentKitV2, AppWorld, SWE-bench-style agent tasks, GUI interaction, and tool-augmented workflows share the same run/output model.
  • External harness integration: Delegate task-batch benchmarks to Harbor, then import trial evidence back into standard GAGE samples, metrics, reports, and raw artifacts.
  • Replayable GameKit runtime: Gomoku, Tic-Tac-Toe, Doudizhu, Mahjong, PettingZoo Space Invaders, Retro Mario, and ViZDoom now emit structured arena traces plus arena_visual sessions.
  • Operational visibility: Runs write summary.json, sample outputs, logs, visual artifacts, and a static report pack so failures can be inspected after the fact.

Design Overview

Core design philosophy: everything is a step, everything is configurable.

Architecture Design

End-to-end flow

Orchestration Design

Step view

Game Arena Design

GameArena runtime core design

AgentKitV2 Design

AgentKit v2 pipeline design

External Harness Design

External Harness (Harbor) pipeline design

Quick Start

1. Installation

# If you are in the mono-repo root:cd gage-eval-main
# Python 3.10+ recommended
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

For Game Arena LLM configs, use the *_openai_gamekit.yaml variants and export OPENAI_API_KEY. The model defaults to gpt-5.4; set GAGE_GAME_ARENA_LLM_MODEL to override it, or set OPENAI_API_BASE for an OpenAI-compatible endpoint.

2. Run a Basic Demo

python run.py \
--config config/run_configs/demo_echo_run_1.yaml \
--output-dir runs \
--run-id demo_echo

3. View Reports

Default output structure:

runs/<run_id>/
events.jsonl
samples.jsonl
summary.json
samples/
<namespace>/
<sample_id>.json
report_pack/
report.html
report_context.json
report_context.md
prompt.txt
diagnostics.json
assets_manifest.json

Open runs/<run_id>/report_pack/report.html for the execution-aware report: primary metrics, key findings, scenario profiles, evidence links, media previews, diagnostics, and reason-code explanations. See Run Reports.

Advanced Configurations

ScenarioConfig ExampleDescription
GameArena Human-vs-AIconfig/custom/doudizhu/doudizhu_human_visual_gamekit.yamlBrowser-controlled Doudizhu match against LLM players
GameArena Pure Human Controlconfig/custom/retro_mario/retro_mario_human_visual_gamekit.yamlBrowser-controlled real-time Retro Mario session
AgentKitV2 Tau2config/custom/manual_e2e/agentkit_v2_tau2_local_lmstudio.yamlNative per-sample local-process Tau2 1-case smoke run
AgentKitV2 SWE-bench Proconfig/custom/manual_e2e/agentkit_v2_swebench_pro_docker_lmstudio_smoke1_qutebrowser.yamlNative Docker-backed SWE-bench Pro smoke run
External Harness Harborconfig/custom/external_harness_kits/harbor_terminal_bench2_lmstudio_1case.yamlDelegates a Terminal-Bench 2.0 task to Harbor and imports results
AgentKitV2 AppWorldconfig/custom/appworld/appworld_official_jsonl.yamlAppWorld sandbox evaluation through the native AgentKitV2 path
Textconfig/custom/aime24/aime2024_chat.yamlAIME, GPQA, Math500, and related text benchmarks
Multimodalconfig/custom/mathvista/chat.yamlMathVista and related multimodal benchmarks
LLM Judgeconfig/custom/examples/single_task_local_judge_qwen.yamlLocal LLM judge example

Roadmap

  • Agent evaluation: Continue hardening AgentKitV2 and External Harness trace import, failure diagnostics, and reproducible live smoke configs.
  • Game Arena expansion: Grow the GameKit catalog and keep browser control, replay, and output contracts consistent.
  • Gage-Client: Add a client tool for configuration management, failure diagnostics, and benchmark onboarding.
  • Distributed inference: Support multi-node task sharding and load balancing for large runs.
  • Benchmark expansion: Continue adding benchmark configs, metrics, and troubleshooting guidance.

Status

This project is in internal validation; APIs, configs, and docs may change rapidly.

About

General AI evaluation and Gauge Engine. A unified evaluation engine for LLMs, MLLMs, audio, and diffusion models.

Topics

Resources

Contributing

Stars

52 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - HiThink-Research/GAGE: General AI evaluation and Gauge Engine. A unified evaluation engine for LLMs, MLLMs, audio, and diffusion models. · GitHub
Skip to content

Repository files navigation

📐 GAGE: General AI evaluation and Gauge Engine

PythonCode StyleLicenseStatus

English · 中文

📧 Contact:zhangrongjunchen@myhexin.com

Overview · Sample Schema · Smart Defaults · Run Reports · Game Arena · Arena Visual Control · AgentKitV2 · External Harness · Benchmark · Contributing · Standards


GAGE is a unified, extensible evaluation framework for large language models, multimodal models, audio models, diffusion models, agents, and game environments. It provides one evaluation engine for datasets, model backends, metrics, arena runtimes, structured outputs, and replayable artifacts.

Game Arena Showcase

Gomoku GameArena demoDoudizhu GameArena demoMahjong GameArena demo

Space Invaders demoMario demoVizDoom demo

Why GAGE?

  • Fast evaluation engine: Run local smoke tests, model-backed jobs, and larger benchmark batches through the same pipeline shape.
  • Unified evaluation surface: Datasets, backends, role adapters, metrics, and output contracts are configured instead of hand-wired per benchmark.
  • Game and agent sandboxing: Game Arena, AgentKitV2, AppWorld, SWE-bench-style agent tasks, GUI interaction, and tool-augmented workflows share the same run/output model.
  • External harness integration: Delegate task-batch benchmarks to Harbor, then import trial evidence back into standard GAGE samples, metrics, reports, and raw artifacts.
  • Replayable GameKit runtime: Gomoku, Tic-Tac-Toe, Doudizhu, Mahjong, PettingZoo Space Invaders, Retro Mario, and ViZDoom now emit structured arena traces plus arena_visual sessions.
  • Operational visibility: Runs write summary.json, sample outputs, logs, visual artifacts, and a static report pack so failures can be inspected after the fact.

Design Overview

Core design philosophy: everything is a step, everything is configurable.

Architecture Design

End-to-end flow

Orchestration Design

Step view

Game Arena Design

GameArena runtime core design

AgentKitV2 Design

AgentKit v2 pipeline design

External Harness Design

External Harness (Harbor) pipeline design

Quick Start

1. Installation

# If you are in the mono-repo root:cd gage-eval-main
# Python 3.10+ recommended
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

For Game Arena LLM configs, use the *_openai_gamekit.yaml variants and export OPENAI_API_KEY. The model defaults to gpt-5.4; set GAGE_GAME_ARENA_LLM_MODEL to override it, or set OPENAI_API_BASE for an OpenAI-compatible endpoint.

2. Run a Basic Demo

python run.py \
--config config/run_configs/demo_echo_run_1.yaml \
--output-dir runs \
--run-id demo_echo

3. View Reports

Default output structure:

runs/<run_id>/
events.jsonl
samples.jsonl
summary.json
samples/
<namespace>/
<sample_id>.json
report_pack/
report.html
report_context.json
report_context.md
prompt.txt
diagnostics.json
assets_manifest.json

Open runs/<run_id>/report_pack/report.html for the execution-aware report: primary metrics, key findings, scenario profiles, evidence links, media previews, diagnostics, and reason-code explanations. See Run Reports.

Advanced Configurations

ScenarioConfig ExampleDescription
GameArena Human-vs-AIconfig/custom/doudizhu/doudizhu_human_visual_gamekit.yamlBrowser-controlled Doudizhu match against LLM players
GameArena Pure Human Controlconfig/custom/retro_mario/retro_mario_human_visual_gamekit.yamlBrowser-controlled real-time Retro Mario session
AgentKitV2 Tau2config/custom/manual_e2e/agentkit_v2_tau2_local_lmstudio.yamlNative per-sample local-process Tau2 1-case smoke run
AgentKitV2 SWE-bench Proconfig/custom/manual_e2e/agentkit_v2_swebench_pro_docker_lmstudio_smoke1_qutebrowser.yamlNative Docker-backed SWE-bench Pro smoke run
External Harness Harborconfig/custom/external_harness_kits/harbor_terminal_bench2_lmstudio_1case.yamlDelegates a Terminal-Bench 2.0 task to Harbor and imports results
AgentKitV2 AppWorldconfig/custom/appworld/appworld_official_jsonl.yamlAppWorld sandbox evaluation through the native AgentKitV2 path
Textconfig/custom/aime24/aime2024_chat.yamlAIME, GPQA, Math500, and related text benchmarks
Multimodalconfig/custom/mathvista/chat.yamlMathVista and related multimodal benchmarks
LLM Judgeconfig/custom/examples/single_task_local_judge_qwen.yamlLocal LLM judge example

Roadmap

  • Agent evaluation: Continue hardening AgentKitV2 and External Harness trace import, failure diagnostics, and reproducible live smoke configs.
  • Game Arena expansion: Grow the GameKit catalog and keep browser control, replay, and output contracts consistent.
  • Gage-Client: Add a client tool for configuration management, failure diagnostics, and benchmark onboarding.
  • Distributed inference: Support multi-node task sharding and load balancing for large runs.
  • Benchmark expansion: Continue adding benchmark configs, metrics, and troubleshooting guidance.

Status

This project is in internal validation; APIs, configs, and docs may change rapidly.

About

General AI evaluation and Gauge Engine. A unified evaluation engine for LLMs, MLLMs, audio, and diffusion models.

Topics

Resources

Contributing

Stars

52 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); GitHub - HiThink-Research/GAGE: General AI evaluation and Gauge Engine. A unified evaluation engine for LLMs, MLLMs, audio, and diffusion models. · GitHub
Skip to content

Repository files navigation

📐 GAGE: General AI evaluation and Gauge Engine

PythonCode StyleLicenseStatus

English · 中文

📧 Contact:zhangrongjunchen@myhexin.com

Overview · Sample Schema · Smart Defaults · Run Reports · Game Arena · Arena Visual Control · AgentKitV2 · External Harness · Benchmark · Contributing · Standards


GAGE is a unified, extensible evaluation framework for large language models, multimodal models, audio models, diffusion models, agents, and game environments. It provides one evaluation engine for datasets, model backends, metrics, arena runtimes, structured outputs, and replayable artifacts.

Game Arena Showcase

Gomoku GameArena demoDoudizhu GameArena demoMahjong GameArena demo

Space Invaders demoMario demoVizDoom demo

Why GAGE?

  • Fast evaluation engine: Run local smoke tests, model-backed jobs, and larger benchmark batches through the same pipeline shape.
  • Unified evaluation surface: Datasets, backends, role adapters, metrics, and output contracts are configured instead of hand-wired per benchmark.
  • Game and agent sandboxing: Game Arena, AgentKitV2, AppWorld, SWE-bench-style agent tasks, GUI interaction, and tool-augmented workflows share the same run/output model.
  • External harness integration: Delegate task-batch benchmarks to Harbor, then import trial evidence back into standard GAGE samples, metrics, reports, and raw artifacts.
  • Replayable GameKit runtime: Gomoku, Tic-Tac-Toe, Doudizhu, Mahjong, PettingZoo Space Invaders, Retro Mario, and ViZDoom now emit structured arena traces plus arena_visual sessions.
  • Operational visibility: Runs write summary.json, sample outputs, logs, visual artifacts, and a static report pack so failures can be inspected after the fact.

Design Overview

Core design philosophy: everything is a step, everything is configurable.

Architecture Design

End-to-end flow

Orchestration Design

Step view

Game Arena Design

GameArena runtime core design

AgentKitV2 Design

AgentKit v2 pipeline design

External Harness Design

External Harness (Harbor) pipeline design

Quick Start

1. Installation

# If you are in the mono-repo root:cd gage-eval-main
# Python 3.10+ recommended
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

For Game Arena LLM configs, use the *_openai_gamekit.yaml variants and export OPENAI_API_KEY. The model defaults to gpt-5.4; set GAGE_GAME_ARENA_LLM_MODEL to override it, or set OPENAI_API_BASE for an OpenAI-compatible endpoint.

2. Run a Basic Demo

python run.py \
--config config/run_configs/demo_echo_run_1.yaml \
--output-dir runs \
--run-id demo_echo

3. View Reports

Default output structure:

runs/<run_id>/
events.jsonl
samples.jsonl
summary.json
samples/
<namespace>/
<sample_id>.json
report_pack/
report.html
report_context.json
report_context.md
prompt.txt
diagnostics.json
assets_manifest.json

Open runs/<run_id>/report_pack/report.html for the execution-aware report: primary metrics, key findings, scenario profiles, evidence links, media previews, diagnostics, and reason-code explanations. See Run Reports.

Advanced Configurations

ScenarioConfig ExampleDescription
GameArena Human-vs-AIconfig/custom/doudizhu/doudizhu_human_visual_gamekit.yamlBrowser-controlled Doudizhu match against LLM players
GameArena Pure Human Controlconfig/custom/retro_mario/retro_mario_human_visual_gamekit.yamlBrowser-controlled real-time Retro Mario session
AgentKitV2 Tau2config/custom/manual_e2e/agentkit_v2_tau2_local_lmstudio.yamlNative per-sample local-process Tau2 1-case smoke run
AgentKitV2 SWE-bench Proconfig/custom/manual_e2e/agentkit_v2_swebench_pro_docker_lmstudio_smoke1_qutebrowser.yamlNative Docker-backed SWE-bench Pro smoke run
External Harness Harborconfig/custom/external_harness_kits/harbor_terminal_bench2_lmstudio_1case.yamlDelegates a Terminal-Bench 2.0 task to Harbor and imports results
AgentKitV2 AppWorldconfig/custom/appworld/appworld_official_jsonl.yamlAppWorld sandbox evaluation through the native AgentKitV2 path
Textconfig/custom/aime24/aime2024_chat.yamlAIME, GPQA, Math500, and related text benchmarks
Multimodalconfig/custom/mathvista/chat.yamlMathVista and related multimodal benchmarks
LLM Judgeconfig/custom/examples/single_task_local_judge_qwen.yamlLocal LLM judge example

Roadmap

  • Agent evaluation: Continue hardening AgentKitV2 and External Harness trace import, failure diagnostics, and reproducible live smoke configs.
  • Game Arena expansion: Grow the GameKit catalog and keep browser control, replay, and output contracts consistent.
  • Gage-Client: Add a client tool for configuration management, failure diagnostics, and benchmark onboarding.
  • Distributed inference: Support multi-node task sharding and load balancing for large runs.
  • Benchmark expansion: Continue adding benchmark configs, metrics, and troubleshooting guidance.

Status

This project is in internal validation; APIs, configs, and docs may change rapidly.

About

General AI evaluation and Gauge Engine. A unified evaluation engine for LLMs, MLLMs, audio, and diffusion models.

Topics

Resources

Contributing

Stars

52 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages