Repository files navigation

Harness Bench

A compact scorer and methodology starter for reviewing normalized run artifacts from agent harnesses.

The current implementation does not run harnesses, guarantee equivalent conditions, or produce a defensible public leaderboard. Its value is narrower: it makes task, rubric, run, cost, latency, and result assumptions explicit enough to inspect and extend.

Start here

ArtifactUse it for
methodology/comparability-contract.mddeciding whether two harness runs support a comparative claim
tasks/benchmark task definitions and starting-state contracts
rubrics/scoring rules
schemas/normalized run-artifact structure
examples/sample_run.jsonfictional example input
harness-bench scorescoring one normalized artifact

What the CLI does

python -m pip install -e .
harness-bench score examples/sample_run.json

The scorer reads a structured artifact and applies the repository’s current weighted rubric. It can support repeatable demonstrations and schema-level review. It does not independently verify that:

  • the task was executed from the correct starting state;
  • two harnesses had equivalent model, context, tools, permissions, or confirmation rules;
  • the supplied latency, cost, or quality values were measured correctly;
  • the evaluator was valid or blind to harness identity;
  • missing and failed runs were retained;
  • the sample supports a broader claim.

Those responsibilities belong to the experiment and evidence package around the scorer.

The comparability problem

A harness comparison changes more than an interface. Results may depend on:

  • harness and version;
  • model and provider version;
  • system, project, and user instructions;
  • repository and retrieval context;
  • tools and effective permissions;
  • approval and confirmation policy;
  • memory, cache, and session behavior;
  • network, filesystem, and execution environment;
  • task fixture and reset procedure;
  • evaluator and rubric;
  • run count and operator intervention.

When several variables change, the result may still be useful, but the claim must be limited accordingly.

See methodology/comparability-contract.md.

Task contract

A benchmark task should define:

  • immutable starting state;
  • user request;
  • allowed and prohibited actions;
  • expected artifact or state change;
  • success and failure rules;
  • timeout and resource budget;
  • evidence required for scoring;
  • cleanup and reset;
  • known ambiguity and valid alternative solutions.

A prose prompt alone is not a reproducible task.

Outcome model

Do not rely on one total score. Report the components and critical failures.

OutcomeExamples
Functional successtests, expected state, valid deliverable
Constraint complianceprohibited tools, files, network, or external actions avoided
Recoverypartial failure detected and state restored
Evidence qualitydiff, sources, logs, and rationale support review
Interaction burdenclarification, confirmation, and human intervention
Efficiencyelapsed time, tool calls, tokens, compute, cost
Robustnessrepeated-run distribution and recurrence of failure

A weighted score should never compensate for a critical prohibited action or silently dropped failure.

Authority and information parity

Comparisons should record whether systems received equivalent:

  • file, network, connector, and code-execution authority;
  • tools and tool versions;
  • data and repository context;
  • tests and reference outputs;
  • hidden or provider-managed instructions where known;
  • confirmation and human-approval behavior;
  • persistent memory and prior context.

A harness with broader authority or richer context is not necessarily reasoning better; it may be receiving a different treatment.

Run design

For repeated evaluation:

  • predefine run count and task set;
  • reset fixtures and relevant state;
  • record sampling / seed policy where available;
  • randomize or counterbalance order when services can drift;
  • retain setup failures, timeouts, refusals, and invalid outputs;
  • separate harness retries from operator rescue;
  • preserve raw artifacts and manifests.

Repeated attempts on one task are clustered observations and should not be presented as broad independent evidence.

Evaluator design

Prefer executable checks for observable state. When judgment is required:

  • define the rubric before seeing outputs;
  • blind reviewers to harness identity where feasible;
  • randomize presentation order;
  • report reviewer disagreement and adjudication;
  • calibrate model judges against human review;
  • record judge model, prompt, and version.

A model judge is part of the measurement system and can favor style, verbosity, order, or model family.

Cost and latency

Specify whether measurements include:

  • setup and environment preparation;
  • confirmations and clarification;
  • retries and recovery;
  • failed runs;
  • cached versus uncached calls;
  • tool and external-service cost;
  • human review time;
  • pricing date.

Cost per successful task is useful only alongside quality, constraint, and recovery outcomes.

Permitted claim levels

EvidenceAppropriate claim
Illustrative rundemonstrates one workflow under stated conditions
Repeated synthetic suiteshows repeatable differences on that versioned suite
Controlled broader studysupports a bounded conclusion for the defined workflow
Uncontrolled artifactsqualitative inspection only; no ranking claim

The current repository supports the first two levels when the surrounding run records are complete. It does not support general provider or product rankings.

Roadmap with evidence value

The most valuable next work is:

  1. versioned system-under-test and task manifests;
  2. repeated-run ingestion with missing/failure preservation;
  3. paired task-level comparison rather than only aggregate ranking;
  4. executable constraint and artifact checks;
  5. variance and uncertainty reporting;
  6. blinded review workflow for subjective rubrics;
  7. synthetic task packs with reset scripts and known alternative solutions.

A larger leaderboard without these controls would make the repository look more impressive while making its conclusions less trustworthy.

Publication safety

Only publish synthetic tasks, fictional outputs, and sanitized run artifacts. Do not publish private prompts, proprietary benchmark tasks, customer or employer data, confidential tool output, credentials, connector details, internal cost accounts, or comparison claims not supported by the recorded method.

Maturity and scope

This is a working prototype scorer plus a comparability methodology. It is not an independent benchmark lab, certified evaluation suite, provider endorsement, or definitive ranking of models, IDEs, or agent harnesses.


Maintained by Sima Bagheri.

About

A starter framework for transparent scoring of synthetic agent-harness run artifacts.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Harness Bench

A compact scorer and methodology starter for reviewing normalized run artifacts from agent harnesses.

The current implementation does not run harnesses, guarantee equivalent conditions, or produce a defensible public leaderboard. Its value is narrower: it makes task, rubric, run, cost, latency, and result assumptions explicit enough to inspect and extend.

Start here

ArtifactUse it for
methodology/comparability-contract.mddeciding whether two harness runs support a comparative claim
tasks/benchmark task definitions and starting-state contracts
rubrics/scoring rules
schemas/normalized run-artifact structure
examples/sample_run.jsonfictional example input
harness-bench scorescoring one normalized artifact

What the CLI does

python -m pip install -e .
harness-bench score examples/sample_run.json

The scorer reads a structured artifact and applies the repository’s current weighted rubric. It can support repeatable demonstrations and schema-level review. It does not independently verify that:

  • the task was executed from the correct starting state;
  • two harnesses had equivalent model, context, tools, permissions, or confirmation rules;
  • the supplied latency, cost, or quality values were measured correctly;
  • the evaluator was valid or blind to harness identity;
  • missing and failed runs were retained;
  • the sample supports a broader claim.

Those responsibilities belong to the experiment and evidence package around the scorer.

The comparability problem

A harness comparison changes more than an interface. Results may depend on:

  • harness and version;
  • model and provider version;
  • system, project, and user instructions;
  • repository and retrieval context;
  • tools and effective permissions;
  • approval and confirmation policy;
  • memory, cache, and session behavior;
  • network, filesystem, and execution environment;
  • task fixture and reset procedure;
  • evaluator and rubric;
  • run count and operator intervention.

When several variables change, the result may still be useful, but the claim must be limited accordingly.

See methodology/comparability-contract.md.

Task contract

A benchmark task should define:

  • immutable starting state;
  • user request;
  • allowed and prohibited actions;
  • expected artifact or state change;
  • success and failure rules;
  • timeout and resource budget;
  • evidence required for scoring;
  • cleanup and reset;
  • known ambiguity and valid alternative solutions.

A prose prompt alone is not a reproducible task.

Outcome model

Do not rely on one total score. Report the components and critical failures.

OutcomeExamples
Functional successtests, expected state, valid deliverable
Constraint complianceprohibited tools, files, network, or external actions avoided
Recoverypartial failure detected and state restored
Evidence qualitydiff, sources, logs, and rationale support review
Interaction burdenclarification, confirmation, and human intervention
Efficiencyelapsed time, tool calls, tokens, compute, cost
Robustnessrepeated-run distribution and recurrence of failure

A weighted score should never compensate for a critical prohibited action or silently dropped failure.

Authority and information parity

Comparisons should record whether systems received equivalent:

  • file, network, connector, and code-execution authority;
  • tools and tool versions;
  • data and repository context;
  • tests and reference outputs;
  • hidden or provider-managed instructions where known;
  • confirmation and human-approval behavior;
  • persistent memory and prior context.

A harness with broader authority or richer context is not necessarily reasoning better; it may be receiving a different treatment.

Run design

For repeated evaluation:

  • predefine run count and task set;
  • reset fixtures and relevant state;
  • record sampling / seed policy where available;
  • randomize or counterbalance order when services can drift;
  • retain setup failures, timeouts, refusals, and invalid outputs;
  • separate harness retries from operator rescue;
  • preserve raw artifacts and manifests.

Repeated attempts on one task are clustered observations and should not be presented as broad independent evidence.

Evaluator design

Prefer executable checks for observable state. When judgment is required:

  • define the rubric before seeing outputs;
  • blind reviewers to harness identity where feasible;
  • randomize presentation order;
  • report reviewer disagreement and adjudication;
  • calibrate model judges against human review;
  • record judge model, prompt, and version.

A model judge is part of the measurement system and can favor style, verbosity, order, or model family.

Cost and latency

Specify whether measurements include:

  • setup and environment preparation;
  • confirmations and clarification;
  • retries and recovery;
  • failed runs;
  • cached versus uncached calls;
  • tool and external-service cost;
  • human review time;
  • pricing date.

Cost per successful task is useful only alongside quality, constraint, and recovery outcomes.

Permitted claim levels

EvidenceAppropriate claim
Illustrative rundemonstrates one workflow under stated conditions
Repeated synthetic suiteshows repeatable differences on that versioned suite
Controlled broader studysupports a bounded conclusion for the defined workflow
Uncontrolled artifactsqualitative inspection only; no ranking claim

The current repository supports the first two levels when the surrounding run records are complete. It does not support general provider or product rankings.

Roadmap with evidence value

The most valuable next work is:

  1. versioned system-under-test and task manifests;
  2. repeated-run ingestion with missing/failure preservation;
  3. paired task-level comparison rather than only aggregate ranking;
  4. executable constraint and artifact checks;
  5. variance and uncertainty reporting;
  6. blinded review workflow for subjective rubrics;
  7. synthetic task packs with reset scripts and known alternative solutions.

A larger leaderboard without these controls would make the repository look more impressive while making its conclusions less trustworthy.

Publication safety

Only publish synthetic tasks, fictional outputs, and sanitized run artifacts. Do not publish private prompts, proprietary benchmark tasks, customer or employer data, confidential tool output, credentials, connector details, internal cost accounts, or comparison claims not supported by the recorded method.

Maturity and scope

This is a working prototype scorer plus a comparability methodology. It is not an independent benchmark lab, certified evaluation suite, provider endorsement, or definitive ranking of models, IDEs, or agent harnesses.


Maintained by Sima Bagheri.

About

A starter framework for transparent scoring of synthetic agent-harness run artifacts.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Harness Bench

A compact scorer and methodology starter for reviewing normalized run artifacts from agent harnesses.

The current implementation does not run harnesses, guarantee equivalent conditions, or produce a defensible public leaderboard. Its value is narrower: it makes task, rubric, run, cost, latency, and result assumptions explicit enough to inspect and extend.

Start here

ArtifactUse it for
methodology/comparability-contract.mddeciding whether two harness runs support a comparative claim
tasks/benchmark task definitions and starting-state contracts
rubrics/scoring rules
schemas/normalized run-artifact structure
examples/sample_run.jsonfictional example input
harness-bench scorescoring one normalized artifact

What the CLI does

python -m pip install -e .
harness-bench score examples/sample_run.json

The scorer reads a structured artifact and applies the repository’s current weighted rubric. It can support repeatable demonstrations and schema-level review. It does not independently verify that:

  • the task was executed from the correct starting state;
  • two harnesses had equivalent model, context, tools, permissions, or confirmation rules;
  • the supplied latency, cost, or quality values were measured correctly;
  • the evaluator was valid or blind to harness identity;
  • missing and failed runs were retained;
  • the sample supports a broader claim.

Those responsibilities belong to the experiment and evidence package around the scorer.

The comparability problem

A harness comparison changes more than an interface. Results may depend on:

  • harness and version;
  • model and provider version;
  • system, project, and user instructions;
  • repository and retrieval context;
  • tools and effective permissions;
  • approval and confirmation policy;
  • memory, cache, and session behavior;
  • network, filesystem, and execution environment;
  • task fixture and reset procedure;
  • evaluator and rubric;
  • run count and operator intervention.

When several variables change, the result may still be useful, but the claim must be limited accordingly.

See methodology/comparability-contract.md.

Task contract

A benchmark task should define:

  • immutable starting state;
  • user request;
  • allowed and prohibited actions;
  • expected artifact or state change;
  • success and failure rules;
  • timeout and resource budget;
  • evidence required for scoring;
  • cleanup and reset;
  • known ambiguity and valid alternative solutions.

A prose prompt alone is not a reproducible task.

Outcome model

Do not rely on one total score. Report the components and critical failures.

OutcomeExamples
Functional successtests, expected state, valid deliverable
Constraint complianceprohibited tools, files, network, or external actions avoided
Recoverypartial failure detected and state restored
Evidence qualitydiff, sources, logs, and rationale support review
Interaction burdenclarification, confirmation, and human intervention
Efficiencyelapsed time, tool calls, tokens, compute, cost
Robustnessrepeated-run distribution and recurrence of failure

A weighted score should never compensate for a critical prohibited action or silently dropped failure.

Authority and information parity

Comparisons should record whether systems received equivalent:

  • file, network, connector, and code-execution authority;
  • tools and tool versions;
  • data and repository context;
  • tests and reference outputs;
  • hidden or provider-managed instructions where known;
  • confirmation and human-approval behavior;
  • persistent memory and prior context.

A harness with broader authority or richer context is not necessarily reasoning better; it may be receiving a different treatment.

Run design

For repeated evaluation:

  • predefine run count and task set;
  • reset fixtures and relevant state;
  • record sampling / seed policy where available;
  • randomize or counterbalance order when services can drift;
  • retain setup failures, timeouts, refusals, and invalid outputs;
  • separate harness retries from operator rescue;
  • preserve raw artifacts and manifests.

Repeated attempts on one task are clustered observations and should not be presented as broad independent evidence.

Evaluator design

Prefer executable checks for observable state. When judgment is required:

  • define the rubric before seeing outputs;
  • blind reviewers to harness identity where feasible;
  • randomize presentation order;
  • report reviewer disagreement and adjudication;
  • calibrate model judges against human review;
  • record judge model, prompt, and version.

A model judge is part of the measurement system and can favor style, verbosity, order, or model family.

Cost and latency

Specify whether measurements include:

  • setup and environment preparation;
  • confirmations and clarification;
  • retries and recovery;
  • failed runs;
  • cached versus uncached calls;
  • tool and external-service cost;
  • human review time;
  • pricing date.

Cost per successful task is useful only alongside quality, constraint, and recovery outcomes.

Permitted claim levels

EvidenceAppropriate claim
Illustrative rundemonstrates one workflow under stated conditions
Repeated synthetic suiteshows repeatable differences on that versioned suite
Controlled broader studysupports a bounded conclusion for the defined workflow
Uncontrolled artifactsqualitative inspection only; no ranking claim

The current repository supports the first two levels when the surrounding run records are complete. It does not support general provider or product rankings.

Roadmap with evidence value

The most valuable next work is:

  1. versioned system-under-test and task manifests;
  2. repeated-run ingestion with missing/failure preservation;
  3. paired task-level comparison rather than only aggregate ranking;
  4. executable constraint and artifact checks;
  5. variance and uncertainty reporting;
  6. blinded review workflow for subjective rubrics;
  7. synthetic task packs with reset scripts and known alternative solutions.

A larger leaderboard without these controls would make the repository look more impressive while making its conclusions less trustworthy.

Publication safety

Only publish synthetic tasks, fictional outputs, and sanitized run artifacts. Do not publish private prompts, proprietary benchmark tasks, customer or employer data, confidential tool output, credentials, connector details, internal cost accounts, or comparison claims not supported by the recorded method.

Maturity and scope

This is a working prototype scorer plus a comparability methodology. It is not an independent benchmark lab, certified evaluation suite, provider endorsement, or definitive ranking of models, IDEs, or agent harnesses.


Maintained by Sima Bagheri.

About

A starter framework for transparent scoring of synthetic agent-harness run artifacts.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Harness Bench

A compact scorer and methodology starter for reviewing normalized run artifacts from agent harnesses.

The current implementation does not run harnesses, guarantee equivalent conditions, or produce a defensible public leaderboard. Its value is narrower: it makes task, rubric, run, cost, latency, and result assumptions explicit enough to inspect and extend.

Start here

ArtifactUse it for
methodology/comparability-contract.mddeciding whether two harness runs support a comparative claim
tasks/benchmark task definitions and starting-state contracts
rubrics/scoring rules
schemas/normalized run-artifact structure
examples/sample_run.jsonfictional example input
harness-bench scorescoring one normalized artifact

What the CLI does

python -m pip install -e .
harness-bench score examples/sample_run.json

The scorer reads a structured artifact and applies the repository’s current weighted rubric. It can support repeatable demonstrations and schema-level review. It does not independently verify that:

  • the task was executed from the correct starting state;
  • two harnesses had equivalent model, context, tools, permissions, or confirmation rules;
  • the supplied latency, cost, or quality values were measured correctly;
  • the evaluator was valid or blind to harness identity;
  • missing and failed runs were retained;
  • the sample supports a broader claim.

Those responsibilities belong to the experiment and evidence package around the scorer.

The comparability problem

A harness comparison changes more than an interface. Results may depend on:

  • harness and version;
  • model and provider version;
  • system, project, and user instructions;
  • repository and retrieval context;
  • tools and effective permissions;
  • approval and confirmation policy;
  • memory, cache, and session behavior;
  • network, filesystem, and execution environment;
  • task fixture and reset procedure;
  • evaluator and rubric;
  • run count and operator intervention.

When several variables change, the result may still be useful, but the claim must be limited accordingly.

See methodology/comparability-contract.md.

Task contract

A benchmark task should define:

  • immutable starting state;
  • user request;
  • allowed and prohibited actions;
  • expected artifact or state change;
  • success and failure rules;
  • timeout and resource budget;
  • evidence required for scoring;
  • cleanup and reset;
  • known ambiguity and valid alternative solutions.

A prose prompt alone is not a reproducible task.

Outcome model

Do not rely on one total score. Report the components and critical failures.

OutcomeExamples
Functional successtests, expected state, valid deliverable
Constraint complianceprohibited tools, files, network, or external actions avoided
Recoverypartial failure detected and state restored
Evidence qualitydiff, sources, logs, and rationale support review
Interaction burdenclarification, confirmation, and human intervention
Efficiencyelapsed time, tool calls, tokens, compute, cost
Robustnessrepeated-run distribution and recurrence of failure

A weighted score should never compensate for a critical prohibited action or silently dropped failure.

Authority and information parity

Comparisons should record whether systems received equivalent:

  • file, network, connector, and code-execution authority;
  • tools and tool versions;
  • data and repository context;
  • tests and reference outputs;
  • hidden or provider-managed instructions where known;
  • confirmation and human-approval behavior;
  • persistent memory and prior context.

A harness with broader authority or richer context is not necessarily reasoning better; it may be receiving a different treatment.

Run design

For repeated evaluation:

  • predefine run count and task set;
  • reset fixtures and relevant state;
  • record sampling / seed policy where available;
  • randomize or counterbalance order when services can drift;
  • retain setup failures, timeouts, refusals, and invalid outputs;
  • separate harness retries from operator rescue;
  • preserve raw artifacts and manifests.

Repeated attempts on one task are clustered observations and should not be presented as broad independent evidence.

Evaluator design

Prefer executable checks for observable state. When judgment is required:

  • define the rubric before seeing outputs;
  • blind reviewers to harness identity where feasible;
  • randomize presentation order;
  • report reviewer disagreement and adjudication;
  • calibrate model judges against human review;
  • record judge model, prompt, and version.

A model judge is part of the measurement system and can favor style, verbosity, order, or model family.

Cost and latency

Specify whether measurements include:

  • setup and environment preparation;
  • confirmations and clarification;
  • retries and recovery;
  • failed runs;
  • cached versus uncached calls;
  • tool and external-service cost;
  • human review time;
  • pricing date.

Cost per successful task is useful only alongside quality, constraint, and recovery outcomes.

Permitted claim levels

EvidenceAppropriate claim
Illustrative rundemonstrates one workflow under stated conditions
Repeated synthetic suiteshows repeatable differences on that versioned suite
Controlled broader studysupports a bounded conclusion for the defined workflow
Uncontrolled artifactsqualitative inspection only; no ranking claim

The current repository supports the first two levels when the surrounding run records are complete. It does not support general provider or product rankings.

Roadmap with evidence value

The most valuable next work is:

  1. versioned system-under-test and task manifests;
  2. repeated-run ingestion with missing/failure preservation;
  3. paired task-level comparison rather than only aggregate ranking;
  4. executable constraint and artifact checks;
  5. variance and uncertainty reporting;
  6. blinded review workflow for subjective rubrics;
  7. synthetic task packs with reset scripts and known alternative solutions.

A larger leaderboard without these controls would make the repository look more impressive while making its conclusions less trustworthy.

Publication safety

Only publish synthetic tasks, fictional outputs, and sanitized run artifacts. Do not publish private prompts, proprietary benchmark tasks, customer or employer data, confidential tool output, credentials, connector details, internal cost accounts, or comparison claims not supported by the recorded method.

Maturity and scope

This is a working prototype scorer plus a comparability methodology. It is not an independent benchmark lab, certified evaluation suite, provider endorsement, or definitive ranking of models, IDEs, or agent harnesses.


Maintained by Sima Bagheri.

About

A starter framework for transparent scoring of synthetic agent-harness run artifacts.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Harness Bench

A compact scorer and methodology starter for reviewing normalized run artifacts from agent harnesses.

The current implementation does not run harnesses, guarantee equivalent conditions, or produce a defensible public leaderboard. Its value is narrower: it makes task, rubric, run, cost, latency, and result assumptions explicit enough to inspect and extend.

Start here

ArtifactUse it for
methodology/comparability-contract.mddeciding whether two harness runs support a comparative claim
tasks/benchmark task definitions and starting-state contracts
rubrics/scoring rules
schemas/normalized run-artifact structure
examples/sample_run.jsonfictional example input
harness-bench scorescoring one normalized artifact

What the CLI does

python -m pip install -e .
harness-bench score examples/sample_run.json

The scorer reads a structured artifact and applies the repository’s current weighted rubric. It can support repeatable demonstrations and schema-level review. It does not independently verify that:

  • the task was executed from the correct starting state;
  • two harnesses had equivalent model, context, tools, permissions, or confirmation rules;
  • the supplied latency, cost, or quality values were measured correctly;
  • the evaluator was valid or blind to harness identity;
  • missing and failed runs were retained;
  • the sample supports a broader claim.

Those responsibilities belong to the experiment and evidence package around the scorer.

The comparability problem

A harness comparison changes more than an interface. Results may depend on:

  • harness and version;
  • model and provider version;
  • system, project, and user instructions;
  • repository and retrieval context;
  • tools and effective permissions;
  • approval and confirmation policy;
  • memory, cache, and session behavior;
  • network, filesystem, and execution environment;
  • task fixture and reset procedure;
  • evaluator and rubric;
  • run count and operator intervention.

When several variables change, the result may still be useful, but the claim must be limited accordingly.

See methodology/comparability-contract.md.

Task contract

A benchmark task should define:

  • immutable starting state;
  • user request;
  • allowed and prohibited actions;
  • expected artifact or state change;
  • success and failure rules;
  • timeout and resource budget;
  • evidence required for scoring;
  • cleanup and reset;
  • known ambiguity and valid alternative solutions.

A prose prompt alone is not a reproducible task.

Outcome model

Do not rely on one total score. Report the components and critical failures.

OutcomeExamples
Functional successtests, expected state, valid deliverable
Constraint complianceprohibited tools, files, network, or external actions avoided
Recoverypartial failure detected and state restored
Evidence qualitydiff, sources, logs, and rationale support review
Interaction burdenclarification, confirmation, and human intervention
Efficiencyelapsed time, tool calls, tokens, compute, cost
Robustnessrepeated-run distribution and recurrence of failure

A weighted score should never compensate for a critical prohibited action or silently dropped failure.

Authority and information parity

Comparisons should record whether systems received equivalent:

  • file, network, connector, and code-execution authority;
  • tools and tool versions;
  • data and repository context;
  • tests and reference outputs;
  • hidden or provider-managed instructions where known;
  • confirmation and human-approval behavior;
  • persistent memory and prior context.

A harness with broader authority or richer context is not necessarily reasoning better; it may be receiving a different treatment.

Run design

For repeated evaluation:

  • predefine run count and task set;
  • reset fixtures and relevant state;
  • record sampling / seed policy where available;
  • randomize or counterbalance order when services can drift;
  • retain setup failures, timeouts, refusals, and invalid outputs;
  • separate harness retries from operator rescue;
  • preserve raw artifacts and manifests.

Repeated attempts on one task are clustered observations and should not be presented as broad independent evidence.

Evaluator design

Prefer executable checks for observable state. When judgment is required:

  • define the rubric before seeing outputs;
  • blind reviewers to harness identity where feasible;
  • randomize presentation order;
  • report reviewer disagreement and adjudication;
  • calibrate model judges against human review;
  • record judge model, prompt, and version.

A model judge is part of the measurement system and can favor style, verbosity, order, or model family.

Cost and latency

Specify whether measurements include:

  • setup and environment preparation;
  • confirmations and clarification;
  • retries and recovery;
  • failed runs;
  • cached versus uncached calls;
  • tool and external-service cost;
  • human review time;
  • pricing date.

Cost per successful task is useful only alongside quality, constraint, and recovery outcomes.

Permitted claim levels

EvidenceAppropriate claim
Illustrative rundemonstrates one workflow under stated conditions
Repeated synthetic suiteshows repeatable differences on that versioned suite
Controlled broader studysupports a bounded conclusion for the defined workflow
Uncontrolled artifactsqualitative inspection only; no ranking claim

The current repository supports the first two levels when the surrounding run records are complete. It does not support general provider or product rankings.

Roadmap with evidence value

The most valuable next work is:

  1. versioned system-under-test and task manifests;
  2. repeated-run ingestion with missing/failure preservation;
  3. paired task-level comparison rather than only aggregate ranking;
  4. executable constraint and artifact checks;
  5. variance and uncertainty reporting;
  6. blinded review workflow for subjective rubrics;
  7. synthetic task packs with reset scripts and known alternative solutions.

A larger leaderboard without these controls would make the repository look more impressive while making its conclusions less trustworthy.

Publication safety

Only publish synthetic tasks, fictional outputs, and sanitized run artifacts. Do not publish private prompts, proprietary benchmark tasks, customer or employer data, confidential tool output, credentials, connector details, internal cost accounts, or comparison claims not supported by the recorded method.

Maturity and scope

This is a working prototype scorer plus a comparability methodology. It is not an independent benchmark lab, certified evaluation suite, provider endorsement, or definitive ranking of models, IDEs, or agent harnesses.


Maintained by Sima Bagheri.

About

A starter framework for transparent scoring of synthetic agent-harness run artifacts.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Harness Bench

A compact scorer and methodology starter for reviewing normalized run artifacts from agent harnesses.

The current implementation does not run harnesses, guarantee equivalent conditions, or produce a defensible public leaderboard. Its value is narrower: it makes task, rubric, run, cost, latency, and result assumptions explicit enough to inspect and extend.

Start here

ArtifactUse it for
methodology/comparability-contract.mddeciding whether two harness runs support a comparative claim
tasks/benchmark task definitions and starting-state contracts
rubrics/scoring rules
schemas/normalized run-artifact structure
examples/sample_run.jsonfictional example input
harness-bench scorescoring one normalized artifact

What the CLI does

python -m pip install -e .
harness-bench score examples/sample_run.json

The scorer reads a structured artifact and applies the repository’s current weighted rubric. It can support repeatable demonstrations and schema-level review. It does not independently verify that:

  • the task was executed from the correct starting state;
  • two harnesses had equivalent model, context, tools, permissions, or confirmation rules;
  • the supplied latency, cost, or quality values were measured correctly;
  • the evaluator was valid or blind to harness identity;
  • missing and failed runs were retained;
  • the sample supports a broader claim.

Those responsibilities belong to the experiment and evidence package around the scorer.

The comparability problem

A harness comparison changes more than an interface. Results may depend on:

  • harness and version;
  • model and provider version;
  • system, project, and user instructions;
  • repository and retrieval context;
  • tools and effective permissions;
  • approval and confirmation policy;
  • memory, cache, and session behavior;
  • network, filesystem, and execution environment;
  • task fixture and reset procedure;
  • evaluator and rubric;
  • run count and operator intervention.

When several variables change, the result may still be useful, but the claim must be limited accordingly.

See methodology/comparability-contract.md.

Task contract

A benchmark task should define:

  • immutable starting state;
  • user request;
  • allowed and prohibited actions;
  • expected artifact or state change;
  • success and failure rules;
  • timeout and resource budget;
  • evidence required for scoring;
  • cleanup and reset;
  • known ambiguity and valid alternative solutions.

A prose prompt alone is not a reproducible task.

Outcome model

Do not rely on one total score. Report the components and critical failures.

OutcomeExamples
Functional successtests, expected state, valid deliverable
Constraint complianceprohibited tools, files, network, or external actions avoided
Recoverypartial failure detected and state restored
Evidence qualitydiff, sources, logs, and rationale support review
Interaction burdenclarification, confirmation, and human intervention
Efficiencyelapsed time, tool calls, tokens, compute, cost
Robustnessrepeated-run distribution and recurrence of failure

A weighted score should never compensate for a critical prohibited action or silently dropped failure.

Authority and information parity

Comparisons should record whether systems received equivalent:

  • file, network, connector, and code-execution authority;
  • tools and tool versions;
  • data and repository context;
  • tests and reference outputs;
  • hidden or provider-managed instructions where known;
  • confirmation and human-approval behavior;
  • persistent memory and prior context.

A harness with broader authority or richer context is not necessarily reasoning better; it may be receiving a different treatment.

Run design

For repeated evaluation:

  • predefine run count and task set;
  • reset fixtures and relevant state;
  • record sampling / seed policy where available;
  • randomize or counterbalance order when services can drift;
  • retain setup failures, timeouts, refusals, and invalid outputs;
  • separate harness retries from operator rescue;
  • preserve raw artifacts and manifests.

Repeated attempts on one task are clustered observations and should not be presented as broad independent evidence.

Evaluator design

Prefer executable checks for observable state. When judgment is required:

  • define the rubric before seeing outputs;
  • blind reviewers to harness identity where feasible;
  • randomize presentation order;
  • report reviewer disagreement and adjudication;
  • calibrate model judges against human review;
  • record judge model, prompt, and version.

A model judge is part of the measurement system and can favor style, verbosity, order, or model family.

Cost and latency

Specify whether measurements include:

  • setup and environment preparation;
  • confirmations and clarification;
  • retries and recovery;
  • failed runs;
  • cached versus uncached calls;
  • tool and external-service cost;
  • human review time;
  • pricing date.

Cost per successful task is useful only alongside quality, constraint, and recovery outcomes.

Permitted claim levels

EvidenceAppropriate claim
Illustrative rundemonstrates one workflow under stated conditions
Repeated synthetic suiteshows repeatable differences on that versioned suite
Controlled broader studysupports a bounded conclusion for the defined workflow
Uncontrolled artifactsqualitative inspection only; no ranking claim

The current repository supports the first two levels when the surrounding run records are complete. It does not support general provider or product rankings.

Roadmap with evidence value

The most valuable next work is:

  1. versioned system-under-test and task manifests;
  2. repeated-run ingestion with missing/failure preservation;
  3. paired task-level comparison rather than only aggregate ranking;
  4. executable constraint and artifact checks;
  5. variance and uncertainty reporting;
  6. blinded review workflow for subjective rubrics;
  7. synthetic task packs with reset scripts and known alternative solutions.

A larger leaderboard without these controls would make the repository look more impressive while making its conclusions less trustworthy.

Publication safety

Only publish synthetic tasks, fictional outputs, and sanitized run artifacts. Do not publish private prompts, proprietary benchmark tasks, customer or employer data, confidential tool output, credentials, connector details, internal cost accounts, or comparison claims not supported by the recorded method.

Maturity and scope

This is a working prototype scorer plus a comparability methodology. It is not an independent benchmark lab, certified evaluation suite, provider endorsement, or definitive ranking of models, IDEs, or agent harnesses.


Maintained by Sima Bagheri.

About

A starter framework for transparent scoring of synthetic agent-harness run artifacts.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Harness Bench

A compact scorer and methodology starter for reviewing normalized run artifacts from agent harnesses.

The current implementation does not run harnesses, guarantee equivalent conditions, or produce a defensible public leaderboard. Its value is narrower: it makes task, rubric, run, cost, latency, and result assumptions explicit enough to inspect and extend.

Start here

ArtifactUse it for
methodology/comparability-contract.mddeciding whether two harness runs support a comparative claim
tasks/benchmark task definitions and starting-state contracts
rubrics/scoring rules
schemas/normalized run-artifact structure
examples/sample_run.jsonfictional example input
harness-bench scorescoring one normalized artifact

What the CLI does

python -m pip install -e .
harness-bench score examples/sample_run.json

The scorer reads a structured artifact and applies the repository’s current weighted rubric. It can support repeatable demonstrations and schema-level review. It does not independently verify that:

  • the task was executed from the correct starting state;
  • two harnesses had equivalent model, context, tools, permissions, or confirmation rules;
  • the supplied latency, cost, or quality values were measured correctly;
  • the evaluator was valid or blind to harness identity;
  • missing and failed runs were retained;
  • the sample supports a broader claim.

Those responsibilities belong to the experiment and evidence package around the scorer.

The comparability problem

A harness comparison changes more than an interface. Results may depend on:

  • harness and version;
  • model and provider version;
  • system, project, and user instructions;
  • repository and retrieval context;
  • tools and effective permissions;
  • approval and confirmation policy;
  • memory, cache, and session behavior;
  • network, filesystem, and execution environment;
  • task fixture and reset procedure;
  • evaluator and rubric;
  • run count and operator intervention.

When several variables change, the result may still be useful, but the claim must be limited accordingly.

See methodology/comparability-contract.md.

Task contract

A benchmark task should define:

  • immutable starting state;
  • user request;
  • allowed and prohibited actions;
  • expected artifact or state change;
  • success and failure rules;
  • timeout and resource budget;
  • evidence required for scoring;
  • cleanup and reset;
  • known ambiguity and valid alternative solutions.

A prose prompt alone is not a reproducible task.

Outcome model

Do not rely on one total score. Report the components and critical failures.

OutcomeExamples
Functional successtests, expected state, valid deliverable
Constraint complianceprohibited tools, files, network, or external actions avoided
Recoverypartial failure detected and state restored
Evidence qualitydiff, sources, logs, and rationale support review
Interaction burdenclarification, confirmation, and human intervention
Efficiencyelapsed time, tool calls, tokens, compute, cost
Robustnessrepeated-run distribution and recurrence of failure

A weighted score should never compensate for a critical prohibited action or silently dropped failure.

Authority and information parity

Comparisons should record whether systems received equivalent:

  • file, network, connector, and code-execution authority;
  • tools and tool versions;
  • data and repository context;
  • tests and reference outputs;
  • hidden or provider-managed instructions where known;
  • confirmation and human-approval behavior;
  • persistent memory and prior context.

A harness with broader authority or richer context is not necessarily reasoning better; it may be receiving a different treatment.

Run design

For repeated evaluation:

  • predefine run count and task set;
  • reset fixtures and relevant state;
  • record sampling / seed policy where available;
  • randomize or counterbalance order when services can drift;
  • retain setup failures, timeouts, refusals, and invalid outputs;
  • separate harness retries from operator rescue;
  • preserve raw artifacts and manifests.

Repeated attempts on one task are clustered observations and should not be presented as broad independent evidence.

Evaluator design

Prefer executable checks for observable state. When judgment is required:

  • define the rubric before seeing outputs;
  • blind reviewers to harness identity where feasible;
  • randomize presentation order;
  • report reviewer disagreement and adjudication;
  • calibrate model judges against human review;
  • record judge model, prompt, and version.

A model judge is part of the measurement system and can favor style, verbosity, order, or model family.

Cost and latency

Specify whether measurements include:

  • setup and environment preparation;
  • confirmations and clarification;
  • retries and recovery;
  • failed runs;
  • cached versus uncached calls;
  • tool and external-service cost;
  • human review time;
  • pricing date.

Cost per successful task is useful only alongside quality, constraint, and recovery outcomes.

Permitted claim levels

EvidenceAppropriate claim
Illustrative rundemonstrates one workflow under stated conditions
Repeated synthetic suiteshows repeatable differences on that versioned suite
Controlled broader studysupports a bounded conclusion for the defined workflow
Uncontrolled artifactsqualitative inspection only; no ranking claim

The current repository supports the first two levels when the surrounding run records are complete. It does not support general provider or product rankings.

Roadmap with evidence value

The most valuable next work is:

  1. versioned system-under-test and task manifests;
  2. repeated-run ingestion with missing/failure preservation;
  3. paired task-level comparison rather than only aggregate ranking;
  4. executable constraint and artifact checks;
  5. variance and uncertainty reporting;
  6. blinded review workflow for subjective rubrics;
  7. synthetic task packs with reset scripts and known alternative solutions.

A larger leaderboard without these controls would make the repository look more impressive while making its conclusions less trustworthy.

Publication safety

Only publish synthetic tasks, fictional outputs, and sanitized run artifacts. Do not publish private prompts, proprietary benchmark tasks, customer or employer data, confidential tool output, credentials, connector details, internal cost accounts, or comparison claims not supported by the recorded method.

Maturity and scope

This is a working prototype scorer plus a comparability methodology. It is not an independent benchmark lab, certified evaluation suite, provider endorsement, or definitive ranking of models, IDEs, or agent harnesses.


Maintained by Sima Bagheri.

About

A starter framework for transparent scoring of synthetic agent-harness run artifacts.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Harness Bench

A compact scorer and methodology starter for reviewing normalized run artifacts from agent harnesses.

The current implementation does not run harnesses, guarantee equivalent conditions, or produce a defensible public leaderboard. Its value is narrower: it makes task, rubric, run, cost, latency, and result assumptions explicit enough to inspect and extend.

Start here

ArtifactUse it for
methodology/comparability-contract.mddeciding whether two harness runs support a comparative claim
tasks/benchmark task definitions and starting-state contracts
rubrics/scoring rules
schemas/normalized run-artifact structure
examples/sample_run.jsonfictional example input
harness-bench scorescoring one normalized artifact

What the CLI does

python -m pip install -e .
harness-bench score examples/sample_run.json

The scorer reads a structured artifact and applies the repository’s current weighted rubric. It can support repeatable demonstrations and schema-level review. It does not independently verify that:

  • the task was executed from the correct starting state;
  • two harnesses had equivalent model, context, tools, permissions, or confirmation rules;
  • the supplied latency, cost, or quality values were measured correctly;
  • the evaluator was valid or blind to harness identity;
  • missing and failed runs were retained;
  • the sample supports a broader claim.

Those responsibilities belong to the experiment and evidence package around the scorer.

The comparability problem

A harness comparison changes more than an interface. Results may depend on:

  • harness and version;
  • model and provider version;
  • system, project, and user instructions;
  • repository and retrieval context;
  • tools and effective permissions;
  • approval and confirmation policy;
  • memory, cache, and session behavior;
  • network, filesystem, and execution environment;
  • task fixture and reset procedure;
  • evaluator and rubric;
  • run count and operator intervention.

When several variables change, the result may still be useful, but the claim must be limited accordingly.

See methodology/comparability-contract.md.

Task contract

A benchmark task should define:

  • immutable starting state;
  • user request;
  • allowed and prohibited actions;
  • expected artifact or state change;
  • success and failure rules;
  • timeout and resource budget;
  • evidence required for scoring;
  • cleanup and reset;
  • known ambiguity and valid alternative solutions.

A prose prompt alone is not a reproducible task.

Outcome model

Do not rely on one total score. Report the components and critical failures.

OutcomeExamples
Functional successtests, expected state, valid deliverable
Constraint complianceprohibited tools, files, network, or external actions avoided
Recoverypartial failure detected and state restored
Evidence qualitydiff, sources, logs, and rationale support review
Interaction burdenclarification, confirmation, and human intervention
Efficiencyelapsed time, tool calls, tokens, compute, cost
Robustnessrepeated-run distribution and recurrence of failure

A weighted score should never compensate for a critical prohibited action or silently dropped failure.

Authority and information parity

Comparisons should record whether systems received equivalent:

  • file, network, connector, and code-execution authority;
  • tools and tool versions;
  • data and repository context;
  • tests and reference outputs;
  • hidden or provider-managed instructions where known;
  • confirmation and human-approval behavior;
  • persistent memory and prior context.

A harness with broader authority or richer context is not necessarily reasoning better; it may be receiving a different treatment.

Run design

For repeated evaluation:

  • predefine run count and task set;
  • reset fixtures and relevant state;
  • record sampling / seed policy where available;
  • randomize or counterbalance order when services can drift;
  • retain setup failures, timeouts, refusals, and invalid outputs;
  • separate harness retries from operator rescue;
  • preserve raw artifacts and manifests.

Repeated attempts on one task are clustered observations and should not be presented as broad independent evidence.

Evaluator design

Prefer executable checks for observable state. When judgment is required:

  • define the rubric before seeing outputs;
  • blind reviewers to harness identity where feasible;
  • randomize presentation order;
  • report reviewer disagreement and adjudication;
  • calibrate model judges against human review;
  • record judge model, prompt, and version.

A model judge is part of the measurement system and can favor style, verbosity, order, or model family.

Cost and latency

Specify whether measurements include:

  • setup and environment preparation;
  • confirmations and clarification;
  • retries and recovery;
  • failed runs;
  • cached versus uncached calls;
  • tool and external-service cost;
  • human review time;
  • pricing date.

Cost per successful task is useful only alongside quality, constraint, and recovery outcomes.

Permitted claim levels

EvidenceAppropriate claim
Illustrative rundemonstrates one workflow under stated conditions
Repeated synthetic suiteshows repeatable differences on that versioned suite
Controlled broader studysupports a bounded conclusion for the defined workflow
Uncontrolled artifactsqualitative inspection only; no ranking claim

The current repository supports the first two levels when the surrounding run records are complete. It does not support general provider or product rankings.

Roadmap with evidence value

The most valuable next work is:

  1. versioned system-under-test and task manifests;
  2. repeated-run ingestion with missing/failure preservation;
  3. paired task-level comparison rather than only aggregate ranking;
  4. executable constraint and artifact checks;
  5. variance and uncertainty reporting;
  6. blinded review workflow for subjective rubrics;
  7. synthetic task packs with reset scripts and known alternative solutions.

A larger leaderboard without these controls would make the repository look more impressive while making its conclusions less trustworthy.

Publication safety

Only publish synthetic tasks, fictional outputs, and sanitized run artifacts. Do not publish private prompts, proprietary benchmark tasks, customer or employer data, confidential tool output, credentials, connector details, internal cost accounts, or comparison claims not supported by the recorded method.

Maturity and scope

This is a working prototype scorer plus a comparability methodology. It is not an independent benchmark lab, certified evaluation suite, provider endorsement, or definitive ranking of models, IDEs, or agent harnesses.


Maintained by Sima Bagheri.

About

A starter framework for transparent scoring of synthetic agent-harness run artifacts.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages