docs(eval): report the three-arm edit-contract comparison - #3158

Draft
Astro-Han wants to merge 1 commit into
mainfrom
docs/eval-edit-contract-arms-report
Draft

docs(eval): report the three-arm edit-contract comparison#3158
Astro-Han wants to merge 1 commit into
mainfrom
docs/eval-edit-contract-arms-report

Conversation

@Astro-Han

Copy link
Copy Markdown
Contributor

Summary

Adds the report and per-task CSV for the three-arm edit-contract experiment on Terminal-Bench 2.1 with deepseek-v4-flash: one harness, one model, one composition, and the file-editing tool contract as the only variable (str_replace_editor baseline vs @deepseek-ai/dsh-tool-fs vs a repo-authored V4A apply_patch plugin).

Headline results, 86 scored tasks per arm:

  • str_replace_editor 56, apply_patch 56, fs 53 — and the run has no power to separate them. Exact McNemar p = 1.00 / 0.63 / 0.65 with 95% CIs of roughly ±11 pp; the minimum detectable effect is ~13 tasks and power against the observed spread is 7%. The report states this as a failure to detect, not a finding of equivalence.
  • The arms disagree on 29/86 tasks; with no same-arm repetition the run cannot decompose that into variance vs a real per-task effect.
  • Costs are computed at DeepSeek's current time-of-day pricing, apportioned peak/off-peak per attempt window: $3.15 / $3.18 / $3.56 per arm (93% off-peak). The harness's own stale flat table would have understated this ~2.4× (fixed separately in fix(eval): update DeepSeek V4 Flash pricing to current published rates #3153).
  • Four operational findings (glibc floor from native modules, CLI double--- parsing, fs.inotify.max_user_instances as a host-wide budget, verifier bandwidth as a scheduling constraint) with their fixes.

Verification

Docs-only. Numbers are derived from the archived run results (5 run directories, later runs superseding earlier cells); the analysis script cross-checks totals against the per-arm token columns. CSV records per-task per-arm outcome, status, seconds, cost, requests and source run.

AI use

Select exactly one:

  • No generative tool made a substantive contribution
  • Generative tooling made a substantive contribution

Tool(s) and scope: Claude Code ran the experiment harness, performed the statistical analysis, and authored the report, CSV and this PR body. The commit carries a Generated-by trailer. A human contributor reviews the final text and owns the merge decision.

Checklist

  • Tests cover the change and fail without it — docs only
  • Lint, format, typecheck and the affected suites pass locally

Does this PR entail a change in behavior?

  • Yes — described under Summary above
  • No

@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Draft detected.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: cb3df9a2-0c56-4757-aea2-220c1ca8e827

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

One harness, one model, one composition, three file-editing plugins. The arms
score 56, 56 and 53 of 86 tasks, and the run has no power to separate them:
with 17 to 22 discordant pairs per comparison its minimum detectable effect is
about 13 tasks, and against the differences actually observed — zero and three
— its power is 3% to 7%. The report says that rather than reporting the three
non-significant p-values as a finding of equivalence. The 95% intervals are the
honest summary: each spans roughly 22 percentage points, so every pair is as
consistent with no difference as with one arm being nine or ten tasks better.
The arms disagree on 29 of 86 tasks. No arm was run against itself, so this
data cannot say how much of that is variance and how much is a real per-task
effect; a same-arm repetition is named as the cheapest experiment that would
make the rest of it interpretable.
Records what the arms differ in, none of it powered: fs nearly doubles the
baseline's verification failures while losing fewer cells to the deadline,
spends 30% more reasoning, and costs 20% more per pass. Records that the
treatment is a tool family rather than a diff format — fs presents four tools
and the least instruction text of the three.
Four harness failures found and fixed during the run are written up as
operational findings, because each is a prerequisite for reproducing it: the
toolchain's base image is a glibc floor for the task images, a task instruction
beginning with a dash needs two argument separators, inotify instances are a
host-wide budget that caps concurrency, and verifier bandwidth is a scheduling
constraint for tasks that build an environment.
Three tasks are unscored. The two torch tasks are reported as unscored rather
than as zeros: an agent that exhausted its budget may still have left a passing
state behind, and the verifier that would have said so never ran.
Generated-by: Claude Code
@M4n5ter
M4n5terforce-pushed the docs/eval-edit-contract-arms-report branch from 20f4778 to b8a83deCompareAugust 26, 2026 10:03
@github-actionsgithub-actionsBot added the effort/M Under 500 readable lines label Aug 27, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

effort/MUnder 500 readable lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

docs(eval): report the three-arm edit-contract comparison - #3158

Draft
Astro-Han wants to merge 1 commit into
mainfrom
docs/eval-edit-contract-arms-report
Draft

docs(eval): report the three-arm edit-contract comparison#3158
Astro-Han wants to merge 1 commit into
mainfrom
docs/eval-edit-contract-arms-report

Conversation

@Astro-Han

Copy link
Copy Markdown
Contributor

Summary

Adds the report and per-task CSV for the three-arm edit-contract experiment on Terminal-Bench 2.1 with deepseek-v4-flash: one harness, one model, one composition, and the file-editing tool contract as the only variable (str_replace_editor baseline vs @deepseek-ai/dsh-tool-fs vs a repo-authored V4A apply_patch plugin).

Headline results, 86 scored tasks per arm:

  • str_replace_editor 56, apply_patch 56, fs 53 — and the run has no power to separate them. Exact McNemar p = 1.00 / 0.63 / 0.65 with 95% CIs of roughly ±11 pp; the minimum detectable effect is ~13 tasks and power against the observed spread is 7%. The report states this as a failure to detect, not a finding of equivalence.
  • The arms disagree on 29/86 tasks; with no same-arm repetition the run cannot decompose that into variance vs a real per-task effect.
  • Costs are computed at DeepSeek's current time-of-day pricing, apportioned peak/off-peak per attempt window: $3.15 / $3.18 / $3.56 per arm (93% off-peak). The harness's own stale flat table would have understated this ~2.4× (fixed separately in fix(eval): update DeepSeek V4 Flash pricing to current published rates #3153).
  • Four operational findings (glibc floor from native modules, CLI double--- parsing, fs.inotify.max_user_instances as a host-wide budget, verifier bandwidth as a scheduling constraint) with their fixes.

Verification

Docs-only. Numbers are derived from the archived run results (5 run directories, later runs superseding earlier cells); the analysis script cross-checks totals against the per-arm token columns. CSV records per-task per-arm outcome, status, seconds, cost, requests and source run.

AI use

Select exactly one:

  • No generative tool made a substantive contribution
  • Generative tooling made a substantive contribution

Tool(s) and scope: Claude Code ran the experiment harness, performed the statistical analysis, and authored the report, CSV and this PR body. The commit carries a Generated-by trailer. A human contributor reviews the final text and owns the merge decision.

Checklist

  • Tests cover the change and fail without it — docs only
  • Lint, format, typecheck and the affected suites pass locally

Does this PR entail a change in behavior?

  • Yes — described under Summary above
  • No

@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Draft detected.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: cb3df9a2-0c56-4757-aea2-220c1ca8e827

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

One harness, one model, one composition, three file-editing plugins. The arms
score 56, 56 and 53 of 86 tasks, and the run has no power to separate them:
with 17 to 22 discordant pairs per comparison its minimum detectable effect is
about 13 tasks, and against the differences actually observed — zero and three
— its power is 3% to 7%. The report says that rather than reporting the three
non-significant p-values as a finding of equivalence. The 95% intervals are the
honest summary: each spans roughly 22 percentage points, so every pair is as
consistent with no difference as with one arm being nine or ten tasks better.
The arms disagree on 29 of 86 tasks. No arm was run against itself, so this
data cannot say how much of that is variance and how much is a real per-task
effect; a same-arm repetition is named as the cheapest experiment that would
make the rest of it interpretable.
Records what the arms differ in, none of it powered: fs nearly doubles the
baseline's verification failures while losing fewer cells to the deadline,
spends 30% more reasoning, and costs 20% more per pass. Records that the
treatment is a tool family rather than a diff format — fs presents four tools
and the least instruction text of the three.
Four harness failures found and fixed during the run are written up as
operational findings, because each is a prerequisite for reproducing it: the
toolchain's base image is a glibc floor for the task images, a task instruction
beginning with a dash needs two argument separators, inotify instances are a
host-wide budget that caps concurrency, and verifier bandwidth is a scheduling
constraint for tasks that build an environment.
Three tasks are unscored. The two torch tasks are reported as unscored rather
than as zeros: an agent that exhausted its budget may still have left a passing
state behind, and the verifier that would have said so never ran.
Generated-by: Claude Code
@M4n5ter
M4n5terforce-pushed the docs/eval-edit-contract-arms-report branch from 20f4778 to b8a83deCompareAugust 26, 2026 10:03
@github-actionsgithub-actionsBot added the effort/M Under 500 readable lines label Aug 27, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

effort/MUnder 500 readable lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

docs(eval): report the three-arm edit-contract comparison - #3158

Draft
Astro-Han wants to merge 1 commit into
mainfrom
docs/eval-edit-contract-arms-report
Draft

docs(eval): report the three-arm edit-contract comparison#3158
Astro-Han wants to merge 1 commit into
mainfrom
docs/eval-edit-contract-arms-report

Conversation

@Astro-Han

Copy link
Copy Markdown
Contributor

Summary

Adds the report and per-task CSV for the three-arm edit-contract experiment on Terminal-Bench 2.1 with deepseek-v4-flash: one harness, one model, one composition, and the file-editing tool contract as the only variable (str_replace_editor baseline vs @deepseek-ai/dsh-tool-fs vs a repo-authored V4A apply_patch plugin).

Headline results, 86 scored tasks per arm:

  • str_replace_editor 56, apply_patch 56, fs 53 — and the run has no power to separate them. Exact McNemar p = 1.00 / 0.63 / 0.65 with 95% CIs of roughly ±11 pp; the minimum detectable effect is ~13 tasks and power against the observed spread is 7%. The report states this as a failure to detect, not a finding of equivalence.
  • The arms disagree on 29/86 tasks; with no same-arm repetition the run cannot decompose that into variance vs a real per-task effect.
  • Costs are computed at DeepSeek's current time-of-day pricing, apportioned peak/off-peak per attempt window: $3.15 / $3.18 / $3.56 per arm (93% off-peak). The harness's own stale flat table would have understated this ~2.4× (fixed separately in fix(eval): update DeepSeek V4 Flash pricing to current published rates #3153).
  • Four operational findings (glibc floor from native modules, CLI double--- parsing, fs.inotify.max_user_instances as a host-wide budget, verifier bandwidth as a scheduling constraint) with their fixes.

Verification

Docs-only. Numbers are derived from the archived run results (5 run directories, later runs superseding earlier cells); the analysis script cross-checks totals against the per-arm token columns. CSV records per-task per-arm outcome, status, seconds, cost, requests and source run.

AI use

Select exactly one:

  • No generative tool made a substantive contribution
  • Generative tooling made a substantive contribution

Tool(s) and scope: Claude Code ran the experiment harness, performed the statistical analysis, and authored the report, CSV and this PR body. The commit carries a Generated-by trailer. A human contributor reviews the final text and owns the merge decision.

Checklist

  • Tests cover the change and fail without it — docs only
  • Lint, format, typecheck and the affected suites pass locally

Does this PR entail a change in behavior?

  • Yes — described under Summary above
  • No

@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Draft detected.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: cb3df9a2-0c56-4757-aea2-220c1ca8e827

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

One harness, one model, one composition, three file-editing plugins. The arms
score 56, 56 and 53 of 86 tasks, and the run has no power to separate them:
with 17 to 22 discordant pairs per comparison its minimum detectable effect is
about 13 tasks, and against the differences actually observed — zero and three
— its power is 3% to 7%. The report says that rather than reporting the three
non-significant p-values as a finding of equivalence. The 95% intervals are the
honest summary: each spans roughly 22 percentage points, so every pair is as
consistent with no difference as with one arm being nine or ten tasks better.
The arms disagree on 29 of 86 tasks. No arm was run against itself, so this
data cannot say how much of that is variance and how much is a real per-task
effect; a same-arm repetition is named as the cheapest experiment that would
make the rest of it interpretable.
Records what the arms differ in, none of it powered: fs nearly doubles the
baseline's verification failures while losing fewer cells to the deadline,
spends 30% more reasoning, and costs 20% more per pass. Records that the
treatment is a tool family rather than a diff format — fs presents four tools
and the least instruction text of the three.
Four harness failures found and fixed during the run are written up as
operational findings, because each is a prerequisite for reproducing it: the
toolchain's base image is a glibc floor for the task images, a task instruction
beginning with a dash needs two argument separators, inotify instances are a
host-wide budget that caps concurrency, and verifier bandwidth is a scheduling
constraint for tasks that build an environment.
Three tasks are unscored. The two torch tasks are reported as unscored rather
than as zeros: an agent that exhausted its budget may still have left a passing
state behind, and the verifier that would have said so never ran.
Generated-by: Claude Code
@M4n5ter
M4n5terforce-pushed the docs/eval-edit-contract-arms-report branch from 20f4778 to b8a83deCompareAugust 26, 2026 10:03
@github-actionsgithub-actionsBot added the effort/M Under 500 readable lines label Aug 27, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

effort/MUnder 500 readable lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

docs(eval): report the three-arm edit-contract comparison - #3158

Draft
Astro-Han wants to merge 1 commit into
mainfrom
docs/eval-edit-contract-arms-report
Draft

docs(eval): report the three-arm edit-contract comparison#3158
Astro-Han wants to merge 1 commit into
mainfrom
docs/eval-edit-contract-arms-report

Conversation

@Astro-Han

Copy link
Copy Markdown
Contributor

Summary

Adds the report and per-task CSV for the three-arm edit-contract experiment on Terminal-Bench 2.1 with deepseek-v4-flash: one harness, one model, one composition, and the file-editing tool contract as the only variable (str_replace_editor baseline vs @deepseek-ai/dsh-tool-fs vs a repo-authored V4A apply_patch plugin).

Headline results, 86 scored tasks per arm:

  • str_replace_editor 56, apply_patch 56, fs 53 — and the run has no power to separate them. Exact McNemar p = 1.00 / 0.63 / 0.65 with 95% CIs of roughly ±11 pp; the minimum detectable effect is ~13 tasks and power against the observed spread is 7%. The report states this as a failure to detect, not a finding of equivalence.
  • The arms disagree on 29/86 tasks; with no same-arm repetition the run cannot decompose that into variance vs a real per-task effect.
  • Costs are computed at DeepSeek's current time-of-day pricing, apportioned peak/off-peak per attempt window: $3.15 / $3.18 / $3.56 per arm (93% off-peak). The harness's own stale flat table would have understated this ~2.4× (fixed separately in fix(eval): update DeepSeek V4 Flash pricing to current published rates #3153).
  • Four operational findings (glibc floor from native modules, CLI double--- parsing, fs.inotify.max_user_instances as a host-wide budget, verifier bandwidth as a scheduling constraint) with their fixes.

Verification

Docs-only. Numbers are derived from the archived run results (5 run directories, later runs superseding earlier cells); the analysis script cross-checks totals against the per-arm token columns. CSV records per-task per-arm outcome, status, seconds, cost, requests and source run.

AI use

Select exactly one:

  • No generative tool made a substantive contribution
  • Generative tooling made a substantive contribution

Tool(s) and scope: Claude Code ran the experiment harness, performed the statistical analysis, and authored the report, CSV and this PR body. The commit carries a Generated-by trailer. A human contributor reviews the final text and owns the merge decision.

Checklist

  • Tests cover the change and fail without it — docs only
  • Lint, format, typecheck and the affected suites pass locally

Does this PR entail a change in behavior?

  • Yes — described under Summary above
  • No

@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Draft detected.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: cb3df9a2-0c56-4757-aea2-220c1ca8e827

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

One harness, one model, one composition, three file-editing plugins. The arms
score 56, 56 and 53 of 86 tasks, and the run has no power to separate them:
with 17 to 22 discordant pairs per comparison its minimum detectable effect is
about 13 tasks, and against the differences actually observed — zero and three
— its power is 3% to 7%. The report says that rather than reporting the three
non-significant p-values as a finding of equivalence. The 95% intervals are the
honest summary: each spans roughly 22 percentage points, so every pair is as
consistent with no difference as with one arm being nine or ten tasks better.
The arms disagree on 29 of 86 tasks. No arm was run against itself, so this
data cannot say how much of that is variance and how much is a real per-task
effect; a same-arm repetition is named as the cheapest experiment that would
make the rest of it interpretable.
Records what the arms differ in, none of it powered: fs nearly doubles the
baseline's verification failures while losing fewer cells to the deadline,
spends 30% more reasoning, and costs 20% more per pass. Records that the
treatment is a tool family rather than a diff format — fs presents four tools
and the least instruction text of the three.
Four harness failures found and fixed during the run are written up as
operational findings, because each is a prerequisite for reproducing it: the
toolchain's base image is a glibc floor for the task images, a task instruction
beginning with a dash needs two argument separators, inotify instances are a
host-wide budget that caps concurrency, and verifier bandwidth is a scheduling
constraint for tasks that build an environment.
Three tasks are unscored. The two torch tasks are reported as unscored rather
than as zeros: an agent that exhausted its budget may still have left a passing
state behind, and the verifier that would have said so never ran.
Generated-by: Claude Code
@M4n5ter
M4n5terforce-pushed the docs/eval-edit-contract-arms-report branch from 20f4778 to b8a83deCompareAugust 26, 2026 10:03
@github-actionsgithub-actionsBot added the effort/M Under 500 readable lines label Aug 27, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

effort/MUnder 500 readable lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

docs(eval): report the three-arm edit-contract comparison - #3158

Draft
Astro-Han wants to merge 1 commit into
mainfrom
docs/eval-edit-contract-arms-report
Draft

docs(eval): report the three-arm edit-contract comparison#3158
Astro-Han wants to merge 1 commit into
mainfrom
docs/eval-edit-contract-arms-report

Conversation

@Astro-Han

Copy link
Copy Markdown
Contributor

Summary

Adds the report and per-task CSV for the three-arm edit-contract experiment on Terminal-Bench 2.1 with deepseek-v4-flash: one harness, one model, one composition, and the file-editing tool contract as the only variable (str_replace_editor baseline vs @deepseek-ai/dsh-tool-fs vs a repo-authored V4A apply_patch plugin).

Headline results, 86 scored tasks per arm:

  • str_replace_editor 56, apply_patch 56, fs 53 — and the run has no power to separate them. Exact McNemar p = 1.00 / 0.63 / 0.65 with 95% CIs of roughly ±11 pp; the minimum detectable effect is ~13 tasks and power against the observed spread is 7%. The report states this as a failure to detect, not a finding of equivalence.
  • The arms disagree on 29/86 tasks; with no same-arm repetition the run cannot decompose that into variance vs a real per-task effect.
  • Costs are computed at DeepSeek's current time-of-day pricing, apportioned peak/off-peak per attempt window: $3.15 / $3.18 / $3.56 per arm (93% off-peak). The harness's own stale flat table would have understated this ~2.4× (fixed separately in fix(eval): update DeepSeek V4 Flash pricing to current published rates #3153).
  • Four operational findings (glibc floor from native modules, CLI double--- parsing, fs.inotify.max_user_instances as a host-wide budget, verifier bandwidth as a scheduling constraint) with their fixes.

Verification

Docs-only. Numbers are derived from the archived run results (5 run directories, later runs superseding earlier cells); the analysis script cross-checks totals against the per-arm token columns. CSV records per-task per-arm outcome, status, seconds, cost, requests and source run.

AI use

Select exactly one:

  • No generative tool made a substantive contribution
  • Generative tooling made a substantive contribution

Tool(s) and scope: Claude Code ran the experiment harness, performed the statistical analysis, and authored the report, CSV and this PR body. The commit carries a Generated-by trailer. A human contributor reviews the final text and owns the merge decision.

Checklist

  • Tests cover the change and fail without it — docs only
  • Lint, format, typecheck and the affected suites pass locally

Does this PR entail a change in behavior?

  • Yes — described under Summary above
  • No

@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Draft detected.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: cb3df9a2-0c56-4757-aea2-220c1ca8e827

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

One harness, one model, one composition, three file-editing plugins. The arms
score 56, 56 and 53 of 86 tasks, and the run has no power to separate them:
with 17 to 22 discordant pairs per comparison its minimum detectable effect is
about 13 tasks, and against the differences actually observed — zero and three
— its power is 3% to 7%. The report says that rather than reporting the three
non-significant p-values as a finding of equivalence. The 95% intervals are the
honest summary: each spans roughly 22 percentage points, so every pair is as
consistent with no difference as with one arm being nine or ten tasks better.
The arms disagree on 29 of 86 tasks. No arm was run against itself, so this
data cannot say how much of that is variance and how much is a real per-task
effect; a same-arm repetition is named as the cheapest experiment that would
make the rest of it interpretable.
Records what the arms differ in, none of it powered: fs nearly doubles the
baseline's verification failures while losing fewer cells to the deadline,
spends 30% more reasoning, and costs 20% more per pass. Records that the
treatment is a tool family rather than a diff format — fs presents four tools
and the least instruction text of the three.
Four harness failures found and fixed during the run are written up as
operational findings, because each is a prerequisite for reproducing it: the
toolchain's base image is a glibc floor for the task images, a task instruction
beginning with a dash needs two argument separators, inotify instances are a
host-wide budget that caps concurrency, and verifier bandwidth is a scheduling
constraint for tasks that build an environment.
Three tasks are unscored. The two torch tasks are reported as unscored rather
than as zeros: an agent that exhausted its budget may still have left a passing
state behind, and the verifier that would have said so never ran.
Generated-by: Claude Code
@M4n5ter
M4n5terforce-pushed the docs/eval-edit-contract-arms-report branch from 20f4778 to b8a83deCompareAugust 26, 2026 10:03
@github-actionsgithub-actionsBot added the effort/M Under 500 readable lines label Aug 27, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

effort/MUnder 500 readable lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

docs(eval): report the three-arm edit-contract comparison - #3158

Draft
Astro-Han wants to merge 1 commit into
mainfrom
docs/eval-edit-contract-arms-report
Draft

docs(eval): report the three-arm edit-contract comparison#3158
Astro-Han wants to merge 1 commit into
mainfrom
docs/eval-edit-contract-arms-report

Conversation

@Astro-Han

Copy link
Copy Markdown
Contributor

Summary

Adds the report and per-task CSV for the three-arm edit-contract experiment on Terminal-Bench 2.1 with deepseek-v4-flash: one harness, one model, one composition, and the file-editing tool contract as the only variable (str_replace_editor baseline vs @deepseek-ai/dsh-tool-fs vs a repo-authored V4A apply_patch plugin).

Headline results, 86 scored tasks per arm:

  • str_replace_editor 56, apply_patch 56, fs 53 — and the run has no power to separate them. Exact McNemar p = 1.00 / 0.63 / 0.65 with 95% CIs of roughly ±11 pp; the minimum detectable effect is ~13 tasks and power against the observed spread is 7%. The report states this as a failure to detect, not a finding of equivalence.
  • The arms disagree on 29/86 tasks; with no same-arm repetition the run cannot decompose that into variance vs a real per-task effect.
  • Costs are computed at DeepSeek's current time-of-day pricing, apportioned peak/off-peak per attempt window: $3.15 / $3.18 / $3.56 per arm (93% off-peak). The harness's own stale flat table would have understated this ~2.4× (fixed separately in fix(eval): update DeepSeek V4 Flash pricing to current published rates #3153).
  • Four operational findings (glibc floor from native modules, CLI double--- parsing, fs.inotify.max_user_instances as a host-wide budget, verifier bandwidth as a scheduling constraint) with their fixes.

Verification

Docs-only. Numbers are derived from the archived run results (5 run directories, later runs superseding earlier cells); the analysis script cross-checks totals against the per-arm token columns. CSV records per-task per-arm outcome, status, seconds, cost, requests and source run.

AI use

Select exactly one:

  • No generative tool made a substantive contribution
  • Generative tooling made a substantive contribution

Tool(s) and scope: Claude Code ran the experiment harness, performed the statistical analysis, and authored the report, CSV and this PR body. The commit carries a Generated-by trailer. A human contributor reviews the final text and owns the merge decision.

Checklist

  • Tests cover the change and fail without it — docs only
  • Lint, format, typecheck and the affected suites pass locally

Does this PR entail a change in behavior?

  • Yes — described under Summary above
  • No

@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Draft detected.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: cb3df9a2-0c56-4757-aea2-220c1ca8e827

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

One harness, one model, one composition, three file-editing plugins. The arms
score 56, 56 and 53 of 86 tasks, and the run has no power to separate them:
with 17 to 22 discordant pairs per comparison its minimum detectable effect is
about 13 tasks, and against the differences actually observed — zero and three
— its power is 3% to 7%. The report says that rather than reporting the three
non-significant p-values as a finding of equivalence. The 95% intervals are the
honest summary: each spans roughly 22 percentage points, so every pair is as
consistent with no difference as with one arm being nine or ten tasks better.
The arms disagree on 29 of 86 tasks. No arm was run against itself, so this
data cannot say how much of that is variance and how much is a real per-task
effect; a same-arm repetition is named as the cheapest experiment that would
make the rest of it interpretable.
Records what the arms differ in, none of it powered: fs nearly doubles the
baseline's verification failures while losing fewer cells to the deadline,
spends 30% more reasoning, and costs 20% more per pass. Records that the
treatment is a tool family rather than a diff format — fs presents four tools
and the least instruction text of the three.
Four harness failures found and fixed during the run are written up as
operational findings, because each is a prerequisite for reproducing it: the
toolchain's base image is a glibc floor for the task images, a task instruction
beginning with a dash needs two argument separators, inotify instances are a
host-wide budget that caps concurrency, and verifier bandwidth is a scheduling
constraint for tasks that build an environment.
Three tasks are unscored. The two torch tasks are reported as unscored rather
than as zeros: an agent that exhausted its budget may still have left a passing
state behind, and the verifier that would have said so never ran.
Generated-by: Claude Code
@M4n5ter
M4n5terforce-pushed the docs/eval-edit-contract-arms-report branch from 20f4778 to b8a83deCompareAugust 26, 2026 10:03
@github-actionsgithub-actionsBot added the effort/M Under 500 readable lines label Aug 27, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

effort/MUnder 500 readable lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

docs(eval): report the three-arm edit-contract comparison - #3158

Draft
Astro-Han wants to merge 1 commit into
mainfrom
docs/eval-edit-contract-arms-report
Draft

docs(eval): report the three-arm edit-contract comparison#3158
Astro-Han wants to merge 1 commit into
mainfrom
docs/eval-edit-contract-arms-report

Conversation

@Astro-Han

Copy link
Copy Markdown
Contributor

Summary

Adds the report and per-task CSV for the three-arm edit-contract experiment on Terminal-Bench 2.1 with deepseek-v4-flash: one harness, one model, one composition, and the file-editing tool contract as the only variable (str_replace_editor baseline vs @deepseek-ai/dsh-tool-fs vs a repo-authored V4A apply_patch plugin).

Headline results, 86 scored tasks per arm:

  • str_replace_editor 56, apply_patch 56, fs 53 — and the run has no power to separate them. Exact McNemar p = 1.00 / 0.63 / 0.65 with 95% CIs of roughly ±11 pp; the minimum detectable effect is ~13 tasks and power against the observed spread is 7%. The report states this as a failure to detect, not a finding of equivalence.
  • The arms disagree on 29/86 tasks; with no same-arm repetition the run cannot decompose that into variance vs a real per-task effect.
  • Costs are computed at DeepSeek's current time-of-day pricing, apportioned peak/off-peak per attempt window: $3.15 / $3.18 / $3.56 per arm (93% off-peak). The harness's own stale flat table would have understated this ~2.4× (fixed separately in fix(eval): update DeepSeek V4 Flash pricing to current published rates #3153).
  • Four operational findings (glibc floor from native modules, CLI double--- parsing, fs.inotify.max_user_instances as a host-wide budget, verifier bandwidth as a scheduling constraint) with their fixes.

Verification

Docs-only. Numbers are derived from the archived run results (5 run directories, later runs superseding earlier cells); the analysis script cross-checks totals against the per-arm token columns. CSV records per-task per-arm outcome, status, seconds, cost, requests and source run.

AI use

Select exactly one:

  • No generative tool made a substantive contribution
  • Generative tooling made a substantive contribution

Tool(s) and scope: Claude Code ran the experiment harness, performed the statistical analysis, and authored the report, CSV and this PR body. The commit carries a Generated-by trailer. A human contributor reviews the final text and owns the merge decision.

Checklist

  • Tests cover the change and fail without it — docs only
  • Lint, format, typecheck and the affected suites pass locally

Does this PR entail a change in behavior?

  • Yes — described under Summary above
  • No

@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Draft detected.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: cb3df9a2-0c56-4757-aea2-220c1ca8e827

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

One harness, one model, one composition, three file-editing plugins. The arms
score 56, 56 and 53 of 86 tasks, and the run has no power to separate them:
with 17 to 22 discordant pairs per comparison its minimum detectable effect is
about 13 tasks, and against the differences actually observed — zero and three
— its power is 3% to 7%. The report says that rather than reporting the three
non-significant p-values as a finding of equivalence. The 95% intervals are the
honest summary: each spans roughly 22 percentage points, so every pair is as
consistent with no difference as with one arm being nine or ten tasks better.
The arms disagree on 29 of 86 tasks. No arm was run against itself, so this
data cannot say how much of that is variance and how much is a real per-task
effect; a same-arm repetition is named as the cheapest experiment that would
make the rest of it interpretable.
Records what the arms differ in, none of it powered: fs nearly doubles the
baseline's verification failures while losing fewer cells to the deadline,
spends 30% more reasoning, and costs 20% more per pass. Records that the
treatment is a tool family rather than a diff format — fs presents four tools
and the least instruction text of the three.
Four harness failures found and fixed during the run are written up as
operational findings, because each is a prerequisite for reproducing it: the
toolchain's base image is a glibc floor for the task images, a task instruction
beginning with a dash needs two argument separators, inotify instances are a
host-wide budget that caps concurrency, and verifier bandwidth is a scheduling
constraint for tasks that build an environment.
Three tasks are unscored. The two torch tasks are reported as unscored rather
than as zeros: an agent that exhausted its budget may still have left a passing
state behind, and the verifier that would have said so never ran.
Generated-by: Claude Code
@M4n5ter
M4n5terforce-pushed the docs/eval-edit-contract-arms-report branch from 20f4778 to b8a83deCompareAugust 26, 2026 10:03
@github-actionsgithub-actionsBot added the effort/M Under 500 readable lines label Aug 27, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

effort/MUnder 500 readable lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

docs(eval): report the three-arm edit-contract comparison - #3158

Draft
Astro-Han wants to merge 1 commit into
mainfrom
docs/eval-edit-contract-arms-report
Draft

docs(eval): report the three-arm edit-contract comparison#3158
Astro-Han wants to merge 1 commit into
mainfrom
docs/eval-edit-contract-arms-report

Conversation

@Astro-Han

Copy link
Copy Markdown
Contributor

Summary

Adds the report and per-task CSV for the three-arm edit-contract experiment on Terminal-Bench 2.1 with deepseek-v4-flash: one harness, one model, one composition, and the file-editing tool contract as the only variable (str_replace_editor baseline vs @deepseek-ai/dsh-tool-fs vs a repo-authored V4A apply_patch plugin).

Headline results, 86 scored tasks per arm:

  • str_replace_editor 56, apply_patch 56, fs 53 — and the run has no power to separate them. Exact McNemar p = 1.00 / 0.63 / 0.65 with 95% CIs of roughly ±11 pp; the minimum detectable effect is ~13 tasks and power against the observed spread is 7%. The report states this as a failure to detect, not a finding of equivalence.
  • The arms disagree on 29/86 tasks; with no same-arm repetition the run cannot decompose that into variance vs a real per-task effect.
  • Costs are computed at DeepSeek's current time-of-day pricing, apportioned peak/off-peak per attempt window: $3.15 / $3.18 / $3.56 per arm (93% off-peak). The harness's own stale flat table would have understated this ~2.4× (fixed separately in fix(eval): update DeepSeek V4 Flash pricing to current published rates #3153).
  • Four operational findings (glibc floor from native modules, CLI double--- parsing, fs.inotify.max_user_instances as a host-wide budget, verifier bandwidth as a scheduling constraint) with their fixes.

Verification

Docs-only. Numbers are derived from the archived run results (5 run directories, later runs superseding earlier cells); the analysis script cross-checks totals against the per-arm token columns. CSV records per-task per-arm outcome, status, seconds, cost, requests and source run.

AI use

Select exactly one:

  • No generative tool made a substantive contribution
  • Generative tooling made a substantive contribution

Tool(s) and scope: Claude Code ran the experiment harness, performed the statistical analysis, and authored the report, CSV and this PR body. The commit carries a Generated-by trailer. A human contributor reviews the final text and owns the merge decision.

Checklist

  • Tests cover the change and fail without it — docs only
  • Lint, format, typecheck and the affected suites pass locally

Does this PR entail a change in behavior?

  • Yes — described under Summary above
  • No

@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Draft detected.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: cb3df9a2-0c56-4757-aea2-220c1ca8e827

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

One harness, one model, one composition, three file-editing plugins. The arms
score 56, 56 and 53 of 86 tasks, and the run has no power to separate them:
with 17 to 22 discordant pairs per comparison its minimum detectable effect is
about 13 tasks, and against the differences actually observed — zero and three
— its power is 3% to 7%. The report says that rather than reporting the three
non-significant p-values as a finding of equivalence. The 95% intervals are the
honest summary: each spans roughly 22 percentage points, so every pair is as
consistent with no difference as with one arm being nine or ten tasks better.
The arms disagree on 29 of 86 tasks. No arm was run against itself, so this
data cannot say how much of that is variance and how much is a real per-task
effect; a same-arm repetition is named as the cheapest experiment that would
make the rest of it interpretable.
Records what the arms differ in, none of it powered: fs nearly doubles the
baseline's verification failures while losing fewer cells to the deadline,
spends 30% more reasoning, and costs 20% more per pass. Records that the
treatment is a tool family rather than a diff format — fs presents four tools
and the least instruction text of the three.
Four harness failures found and fixed during the run are written up as
operational findings, because each is a prerequisite for reproducing it: the
toolchain's base image is a glibc floor for the task images, a task instruction
beginning with a dash needs two argument separators, inotify instances are a
host-wide budget that caps concurrency, and verifier bandwidth is a scheduling
constraint for tasks that build an environment.
Three tasks are unscored. The two torch tasks are reported as unscored rather
than as zeros: an agent that exhausted its budget may still have left a passing
state behind, and the verifier that would have said so never ran.
Generated-by: Claude Code
@M4n5ter
M4n5terforce-pushed the docs/eval-edit-contract-arms-report branch from 20f4778 to b8a83deCompareAugust 26, 2026 10:03
@github-actionsgithub-actionsBot added the effort/M Under 500 readable lines label Aug 27, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

effort/MUnder 500 readable lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han