refactor(headless): move the benchmark identity out of the build a graded arm mounts - #2303

Merged
Astro-Han merged 4 commits into
mainfrom
refactor/benchmark-identity-out-of-dist
Aug 6, 2026
Merged

refactor(headless): move the benchmark identity out of the build a graded arm mounts#2303
Astro-Han merged 4 commits into
mainfrom
refactor/benchmark-identity-out-of-dist

Conversation

@Astro-Han

Copy link
Copy Markdown
Contributor

What

A graded arm executes packages/headless/dist, so anything compiled into a shipped module is readable from inside the container being scored. harness-ab-manifest.ts carried each benchmark's pinned revision, task-tree fingerprint and full task list — 89 Terminal-Bench ids and 113 DeepSWE ids — and every one of them shipped.

That is the exact material of the #1970 incident: an arm read the pinned revision out of the mount, fetched the task's reference solution from the public repository at that revision, and recorded a pass that measured retrieval.

#2298 narrowed the mount so docs/eval and src no longer reach the container. It could not close this one, because the mount is a list of paths and the identity was inside a path the container legitimately needs.

How

  • The data moves to packages/headless/harbor/benchmark-identity.json, loaded by harbor/benchmark-identity.mjs. Neither is mounted by any arm — agent-repo-mount.ts names exactly one file out of harbor/, and it is run-cell.mjs.
  • harness-ab-manifest.ts keeps the assertions and loses the answers: assertTerminalBench21TaskSet and its three siblings become assertFrozenTaskSet(label, expected, actual) and assertFrozenTaskTreeFingerprint(label, expected, actual). Error messages are unchanged.
  • Each entry in HARNESS_BENCHMARK_PROFILES now carries its own taskTreeFingerprint. With that, the DeepSWE-full and Terminal-Bench task-source branches were the same code under different names and collapse into one; only the subset-30 profile still differs, because its tree carries more tasks than the frozen set.
  • No product module was affected — the identity was only ever read by the two host-side harbor/*.mjs entrypoints and by tests.

Verification

A new test in harness-ab-manifest.test.ts reads every shipped dist/*.js and asserts it contains no revision, no upstream URL and none of the 202 task ids. This is the first check on what the mounted files contain rather than on which paths are declared, and it immediately found a real leak: harbor-smoke-config.ts hardcoded '*sqlite-with-gcov' — a Terminal-Bench 2.1 task id — as a fallback the smoke manifest was already supplying. Removed, and it now fails loudly instead.

Red/green confirmed: restoring that literal fails the new test.

Loading the real entrypoint reproduces all three profiles byte-for-byte:

terminal-bench-2.1 d49e28f1 89 sha256:456826a
deep-swe-1.1 6db64a40 30 sha256:508aedc
deep-swe-1.1-full 6db64a40 113 sha256:973091a

npm run -w @maka/headless test — 1422 pass, 0 fail. Lint and format clean.

Known residue

Compiled tests under dist/__tests__ are still mounted and still mention individual task ids in fixtures. Closing that means splitting the test build out of dist, which is a cross-cutting change to every workspace's tsconfig and the test runner; a fixture id is also not the benchmark's identity in the way a pinned revision and a complete task list are. Left for its own change.

dist/terminal-bench-adapter.js remains readable and remains in the graded closure. It is genuinely needed to run, and what it reveals — a verifier grades by exit code — is generic harness mechanics rather than any task's answer.

@Astro-Han
Astro-Hanforce-pushed the refactor/benchmark-identity-out-of-dist branch from 1a563c6 to 2a08231CompareAugust 6, 2026 02:12
@Astro-Han
Astro-Han changed the base branch from main to fix/harness-maka-repo-mount-scopeAugust 6, 2026 02:12
@Astro-Han
Astro-Hanforce-pushed the refactor/benchmark-identity-out-of-dist branch 2 times, most recently from e79c5ba to 683e1c9CompareAugust 6, 2026 06:04
@Astro-Han
Astro-Han changed the base branch from fix/harness-maka-repo-mount-scope to mainAugust 6, 2026 06:04
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
Review of the first cut found the comments claiming more than the test held.
The content check read only the top level of dist while the whole directory
is mounted, and described the compiled tests as excluded when they are not.
It now scans dist recursively and states two boundaries, because they are two
different things: the retrieval keys — revision, upstream URL, tree
fingerprint — must be absent from every mounted file, while a single task id
is not a retrieval key and is only held out of the shipped modules, where a
complete list could accumulate.
No mount list may declare the identity file, the Maka one included, and the
DeepSWE cache path derives its revision instead of spelling a second copy.
…t forgot
The scan filtered to `.js`, so a source map — whose `sourcesContent` carries
the whole of `src` — would have walked straight past it. And the retrieval
keys named every task-tree fingerprint but one.
@Astro-Han
Astro-Hanforce-pushed the refactor/benchmark-identity-out-of-dist branch from 683e1c9 to d0e7cf3CompareAugust 6, 2026 06:11
@Astro-Han
Astro-Han marked this pull request as ready for review August 6, 2026 06:25
@Astro-Han
Astro-Han merged commit 43b6694 into mainAug 6, 2026
11 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

refactor(headless): move the benchmark identity out of the build a graded arm mounts - #2303

Merged
Astro-Han merged 4 commits into
mainfrom
refactor/benchmark-identity-out-of-dist
Aug 6, 2026
Merged

refactor(headless): move the benchmark identity out of the build a graded arm mounts#2303
Astro-Han merged 4 commits into
mainfrom
refactor/benchmark-identity-out-of-dist

Conversation

@Astro-Han

Copy link
Copy Markdown
Contributor

What

A graded arm executes packages/headless/dist, so anything compiled into a shipped module is readable from inside the container being scored. harness-ab-manifest.ts carried each benchmark's pinned revision, task-tree fingerprint and full task list — 89 Terminal-Bench ids and 113 DeepSWE ids — and every one of them shipped.

That is the exact material of the #1970 incident: an arm read the pinned revision out of the mount, fetched the task's reference solution from the public repository at that revision, and recorded a pass that measured retrieval.

#2298 narrowed the mount so docs/eval and src no longer reach the container. It could not close this one, because the mount is a list of paths and the identity was inside a path the container legitimately needs.

How

  • The data moves to packages/headless/harbor/benchmark-identity.json, loaded by harbor/benchmark-identity.mjs. Neither is mounted by any arm — agent-repo-mount.ts names exactly one file out of harbor/, and it is run-cell.mjs.
  • harness-ab-manifest.ts keeps the assertions and loses the answers: assertTerminalBench21TaskSet and its three siblings become assertFrozenTaskSet(label, expected, actual) and assertFrozenTaskTreeFingerprint(label, expected, actual). Error messages are unchanged.
  • Each entry in HARNESS_BENCHMARK_PROFILES now carries its own taskTreeFingerprint. With that, the DeepSWE-full and Terminal-Bench task-source branches were the same code under different names and collapse into one; only the subset-30 profile still differs, because its tree carries more tasks than the frozen set.
  • No product module was affected — the identity was only ever read by the two host-side harbor/*.mjs entrypoints and by tests.

Verification

A new test in harness-ab-manifest.test.ts reads every shipped dist/*.js and asserts it contains no revision, no upstream URL and none of the 202 task ids. This is the first check on what the mounted files contain rather than on which paths are declared, and it immediately found a real leak: harbor-smoke-config.ts hardcoded '*sqlite-with-gcov' — a Terminal-Bench 2.1 task id — as a fallback the smoke manifest was already supplying. Removed, and it now fails loudly instead.

Red/green confirmed: restoring that literal fails the new test.

Loading the real entrypoint reproduces all three profiles byte-for-byte:

terminal-bench-2.1 d49e28f1 89 sha256:456826a
deep-swe-1.1 6db64a40 30 sha256:508aedc
deep-swe-1.1-full 6db64a40 113 sha256:973091a

npm run -w @maka/headless test — 1422 pass, 0 fail. Lint and format clean.

Known residue

Compiled tests under dist/__tests__ are still mounted and still mention individual task ids in fixtures. Closing that means splitting the test build out of dist, which is a cross-cutting change to every workspace's tsconfig and the test runner; a fixture id is also not the benchmark's identity in the way a pinned revision and a complete task list are. Left for its own change.

dist/terminal-bench-adapter.js remains readable and remains in the graded closure. It is genuinely needed to run, and what it reveals — a verifier grades by exit code — is generic harness mechanics rather than any task's answer.

@Astro-Han
Astro-Hanforce-pushed the refactor/benchmark-identity-out-of-dist branch from 1a563c6 to 2a08231CompareAugust 6, 2026 02:12
@Astro-Han
Astro-Han changed the base branch from main to fix/harness-maka-repo-mount-scopeAugust 6, 2026 02:12
@Astro-Han
Astro-Hanforce-pushed the refactor/benchmark-identity-out-of-dist branch 2 times, most recently from e79c5ba to 683e1c9CompareAugust 6, 2026 06:04
@Astro-Han
Astro-Han changed the base branch from fix/harness-maka-repo-mount-scope to mainAugust 6, 2026 06:04
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
Review of the first cut found the comments claiming more than the test held.
The content check read only the top level of dist while the whole directory
is mounted, and described the compiled tests as excluded when they are not.
It now scans dist recursively and states two boundaries, because they are two
different things: the retrieval keys — revision, upstream URL, tree
fingerprint — must be absent from every mounted file, while a single task id
is not a retrieval key and is only held out of the shipped modules, where a
complete list could accumulate.
No mount list may declare the identity file, the Maka one included, and the
DeepSWE cache path derives its revision instead of spelling a second copy.
…t forgot
The scan filtered to `.js`, so a source map — whose `sourcesContent` carries
the whole of `src` — would have walked straight past it. And the retrieval
keys named every task-tree fingerprint but one.
@Astro-Han
Astro-Hanforce-pushed the refactor/benchmark-identity-out-of-dist branch from 683e1c9 to d0e7cf3CompareAugust 6, 2026 06:11
@Astro-Han
Astro-Han marked this pull request as ready for review August 6, 2026 06:25
@Astro-Han
Astro-Han merged commit 43b6694 into mainAug 6, 2026
11 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

refactor(headless): move the benchmark identity out of the build a graded arm mounts - #2303

Merged
Astro-Han merged 4 commits into
mainfrom
refactor/benchmark-identity-out-of-dist
Aug 6, 2026
Merged

refactor(headless): move the benchmark identity out of the build a graded arm mounts#2303
Astro-Han merged 4 commits into
mainfrom
refactor/benchmark-identity-out-of-dist

Conversation

@Astro-Han

Copy link
Copy Markdown
Contributor

What

A graded arm executes packages/headless/dist, so anything compiled into a shipped module is readable from inside the container being scored. harness-ab-manifest.ts carried each benchmark's pinned revision, task-tree fingerprint and full task list — 89 Terminal-Bench ids and 113 DeepSWE ids — and every one of them shipped.

That is the exact material of the #1970 incident: an arm read the pinned revision out of the mount, fetched the task's reference solution from the public repository at that revision, and recorded a pass that measured retrieval.

#2298 narrowed the mount so docs/eval and src no longer reach the container. It could not close this one, because the mount is a list of paths and the identity was inside a path the container legitimately needs.

How

  • The data moves to packages/headless/harbor/benchmark-identity.json, loaded by harbor/benchmark-identity.mjs. Neither is mounted by any arm — agent-repo-mount.ts names exactly one file out of harbor/, and it is run-cell.mjs.
  • harness-ab-manifest.ts keeps the assertions and loses the answers: assertTerminalBench21TaskSet and its three siblings become assertFrozenTaskSet(label, expected, actual) and assertFrozenTaskTreeFingerprint(label, expected, actual). Error messages are unchanged.
  • Each entry in HARNESS_BENCHMARK_PROFILES now carries its own taskTreeFingerprint. With that, the DeepSWE-full and Terminal-Bench task-source branches were the same code under different names and collapse into one; only the subset-30 profile still differs, because its tree carries more tasks than the frozen set.
  • No product module was affected — the identity was only ever read by the two host-side harbor/*.mjs entrypoints and by tests.

Verification

A new test in harness-ab-manifest.test.ts reads every shipped dist/*.js and asserts it contains no revision, no upstream URL and none of the 202 task ids. This is the first check on what the mounted files contain rather than on which paths are declared, and it immediately found a real leak: harbor-smoke-config.ts hardcoded '*sqlite-with-gcov' — a Terminal-Bench 2.1 task id — as a fallback the smoke manifest was already supplying. Removed, and it now fails loudly instead.

Red/green confirmed: restoring that literal fails the new test.

Loading the real entrypoint reproduces all three profiles byte-for-byte:

terminal-bench-2.1 d49e28f1 89 sha256:456826a
deep-swe-1.1 6db64a40 30 sha256:508aedc
deep-swe-1.1-full 6db64a40 113 sha256:973091a

npm run -w @maka/headless test — 1422 pass, 0 fail. Lint and format clean.

Known residue

Compiled tests under dist/__tests__ are still mounted and still mention individual task ids in fixtures. Closing that means splitting the test build out of dist, which is a cross-cutting change to every workspace's tsconfig and the test runner; a fixture id is also not the benchmark's identity in the way a pinned revision and a complete task list are. Left for its own change.

dist/terminal-bench-adapter.js remains readable and remains in the graded closure. It is genuinely needed to run, and what it reveals — a verifier grades by exit code — is generic harness mechanics rather than any task's answer.

@Astro-Han
Astro-Hanforce-pushed the refactor/benchmark-identity-out-of-dist branch from 1a563c6 to 2a08231CompareAugust 6, 2026 02:12
@Astro-Han
Astro-Han changed the base branch from main to fix/harness-maka-repo-mount-scopeAugust 6, 2026 02:12
@Astro-Han
Astro-Hanforce-pushed the refactor/benchmark-identity-out-of-dist branch 2 times, most recently from e79c5ba to 683e1c9CompareAugust 6, 2026 06:04
@Astro-Han
Astro-Han changed the base branch from fix/harness-maka-repo-mount-scope to mainAugust 6, 2026 06:04
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
Review of the first cut found the comments claiming more than the test held.
The content check read only the top level of dist while the whole directory
is mounted, and described the compiled tests as excluded when they are not.
It now scans dist recursively and states two boundaries, because they are two
different things: the retrieval keys — revision, upstream URL, tree
fingerprint — must be absent from every mounted file, while a single task id
is not a retrieval key and is only held out of the shipped modules, where a
complete list could accumulate.
No mount list may declare the identity file, the Maka one included, and the
DeepSWE cache path derives its revision instead of spelling a second copy.
…t forgot
The scan filtered to `.js`, so a source map — whose `sourcesContent` carries
the whole of `src` — would have walked straight past it. And the retrieval
keys named every task-tree fingerprint but one.
@Astro-Han
Astro-Hanforce-pushed the refactor/benchmark-identity-out-of-dist branch from 683e1c9 to d0e7cf3CompareAugust 6, 2026 06:11
@Astro-Han
Astro-Han marked this pull request as ready for review August 6, 2026 06:25
@Astro-Han
Astro-Han merged commit 43b6694 into mainAug 6, 2026
11 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

refactor(headless): move the benchmark identity out of the build a graded arm mounts - #2303

Merged
Astro-Han merged 4 commits into
mainfrom
refactor/benchmark-identity-out-of-dist
Aug 6, 2026
Merged

refactor(headless): move the benchmark identity out of the build a graded arm mounts#2303
Astro-Han merged 4 commits into
mainfrom
refactor/benchmark-identity-out-of-dist

Conversation

@Astro-Han

Copy link
Copy Markdown
Contributor

What

A graded arm executes packages/headless/dist, so anything compiled into a shipped module is readable from inside the container being scored. harness-ab-manifest.ts carried each benchmark's pinned revision, task-tree fingerprint and full task list — 89 Terminal-Bench ids and 113 DeepSWE ids — and every one of them shipped.

That is the exact material of the #1970 incident: an arm read the pinned revision out of the mount, fetched the task's reference solution from the public repository at that revision, and recorded a pass that measured retrieval.

#2298 narrowed the mount so docs/eval and src no longer reach the container. It could not close this one, because the mount is a list of paths and the identity was inside a path the container legitimately needs.

How

  • The data moves to packages/headless/harbor/benchmark-identity.json, loaded by harbor/benchmark-identity.mjs. Neither is mounted by any arm — agent-repo-mount.ts names exactly one file out of harbor/, and it is run-cell.mjs.
  • harness-ab-manifest.ts keeps the assertions and loses the answers: assertTerminalBench21TaskSet and its three siblings become assertFrozenTaskSet(label, expected, actual) and assertFrozenTaskTreeFingerprint(label, expected, actual). Error messages are unchanged.
  • Each entry in HARNESS_BENCHMARK_PROFILES now carries its own taskTreeFingerprint. With that, the DeepSWE-full and Terminal-Bench task-source branches were the same code under different names and collapse into one; only the subset-30 profile still differs, because its tree carries more tasks than the frozen set.
  • No product module was affected — the identity was only ever read by the two host-side harbor/*.mjs entrypoints and by tests.

Verification

A new test in harness-ab-manifest.test.ts reads every shipped dist/*.js and asserts it contains no revision, no upstream URL and none of the 202 task ids. This is the first check on what the mounted files contain rather than on which paths are declared, and it immediately found a real leak: harbor-smoke-config.ts hardcoded '*sqlite-with-gcov' — a Terminal-Bench 2.1 task id — as a fallback the smoke manifest was already supplying. Removed, and it now fails loudly instead.

Red/green confirmed: restoring that literal fails the new test.

Loading the real entrypoint reproduces all three profiles byte-for-byte:

terminal-bench-2.1 d49e28f1 89 sha256:456826a
deep-swe-1.1 6db64a40 30 sha256:508aedc
deep-swe-1.1-full 6db64a40 113 sha256:973091a

npm run -w @maka/headless test — 1422 pass, 0 fail. Lint and format clean.

Known residue

Compiled tests under dist/__tests__ are still mounted and still mention individual task ids in fixtures. Closing that means splitting the test build out of dist, which is a cross-cutting change to every workspace's tsconfig and the test runner; a fixture id is also not the benchmark's identity in the way a pinned revision and a complete task list are. Left for its own change.

dist/terminal-bench-adapter.js remains readable and remains in the graded closure. It is genuinely needed to run, and what it reveals — a verifier grades by exit code — is generic harness mechanics rather than any task's answer.

@Astro-Han
Astro-Hanforce-pushed the refactor/benchmark-identity-out-of-dist branch from 1a563c6 to 2a08231CompareAugust 6, 2026 02:12
@Astro-Han
Astro-Han changed the base branch from main to fix/harness-maka-repo-mount-scopeAugust 6, 2026 02:12
@Astro-Han
Astro-Hanforce-pushed the refactor/benchmark-identity-out-of-dist branch 2 times, most recently from e79c5ba to 683e1c9CompareAugust 6, 2026 06:04
@Astro-Han
Astro-Han changed the base branch from fix/harness-maka-repo-mount-scope to mainAugust 6, 2026 06:04
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
Review of the first cut found the comments claiming more than the test held.
The content check read only the top level of dist while the whole directory
is mounted, and described the compiled tests as excluded when they are not.
It now scans dist recursively and states two boundaries, because they are two
different things: the retrieval keys — revision, upstream URL, tree
fingerprint — must be absent from every mounted file, while a single task id
is not a retrieval key and is only held out of the shipped modules, where a
complete list could accumulate.
No mount list may declare the identity file, the Maka one included, and the
DeepSWE cache path derives its revision instead of spelling a second copy.
…t forgot
The scan filtered to `.js`, so a source map — whose `sourcesContent` carries
the whole of `src` — would have walked straight past it. And the retrieval
keys named every task-tree fingerprint but one.
@Astro-Han
Astro-Hanforce-pushed the refactor/benchmark-identity-out-of-dist branch from 683e1c9 to d0e7cf3CompareAugust 6, 2026 06:11
@Astro-Han
Astro-Han marked this pull request as ready for review August 6, 2026 06:25
@Astro-Han
Astro-Han merged commit 43b6694 into mainAug 6, 2026
11 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

refactor(headless): move the benchmark identity out of the build a graded arm mounts - #2303

Merged
Astro-Han merged 4 commits into
mainfrom
refactor/benchmark-identity-out-of-dist
Aug 6, 2026
Merged

refactor(headless): move the benchmark identity out of the build a graded arm mounts#2303
Astro-Han merged 4 commits into
mainfrom
refactor/benchmark-identity-out-of-dist

Conversation

@Astro-Han

Copy link
Copy Markdown
Contributor

What

A graded arm executes packages/headless/dist, so anything compiled into a shipped module is readable from inside the container being scored. harness-ab-manifest.ts carried each benchmark's pinned revision, task-tree fingerprint and full task list — 89 Terminal-Bench ids and 113 DeepSWE ids — and every one of them shipped.

That is the exact material of the #1970 incident: an arm read the pinned revision out of the mount, fetched the task's reference solution from the public repository at that revision, and recorded a pass that measured retrieval.

#2298 narrowed the mount so docs/eval and src no longer reach the container. It could not close this one, because the mount is a list of paths and the identity was inside a path the container legitimately needs.

How

  • The data moves to packages/headless/harbor/benchmark-identity.json, loaded by harbor/benchmark-identity.mjs. Neither is mounted by any arm — agent-repo-mount.ts names exactly one file out of harbor/, and it is run-cell.mjs.
  • harness-ab-manifest.ts keeps the assertions and loses the answers: assertTerminalBench21TaskSet and its three siblings become assertFrozenTaskSet(label, expected, actual) and assertFrozenTaskTreeFingerprint(label, expected, actual). Error messages are unchanged.
  • Each entry in HARNESS_BENCHMARK_PROFILES now carries its own taskTreeFingerprint. With that, the DeepSWE-full and Terminal-Bench task-source branches were the same code under different names and collapse into one; only the subset-30 profile still differs, because its tree carries more tasks than the frozen set.
  • No product module was affected — the identity was only ever read by the two host-side harbor/*.mjs entrypoints and by tests.

Verification

A new test in harness-ab-manifest.test.ts reads every shipped dist/*.js and asserts it contains no revision, no upstream URL and none of the 202 task ids. This is the first check on what the mounted files contain rather than on which paths are declared, and it immediately found a real leak: harbor-smoke-config.ts hardcoded '*sqlite-with-gcov' — a Terminal-Bench 2.1 task id — as a fallback the smoke manifest was already supplying. Removed, and it now fails loudly instead.

Red/green confirmed: restoring that literal fails the new test.

Loading the real entrypoint reproduces all three profiles byte-for-byte:

terminal-bench-2.1 d49e28f1 89 sha256:456826a
deep-swe-1.1 6db64a40 30 sha256:508aedc
deep-swe-1.1-full 6db64a40 113 sha256:973091a

npm run -w @maka/headless test — 1422 pass, 0 fail. Lint and format clean.

Known residue

Compiled tests under dist/__tests__ are still mounted and still mention individual task ids in fixtures. Closing that means splitting the test build out of dist, which is a cross-cutting change to every workspace's tsconfig and the test runner; a fixture id is also not the benchmark's identity in the way a pinned revision and a complete task list are. Left for its own change.

dist/terminal-bench-adapter.js remains readable and remains in the graded closure. It is genuinely needed to run, and what it reveals — a verifier grades by exit code — is generic harness mechanics rather than any task's answer.

@Astro-Han
Astro-Hanforce-pushed the refactor/benchmark-identity-out-of-dist branch from 1a563c6 to 2a08231CompareAugust 6, 2026 02:12
@Astro-Han
Astro-Han changed the base branch from main to fix/harness-maka-repo-mount-scopeAugust 6, 2026 02:12
@Astro-Han
Astro-Hanforce-pushed the refactor/benchmark-identity-out-of-dist branch 2 times, most recently from e79c5ba to 683e1c9CompareAugust 6, 2026 06:04
@Astro-Han
Astro-Han changed the base branch from fix/harness-maka-repo-mount-scope to mainAugust 6, 2026 06:04
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
Review of the first cut found the comments claiming more than the test held.
The content check read only the top level of dist while the whole directory
is mounted, and described the compiled tests as excluded when they are not.
It now scans dist recursively and states two boundaries, because they are two
different things: the retrieval keys — revision, upstream URL, tree
fingerprint — must be absent from every mounted file, while a single task id
is not a retrieval key and is only held out of the shipped modules, where a
complete list could accumulate.
No mount list may declare the identity file, the Maka one included, and the
DeepSWE cache path derives its revision instead of spelling a second copy.
…t forgot
The scan filtered to `.js`, so a source map — whose `sourcesContent` carries
the whole of `src` — would have walked straight past it. And the retrieval
keys named every task-tree fingerprint but one.
@Astro-Han
Astro-Hanforce-pushed the refactor/benchmark-identity-out-of-dist branch from 683e1c9 to d0e7cf3CompareAugust 6, 2026 06:11
@Astro-Han
Astro-Han marked this pull request as ready for review August 6, 2026 06:25
@Astro-Han
Astro-Han merged commit 43b6694 into mainAug 6, 2026
11 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

refactor(headless): move the benchmark identity out of the build a graded arm mounts - #2303

Merged
Astro-Han merged 4 commits into
mainfrom
refactor/benchmark-identity-out-of-dist
Aug 6, 2026
Merged

refactor(headless): move the benchmark identity out of the build a graded arm mounts#2303
Astro-Han merged 4 commits into
mainfrom
refactor/benchmark-identity-out-of-dist

Conversation

@Astro-Han

Copy link
Copy Markdown
Contributor

What

A graded arm executes packages/headless/dist, so anything compiled into a shipped module is readable from inside the container being scored. harness-ab-manifest.ts carried each benchmark's pinned revision, task-tree fingerprint and full task list — 89 Terminal-Bench ids and 113 DeepSWE ids — and every one of them shipped.

That is the exact material of the #1970 incident: an arm read the pinned revision out of the mount, fetched the task's reference solution from the public repository at that revision, and recorded a pass that measured retrieval.

#2298 narrowed the mount so docs/eval and src no longer reach the container. It could not close this one, because the mount is a list of paths and the identity was inside a path the container legitimately needs.

How

  • The data moves to packages/headless/harbor/benchmark-identity.json, loaded by harbor/benchmark-identity.mjs. Neither is mounted by any arm — agent-repo-mount.ts names exactly one file out of harbor/, and it is run-cell.mjs.
  • harness-ab-manifest.ts keeps the assertions and loses the answers: assertTerminalBench21TaskSet and its three siblings become assertFrozenTaskSet(label, expected, actual) and assertFrozenTaskTreeFingerprint(label, expected, actual). Error messages are unchanged.
  • Each entry in HARNESS_BENCHMARK_PROFILES now carries its own taskTreeFingerprint. With that, the DeepSWE-full and Terminal-Bench task-source branches were the same code under different names and collapse into one; only the subset-30 profile still differs, because its tree carries more tasks than the frozen set.
  • No product module was affected — the identity was only ever read by the two host-side harbor/*.mjs entrypoints and by tests.

Verification

A new test in harness-ab-manifest.test.ts reads every shipped dist/*.js and asserts it contains no revision, no upstream URL and none of the 202 task ids. This is the first check on what the mounted files contain rather than on which paths are declared, and it immediately found a real leak: harbor-smoke-config.ts hardcoded '*sqlite-with-gcov' — a Terminal-Bench 2.1 task id — as a fallback the smoke manifest was already supplying. Removed, and it now fails loudly instead.

Red/green confirmed: restoring that literal fails the new test.

Loading the real entrypoint reproduces all three profiles byte-for-byte:

terminal-bench-2.1 d49e28f1 89 sha256:456826a
deep-swe-1.1 6db64a40 30 sha256:508aedc
deep-swe-1.1-full 6db64a40 113 sha256:973091a

npm run -w @maka/headless test — 1422 pass, 0 fail. Lint and format clean.

Known residue

Compiled tests under dist/__tests__ are still mounted and still mention individual task ids in fixtures. Closing that means splitting the test build out of dist, which is a cross-cutting change to every workspace's tsconfig and the test runner; a fixture id is also not the benchmark's identity in the way a pinned revision and a complete task list are. Left for its own change.

dist/terminal-bench-adapter.js remains readable and remains in the graded closure. It is genuinely needed to run, and what it reveals — a verifier grades by exit code — is generic harness mechanics rather than any task's answer.

@Astro-Han
Astro-Hanforce-pushed the refactor/benchmark-identity-out-of-dist branch from 1a563c6 to 2a08231CompareAugust 6, 2026 02:12
@Astro-Han
Astro-Han changed the base branch from main to fix/harness-maka-repo-mount-scopeAugust 6, 2026 02:12
@Astro-Han
Astro-Hanforce-pushed the refactor/benchmark-identity-out-of-dist branch 2 times, most recently from e79c5ba to 683e1c9CompareAugust 6, 2026 06:04
@Astro-Han
Astro-Han changed the base branch from fix/harness-maka-repo-mount-scope to mainAugust 6, 2026 06:04
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
Review of the first cut found the comments claiming more than the test held.
The content check read only the top level of dist while the whole directory
is mounted, and described the compiled tests as excluded when they are not.
It now scans dist recursively and states two boundaries, because they are two
different things: the retrieval keys — revision, upstream URL, tree
fingerprint — must be absent from every mounted file, while a single task id
is not a retrieval key and is only held out of the shipped modules, where a
complete list could accumulate.
No mount list may declare the identity file, the Maka one included, and the
DeepSWE cache path derives its revision instead of spelling a second copy.
…t forgot
The scan filtered to `.js`, so a source map — whose `sourcesContent` carries
the whole of `src` — would have walked straight past it. And the retrieval
keys named every task-tree fingerprint but one.
@Astro-Han
Astro-Hanforce-pushed the refactor/benchmark-identity-out-of-dist branch from 683e1c9 to d0e7cf3CompareAugust 6, 2026 06:11
@Astro-Han
Astro-Han marked this pull request as ready for review August 6, 2026 06:25
@Astro-Han
Astro-Han merged commit 43b6694 into mainAug 6, 2026
11 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

refactor(headless): move the benchmark identity out of the build a graded arm mounts - #2303

Merged
Astro-Han merged 4 commits into
mainfrom
refactor/benchmark-identity-out-of-dist
Aug 6, 2026
Merged

refactor(headless): move the benchmark identity out of the build a graded arm mounts#2303
Astro-Han merged 4 commits into
mainfrom
refactor/benchmark-identity-out-of-dist

Conversation

@Astro-Han

Copy link
Copy Markdown
Contributor

What

A graded arm executes packages/headless/dist, so anything compiled into a shipped module is readable from inside the container being scored. harness-ab-manifest.ts carried each benchmark's pinned revision, task-tree fingerprint and full task list — 89 Terminal-Bench ids and 113 DeepSWE ids — and every one of them shipped.

That is the exact material of the #1970 incident: an arm read the pinned revision out of the mount, fetched the task's reference solution from the public repository at that revision, and recorded a pass that measured retrieval.

#2298 narrowed the mount so docs/eval and src no longer reach the container. It could not close this one, because the mount is a list of paths and the identity was inside a path the container legitimately needs.

How

  • The data moves to packages/headless/harbor/benchmark-identity.json, loaded by harbor/benchmark-identity.mjs. Neither is mounted by any arm — agent-repo-mount.ts names exactly one file out of harbor/, and it is run-cell.mjs.
  • harness-ab-manifest.ts keeps the assertions and loses the answers: assertTerminalBench21TaskSet and its three siblings become assertFrozenTaskSet(label, expected, actual) and assertFrozenTaskTreeFingerprint(label, expected, actual). Error messages are unchanged.
  • Each entry in HARNESS_BENCHMARK_PROFILES now carries its own taskTreeFingerprint. With that, the DeepSWE-full and Terminal-Bench task-source branches were the same code under different names and collapse into one; only the subset-30 profile still differs, because its tree carries more tasks than the frozen set.
  • No product module was affected — the identity was only ever read by the two host-side harbor/*.mjs entrypoints and by tests.

Verification

A new test in harness-ab-manifest.test.ts reads every shipped dist/*.js and asserts it contains no revision, no upstream URL and none of the 202 task ids. This is the first check on what the mounted files contain rather than on which paths are declared, and it immediately found a real leak: harbor-smoke-config.ts hardcoded '*sqlite-with-gcov' — a Terminal-Bench 2.1 task id — as a fallback the smoke manifest was already supplying. Removed, and it now fails loudly instead.

Red/green confirmed: restoring that literal fails the new test.

Loading the real entrypoint reproduces all three profiles byte-for-byte:

terminal-bench-2.1 d49e28f1 89 sha256:456826a
deep-swe-1.1 6db64a40 30 sha256:508aedc
deep-swe-1.1-full 6db64a40 113 sha256:973091a

npm run -w @maka/headless test — 1422 pass, 0 fail. Lint and format clean.

Known residue

Compiled tests under dist/__tests__ are still mounted and still mention individual task ids in fixtures. Closing that means splitting the test build out of dist, which is a cross-cutting change to every workspace's tsconfig and the test runner; a fixture id is also not the benchmark's identity in the way a pinned revision and a complete task list are. Left for its own change.

dist/terminal-bench-adapter.js remains readable and remains in the graded closure. It is genuinely needed to run, and what it reveals — a verifier grades by exit code — is generic harness mechanics rather than any task's answer.

@Astro-Han
Astro-Hanforce-pushed the refactor/benchmark-identity-out-of-dist branch from 1a563c6 to 2a08231CompareAugust 6, 2026 02:12
@Astro-Han
Astro-Han changed the base branch from main to fix/harness-maka-repo-mount-scopeAugust 6, 2026 02:12
@Astro-Han
Astro-Hanforce-pushed the refactor/benchmark-identity-out-of-dist branch 2 times, most recently from e79c5ba to 683e1c9CompareAugust 6, 2026 06:04
@Astro-Han
Astro-Han changed the base branch from fix/harness-maka-repo-mount-scope to mainAugust 6, 2026 06:04
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
Review of the first cut found the comments claiming more than the test held.
The content check read only the top level of dist while the whole directory
is mounted, and described the compiled tests as excluded when they are not.
It now scans dist recursively and states two boundaries, because they are two
different things: the retrieval keys — revision, upstream URL, tree
fingerprint — must be absent from every mounted file, while a single task id
is not a retrieval key and is only held out of the shipped modules, where a
complete list could accumulate.
No mount list may declare the identity file, the Maka one included, and the
DeepSWE cache path derives its revision instead of spelling a second copy.
…t forgot
The scan filtered to `.js`, so a source map — whose `sourcesContent` carries
the whole of `src` — would have walked straight past it. And the retrieval
keys named every task-tree fingerprint but one.
@Astro-Han
Astro-Hanforce-pushed the refactor/benchmark-identity-out-of-dist branch from 683e1c9 to d0e7cf3CompareAugust 6, 2026 06:11
@Astro-Han
Astro-Han marked this pull request as ready for review August 6, 2026 06:25
@Astro-Han
Astro-Han merged commit 43b6694 into mainAug 6, 2026
11 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

refactor(headless): move the benchmark identity out of the build a graded arm mounts - #2303

Merged
Astro-Han merged 4 commits into
mainfrom
refactor/benchmark-identity-out-of-dist
Aug 6, 2026
Merged

refactor(headless): move the benchmark identity out of the build a graded arm mounts#2303
Astro-Han merged 4 commits into
mainfrom
refactor/benchmark-identity-out-of-dist

Conversation

@Astro-Han

Copy link
Copy Markdown
Contributor

What

A graded arm executes packages/headless/dist, so anything compiled into a shipped module is readable from inside the container being scored. harness-ab-manifest.ts carried each benchmark's pinned revision, task-tree fingerprint and full task list — 89 Terminal-Bench ids and 113 DeepSWE ids — and every one of them shipped.

That is the exact material of the #1970 incident: an arm read the pinned revision out of the mount, fetched the task's reference solution from the public repository at that revision, and recorded a pass that measured retrieval.

#2298 narrowed the mount so docs/eval and src no longer reach the container. It could not close this one, because the mount is a list of paths and the identity was inside a path the container legitimately needs.

How

  • The data moves to packages/headless/harbor/benchmark-identity.json, loaded by harbor/benchmark-identity.mjs. Neither is mounted by any arm — agent-repo-mount.ts names exactly one file out of harbor/, and it is run-cell.mjs.
  • harness-ab-manifest.ts keeps the assertions and loses the answers: assertTerminalBench21TaskSet and its three siblings become assertFrozenTaskSet(label, expected, actual) and assertFrozenTaskTreeFingerprint(label, expected, actual). Error messages are unchanged.
  • Each entry in HARNESS_BENCHMARK_PROFILES now carries its own taskTreeFingerprint. With that, the DeepSWE-full and Terminal-Bench task-source branches were the same code under different names and collapse into one; only the subset-30 profile still differs, because its tree carries more tasks than the frozen set.
  • No product module was affected — the identity was only ever read by the two host-side harbor/*.mjs entrypoints and by tests.

Verification

A new test in harness-ab-manifest.test.ts reads every shipped dist/*.js and asserts it contains no revision, no upstream URL and none of the 202 task ids. This is the first check on what the mounted files contain rather than on which paths are declared, and it immediately found a real leak: harbor-smoke-config.ts hardcoded '*sqlite-with-gcov' — a Terminal-Bench 2.1 task id — as a fallback the smoke manifest was already supplying. Removed, and it now fails loudly instead.

Red/green confirmed: restoring that literal fails the new test.

Loading the real entrypoint reproduces all three profiles byte-for-byte:

terminal-bench-2.1 d49e28f1 89 sha256:456826a
deep-swe-1.1 6db64a40 30 sha256:508aedc
deep-swe-1.1-full 6db64a40 113 sha256:973091a

npm run -w @maka/headless test — 1422 pass, 0 fail. Lint and format clean.

Known residue

Compiled tests under dist/__tests__ are still mounted and still mention individual task ids in fixtures. Closing that means splitting the test build out of dist, which is a cross-cutting change to every workspace's tsconfig and the test runner; a fixture id is also not the benchmark's identity in the way a pinned revision and a complete task list are. Left for its own change.

dist/terminal-bench-adapter.js remains readable and remains in the graded closure. It is genuinely needed to run, and what it reveals — a verifier grades by exit code — is generic harness mechanics rather than any task's answer.

@Astro-Han
Astro-Hanforce-pushed the refactor/benchmark-identity-out-of-dist branch from 1a563c6 to 2a08231CompareAugust 6, 2026 02:12
@Astro-Han
Astro-Han changed the base branch from main to fix/harness-maka-repo-mount-scopeAugust 6, 2026 02:12
@Astro-Han
Astro-Hanforce-pushed the refactor/benchmark-identity-out-of-dist branch 2 times, most recently from e79c5ba to 683e1c9CompareAugust 6, 2026 06:04
@Astro-Han
Astro-Han changed the base branch from fix/harness-maka-repo-mount-scope to mainAugust 6, 2026 06:04
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
Review of the first cut found the comments claiming more than the test held.
The content check read only the top level of dist while the whole directory
is mounted, and described the compiled tests as excluded when they are not.
It now scans dist recursively and states two boundaries, because they are two
different things: the retrieval keys — revision, upstream URL, tree
fingerprint — must be absent from every mounted file, while a single task id
is not a retrieval key and is only held out of the shipped modules, where a
complete list could accumulate.
No mount list may declare the identity file, the Maka one included, and the
DeepSWE cache path derives its revision instead of spelling a second copy.
…t forgot
The scan filtered to `.js`, so a source map — whose `sourcesContent` carries
the whole of `src` — would have walked straight past it. And the retrieval
keys named every task-tree fingerprint but one.
@Astro-Han
Astro-Hanforce-pushed the refactor/benchmark-identity-out-of-dist branch from 683e1c9 to d0e7cf3CompareAugust 6, 2026 06:11
@Astro-Han
Astro-Han marked this pull request as ready for review August 6, 2026 06:25
@Astro-Han
Astro-Han merged commit 43b6694 into mainAug 6, 2026
11 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han