fix(headless): mount Maka's build outputs into the task container, not the repo root - #2298

Merged
Astro-Han merged 4 commits into
mainfrom
fix/harness-maka-repo-mount-scope
Aug 6, 2026
Merged

fix(headless): mount Maka's build outputs into the task container, not the repo root#2298
Astro-Han merged 4 commits into
mainfrom
fix/harness-maka-repo-mount-scope

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Aug 6, 2026

Copy link
Copy Markdown
Contributor

Summary

The Maka arm received the repo root at /opt/maka-agent, exempt from the declared-mount rule every competitor arm follows. The exemption's premise was that Maka runs our own runtime — but the model does not know it is ours, and it reads what is reachable.

In the #2245 two-arm run it did exactly that. On extract-elf it ran grep -rl "extract-elf|extract memory values" /opt/maka-agent and read docs/eval/terminal-bench-2.1-*.csv — this benchmark's own per-task pass/fail table from earlier runs. On db-wal-recovery it read packages/headless/src/terminal-bench-adapter.ts, the verifier source. Both cells passed. Neither file is one the runtime executes; both were reachable only because the mount was the root.

The container runs dist, so the arm now gets dependencies, workspace manifests, each workspace's build output, and the single harbor entrypoint the adapter invokes. Sources, docs/, repo history and the harness manifest source are gone from inside it.

Refs #2245

Verification

  • npm run -w @maka/headless test — 1425 pass, 0 fail. Three existing tests asserted the root mount and now assert the narrowed shape.
  • npm run lint, npm run format — clean
  • Ran the container for real, since the contract tests only check that the declared list is self-consistent — they cannot tell whether Maka still starts:
docker run --rm $(mounts from makaRepoPaths()) node:24-slim \
node /opt/maka-agent/packages/headless/dist/cli.js --help

The first attempt failed with Cannot find package '@slack/socket-mode': npm keeps a private node_modules in any workspace whose version pin it could not hoist, and the list only had the root one. That is exactly the class of failure the contract tests cannot see. With those added, the CLI starts and harbor --help resolves.

Reachability from inside the same container, after the change:

pathbeforeafter
docs/ (per-task results)readableblocked
packages/headless/src/ (verifier source)readableblocked
.git/readableblocked
maka-eval/readableblocked
packages/headless/harbor/run-harness-ab.mjsreadableblocked
packages/headless/dist/harness-ab-manifest.jsreadablereadable

Known gap — now closed by #2303

The last row is not fixed here. harness-ab-manifest.ts compiles TERMINAL_BENCH_2_1_REVISION, the upstream repository URL, the task-tree fingerprint and all 89 task ids into packages/headless/dist, which the container needs as a directory. Fixing it means moving benchmark identity out of a compiled runtime package, which is a data-layout change rather than a mount change; #2303 does that, and adds the content-level check this PR's path-level assertions structurally cannot make.

Review round 2

The suite checked that an absent workspace dependency tree is dropped. The incident was the opposite direction: packages/runtime/node_modules exists — npm could not hoist @slack/socket-mode past a version conflict — and omitting it resolved on the host, then failed inside the container on the first import that needed it. Only a live container run caught that.

declares every workspace dependency tree npm could not hoist away now asks the repo directly, and fails on the pre-fix list. The renderer-only exclusion and the workspace list are also derived from the root manifest in one place instead of two.

…t the repo root
The Maka arm was exempt from the declared-mount rule on the theory that
it runs our own runtime. The model does not know it is ours. In the
#2245 two-arm run it grepped the mount and read
`docs/eval/terminal-bench-2.1-*.csv` — this benchmark's own per-task
pass/fail table from earlier runs — and, on another task, the verifier
source at `packages/headless/src/terminal-bench-adapter.ts`.
The container executes `dist`, so it now receives dependencies,
workspace manifests, build outputs and the one harbor entrypoint the
adapter runs. Sources, docs, repo history and the harness manifest are
no longer reachable from inside it.
Workspace-private `node_modules` are declared but dropped when absent:
which ones exist depends on what npm hoisted, and Docker materialises a
missing bind source as a new empty directory inside the repo.
The suite checked that an absent workspace dependency tree is dropped. The
incident was the opposite: `packages/runtime/node_modules` exists — npm could
not hoist `@slack/socket-mode` past a version conflict — and omitting it
resolved on the host, then failed in the container on the first import that
needed it. Only a live container run caught that. Ask the repo instead.
…obes
`install()` tests for both `run-cell.mjs` and `run-host-cell.mjs` before
either can run, in every cell-mode branch. Only the first was declared, so
the narrowed mount aborted the trial at install rather than degrading it —
a regression against the repo-root mount this PR replaces, and one the
container smoke check missed because it exercised loading, not installing.
The competitor arms already have the check that would have caught it: read
the adapter for what it names at a container path and assert the mount
declares it. Maka had no equivalent, which is why the omission shipped. It
does now — `maka_agent.py` joins path segments onto the mount root instead
of writing literals, so the test matches the join.
Also asserts the sources-and-records invariant against the mounts actually
produced rather than the declared list: a declaration naming no forbidden
path still leaks every one of them if the builder mounts the root anyway.
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
The list lived twice — once implied by the mount declaration, once spelled in
the test that cross-checks it against the root manifest. Two copies of an
exclusion drift into a workspace nobody mounts.
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
@Astro-Han
Astro-Han marked this pull request as ready for review August 6, 2026 06:02
@Astro-Han
Astro-Han merged commit 474afee into mainAug 6, 2026
11 checks passed
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts (#2303)
* refactor(headless): move the benchmark identity out of the build a graded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
* refactor(headless): hold the identity boundary where it actually is
Review of the first cut found the comments claiming more than the test held.
The content check read only the top level of dist while the whole directory
is mounted, and described the compiled tests as excluded when they are not.
It now scans dist recursively and states two boundaries, because they are two
different things: the retrieval keys — revision, upstream URL, tree
fingerprint — must be absent from every mounted file, while a single task id
is not a retrieval key and is only held out of the shipped modules, where a
complete list could accumulate.
No mount list may declare the identity file, the Maka one included, and the
DeepSWE cache path derives its revision instead of spelling a second copy.
* test(headless): check every mounted file, and the fingerprint the list forgot
The scan filtered to `.js`, so a source map — whose `sourcesContent` carries
the whole of `src` — would have walked straight past it. And the retrieval
keys named every task-tree fingerprint but one.
* refactor(headless): drop a benchmark task id from a retry comment
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

fix(headless): mount Maka's build outputs into the task container, not the repo root - #2298

Merged
Astro-Han merged 4 commits into
mainfrom
fix/harness-maka-repo-mount-scope
Aug 6, 2026
Merged

fix(headless): mount Maka's build outputs into the task container, not the repo root#2298
Astro-Han merged 4 commits into
mainfrom
fix/harness-maka-repo-mount-scope

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Aug 6, 2026

Copy link
Copy Markdown
Contributor

Summary

The Maka arm received the repo root at /opt/maka-agent, exempt from the declared-mount rule every competitor arm follows. The exemption's premise was that Maka runs our own runtime — but the model does not know it is ours, and it reads what is reachable.

In the #2245 two-arm run it did exactly that. On extract-elf it ran grep -rl "extract-elf|extract memory values" /opt/maka-agent and read docs/eval/terminal-bench-2.1-*.csv — this benchmark's own per-task pass/fail table from earlier runs. On db-wal-recovery it read packages/headless/src/terminal-bench-adapter.ts, the verifier source. Both cells passed. Neither file is one the runtime executes; both were reachable only because the mount was the root.

The container runs dist, so the arm now gets dependencies, workspace manifests, each workspace's build output, and the single harbor entrypoint the adapter invokes. Sources, docs/, repo history and the harness manifest source are gone from inside it.

Refs #2245

Verification

  • npm run -w @maka/headless test — 1425 pass, 0 fail. Three existing tests asserted the root mount and now assert the narrowed shape.
  • npm run lint, npm run format — clean
  • Ran the container for real, since the contract tests only check that the declared list is self-consistent — they cannot tell whether Maka still starts:
docker run --rm $(mounts from makaRepoPaths()) node:24-slim \
node /opt/maka-agent/packages/headless/dist/cli.js --help

The first attempt failed with Cannot find package '@slack/socket-mode': npm keeps a private node_modules in any workspace whose version pin it could not hoist, and the list only had the root one. That is exactly the class of failure the contract tests cannot see. With those added, the CLI starts and harbor --help resolves.

Reachability from inside the same container, after the change:

pathbeforeafter
docs/ (per-task results)readableblocked
packages/headless/src/ (verifier source)readableblocked
.git/readableblocked
maka-eval/readableblocked
packages/headless/harbor/run-harness-ab.mjsreadableblocked
packages/headless/dist/harness-ab-manifest.jsreadablereadable

Known gap — now closed by #2303

The last row is not fixed here. harness-ab-manifest.ts compiles TERMINAL_BENCH_2_1_REVISION, the upstream repository URL, the task-tree fingerprint and all 89 task ids into packages/headless/dist, which the container needs as a directory. Fixing it means moving benchmark identity out of a compiled runtime package, which is a data-layout change rather than a mount change; #2303 does that, and adds the content-level check this PR's path-level assertions structurally cannot make.

Review round 2

The suite checked that an absent workspace dependency tree is dropped. The incident was the opposite direction: packages/runtime/node_modules exists — npm could not hoist @slack/socket-mode past a version conflict — and omitting it resolved on the host, then failed inside the container on the first import that needed it. Only a live container run caught that.

declares every workspace dependency tree npm could not hoist away now asks the repo directly, and fails on the pre-fix list. The renderer-only exclusion and the workspace list are also derived from the root manifest in one place instead of two.

…t the repo root
The Maka arm was exempt from the declared-mount rule on the theory that
it runs our own runtime. The model does not know it is ours. In the
#2245 two-arm run it grepped the mount and read
`docs/eval/terminal-bench-2.1-*.csv` — this benchmark's own per-task
pass/fail table from earlier runs — and, on another task, the verifier
source at `packages/headless/src/terminal-bench-adapter.ts`.
The container executes `dist`, so it now receives dependencies,
workspace manifests, build outputs and the one harbor entrypoint the
adapter runs. Sources, docs, repo history and the harness manifest are
no longer reachable from inside it.
Workspace-private `node_modules` are declared but dropped when absent:
which ones exist depends on what npm hoisted, and Docker materialises a
missing bind source as a new empty directory inside the repo.
The suite checked that an absent workspace dependency tree is dropped. The
incident was the opposite: `packages/runtime/node_modules` exists — npm could
not hoist `@slack/socket-mode` past a version conflict — and omitting it
resolved on the host, then failed in the container on the first import that
needed it. Only a live container run caught that. Ask the repo instead.
…obes
`install()` tests for both `run-cell.mjs` and `run-host-cell.mjs` before
either can run, in every cell-mode branch. Only the first was declared, so
the narrowed mount aborted the trial at install rather than degrading it —
a regression against the repo-root mount this PR replaces, and one the
container smoke check missed because it exercised loading, not installing.
The competitor arms already have the check that would have caught it: read
the adapter for what it names at a container path and assert the mount
declares it. Maka had no equivalent, which is why the omission shipped. It
does now — `maka_agent.py` joins path segments onto the mount root instead
of writing literals, so the test matches the join.
Also asserts the sources-and-records invariant against the mounts actually
produced rather than the declared list: a declaration naming no forbidden
path still leaks every one of them if the builder mounts the root anyway.
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
The list lived twice — once implied by the mount declaration, once spelled in
the test that cross-checks it against the root manifest. Two copies of an
exclusion drift into a workspace nobody mounts.
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
@Astro-Han
Astro-Han marked this pull request as ready for review August 6, 2026 06:02
@Astro-Han
Astro-Han merged commit 474afee into mainAug 6, 2026
11 checks passed
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts (#2303)
* refactor(headless): move the benchmark identity out of the build a graded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
* refactor(headless): hold the identity boundary where it actually is
Review of the first cut found the comments claiming more than the test held.
The content check read only the top level of dist while the whole directory
is mounted, and described the compiled tests as excluded when they are not.
It now scans dist recursively and states two boundaries, because they are two
different things: the retrieval keys — revision, upstream URL, tree
fingerprint — must be absent from every mounted file, while a single task id
is not a retrieval key and is only held out of the shipped modules, where a
complete list could accumulate.
No mount list may declare the identity file, the Maka one included, and the
DeepSWE cache path derives its revision instead of spelling a second copy.
* test(headless): check every mounted file, and the fingerprint the list forgot
The scan filtered to `.js`, so a source map — whose `sourcesContent` carries
the whole of `src` — would have walked straight past it. And the retrieval
keys named every task-tree fingerprint but one.
* refactor(headless): drop a benchmark task id from a retry comment
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(headless): mount Maka's build outputs into the task container, not the repo root - #2298

Merged
Astro-Han merged 4 commits into
mainfrom
fix/harness-maka-repo-mount-scope
Aug 6, 2026
Merged

fix(headless): mount Maka's build outputs into the task container, not the repo root#2298
Astro-Han merged 4 commits into
mainfrom
fix/harness-maka-repo-mount-scope

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Aug 6, 2026

Copy link
Copy Markdown
Contributor

Summary

The Maka arm received the repo root at /opt/maka-agent, exempt from the declared-mount rule every competitor arm follows. The exemption's premise was that Maka runs our own runtime — but the model does not know it is ours, and it reads what is reachable.

In the #2245 two-arm run it did exactly that. On extract-elf it ran grep -rl "extract-elf|extract memory values" /opt/maka-agent and read docs/eval/terminal-bench-2.1-*.csv — this benchmark's own per-task pass/fail table from earlier runs. On db-wal-recovery it read packages/headless/src/terminal-bench-adapter.ts, the verifier source. Both cells passed. Neither file is one the runtime executes; both were reachable only because the mount was the root.

The container runs dist, so the arm now gets dependencies, workspace manifests, each workspace's build output, and the single harbor entrypoint the adapter invokes. Sources, docs/, repo history and the harness manifest source are gone from inside it.

Refs #2245

Verification

  • npm run -w @maka/headless test — 1425 pass, 0 fail. Three existing tests asserted the root mount and now assert the narrowed shape.
  • npm run lint, npm run format — clean
  • Ran the container for real, since the contract tests only check that the declared list is self-consistent — they cannot tell whether Maka still starts:
docker run --rm $(mounts from makaRepoPaths()) node:24-slim \
node /opt/maka-agent/packages/headless/dist/cli.js --help

The first attempt failed with Cannot find package '@slack/socket-mode': npm keeps a private node_modules in any workspace whose version pin it could not hoist, and the list only had the root one. That is exactly the class of failure the contract tests cannot see. With those added, the CLI starts and harbor --help resolves.

Reachability from inside the same container, after the change:

pathbeforeafter
docs/ (per-task results)readableblocked
packages/headless/src/ (verifier source)readableblocked
.git/readableblocked
maka-eval/readableblocked
packages/headless/harbor/run-harness-ab.mjsreadableblocked
packages/headless/dist/harness-ab-manifest.jsreadablereadable

Known gap — now closed by #2303

The last row is not fixed here. harness-ab-manifest.ts compiles TERMINAL_BENCH_2_1_REVISION, the upstream repository URL, the task-tree fingerprint and all 89 task ids into packages/headless/dist, which the container needs as a directory. Fixing it means moving benchmark identity out of a compiled runtime package, which is a data-layout change rather than a mount change; #2303 does that, and adds the content-level check this PR's path-level assertions structurally cannot make.

Review round 2

The suite checked that an absent workspace dependency tree is dropped. The incident was the opposite direction: packages/runtime/node_modules exists — npm could not hoist @slack/socket-mode past a version conflict — and omitting it resolved on the host, then failed inside the container on the first import that needed it. Only a live container run caught that.

declares every workspace dependency tree npm could not hoist away now asks the repo directly, and fails on the pre-fix list. The renderer-only exclusion and the workspace list are also derived from the root manifest in one place instead of two.

…t the repo root
The Maka arm was exempt from the declared-mount rule on the theory that
it runs our own runtime. The model does not know it is ours. In the
#2245 two-arm run it grepped the mount and read
`docs/eval/terminal-bench-2.1-*.csv` — this benchmark's own per-task
pass/fail table from earlier runs — and, on another task, the verifier
source at `packages/headless/src/terminal-bench-adapter.ts`.
The container executes `dist`, so it now receives dependencies,
workspace manifests, build outputs and the one harbor entrypoint the
adapter runs. Sources, docs, repo history and the harness manifest are
no longer reachable from inside it.
Workspace-private `node_modules` are declared but dropped when absent:
which ones exist depends on what npm hoisted, and Docker materialises a
missing bind source as a new empty directory inside the repo.
The suite checked that an absent workspace dependency tree is dropped. The
incident was the opposite: `packages/runtime/node_modules` exists — npm could
not hoist `@slack/socket-mode` past a version conflict — and omitting it
resolved on the host, then failed in the container on the first import that
needed it. Only a live container run caught that. Ask the repo instead.
…obes
`install()` tests for both `run-cell.mjs` and `run-host-cell.mjs` before
either can run, in every cell-mode branch. Only the first was declared, so
the narrowed mount aborted the trial at install rather than degrading it —
a regression against the repo-root mount this PR replaces, and one the
container smoke check missed because it exercised loading, not installing.
The competitor arms already have the check that would have caught it: read
the adapter for what it names at a container path and assert the mount
declares it. Maka had no equivalent, which is why the omission shipped. It
does now — `maka_agent.py` joins path segments onto the mount root instead
of writing literals, so the test matches the join.
Also asserts the sources-and-records invariant against the mounts actually
produced rather than the declared list: a declaration naming no forbidden
path still leaks every one of them if the builder mounts the root anyway.
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
The list lived twice — once implied by the mount declaration, once spelled in
the test that cross-checks it against the root manifest. Two copies of an
exclusion drift into a workspace nobody mounts.
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
@Astro-Han
Astro-Han marked this pull request as ready for review August 6, 2026 06:02
@Astro-Han
Astro-Han merged commit 474afee into mainAug 6, 2026
11 checks passed
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts (#2303)
* refactor(headless): move the benchmark identity out of the build a graded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
* refactor(headless): hold the identity boundary where it actually is
Review of the first cut found the comments claiming more than the test held.
The content check read only the top level of dist while the whole directory
is mounted, and described the compiled tests as excluded when they are not.
It now scans dist recursively and states two boundaries, because they are two
different things: the retrieval keys — revision, upstream URL, tree
fingerprint — must be absent from every mounted file, while a single task id
is not a retrieval key and is only held out of the shipped modules, where a
complete list could accumulate.
No mount list may declare the identity file, the Maka one included, and the
DeepSWE cache path derives its revision instead of spelling a second copy.
* test(headless): check every mounted file, and the fingerprint the list forgot
The scan filtered to `.js`, so a source map — whose `sourcesContent` carries
the whole of `src` — would have walked straight past it. And the retrieval
keys named every task-tree fingerprint but one.
* refactor(headless): drop a benchmark task id from a retry comment
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(headless): mount Maka's build outputs into the task container, not the repo root - #2298

Merged
Astro-Han merged 4 commits into
mainfrom
fix/harness-maka-repo-mount-scope
Aug 6, 2026
Merged

fix(headless): mount Maka's build outputs into the task container, not the repo root#2298
Astro-Han merged 4 commits into
mainfrom
fix/harness-maka-repo-mount-scope

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Aug 6, 2026

Copy link
Copy Markdown
Contributor

Summary

The Maka arm received the repo root at /opt/maka-agent, exempt from the declared-mount rule every competitor arm follows. The exemption's premise was that Maka runs our own runtime — but the model does not know it is ours, and it reads what is reachable.

In the #2245 two-arm run it did exactly that. On extract-elf it ran grep -rl "extract-elf|extract memory values" /opt/maka-agent and read docs/eval/terminal-bench-2.1-*.csv — this benchmark's own per-task pass/fail table from earlier runs. On db-wal-recovery it read packages/headless/src/terminal-bench-adapter.ts, the verifier source. Both cells passed. Neither file is one the runtime executes; both were reachable only because the mount was the root.

The container runs dist, so the arm now gets dependencies, workspace manifests, each workspace's build output, and the single harbor entrypoint the adapter invokes. Sources, docs/, repo history and the harness manifest source are gone from inside it.

Refs #2245

Verification

  • npm run -w @maka/headless test — 1425 pass, 0 fail. Three existing tests asserted the root mount and now assert the narrowed shape.
  • npm run lint, npm run format — clean
  • Ran the container for real, since the contract tests only check that the declared list is self-consistent — they cannot tell whether Maka still starts:
docker run --rm $(mounts from makaRepoPaths()) node:24-slim \
node /opt/maka-agent/packages/headless/dist/cli.js --help

The first attempt failed with Cannot find package '@slack/socket-mode': npm keeps a private node_modules in any workspace whose version pin it could not hoist, and the list only had the root one. That is exactly the class of failure the contract tests cannot see. With those added, the CLI starts and harbor --help resolves.

Reachability from inside the same container, after the change:

pathbeforeafter
docs/ (per-task results)readableblocked
packages/headless/src/ (verifier source)readableblocked
.git/readableblocked
maka-eval/readableblocked
packages/headless/harbor/run-harness-ab.mjsreadableblocked
packages/headless/dist/harness-ab-manifest.jsreadablereadable

Known gap — now closed by #2303

The last row is not fixed here. harness-ab-manifest.ts compiles TERMINAL_BENCH_2_1_REVISION, the upstream repository URL, the task-tree fingerprint and all 89 task ids into packages/headless/dist, which the container needs as a directory. Fixing it means moving benchmark identity out of a compiled runtime package, which is a data-layout change rather than a mount change; #2303 does that, and adds the content-level check this PR's path-level assertions structurally cannot make.

Review round 2

The suite checked that an absent workspace dependency tree is dropped. The incident was the opposite direction: packages/runtime/node_modules exists — npm could not hoist @slack/socket-mode past a version conflict — and omitting it resolved on the host, then failed inside the container on the first import that needed it. Only a live container run caught that.

declares every workspace dependency tree npm could not hoist away now asks the repo directly, and fails on the pre-fix list. The renderer-only exclusion and the workspace list are also derived from the root manifest in one place instead of two.

…t the repo root
The Maka arm was exempt from the declared-mount rule on the theory that
it runs our own runtime. The model does not know it is ours. In the
#2245 two-arm run it grepped the mount and read
`docs/eval/terminal-bench-2.1-*.csv` — this benchmark's own per-task
pass/fail table from earlier runs — and, on another task, the verifier
source at `packages/headless/src/terminal-bench-adapter.ts`.
The container executes `dist`, so it now receives dependencies,
workspace manifests, build outputs and the one harbor entrypoint the
adapter runs. Sources, docs, repo history and the harness manifest are
no longer reachable from inside it.
Workspace-private `node_modules` are declared but dropped when absent:
which ones exist depends on what npm hoisted, and Docker materialises a
missing bind source as a new empty directory inside the repo.
The suite checked that an absent workspace dependency tree is dropped. The
incident was the opposite: `packages/runtime/node_modules` exists — npm could
not hoist `@slack/socket-mode` past a version conflict — and omitting it
resolved on the host, then failed in the container on the first import that
needed it. Only a live container run caught that. Ask the repo instead.
…obes
`install()` tests for both `run-cell.mjs` and `run-host-cell.mjs` before
either can run, in every cell-mode branch. Only the first was declared, so
the narrowed mount aborted the trial at install rather than degrading it —
a regression against the repo-root mount this PR replaces, and one the
container smoke check missed because it exercised loading, not installing.
The competitor arms already have the check that would have caught it: read
the adapter for what it names at a container path and assert the mount
declares it. Maka had no equivalent, which is why the omission shipped. It
does now — `maka_agent.py` joins path segments onto the mount root instead
of writing literals, so the test matches the join.
Also asserts the sources-and-records invariant against the mounts actually
produced rather than the declared list: a declaration naming no forbidden
path still leaks every one of them if the builder mounts the root anyway.
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
The list lived twice — once implied by the mount declaration, once spelled in
the test that cross-checks it against the root manifest. Two copies of an
exclusion drift into a workspace nobody mounts.
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
@Astro-Han
Astro-Han marked this pull request as ready for review August 6, 2026 06:02
@Astro-Han
Astro-Han merged commit 474afee into mainAug 6, 2026
11 checks passed
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts (#2303)
* refactor(headless): move the benchmark identity out of the build a graded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
* refactor(headless): hold the identity boundary where it actually is
Review of the first cut found the comments claiming more than the test held.
The content check read only the top level of dist while the whole directory
is mounted, and described the compiled tests as excluded when they are not.
It now scans dist recursively and states two boundaries, because they are two
different things: the retrieval keys — revision, upstream URL, tree
fingerprint — must be absent from every mounted file, while a single task id
is not a retrieval key and is only held out of the shipped modules, where a
complete list could accumulate.
No mount list may declare the identity file, the Maka one included, and the
DeepSWE cache path derives its revision instead of spelling a second copy.
* test(headless): check every mounted file, and the fingerprint the list forgot
The scan filtered to `.js`, so a source map — whose `sourcesContent` carries
the whole of `src` — would have walked straight past it. And the retrieval
keys named every task-tree fingerprint but one.
* refactor(headless): drop a benchmark task id from a retry comment
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

fix(headless): mount Maka's build outputs into the task container, not the repo root - #2298

Merged
Astro-Han merged 4 commits into
mainfrom
fix/harness-maka-repo-mount-scope
Aug 6, 2026
Merged

fix(headless): mount Maka's build outputs into the task container, not the repo root#2298
Astro-Han merged 4 commits into
mainfrom
fix/harness-maka-repo-mount-scope

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Aug 6, 2026

Copy link
Copy Markdown
Contributor

Summary

The Maka arm received the repo root at /opt/maka-agent, exempt from the declared-mount rule every competitor arm follows. The exemption's premise was that Maka runs our own runtime — but the model does not know it is ours, and it reads what is reachable.

In the #2245 two-arm run it did exactly that. On extract-elf it ran grep -rl "extract-elf|extract memory values" /opt/maka-agent and read docs/eval/terminal-bench-2.1-*.csv — this benchmark's own per-task pass/fail table from earlier runs. On db-wal-recovery it read packages/headless/src/terminal-bench-adapter.ts, the verifier source. Both cells passed. Neither file is one the runtime executes; both were reachable only because the mount was the root.

The container runs dist, so the arm now gets dependencies, workspace manifests, each workspace's build output, and the single harbor entrypoint the adapter invokes. Sources, docs/, repo history and the harness manifest source are gone from inside it.

Refs #2245

Verification

  • npm run -w @maka/headless test — 1425 pass, 0 fail. Three existing tests asserted the root mount and now assert the narrowed shape.
  • npm run lint, npm run format — clean
  • Ran the container for real, since the contract tests only check that the declared list is self-consistent — they cannot tell whether Maka still starts:
docker run --rm $(mounts from makaRepoPaths()) node:24-slim \
node /opt/maka-agent/packages/headless/dist/cli.js --help

The first attempt failed with Cannot find package '@slack/socket-mode': npm keeps a private node_modules in any workspace whose version pin it could not hoist, and the list only had the root one. That is exactly the class of failure the contract tests cannot see. With those added, the CLI starts and harbor --help resolves.

Reachability from inside the same container, after the change:

pathbeforeafter
docs/ (per-task results)readableblocked
packages/headless/src/ (verifier source)readableblocked
.git/readableblocked
maka-eval/readableblocked
packages/headless/harbor/run-harness-ab.mjsreadableblocked
packages/headless/dist/harness-ab-manifest.jsreadablereadable

Known gap — now closed by #2303

The last row is not fixed here. harness-ab-manifest.ts compiles TERMINAL_BENCH_2_1_REVISION, the upstream repository URL, the task-tree fingerprint and all 89 task ids into packages/headless/dist, which the container needs as a directory. Fixing it means moving benchmark identity out of a compiled runtime package, which is a data-layout change rather than a mount change; #2303 does that, and adds the content-level check this PR's path-level assertions structurally cannot make.

Review round 2

The suite checked that an absent workspace dependency tree is dropped. The incident was the opposite direction: packages/runtime/node_modules exists — npm could not hoist @slack/socket-mode past a version conflict — and omitting it resolved on the host, then failed inside the container on the first import that needed it. Only a live container run caught that.

declares every workspace dependency tree npm could not hoist away now asks the repo directly, and fails on the pre-fix list. The renderer-only exclusion and the workspace list are also derived from the root manifest in one place instead of two.

…t the repo root
The Maka arm was exempt from the declared-mount rule on the theory that
it runs our own runtime. The model does not know it is ours. In the
#2245 two-arm run it grepped the mount and read
`docs/eval/terminal-bench-2.1-*.csv` — this benchmark's own per-task
pass/fail table from earlier runs — and, on another task, the verifier
source at `packages/headless/src/terminal-bench-adapter.ts`.
The container executes `dist`, so it now receives dependencies,
workspace manifests, build outputs and the one harbor entrypoint the
adapter runs. Sources, docs, repo history and the harness manifest are
no longer reachable from inside it.
Workspace-private `node_modules` are declared but dropped when absent:
which ones exist depends on what npm hoisted, and Docker materialises a
missing bind source as a new empty directory inside the repo.
The suite checked that an absent workspace dependency tree is dropped. The
incident was the opposite: `packages/runtime/node_modules` exists — npm could
not hoist `@slack/socket-mode` past a version conflict — and omitting it
resolved on the host, then failed in the container on the first import that
needed it. Only a live container run caught that. Ask the repo instead.
…obes
`install()` tests for both `run-cell.mjs` and `run-host-cell.mjs` before
either can run, in every cell-mode branch. Only the first was declared, so
the narrowed mount aborted the trial at install rather than degrading it —
a regression against the repo-root mount this PR replaces, and one the
container smoke check missed because it exercised loading, not installing.
The competitor arms already have the check that would have caught it: read
the adapter for what it names at a container path and assert the mount
declares it. Maka had no equivalent, which is why the omission shipped. It
does now — `maka_agent.py` joins path segments onto the mount root instead
of writing literals, so the test matches the join.
Also asserts the sources-and-records invariant against the mounts actually
produced rather than the declared list: a declaration naming no forbidden
path still leaks every one of them if the builder mounts the root anyway.
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
The list lived twice — once implied by the mount declaration, once spelled in
the test that cross-checks it against the root manifest. Two copies of an
exclusion drift into a workspace nobody mounts.
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
@Astro-Han
Astro-Han marked this pull request as ready for review August 6, 2026 06:02
@Astro-Han
Astro-Han merged commit 474afee into mainAug 6, 2026
11 checks passed
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts (#2303)
* refactor(headless): move the benchmark identity out of the build a graded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
* refactor(headless): hold the identity boundary where it actually is
Review of the first cut found the comments claiming more than the test held.
The content check read only the top level of dist while the whole directory
is mounted, and described the compiled tests as excluded when they are not.
It now scans dist recursively and states two boundaries, because they are two
different things: the retrieval keys — revision, upstream URL, tree
fingerprint — must be absent from every mounted file, while a single task id
is not a retrieval key and is only held out of the shipped modules, where a
complete list could accumulate.
No mount list may declare the identity file, the Maka one included, and the
DeepSWE cache path derives its revision instead of spelling a second copy.
* test(headless): check every mounted file, and the fingerprint the list forgot
The scan filtered to `.js`, so a source map — whose `sourcesContent` carries
the whole of `src` — would have walked straight past it. And the retrieval
keys named every task-tree fingerprint but one.
* refactor(headless): drop a benchmark task id from a retry comment
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(headless): mount Maka's build outputs into the task container, not the repo root - #2298

Merged
Astro-Han merged 4 commits into
mainfrom
fix/harness-maka-repo-mount-scope
Aug 6, 2026
Merged

fix(headless): mount Maka's build outputs into the task container, not the repo root#2298
Astro-Han merged 4 commits into
mainfrom
fix/harness-maka-repo-mount-scope

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Aug 6, 2026

Copy link
Copy Markdown
Contributor

Summary

The Maka arm received the repo root at /opt/maka-agent, exempt from the declared-mount rule every competitor arm follows. The exemption's premise was that Maka runs our own runtime — but the model does not know it is ours, and it reads what is reachable.

In the #2245 two-arm run it did exactly that. On extract-elf it ran grep -rl "extract-elf|extract memory values" /opt/maka-agent and read docs/eval/terminal-bench-2.1-*.csv — this benchmark's own per-task pass/fail table from earlier runs. On db-wal-recovery it read packages/headless/src/terminal-bench-adapter.ts, the verifier source. Both cells passed. Neither file is one the runtime executes; both were reachable only because the mount was the root.

The container runs dist, so the arm now gets dependencies, workspace manifests, each workspace's build output, and the single harbor entrypoint the adapter invokes. Sources, docs/, repo history and the harness manifest source are gone from inside it.

Refs #2245

Verification

  • npm run -w @maka/headless test — 1425 pass, 0 fail. Three existing tests asserted the root mount and now assert the narrowed shape.
  • npm run lint, npm run format — clean
  • Ran the container for real, since the contract tests only check that the declared list is self-consistent — they cannot tell whether Maka still starts:
docker run --rm $(mounts from makaRepoPaths()) node:24-slim \
node /opt/maka-agent/packages/headless/dist/cli.js --help

The first attempt failed with Cannot find package '@slack/socket-mode': npm keeps a private node_modules in any workspace whose version pin it could not hoist, and the list only had the root one. That is exactly the class of failure the contract tests cannot see. With those added, the CLI starts and harbor --help resolves.

Reachability from inside the same container, after the change:

pathbeforeafter
docs/ (per-task results)readableblocked
packages/headless/src/ (verifier source)readableblocked
.git/readableblocked
maka-eval/readableblocked
packages/headless/harbor/run-harness-ab.mjsreadableblocked
packages/headless/dist/harness-ab-manifest.jsreadablereadable

Known gap — now closed by #2303

The last row is not fixed here. harness-ab-manifest.ts compiles TERMINAL_BENCH_2_1_REVISION, the upstream repository URL, the task-tree fingerprint and all 89 task ids into packages/headless/dist, which the container needs as a directory. Fixing it means moving benchmark identity out of a compiled runtime package, which is a data-layout change rather than a mount change; #2303 does that, and adds the content-level check this PR's path-level assertions structurally cannot make.

Review round 2

The suite checked that an absent workspace dependency tree is dropped. The incident was the opposite direction: packages/runtime/node_modules exists — npm could not hoist @slack/socket-mode past a version conflict — and omitting it resolved on the host, then failed inside the container on the first import that needed it. Only a live container run caught that.

declares every workspace dependency tree npm could not hoist away now asks the repo directly, and fails on the pre-fix list. The renderer-only exclusion and the workspace list are also derived from the root manifest in one place instead of two.

…t the repo root
The Maka arm was exempt from the declared-mount rule on the theory that
it runs our own runtime. The model does not know it is ours. In the
#2245 two-arm run it grepped the mount and read
`docs/eval/terminal-bench-2.1-*.csv` — this benchmark's own per-task
pass/fail table from earlier runs — and, on another task, the verifier
source at `packages/headless/src/terminal-bench-adapter.ts`.
The container executes `dist`, so it now receives dependencies,
workspace manifests, build outputs and the one harbor entrypoint the
adapter runs. Sources, docs, repo history and the harness manifest are
no longer reachable from inside it.
Workspace-private `node_modules` are declared but dropped when absent:
which ones exist depends on what npm hoisted, and Docker materialises a
missing bind source as a new empty directory inside the repo.
The suite checked that an absent workspace dependency tree is dropped. The
incident was the opposite: `packages/runtime/node_modules` exists — npm could
not hoist `@slack/socket-mode` past a version conflict — and omitting it
resolved on the host, then failed in the container on the first import that
needed it. Only a live container run caught that. Ask the repo instead.
…obes
`install()` tests for both `run-cell.mjs` and `run-host-cell.mjs` before
either can run, in every cell-mode branch. Only the first was declared, so
the narrowed mount aborted the trial at install rather than degrading it —
a regression against the repo-root mount this PR replaces, and one the
container smoke check missed because it exercised loading, not installing.
The competitor arms already have the check that would have caught it: read
the adapter for what it names at a container path and assert the mount
declares it. Maka had no equivalent, which is why the omission shipped. It
does now — `maka_agent.py` joins path segments onto the mount root instead
of writing literals, so the test matches the join.
Also asserts the sources-and-records invariant against the mounts actually
produced rather than the declared list: a declaration naming no forbidden
path still leaks every one of them if the builder mounts the root anyway.
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
The list lived twice — once implied by the mount declaration, once spelled in
the test that cross-checks it against the root manifest. Two copies of an
exclusion drift into a workspace nobody mounts.
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
@Astro-Han
Astro-Han marked this pull request as ready for review August 6, 2026 06:02
@Astro-Han
Astro-Han merged commit 474afee into mainAug 6, 2026
11 checks passed
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts (#2303)
* refactor(headless): move the benchmark identity out of the build a graded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
* refactor(headless): hold the identity boundary where it actually is
Review of the first cut found the comments claiming more than the test held.
The content check read only the top level of dist while the whole directory
is mounted, and described the compiled tests as excluded when they are not.
It now scans dist recursively and states two boundaries, because they are two
different things: the retrieval keys — revision, upstream URL, tree
fingerprint — must be absent from every mounted file, while a single task id
is not a retrieval key and is only held out of the shipped modules, where a
complete list could accumulate.
No mount list may declare the identity file, the Maka one included, and the
DeepSWE cache path derives its revision instead of spelling a second copy.
* test(headless): check every mounted file, and the fingerprint the list forgot
The scan filtered to `.js`, so a source map — whose `sourcesContent` carries
the whole of `src` — would have walked straight past it. And the retrieval
keys named every task-tree fingerprint but one.
* refactor(headless): drop a benchmark task id from a retry comment
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(headless): mount Maka's build outputs into the task container, not the repo root - #2298

Merged
Astro-Han merged 4 commits into
mainfrom
fix/harness-maka-repo-mount-scope
Aug 6, 2026
Merged

fix(headless): mount Maka's build outputs into the task container, not the repo root#2298
Astro-Han merged 4 commits into
mainfrom
fix/harness-maka-repo-mount-scope

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Aug 6, 2026

Copy link
Copy Markdown
Contributor

Summary

The Maka arm received the repo root at /opt/maka-agent, exempt from the declared-mount rule every competitor arm follows. The exemption's premise was that Maka runs our own runtime — but the model does not know it is ours, and it reads what is reachable.

In the #2245 two-arm run it did exactly that. On extract-elf it ran grep -rl "extract-elf|extract memory values" /opt/maka-agent and read docs/eval/terminal-bench-2.1-*.csv — this benchmark's own per-task pass/fail table from earlier runs. On db-wal-recovery it read packages/headless/src/terminal-bench-adapter.ts, the verifier source. Both cells passed. Neither file is one the runtime executes; both were reachable only because the mount was the root.

The container runs dist, so the arm now gets dependencies, workspace manifests, each workspace's build output, and the single harbor entrypoint the adapter invokes. Sources, docs/, repo history and the harness manifest source are gone from inside it.

Refs #2245

Verification

  • npm run -w @maka/headless test — 1425 pass, 0 fail. Three existing tests asserted the root mount and now assert the narrowed shape.
  • npm run lint, npm run format — clean
  • Ran the container for real, since the contract tests only check that the declared list is self-consistent — they cannot tell whether Maka still starts:
docker run --rm $(mounts from makaRepoPaths()) node:24-slim \
node /opt/maka-agent/packages/headless/dist/cli.js --help

The first attempt failed with Cannot find package '@slack/socket-mode': npm keeps a private node_modules in any workspace whose version pin it could not hoist, and the list only had the root one. That is exactly the class of failure the contract tests cannot see. With those added, the CLI starts and harbor --help resolves.

Reachability from inside the same container, after the change:

pathbeforeafter
docs/ (per-task results)readableblocked
packages/headless/src/ (verifier source)readableblocked
.git/readableblocked
maka-eval/readableblocked
packages/headless/harbor/run-harness-ab.mjsreadableblocked
packages/headless/dist/harness-ab-manifest.jsreadablereadable

Known gap — now closed by #2303

The last row is not fixed here. harness-ab-manifest.ts compiles TERMINAL_BENCH_2_1_REVISION, the upstream repository URL, the task-tree fingerprint and all 89 task ids into packages/headless/dist, which the container needs as a directory. Fixing it means moving benchmark identity out of a compiled runtime package, which is a data-layout change rather than a mount change; #2303 does that, and adds the content-level check this PR's path-level assertions structurally cannot make.

Review round 2

The suite checked that an absent workspace dependency tree is dropped. The incident was the opposite direction: packages/runtime/node_modules exists — npm could not hoist @slack/socket-mode past a version conflict — and omitting it resolved on the host, then failed inside the container on the first import that needed it. Only a live container run caught that.

declares every workspace dependency tree npm could not hoist away now asks the repo directly, and fails on the pre-fix list. The renderer-only exclusion and the workspace list are also derived from the root manifest in one place instead of two.

…t the repo root
The Maka arm was exempt from the declared-mount rule on the theory that
it runs our own runtime. The model does not know it is ours. In the
#2245 two-arm run it grepped the mount and read
`docs/eval/terminal-bench-2.1-*.csv` — this benchmark's own per-task
pass/fail table from earlier runs — and, on another task, the verifier
source at `packages/headless/src/terminal-bench-adapter.ts`.
The container executes `dist`, so it now receives dependencies,
workspace manifests, build outputs and the one harbor entrypoint the
adapter runs. Sources, docs, repo history and the harness manifest are
no longer reachable from inside it.
Workspace-private `node_modules` are declared but dropped when absent:
which ones exist depends on what npm hoisted, and Docker materialises a
missing bind source as a new empty directory inside the repo.
The suite checked that an absent workspace dependency tree is dropped. The
incident was the opposite: `packages/runtime/node_modules` exists — npm could
not hoist `@slack/socket-mode` past a version conflict — and omitting it
resolved on the host, then failed in the container on the first import that
needed it. Only a live container run caught that. Ask the repo instead.
…obes
`install()` tests for both `run-cell.mjs` and `run-host-cell.mjs` before
either can run, in every cell-mode branch. Only the first was declared, so
the narrowed mount aborted the trial at install rather than degrading it —
a regression against the repo-root mount this PR replaces, and one the
container smoke check missed because it exercised loading, not installing.
The competitor arms already have the check that would have caught it: read
the adapter for what it names at a container path and assert the mount
declares it. Maka had no equivalent, which is why the omission shipped. It
does now — `maka_agent.py` joins path segments onto the mount root instead
of writing literals, so the test matches the join.
Also asserts the sources-and-records invariant against the mounts actually
produced rather than the declared list: a declaration naming no forbidden
path still leaks every one of them if the builder mounts the root anyway.
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
The list lived twice — once implied by the mount declaration, once spelled in
the test that cross-checks it against the root manifest. Two copies of an
exclusion drift into a workspace nobody mounts.
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
@Astro-Han
Astro-Han marked this pull request as ready for review August 6, 2026 06:02
@Astro-Han
Astro-Han merged commit 474afee into mainAug 6, 2026
11 checks passed
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts (#2303)
* refactor(headless): move the benchmark identity out of the build a graded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
* refactor(headless): hold the identity boundary where it actually is
Review of the first cut found the comments claiming more than the test held.
The content check read only the top level of dist while the whole directory
is mounted, and described the compiled tests as excluded when they are not.
It now scans dist recursively and states two boundaries, because they are two
different things: the retrieval keys — revision, upstream URL, tree
fingerprint — must be absent from every mounted file, while a single task id
is not a retrieval key and is only held out of the shipped modules, where a
complete list could accumulate.
No mount list may declare the identity file, the Maka one included, and the
DeepSWE cache path derives its revision instead of spelling a second copy.
* test(headless): check every mounted file, and the fingerprint the list forgot
The scan filtered to `.js`, so a source map — whose `sourcesContent` carries
the whole of `src` — would have walked straight past it. And the retrieval
keys named every task-tree fingerprint but one.
* refactor(headless): drop a benchmark task id from a retry comment
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

fix(headless): mount Maka's build outputs into the task container, not the repo root - #2298

Merged
Astro-Han merged 4 commits into
mainfrom
fix/harness-maka-repo-mount-scope
Aug 6, 2026
Merged

fix(headless): mount Maka's build outputs into the task container, not the repo root#2298
Astro-Han merged 4 commits into
mainfrom
fix/harness-maka-repo-mount-scope

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Aug 6, 2026

Copy link
Copy Markdown
Contributor

Summary

The Maka arm received the repo root at /opt/maka-agent, exempt from the declared-mount rule every competitor arm follows. The exemption's premise was that Maka runs our own runtime — but the model does not know it is ours, and it reads what is reachable.

In the #2245 two-arm run it did exactly that. On extract-elf it ran grep -rl "extract-elf|extract memory values" /opt/maka-agent and read docs/eval/terminal-bench-2.1-*.csv — this benchmark's own per-task pass/fail table from earlier runs. On db-wal-recovery it read packages/headless/src/terminal-bench-adapter.ts, the verifier source. Both cells passed. Neither file is one the runtime executes; both were reachable only because the mount was the root.

The container runs dist, so the arm now gets dependencies, workspace manifests, each workspace's build output, and the single harbor entrypoint the adapter invokes. Sources, docs/, repo history and the harness manifest source are gone from inside it.

Refs #2245

Verification

  • npm run -w @maka/headless test — 1425 pass, 0 fail. Three existing tests asserted the root mount and now assert the narrowed shape.
  • npm run lint, npm run format — clean
  • Ran the container for real, since the contract tests only check that the declared list is self-consistent — they cannot tell whether Maka still starts:
docker run --rm $(mounts from makaRepoPaths()) node:24-slim \
node /opt/maka-agent/packages/headless/dist/cli.js --help

The first attempt failed with Cannot find package '@slack/socket-mode': npm keeps a private node_modules in any workspace whose version pin it could not hoist, and the list only had the root one. That is exactly the class of failure the contract tests cannot see. With those added, the CLI starts and harbor --help resolves.

Reachability from inside the same container, after the change:

pathbeforeafter
docs/ (per-task results)readableblocked
packages/headless/src/ (verifier source)readableblocked
.git/readableblocked
maka-eval/readableblocked
packages/headless/harbor/run-harness-ab.mjsreadableblocked
packages/headless/dist/harness-ab-manifest.jsreadablereadable

Known gap — now closed by #2303

The last row is not fixed here. harness-ab-manifest.ts compiles TERMINAL_BENCH_2_1_REVISION, the upstream repository URL, the task-tree fingerprint and all 89 task ids into packages/headless/dist, which the container needs as a directory. Fixing it means moving benchmark identity out of a compiled runtime package, which is a data-layout change rather than a mount change; #2303 does that, and adds the content-level check this PR's path-level assertions structurally cannot make.

Review round 2

The suite checked that an absent workspace dependency tree is dropped. The incident was the opposite direction: packages/runtime/node_modules exists — npm could not hoist @slack/socket-mode past a version conflict — and omitting it resolved on the host, then failed inside the container on the first import that needed it. Only a live container run caught that.

declares every workspace dependency tree npm could not hoist away now asks the repo directly, and fails on the pre-fix list. The renderer-only exclusion and the workspace list are also derived from the root manifest in one place instead of two.

…t the repo root
The Maka arm was exempt from the declared-mount rule on the theory that
it runs our own runtime. The model does not know it is ours. In the
#2245 two-arm run it grepped the mount and read
`docs/eval/terminal-bench-2.1-*.csv` — this benchmark's own per-task
pass/fail table from earlier runs — and, on another task, the verifier
source at `packages/headless/src/terminal-bench-adapter.ts`.
The container executes `dist`, so it now receives dependencies,
workspace manifests, build outputs and the one harbor entrypoint the
adapter runs. Sources, docs, repo history and the harness manifest are
no longer reachable from inside it.
Workspace-private `node_modules` are declared but dropped when absent:
which ones exist depends on what npm hoisted, and Docker materialises a
missing bind source as a new empty directory inside the repo.
The suite checked that an absent workspace dependency tree is dropped. The
incident was the opposite: `packages/runtime/node_modules` exists — npm could
not hoist `@slack/socket-mode` past a version conflict — and omitting it
resolved on the host, then failed in the container on the first import that
needed it. Only a live container run caught that. Ask the repo instead.
…obes
`install()` tests for both `run-cell.mjs` and `run-host-cell.mjs` before
either can run, in every cell-mode branch. Only the first was declared, so
the narrowed mount aborted the trial at install rather than degrading it —
a regression against the repo-root mount this PR replaces, and one the
container smoke check missed because it exercised loading, not installing.
The competitor arms already have the check that would have caught it: read
the adapter for what it names at a container path and assert the mount
declares it. Maka had no equivalent, which is why the omission shipped. It
does now — `maka_agent.py` joins path segments onto the mount root instead
of writing literals, so the test matches the join.
Also asserts the sources-and-records invariant against the mounts actually
produced rather than the declared list: a declaration naming no forbidden
path still leaks every one of them if the builder mounts the root anyway.
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
The list lived twice — once implied by the mount declaration, once spelled in
the test that cross-checks it against the root manifest. Two copies of an
exclusion drift into a workspace nobody mounts.
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
@Astro-Han
Astro-Han marked this pull request as ready for review August 6, 2026 06:02
@Astro-Han
Astro-Han merged commit 474afee into mainAug 6, 2026
11 checks passed
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
Astro-Han added a commit that referenced this pull request Aug 6, 2026
…aded arm mounts (#2303)
* refactor(headless): move the benchmark identity out of the build a graded arm mounts
An arm executes this package's dist, so a benchmark's revision, upstream URL
and task list compiled into a shipped module are readable from inside the
container being scored — which is how the #1970 run turned a task into a
lookup at the pinned revision. Narrowing the mount (#2298) closed docs/eval
and src; the identity itself still shipped in harness-ab-manifest.js, and the
path-level mount checks cannot see it because they inspect declared paths,
never what the mounted files contain.
The data moves to harbor/benchmark-identity.json, which nothing mounts, and
the assertions in src become the comparison without the answer. A content
test over the shipped dist now holds the invariant directly; it caught a
Terminal-Bench task id hardcoded as a smoke default that the manifest was
already supplying.
Collapses two identical task-source branches: with each profile carrying its
own frozen fingerprint, the full-tree profiles differ only in their label.
* refactor(headless): hold the identity boundary where it actually is
Review of the first cut found the comments claiming more than the test held.
The content check read only the top level of dist while the whole directory
is mounted, and described the compiled tests as excluded when they are not.
It now scans dist recursively and states two boundaries, because they are two
different things: the retrieval keys — revision, upstream URL, tree
fingerprint — must be absent from every mounted file, while a single task id
is not a retrieval key and is only held out of the shipped modules, where a
complete list could accumulate.
No mount list may declare the identity file, the Maka one included, and the
DeepSWE cache path derives its revision instead of spelling a second copy.
* test(headless): check every mounted file, and the fingerprint the list forgot
The scan filtered to `.js`, so a source map — whose `sourcesContent` carries
the whole of `src` — would have walked straight past it. And the retrieval
keys named every task-tree fingerprint but one.
* refactor(headless): drop a benchmark task id from a retry comment
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han