Skip to content

feat(lab): @maka/lab — headless agent experiment lab (MVP, #31) - #36

Closed
Astro-Han wants to merge 6 commits into
mainfrom
claude/lab
Closed

feat(lab): @maka/lab — headless agent experiment lab (MVP, #31)#36
Astro-Han wants to merge 6 commits into
mainfrom
claude/lab

Conversation

@Astro-Han

Copy link
Copy Markdown
Contributor

Implements the MVP from #31. Independent of #32 / #33 — the lab runs on the FakeBackend with no credentials, and the real backend reads keys from an env var, not the credential store.

What & why

A headless lab for measuring an agent configuration: run a Config × Task grid in isolated sandboxes, capture each trajectory, score it with the task's own test command, and compare. Vary the model/connection (Config), hold the work (Task), read pass/fail off the table.

Config × Task → throwaway workspace → headless agent run → trajectory
↓
ResultRecord (JSONL) ← verification command

What's in this PR

New package @maka/lab (added to the workspaces list):

  • contractsTask (instruction + fixture workspace + verification command), Config (backend / connection / model), ResultRecord (one canonical row per run). Minimal on purpose; kept in @maka/lab for now, promotable to @maka/core once a second consumer exists.
  • sandbox — copies a task fixture into a throwaway dir; the source is never mutated and runs never bleed into each other. Isolation, not asking, is the safety model.
  • evaluator — runs the task's verification command (exit 0 = pass) in the throwaway workspace after the agent finishes, so a config can't grade itself; output-capped + timeout-killed.
  • runner — drives one headless turn through SessionManager, captures the trajectory via runtimeInvocationObserver, and auto-approves permission prompts (a benchmark has no human to confirm; the throwaway workspace is the boundary).
  • matrix — the full Config × Task cross product; a thrown run becomes a failed cell instead of aborting the grid.
  • results — ResultRecord JSONL is canonical truth; a git-diffable markdown comparison table is derived from it.
  • backends — the two concrete wirings, kept out of the engine: fake (deterministic) and ai-sdk (real model; key from a named env var, minimal AiSdkBackend = model + builtin tools + execute-mode permission).
  • CLImaka-lab run <spec.json> [--out <dir>] and maka-lab compare <results.jsonl>.
  • examplesdemo.spec.json (fake, no key) and fix-add.spec.json (a real coding task).

Live validation

The real ai-sdk backend was smoke-tested end-to-end on DeepSeek (deepseek-chat) with examples/fix-add (a buggy add() + a failing node:test):

  • completed, passed, exit 0, a 124-event trajectory using Edit / Write / Bash
  • the source fixture stayed buggy — the agent only edited the throwaway copy, proving sandbox isolation + headless permission auto-approve work against a real model

Testing

15 deterministic tests (FakeBackend), all green: CLI end-to-end smoke (spawns the built bin → results.jsonl + comparison.mdcompare), matrix cross-product + failed-cell, JSONL round-trip, table rendering, sandbox isolation, evaluator pass/fail/timeout, walking-skeleton e2e. Plus the live DeepSeek smoke above.

Out of scope (deliberate — pure additions later)

Parallel matrix execution, LLM/rule evaluators (today: command exit code), Docker / network isolation, SWE-bench pack ingestion, a richer report than the markdown grid, and promoting the contracts into @maka/core.

#31)
First end-to-end slice of the agent experiment lab: run one Config
against one Task in an isolated sandbox, capture the trajectory, score
it, and return a ResultRecord. De-risks the integration (headless
SessionManager run + sandbox + evaluator composing into one real run);
matrix/compare/CLI and real-model backends are later additions.
- contracts.ts: Task (instruction + fixture workspace + verification
command), Config (backend/connection/model), ResultRecord. Minimal;
systemPrompt/Execution/SWE-bench-pack ingestion deferred.
- sandbox.ts: prepareWorkspace copies the fixture to a throwaway dir so
the agent never mutates the source (isolation, not asking).
- evaluator.ts: runVerification runs the Task's test command in the
throwaway workspace after the agent finishes (config can't grade
itself), with output cap + timeout kill.
- runner.ts: runExperiment wires a pure-Node SessionManager (injected
registerBackends keeps model/credential wiring out of the lab core),
drives one turn, captures the InvocationResult trajectory via
runtimeInvocationObserver, scores it, returns a ResultRecord.
- tests: FakeBackend e2e (pass + fail fixtures, trajectory persisted as
runtime-events.jsonl), sandbox isolation, evaluator pass/fail/timeout.
Independent of the credential migration (#32/#33): the skeleton runs on
FakeBackend and needs no credentials.
Completes the MVP loop on top of the walking skeleton: run a grid of
Configs × Tasks, persist canonical results, and compare.
- matrix.ts: runMatrix runs the full cross product (sequential; a thrown
run becomes a failed cell instead of aborting the grid) with a
per-cell onResult callback.
- results.ts: ResultRecord JSONL is canonical truth; toComparisonTable
derives a git-diffable markdown grid (tasks × configs, ✅/❌/⚠️,
pass-rate footer). ResultRecord gains an optional `error`.
- backends.ts: the two concrete backend wirings, kept out of the engine.
registerFakeBackend (deterministic). registerAiSdkBackend resolves a
Config's slug against spec connections, reads the API key from a named
env var (no secrets at rest), and wires a minimal AiSdkBackend (model +
builtin tools + execute-mode permission). Telemetry/artifact/synthesis
hooks omitted — a benchmark scores via the verification command.
- runner.ts: the drain loop now auto-approves permission requests — a
headless benchmark has no human to confirm; throwaway-workspace
isolation is the safety net.
- cli.ts: `maka-lab run <spec.json> [--out <dir>]` and
`maka-lab compare <results.jsonl>`; task fixtures resolve relative to
the spec. Exposed as the `maka-lab` bin.
- tests (15 total): CLI end-to-end smoke (spawns the built bin on a fake
spec → results.jsonl + comparison.md → compare), matrix cross-product
+ failed-cell, JSONL round-trip, table rendering.
The real ai-sdk backend is typecheck-verified but not unit-tested (a live
model call is non-deterministic and costs money); it needs a live smoke
test with a real API key. Everything else is deterministic and green.
README documents the Config × Task model, the CLI (run/compare), the
spec shape (incl. a real ai-sdk connection with apiKeyEnv), the two
backends, and what's deliberately out of MVP scope. examples/demo.spec.json
+ examples/demo/marker.txt run green on the fake backend with no API key:
maka-lab run examples/demo.spec.json --out /tmp/maka-lab-demo
examples/fix-add — a buggy add() with a failing node:test. A real model
run must read the files, fix src.mjs, and turn `node --test` green.
Live smoke confirmed end-to-end on DeepSeek (deepseek-chat): completed,
passed, exit 0, 124-event trajectory using Edit/Write/Bash; the source
fixture stayed buggy (the agent only touched the throwaway copy), proving
sandbox isolation + headless permission auto-approve work against a real
backend. README points at it as the canonical real-run example.
…r, table (#36)
- runner: stop blanket-approving permission prompts. Allow ordinary tool
use, DENY dangerous categories (fs_destructive / git_destructive /
privileged / browser) by default — the workspace is a copy, not a jail,
so a tool can still escape via absolute paths/network. Opt in with
allowDangerousTools (real sandbox only). Corrects the misleading
"workspace is the safety net" comment. Also: a run that didn't complete
can no longer read as passed.
- evaluator: spawn detached + SIGKILL the process group on timeout, so
backgrounded grandchildren die too (a plain child.kill leaked them).
- sandbox: reject fixture symlinks (fs.cp preserved them verbatim → escape
to source/host) and clean up the temp dir if the copy fails (it was
leaked before the runner's finally registered).
- results: cellKey was separated by a NUL byte, making results.ts a binary
file to git — switch to a JSON key. Render a failed run as ⚠️ (distinct
from ❌ verification-failed) and exclude it from the pass count. Escape
`|`/newline in ids so they can't break the table.
- cli: reject unknown flags and require a value after --out.
- README: honest permission/safety wording.
+5 tests (symlink reject, failed-run render, id escaping, unknown flag,
missing flag value); 20 total green. Cross-process credential lock and a
real container sandbox remain deliberately out of scope.
The throwaway workspace stops a run from mutating the source fixture, but
it is NOT a security sandbox and a config could rewrite its own test to
pass. This addresses the two open review P1s without ripping out the real
backend (kept usable for your-own-models-on-trusted-tasks):
- Clean-room grading (verification integrity): Task.verification.protectedPaths
lists the grading assets; the runner restores them from the pristine
fixture AFTER the agent finishes and BEFORE the verification command runs,
so a model that rewrote its own test has that edit reverted. Each path is
removed first (drops an agent-planted symlink) and rejected if it escapes
the workspace. examples/fix-add now protects test.mjs.
- Honest docs (host isolation): the README states plainly that a real run
executes tool calls on your machine with your privileges — same exposure
as running Maka directly — so run only models/tasks you trust. Per-run
container isolation (mount workspace only, env allowlist, network policy)
is the named next hardening; allowDangerousTools is for inside it.
Tests (+4, 24 total green): restoreProtectedPaths unit (reverts protected,
keeps the rest, rejects escapes) + a malicious TamperBackend integration
proving protectedPaths reverts the cheat (passed=false) while the unguarded
task lets the cheat win (passed=true) — the guard is load-bearing. The
integration grades via `node check.mjs` (exit-code), not `node --test`, so
the verification child doesn't collide with the lab's own test runner.
@Astro-Han
Astro-Han deleted the claude/lab branch June 17, 2026 12:51
jackwener pushed a commit that referenced this pull request Jun 21, 2026
…r, table (#36)
- runner: stop blanket-approving permission prompts. Allow ordinary tool
use, DENY dangerous categories (fs_destructive / git_destructive /
privileged / browser) by default — the workspace is a copy, not a jail,
so a tool can still escape via absolute paths/network. Opt in with
allowDangerousTools (real sandbox only). Corrects the misleading
"workspace is the safety net" comment. Also: a run that didn't complete
can no longer read as passed.
- evaluator: spawn detached + SIGKILL the process group on timeout, so
backgrounded grandchildren die too (a plain child.kill leaked them).
- sandbox: reject fixture symlinks (fs.cp preserved them verbatim → escape
to source/host) and clean up the temp dir if the copy fails (it was
leaked before the runner's finally registered).
- results: cellKey was separated by a NUL byte, making results.ts a binary
file to git — switch to a JSON key. Render a failed run as ⚠️ (distinct
from ❌ verification-failed) and exclude it from the pass count. Escape
`|`/newline in ids so they can't break the table.
- cli: reject unknown flags and require a value after --out.
- README: honest permission/safety wording.
+5 tests (symlink reject, failed-run render, id escaping, unknown flag,
missing flag value); 20 total green. Cross-process credential lock and a
real container sandbox remain deliberately out of scope.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
The throwaway workspace stops a run from mutating the source fixture, but
it is NOT a security sandbox and a config could rewrite its own test to
pass. This addresses the two open review P1s without ripping out the real
backend (kept usable for your-own-models-on-trusted-tasks):
- Clean-room grading (verification integrity): Task.verification.protectedPaths
lists the grading assets; the runner restores them from the pristine
fixture AFTER the agent finishes and BEFORE the verification command runs,
so a model that rewrote its own test has that edit reverted. Each path is
removed first (drops an agent-planted symlink) and rejected if it escapes
the workspace. examples/fix-add now protects test.mjs.
- Honest docs (host isolation): the README states plainly that a real run
executes tool calls on your machine with your privileges — same exposure
as running Maka directly — so run only models/tasks you trust. Per-run
container isolation (mount workspace only, env allowlist, network policy)
is the named next hardening; allowDangerousTools is for inside it.
Tests (+4, 24 total green): restoreProtectedPaths unit (reverts protected,
keeps the rest, rejects escapes) + a malicious TamperBackend integration
proving protectedPaths reverts the cheat (passed=false) while the unguarded
task lets the cheat win (passed=true) — the guard is load-bearing. The
integration grades via `node check.mjs` (exit-code), not `node --test`, so
the verification child doesn't collide with the lab's own test runner.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
…r, table (#36)
- runner: stop blanket-approving permission prompts. Allow ordinary tool
use, DENY dangerous categories (fs_destructive / git_destructive /
privileged / browser) by default — the workspace is a copy, not a jail,
so a tool can still escape via absolute paths/network. Opt in with
allowDangerousTools (real sandbox only). Corrects the misleading
"workspace is the safety net" comment. Also: a run that didn't complete
can no longer read as passed.
- evaluator: spawn detached + SIGKILL the process group on timeout, so
backgrounded grandchildren die too (a plain child.kill leaked them).
- sandbox: reject fixture symlinks (fs.cp preserved them verbatim → escape
to source/host) and clean up the temp dir if the copy fails (it was
leaked before the runner's finally registered).
- results: cellKey was separated by a NUL byte, making results.ts a binary
file to git — switch to a JSON key. Render a failed run as ⚠️ (distinct
from ❌ verification-failed) and exclude it from the pass count. Escape
`|`/newline in ids so they can't break the table.
- cli: reject unknown flags and require a value after --out.
- README: honest permission/safety wording.
+5 tests (symlink reject, failed-run render, id escaping, unknown flag,
missing flag value); 20 total green. Cross-process credential lock and a
real container sandbox remain deliberately out of scope.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
The throwaway workspace stops a run from mutating the source fixture, but
it is NOT a security sandbox and a config could rewrite its own test to
pass. This addresses the two open review P1s without ripping out the real
backend (kept usable for your-own-models-on-trusted-tasks):
- Clean-room grading (verification integrity): Task.verification.protectedPaths
lists the grading assets; the runner restores them from the pristine
fixture AFTER the agent finishes and BEFORE the verification command runs,
so a model that rewrote its own test has that edit reverted. Each path is
removed first (drops an agent-planted symlink) and rejected if it escapes
the workspace. examples/fix-add now protects test.mjs.
- Honest docs (host isolation): the README states plainly that a real run
executes tool calls on your machine with your privileges — same exposure
as running Maka directly — so run only models/tasks you trust. Per-run
container isolation (mount workspace only, env allowlist, network policy)
is the named next hardening; allowDangerousTools is for inside it.
Tests (+4, 24 total green): restoreProtectedPaths unit (reverts protected,
keeps the rest, rejects escapes) + a malicious TamperBackend integration
proving protectedPaths reverts the cheat (passed=false) while the unguarded
task lets the cheat win (passed=true) — the guard is load-bearing. The
integration grades via `node check.mjs` (exit-code), not `node --test`, so
the verification child doesn't collide with the lab's own test runner.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
…r, table (#36)
- runner: stop blanket-approving permission prompts. Allow ordinary tool
use, DENY dangerous categories (fs_destructive / git_destructive /
privileged / browser) by default — the workspace is a copy, not a jail,
so a tool can still escape via absolute paths/network. Opt in with
allowDangerousTools (real sandbox only). Corrects the misleading
"workspace is the safety net" comment. Also: a run that didn't complete
can no longer read as passed.
- evaluator: spawn detached + SIGKILL the process group on timeout, so
backgrounded grandchildren die too (a plain child.kill leaked them).
- sandbox: reject fixture symlinks (fs.cp preserved them verbatim → escape
to source/host) and clean up the temp dir if the copy fails (it was
leaked before the runner's finally registered).
- results: cellKey was separated by a NUL byte, making results.ts a binary
file to git — switch to a JSON key. Render a failed run as ⚠️ (distinct
from ❌ verification-failed) and exclude it from the pass count. Escape
`|`/newline in ids so they can't break the table.
- cli: reject unknown flags and require a value after --out.
- README: honest permission/safety wording.
+5 tests (symlink reject, failed-run render, id escaping, unknown flag,
missing flag value); 20 total green. Cross-process credential lock and a
real container sandbox remain deliberately out of scope.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
The throwaway workspace stops a run from mutating the source fixture, but
it is NOT a security sandbox and a config could rewrite its own test to
pass. This addresses the two open review P1s without ripping out the real
backend (kept usable for your-own-models-on-trusted-tasks):
- Clean-room grading (verification integrity): Task.verification.protectedPaths
lists the grading assets; the runner restores them from the pristine
fixture AFTER the agent finishes and BEFORE the verification command runs,
so a model that rewrote its own test has that edit reverted. Each path is
removed first (drops an agent-planted symlink) and rejected if it escapes
the workspace. examples/fix-add now protects test.mjs.
- Honest docs (host isolation): the README states plainly that a real run
executes tool calls on your machine with your privileges — same exposure
as running Maka directly — so run only models/tasks you trust. Per-run
container isolation (mount workspace only, env allowlist, network policy)
is the named next hardening; allowDangerousTools is for inside it.
Tests (+4, 24 total green): restoreProtectedPaths unit (reverts protected,
keeps the rest, rejects escapes) + a malicious TamperBackend integration
proving protectedPaths reverts the cheat (passed=false) while the unguarded
task lets the cheat win (passed=true) — the guard is load-bearing. The
integration grades via `node check.mjs` (exit-code), not `node --test`, so
the verification child doesn't collide with the lab's own test runner.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
feat(lab): @maka/lab — headless agent experiment lab (MVP, #31) by Astro-Han · Pull Request #36 · apache/maka · GitHub
Skip to content

feat(lab): @maka/lab — headless agent experiment lab (MVP, #31) - #36

Closed
Astro-Han wants to merge 6 commits into
mainfrom
claude/lab
Closed

feat(lab): @maka/lab — headless agent experiment lab (MVP, #31)#36
Astro-Han wants to merge 6 commits into
mainfrom
claude/lab

Conversation

@Astro-Han

Copy link
Copy Markdown
Contributor

Implements the MVP from #31. Independent of #32 / #33 — the lab runs on the FakeBackend with no credentials, and the real backend reads keys from an env var, not the credential store.

What & why

A headless lab for measuring an agent configuration: run a Config × Task grid in isolated sandboxes, capture each trajectory, score it with the task's own test command, and compare. Vary the model/connection (Config), hold the work (Task), read pass/fail off the table.

Config × Task → throwaway workspace → headless agent run → trajectory
↓
ResultRecord (JSONL) ← verification command

What's in this PR

New package @maka/lab (added to the workspaces list):

  • contractsTask (instruction + fixture workspace + verification command), Config (backend / connection / model), ResultRecord (one canonical row per run). Minimal on purpose; kept in @maka/lab for now, promotable to @maka/core once a second consumer exists.
  • sandbox — copies a task fixture into a throwaway dir; the source is never mutated and runs never bleed into each other. Isolation, not asking, is the safety model.
  • evaluator — runs the task's verification command (exit 0 = pass) in the throwaway workspace after the agent finishes, so a config can't grade itself; output-capped + timeout-killed.
  • runner — drives one headless turn through SessionManager, captures the trajectory via runtimeInvocationObserver, and auto-approves permission prompts (a benchmark has no human to confirm; the throwaway workspace is the boundary).
  • matrix — the full Config × Task cross product; a thrown run becomes a failed cell instead of aborting the grid.
  • results — ResultRecord JSONL is canonical truth; a git-diffable markdown comparison table is derived from it.
  • backends — the two concrete wirings, kept out of the engine: fake (deterministic) and ai-sdk (real model; key from a named env var, minimal AiSdkBackend = model + builtin tools + execute-mode permission).
  • CLImaka-lab run <spec.json> [--out <dir>] and maka-lab compare <results.jsonl>.
  • examplesdemo.spec.json (fake, no key) and fix-add.spec.json (a real coding task).

Live validation

The real ai-sdk backend was smoke-tested end-to-end on DeepSeek (deepseek-chat) with examples/fix-add (a buggy add() + a failing node:test):

  • completed, passed, exit 0, a 124-event trajectory using Edit / Write / Bash
  • the source fixture stayed buggy — the agent only edited the throwaway copy, proving sandbox isolation + headless permission auto-approve work against a real model

Testing

15 deterministic tests (FakeBackend), all green: CLI end-to-end smoke (spawns the built bin → results.jsonl + comparison.mdcompare), matrix cross-product + failed-cell, JSONL round-trip, table rendering, sandbox isolation, evaluator pass/fail/timeout, walking-skeleton e2e. Plus the live DeepSeek smoke above.

Out of scope (deliberate — pure additions later)

Parallel matrix execution, LLM/rule evaluators (today: command exit code), Docker / network isolation, SWE-bench pack ingestion, a richer report than the markdown grid, and promoting the contracts into @maka/core.

#31)
First end-to-end slice of the agent experiment lab: run one Config
against one Task in an isolated sandbox, capture the trajectory, score
it, and return a ResultRecord. De-risks the integration (headless
SessionManager run + sandbox + evaluator composing into one real run);
matrix/compare/CLI and real-model backends are later additions.
- contracts.ts: Task (instruction + fixture workspace + verification
command), Config (backend/connection/model), ResultRecord. Minimal;
systemPrompt/Execution/SWE-bench-pack ingestion deferred.
- sandbox.ts: prepareWorkspace copies the fixture to a throwaway dir so
the agent never mutates the source (isolation, not asking).
- evaluator.ts: runVerification runs the Task's test command in the
throwaway workspace after the agent finishes (config can't grade
itself), with output cap + timeout kill.
- runner.ts: runExperiment wires a pure-Node SessionManager (injected
registerBackends keeps model/credential wiring out of the lab core),
drives one turn, captures the InvocationResult trajectory via
runtimeInvocationObserver, scores it, returns a ResultRecord.
- tests: FakeBackend e2e (pass + fail fixtures, trajectory persisted as
runtime-events.jsonl), sandbox isolation, evaluator pass/fail/timeout.
Independent of the credential migration (#32/#33): the skeleton runs on
FakeBackend and needs no credentials.
Completes the MVP loop on top of the walking skeleton: run a grid of
Configs × Tasks, persist canonical results, and compare.
- matrix.ts: runMatrix runs the full cross product (sequential; a thrown
run becomes a failed cell instead of aborting the grid) with a
per-cell onResult callback.
- results.ts: ResultRecord JSONL is canonical truth; toComparisonTable
derives a git-diffable markdown grid (tasks × configs, ✅/❌/⚠️,
pass-rate footer). ResultRecord gains an optional `error`.
- backends.ts: the two concrete backend wirings, kept out of the engine.
registerFakeBackend (deterministic). registerAiSdkBackend resolves a
Config's slug against spec connections, reads the API key from a named
env var (no secrets at rest), and wires a minimal AiSdkBackend (model +
builtin tools + execute-mode permission). Telemetry/artifact/synthesis
hooks omitted — a benchmark scores via the verification command.
- runner.ts: the drain loop now auto-approves permission requests — a
headless benchmark has no human to confirm; throwaway-workspace
isolation is the safety net.
- cli.ts: `maka-lab run <spec.json> [--out <dir>]` and
`maka-lab compare <results.jsonl>`; task fixtures resolve relative to
the spec. Exposed as the `maka-lab` bin.
- tests (15 total): CLI end-to-end smoke (spawns the built bin on a fake
spec → results.jsonl + comparison.md → compare), matrix cross-product
+ failed-cell, JSONL round-trip, table rendering.
The real ai-sdk backend is typecheck-verified but not unit-tested (a live
model call is non-deterministic and costs money); it needs a live smoke
test with a real API key. Everything else is deterministic and green.
README documents the Config × Task model, the CLI (run/compare), the
spec shape (incl. a real ai-sdk connection with apiKeyEnv), the two
backends, and what's deliberately out of MVP scope. examples/demo.spec.json
+ examples/demo/marker.txt run green on the fake backend with no API key:
maka-lab run examples/demo.spec.json --out /tmp/maka-lab-demo
examples/fix-add — a buggy add() with a failing node:test. A real model
run must read the files, fix src.mjs, and turn `node --test` green.
Live smoke confirmed end-to-end on DeepSeek (deepseek-chat): completed,
passed, exit 0, 124-event trajectory using Edit/Write/Bash; the source
fixture stayed buggy (the agent only touched the throwaway copy), proving
sandbox isolation + headless permission auto-approve work against a real
backend. README points at it as the canonical real-run example.
…r, table (#36)
- runner: stop blanket-approving permission prompts. Allow ordinary tool
use, DENY dangerous categories (fs_destructive / git_destructive /
privileged / browser) by default — the workspace is a copy, not a jail,
so a tool can still escape via absolute paths/network. Opt in with
allowDangerousTools (real sandbox only). Corrects the misleading
"workspace is the safety net" comment. Also: a run that didn't complete
can no longer read as passed.
- evaluator: spawn detached + SIGKILL the process group on timeout, so
backgrounded grandchildren die too (a plain child.kill leaked them).
- sandbox: reject fixture symlinks (fs.cp preserved them verbatim → escape
to source/host) and clean up the temp dir if the copy fails (it was
leaked before the runner's finally registered).
- results: cellKey was separated by a NUL byte, making results.ts a binary
file to git — switch to a JSON key. Render a failed run as ⚠️ (distinct
from ❌ verification-failed) and exclude it from the pass count. Escape
`|`/newline in ids so they can't break the table.
- cli: reject unknown flags and require a value after --out.
- README: honest permission/safety wording.
+5 tests (symlink reject, failed-run render, id escaping, unknown flag,
missing flag value); 20 total green. Cross-process credential lock and a
real container sandbox remain deliberately out of scope.
The throwaway workspace stops a run from mutating the source fixture, but
it is NOT a security sandbox and a config could rewrite its own test to
pass. This addresses the two open review P1s without ripping out the real
backend (kept usable for your-own-models-on-trusted-tasks):
- Clean-room grading (verification integrity): Task.verification.protectedPaths
lists the grading assets; the runner restores them from the pristine
fixture AFTER the agent finishes and BEFORE the verification command runs,
so a model that rewrote its own test has that edit reverted. Each path is
removed first (drops an agent-planted symlink) and rejected if it escapes
the workspace. examples/fix-add now protects test.mjs.
- Honest docs (host isolation): the README states plainly that a real run
executes tool calls on your machine with your privileges — same exposure
as running Maka directly — so run only models/tasks you trust. Per-run
container isolation (mount workspace only, env allowlist, network policy)
is the named next hardening; allowDangerousTools is for inside it.
Tests (+4, 24 total green): restoreProtectedPaths unit (reverts protected,
keeps the rest, rejects escapes) + a malicious TamperBackend integration
proving protectedPaths reverts the cheat (passed=false) while the unguarded
task lets the cheat win (passed=true) — the guard is load-bearing. The
integration grades via `node check.mjs` (exit-code), not `node --test`, so
the verification child doesn't collide with the lab's own test runner.
@Astro-Han
Astro-Han deleted the claude/lab branch June 17, 2026 12:51
jackwener pushed a commit that referenced this pull request Jun 21, 2026
…r, table (#36)
- runner: stop blanket-approving permission prompts. Allow ordinary tool
use, DENY dangerous categories (fs_destructive / git_destructive /
privileged / browser) by default — the workspace is a copy, not a jail,
so a tool can still escape via absolute paths/network. Opt in with
allowDangerousTools (real sandbox only). Corrects the misleading
"workspace is the safety net" comment. Also: a run that didn't complete
can no longer read as passed.
- evaluator: spawn detached + SIGKILL the process group on timeout, so
backgrounded grandchildren die too (a plain child.kill leaked them).
- sandbox: reject fixture symlinks (fs.cp preserved them verbatim → escape
to source/host) and clean up the temp dir if the copy fails (it was
leaked before the runner's finally registered).
- results: cellKey was separated by a NUL byte, making results.ts a binary
file to git — switch to a JSON key. Render a failed run as ⚠️ (distinct
from ❌ verification-failed) and exclude it from the pass count. Escape
`|`/newline in ids so they can't break the table.
- cli: reject unknown flags and require a value after --out.
- README: honest permission/safety wording.
+5 tests (symlink reject, failed-run render, id escaping, unknown flag,
missing flag value); 20 total green. Cross-process credential lock and a
real container sandbox remain deliberately out of scope.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
The throwaway workspace stops a run from mutating the source fixture, but
it is NOT a security sandbox and a config could rewrite its own test to
pass. This addresses the two open review P1s without ripping out the real
backend (kept usable for your-own-models-on-trusted-tasks):
- Clean-room grading (verification integrity): Task.verification.protectedPaths
lists the grading assets; the runner restores them from the pristine
fixture AFTER the agent finishes and BEFORE the verification command runs,
so a model that rewrote its own test has that edit reverted. Each path is
removed first (drops an agent-planted symlink) and rejected if it escapes
the workspace. examples/fix-add now protects test.mjs.
- Honest docs (host isolation): the README states plainly that a real run
executes tool calls on your machine with your privileges — same exposure
as running Maka directly — so run only models/tasks you trust. Per-run
container isolation (mount workspace only, env allowlist, network policy)
is the named next hardening; allowDangerousTools is for inside it.
Tests (+4, 24 total green): restoreProtectedPaths unit (reverts protected,
keeps the rest, rejects escapes) + a malicious TamperBackend integration
proving protectedPaths reverts the cheat (passed=false) while the unguarded
task lets the cheat win (passed=true) — the guard is load-bearing. The
integration grades via `node check.mjs` (exit-code), not `node --test`, so
the verification child doesn't collide with the lab's own test runner.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
…r, table (#36)
- runner: stop blanket-approving permission prompts. Allow ordinary tool
use, DENY dangerous categories (fs_destructive / git_destructive /
privileged / browser) by default — the workspace is a copy, not a jail,
so a tool can still escape via absolute paths/network. Opt in with
allowDangerousTools (real sandbox only). Corrects the misleading
"workspace is the safety net" comment. Also: a run that didn't complete
can no longer read as passed.
- evaluator: spawn detached + SIGKILL the process group on timeout, so
backgrounded grandchildren die too (a plain child.kill leaked them).
- sandbox: reject fixture symlinks (fs.cp preserved them verbatim → escape
to source/host) and clean up the temp dir if the copy fails (it was
leaked before the runner's finally registered).
- results: cellKey was separated by a NUL byte, making results.ts a binary
file to git — switch to a JSON key. Render a failed run as ⚠️ (distinct
from ❌ verification-failed) and exclude it from the pass count. Escape
`|`/newline in ids so they can't break the table.
- cli: reject unknown flags and require a value after --out.
- README: honest permission/safety wording.
+5 tests (symlink reject, failed-run render, id escaping, unknown flag,
missing flag value); 20 total green. Cross-process credential lock and a
real container sandbox remain deliberately out of scope.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
The throwaway workspace stops a run from mutating the source fixture, but
it is NOT a security sandbox and a config could rewrite its own test to
pass. This addresses the two open review P1s without ripping out the real
backend (kept usable for your-own-models-on-trusted-tasks):
- Clean-room grading (verification integrity): Task.verification.protectedPaths
lists the grading assets; the runner restores them from the pristine
fixture AFTER the agent finishes and BEFORE the verification command runs,
so a model that rewrote its own test has that edit reverted. Each path is
removed first (drops an agent-planted symlink) and rejected if it escapes
the workspace. examples/fix-add now protects test.mjs.
- Honest docs (host isolation): the README states plainly that a real run
executes tool calls on your machine with your privileges — same exposure
as running Maka directly — so run only models/tasks you trust. Per-run
container isolation (mount workspace only, env allowlist, network policy)
is the named next hardening; allowDangerousTools is for inside it.
Tests (+4, 24 total green): restoreProtectedPaths unit (reverts protected,
keeps the rest, rejects escapes) + a malicious TamperBackend integration
proving protectedPaths reverts the cheat (passed=false) while the unguarded
task lets the cheat win (passed=true) — the guard is load-bearing. The
integration grades via `node check.mjs` (exit-code), not `node --test`, so
the verification child doesn't collide with the lab's own test runner.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
…r, table (#36)
- runner: stop blanket-approving permission prompts. Allow ordinary tool
use, DENY dangerous categories (fs_destructive / git_destructive /
privileged / browser) by default — the workspace is a copy, not a jail,
so a tool can still escape via absolute paths/network. Opt in with
allowDangerousTools (real sandbox only). Corrects the misleading
"workspace is the safety net" comment. Also: a run that didn't complete
can no longer read as passed.
- evaluator: spawn detached + SIGKILL the process group on timeout, so
backgrounded grandchildren die too (a plain child.kill leaked them).
- sandbox: reject fixture symlinks (fs.cp preserved them verbatim → escape
to source/host) and clean up the temp dir if the copy fails (it was
leaked before the runner's finally registered).
- results: cellKey was separated by a NUL byte, making results.ts a binary
file to git — switch to a JSON key. Render a failed run as ⚠️ (distinct
from ❌ verification-failed) and exclude it from the pass count. Escape
`|`/newline in ids so they can't break the table.
- cli: reject unknown flags and require a value after --out.
- README: honest permission/safety wording.
+5 tests (symlink reject, failed-run render, id escaping, unknown flag,
missing flag value); 20 total green. Cross-process credential lock and a
real container sandbox remain deliberately out of scope.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
The throwaway workspace stops a run from mutating the source fixture, but
it is NOT a security sandbox and a config could rewrite its own test to
pass. This addresses the two open review P1s without ripping out the real
backend (kept usable for your-own-models-on-trusted-tasks):
- Clean-room grading (verification integrity): Task.verification.protectedPaths
lists the grading assets; the runner restores them from the pristine
fixture AFTER the agent finishes and BEFORE the verification command runs,
so a model that rewrote its own test has that edit reverted. Each path is
removed first (drops an agent-planted symlink) and rejected if it escapes
the workspace. examples/fix-add now protects test.mjs.
- Honest docs (host isolation): the README states plainly that a real run
executes tool calls on your machine with your privileges — same exposure
as running Maka directly — so run only models/tasks you trust. Per-run
container isolation (mount workspace only, env allowlist, network policy)
is the named next hardening; allowDangerousTools is for inside it.
Tests (+4, 24 total green): restoreProtectedPaths unit (reverts protected,
keeps the rest, rejects escapes) + a malicious TamperBackend integration
proving protectedPaths reverts the cheat (passed=false) while the unguarded
task lets the cheat win (passed=true) — the guard is load-bearing. The
integration grades via `node check.mjs` (exit-code), not `node --test`, so
the verification child doesn't collide with the lab's own test runner.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' feat(lab): @maka/lab — headless agent experiment lab (MVP, #31) by Astro-Han · Pull Request #36 · apache/maka · GitHub
Skip to content

feat(lab): @maka/lab — headless agent experiment lab (MVP, #31) - #36

Closed
Astro-Han wants to merge 6 commits into
mainfrom
claude/lab
Closed

feat(lab): @maka/lab — headless agent experiment lab (MVP, #31)#36
Astro-Han wants to merge 6 commits into
mainfrom
claude/lab

Conversation

@Astro-Han

Copy link
Copy Markdown
Contributor

Implements the MVP from #31. Independent of #32 / #33 — the lab runs on the FakeBackend with no credentials, and the real backend reads keys from an env var, not the credential store.

What & why

A headless lab for measuring an agent configuration: run a Config × Task grid in isolated sandboxes, capture each trajectory, score it with the task's own test command, and compare. Vary the model/connection (Config), hold the work (Task), read pass/fail off the table.

Config × Task → throwaway workspace → headless agent run → trajectory
↓
ResultRecord (JSONL) ← verification command

What's in this PR

New package @maka/lab (added to the workspaces list):

  • contractsTask (instruction + fixture workspace + verification command), Config (backend / connection / model), ResultRecord (one canonical row per run). Minimal on purpose; kept in @maka/lab for now, promotable to @maka/core once a second consumer exists.
  • sandbox — copies a task fixture into a throwaway dir; the source is never mutated and runs never bleed into each other. Isolation, not asking, is the safety model.
  • evaluator — runs the task's verification command (exit 0 = pass) in the throwaway workspace after the agent finishes, so a config can't grade itself; output-capped + timeout-killed.
  • runner — drives one headless turn through SessionManager, captures the trajectory via runtimeInvocationObserver, and auto-approves permission prompts (a benchmark has no human to confirm; the throwaway workspace is the boundary).
  • matrix — the full Config × Task cross product; a thrown run becomes a failed cell instead of aborting the grid.
  • results — ResultRecord JSONL is canonical truth; a git-diffable markdown comparison table is derived from it.
  • backends — the two concrete wirings, kept out of the engine: fake (deterministic) and ai-sdk (real model; key from a named env var, minimal AiSdkBackend = model + builtin tools + execute-mode permission).
  • CLImaka-lab run <spec.json> [--out <dir>] and maka-lab compare <results.jsonl>.
  • examplesdemo.spec.json (fake, no key) and fix-add.spec.json (a real coding task).

Live validation

The real ai-sdk backend was smoke-tested end-to-end on DeepSeek (deepseek-chat) with examples/fix-add (a buggy add() + a failing node:test):

  • completed, passed, exit 0, a 124-event trajectory using Edit / Write / Bash
  • the source fixture stayed buggy — the agent only edited the throwaway copy, proving sandbox isolation + headless permission auto-approve work against a real model

Testing

15 deterministic tests (FakeBackend), all green: CLI end-to-end smoke (spawns the built bin → results.jsonl + comparison.mdcompare), matrix cross-product + failed-cell, JSONL round-trip, table rendering, sandbox isolation, evaluator pass/fail/timeout, walking-skeleton e2e. Plus the live DeepSeek smoke above.

Out of scope (deliberate — pure additions later)

Parallel matrix execution, LLM/rule evaluators (today: command exit code), Docker / network isolation, SWE-bench pack ingestion, a richer report than the markdown grid, and promoting the contracts into @maka/core.

#31)
First end-to-end slice of the agent experiment lab: run one Config
against one Task in an isolated sandbox, capture the trajectory, score
it, and return a ResultRecord. De-risks the integration (headless
SessionManager run + sandbox + evaluator composing into one real run);
matrix/compare/CLI and real-model backends are later additions.
- contracts.ts: Task (instruction + fixture workspace + verification
command), Config (backend/connection/model), ResultRecord. Minimal;
systemPrompt/Execution/SWE-bench-pack ingestion deferred.
- sandbox.ts: prepareWorkspace copies the fixture to a throwaway dir so
the agent never mutates the source (isolation, not asking).
- evaluator.ts: runVerification runs the Task's test command in the
throwaway workspace after the agent finishes (config can't grade
itself), with output cap + timeout kill.
- runner.ts: runExperiment wires a pure-Node SessionManager (injected
registerBackends keeps model/credential wiring out of the lab core),
drives one turn, captures the InvocationResult trajectory via
runtimeInvocationObserver, scores it, returns a ResultRecord.
- tests: FakeBackend e2e (pass + fail fixtures, trajectory persisted as
runtime-events.jsonl), sandbox isolation, evaluator pass/fail/timeout.
Independent of the credential migration (#32/#33): the skeleton runs on
FakeBackend and needs no credentials.
Completes the MVP loop on top of the walking skeleton: run a grid of
Configs × Tasks, persist canonical results, and compare.
- matrix.ts: runMatrix runs the full cross product (sequential; a thrown
run becomes a failed cell instead of aborting the grid) with a
per-cell onResult callback.
- results.ts: ResultRecord JSONL is canonical truth; toComparisonTable
derives a git-diffable markdown grid (tasks × configs, ✅/❌/⚠️,
pass-rate footer). ResultRecord gains an optional `error`.
- backends.ts: the two concrete backend wirings, kept out of the engine.
registerFakeBackend (deterministic). registerAiSdkBackend resolves a
Config's slug against spec connections, reads the API key from a named
env var (no secrets at rest), and wires a minimal AiSdkBackend (model +
builtin tools + execute-mode permission). Telemetry/artifact/synthesis
hooks omitted — a benchmark scores via the verification command.
- runner.ts: the drain loop now auto-approves permission requests — a
headless benchmark has no human to confirm; throwaway-workspace
isolation is the safety net.
- cli.ts: `maka-lab run <spec.json> [--out <dir>]` and
`maka-lab compare <results.jsonl>`; task fixtures resolve relative to
the spec. Exposed as the `maka-lab` bin.
- tests (15 total): CLI end-to-end smoke (spawns the built bin on a fake
spec → results.jsonl + comparison.md → compare), matrix cross-product
+ failed-cell, JSONL round-trip, table rendering.
The real ai-sdk backend is typecheck-verified but not unit-tested (a live
model call is non-deterministic and costs money); it needs a live smoke
test with a real API key. Everything else is deterministic and green.
README documents the Config × Task model, the CLI (run/compare), the
spec shape (incl. a real ai-sdk connection with apiKeyEnv), the two
backends, and what's deliberately out of MVP scope. examples/demo.spec.json
+ examples/demo/marker.txt run green on the fake backend with no API key:
maka-lab run examples/demo.spec.json --out /tmp/maka-lab-demo
examples/fix-add — a buggy add() with a failing node:test. A real model
run must read the files, fix src.mjs, and turn `node --test` green.
Live smoke confirmed end-to-end on DeepSeek (deepseek-chat): completed,
passed, exit 0, 124-event trajectory using Edit/Write/Bash; the source
fixture stayed buggy (the agent only touched the throwaway copy), proving
sandbox isolation + headless permission auto-approve work against a real
backend. README points at it as the canonical real-run example.
…r, table (#36)
- runner: stop blanket-approving permission prompts. Allow ordinary tool
use, DENY dangerous categories (fs_destructive / git_destructive /
privileged / browser) by default — the workspace is a copy, not a jail,
so a tool can still escape via absolute paths/network. Opt in with
allowDangerousTools (real sandbox only). Corrects the misleading
"workspace is the safety net" comment. Also: a run that didn't complete
can no longer read as passed.
- evaluator: spawn detached + SIGKILL the process group on timeout, so
backgrounded grandchildren die too (a plain child.kill leaked them).
- sandbox: reject fixture symlinks (fs.cp preserved them verbatim → escape
to source/host) and clean up the temp dir if the copy fails (it was
leaked before the runner's finally registered).
- results: cellKey was separated by a NUL byte, making results.ts a binary
file to git — switch to a JSON key. Render a failed run as ⚠️ (distinct
from ❌ verification-failed) and exclude it from the pass count. Escape
`|`/newline in ids so they can't break the table.
- cli: reject unknown flags and require a value after --out.
- README: honest permission/safety wording.
+5 tests (symlink reject, failed-run render, id escaping, unknown flag,
missing flag value); 20 total green. Cross-process credential lock and a
real container sandbox remain deliberately out of scope.
The throwaway workspace stops a run from mutating the source fixture, but
it is NOT a security sandbox and a config could rewrite its own test to
pass. This addresses the two open review P1s without ripping out the real
backend (kept usable for your-own-models-on-trusted-tasks):
- Clean-room grading (verification integrity): Task.verification.protectedPaths
lists the grading assets; the runner restores them from the pristine
fixture AFTER the agent finishes and BEFORE the verification command runs,
so a model that rewrote its own test has that edit reverted. Each path is
removed first (drops an agent-planted symlink) and rejected if it escapes
the workspace. examples/fix-add now protects test.mjs.
- Honest docs (host isolation): the README states plainly that a real run
executes tool calls on your machine with your privileges — same exposure
as running Maka directly — so run only models/tasks you trust. Per-run
container isolation (mount workspace only, env allowlist, network policy)
is the named next hardening; allowDangerousTools is for inside it.
Tests (+4, 24 total green): restoreProtectedPaths unit (reverts protected,
keeps the rest, rejects escapes) + a malicious TamperBackend integration
proving protectedPaths reverts the cheat (passed=false) while the unguarded
task lets the cheat win (passed=true) — the guard is load-bearing. The
integration grades via `node check.mjs` (exit-code), not `node --test`, so
the verification child doesn't collide with the lab's own test runner.
@Astro-Han
Astro-Han deleted the claude/lab branch June 17, 2026 12:51
jackwener pushed a commit that referenced this pull request Jun 21, 2026
…r, table (#36)
- runner: stop blanket-approving permission prompts. Allow ordinary tool
use, DENY dangerous categories (fs_destructive / git_destructive /
privileged / browser) by default — the workspace is a copy, not a jail,
so a tool can still escape via absolute paths/network. Opt in with
allowDangerousTools (real sandbox only). Corrects the misleading
"workspace is the safety net" comment. Also: a run that didn't complete
can no longer read as passed.
- evaluator: spawn detached + SIGKILL the process group on timeout, so
backgrounded grandchildren die too (a plain child.kill leaked them).
- sandbox: reject fixture symlinks (fs.cp preserved them verbatim → escape
to source/host) and clean up the temp dir if the copy fails (it was
leaked before the runner's finally registered).
- results: cellKey was separated by a NUL byte, making results.ts a binary
file to git — switch to a JSON key. Render a failed run as ⚠️ (distinct
from ❌ verification-failed) and exclude it from the pass count. Escape
`|`/newline in ids so they can't break the table.
- cli: reject unknown flags and require a value after --out.
- README: honest permission/safety wording.
+5 tests (symlink reject, failed-run render, id escaping, unknown flag,
missing flag value); 20 total green. Cross-process credential lock and a
real container sandbox remain deliberately out of scope.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
The throwaway workspace stops a run from mutating the source fixture, but
it is NOT a security sandbox and a config could rewrite its own test to
pass. This addresses the two open review P1s without ripping out the real
backend (kept usable for your-own-models-on-trusted-tasks):
- Clean-room grading (verification integrity): Task.verification.protectedPaths
lists the grading assets; the runner restores them from the pristine
fixture AFTER the agent finishes and BEFORE the verification command runs,
so a model that rewrote its own test has that edit reverted. Each path is
removed first (drops an agent-planted symlink) and rejected if it escapes
the workspace. examples/fix-add now protects test.mjs.
- Honest docs (host isolation): the README states plainly that a real run
executes tool calls on your machine with your privileges — same exposure
as running Maka directly — so run only models/tasks you trust. Per-run
container isolation (mount workspace only, env allowlist, network policy)
is the named next hardening; allowDangerousTools is for inside it.
Tests (+4, 24 total green): restoreProtectedPaths unit (reverts protected,
keeps the rest, rejects escapes) + a malicious TamperBackend integration
proving protectedPaths reverts the cheat (passed=false) while the unguarded
task lets the cheat win (passed=true) — the guard is load-bearing. The
integration grades via `node check.mjs` (exit-code), not `node --test`, so
the verification child doesn't collide with the lab's own test runner.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
…r, table (#36)
- runner: stop blanket-approving permission prompts. Allow ordinary tool
use, DENY dangerous categories (fs_destructive / git_destructive /
privileged / browser) by default — the workspace is a copy, not a jail,
so a tool can still escape via absolute paths/network. Opt in with
allowDangerousTools (real sandbox only). Corrects the misleading
"workspace is the safety net" comment. Also: a run that didn't complete
can no longer read as passed.
- evaluator: spawn detached + SIGKILL the process group on timeout, so
backgrounded grandchildren die too (a plain child.kill leaked them).
- sandbox: reject fixture symlinks (fs.cp preserved them verbatim → escape
to source/host) and clean up the temp dir if the copy fails (it was
leaked before the runner's finally registered).
- results: cellKey was separated by a NUL byte, making results.ts a binary
file to git — switch to a JSON key. Render a failed run as ⚠️ (distinct
from ❌ verification-failed) and exclude it from the pass count. Escape
`|`/newline in ids so they can't break the table.
- cli: reject unknown flags and require a value after --out.
- README: honest permission/safety wording.
+5 tests (symlink reject, failed-run render, id escaping, unknown flag,
missing flag value); 20 total green. Cross-process credential lock and a
real container sandbox remain deliberately out of scope.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
The throwaway workspace stops a run from mutating the source fixture, but
it is NOT a security sandbox and a config could rewrite its own test to
pass. This addresses the two open review P1s without ripping out the real
backend (kept usable for your-own-models-on-trusted-tasks):
- Clean-room grading (verification integrity): Task.verification.protectedPaths
lists the grading assets; the runner restores them from the pristine
fixture AFTER the agent finishes and BEFORE the verification command runs,
so a model that rewrote its own test has that edit reverted. Each path is
removed first (drops an agent-planted symlink) and rejected if it escapes
the workspace. examples/fix-add now protects test.mjs.
- Honest docs (host isolation): the README states plainly that a real run
executes tool calls on your machine with your privileges — same exposure
as running Maka directly — so run only models/tasks you trust. Per-run
container isolation (mount workspace only, env allowlist, network policy)
is the named next hardening; allowDangerousTools is for inside it.
Tests (+4, 24 total green): restoreProtectedPaths unit (reverts protected,
keeps the rest, rejects escapes) + a malicious TamperBackend integration
proving protectedPaths reverts the cheat (passed=false) while the unguarded
task lets the cheat win (passed=true) — the guard is load-bearing. The
integration grades via `node check.mjs` (exit-code), not `node --test`, so
the verification child doesn't collide with the lab's own test runner.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
…r, table (#36)
- runner: stop blanket-approving permission prompts. Allow ordinary tool
use, DENY dangerous categories (fs_destructive / git_destructive /
privileged / browser) by default — the workspace is a copy, not a jail,
so a tool can still escape via absolute paths/network. Opt in with
allowDangerousTools (real sandbox only). Corrects the misleading
"workspace is the safety net" comment. Also: a run that didn't complete
can no longer read as passed.
- evaluator: spawn detached + SIGKILL the process group on timeout, so
backgrounded grandchildren die too (a plain child.kill leaked them).
- sandbox: reject fixture symlinks (fs.cp preserved them verbatim → escape
to source/host) and clean up the temp dir if the copy fails (it was
leaked before the runner's finally registered).
- results: cellKey was separated by a NUL byte, making results.ts a binary
file to git — switch to a JSON key. Render a failed run as ⚠️ (distinct
from ❌ verification-failed) and exclude it from the pass count. Escape
`|`/newline in ids so they can't break the table.
- cli: reject unknown flags and require a value after --out.
- README: honest permission/safety wording.
+5 tests (symlink reject, failed-run render, id escaping, unknown flag,
missing flag value); 20 total green. Cross-process credential lock and a
real container sandbox remain deliberately out of scope.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
The throwaway workspace stops a run from mutating the source fixture, but
it is NOT a security sandbox and a config could rewrite its own test to
pass. This addresses the two open review P1s without ripping out the real
backend (kept usable for your-own-models-on-trusted-tasks):
- Clean-room grading (verification integrity): Task.verification.protectedPaths
lists the grading assets; the runner restores them from the pristine
fixture AFTER the agent finishes and BEFORE the verification command runs,
so a model that rewrote its own test has that edit reverted. Each path is
removed first (drops an agent-planted symlink) and rejected if it escapes
the workspace. examples/fix-add now protects test.mjs.
- Honest docs (host isolation): the README states plainly that a real run
executes tool calls on your machine with your privileges — same exposure
as running Maka directly — so run only models/tasks you trust. Per-run
container isolation (mount workspace only, env allowlist, network policy)
is the named next hardening; allowDangerousTools is for inside it.
Tests (+4, 24 total green): restoreProtectedPaths unit (reverts protected,
keeps the rest, rejects escapes) + a malicious TamperBackend integration
proving protectedPaths reverts the cheat (passed=false) while the unguarded
task lets the cheat win (passed=true) — the guard is load-bearing. The
integration grades via `node check.mjs` (exit-code), not `node --test`, so
the verification child doesn't collide with the lab's own test runner.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' feat(lab): @maka/lab — headless agent experiment lab (MVP, #31) by Astro-Han · Pull Request #36 · apache/maka · GitHub
Skip to content

feat(lab): @maka/lab — headless agent experiment lab (MVP, #31) - #36

Closed
Astro-Han wants to merge 6 commits into
mainfrom
claude/lab
Closed

feat(lab): @maka/lab — headless agent experiment lab (MVP, #31)#36
Astro-Han wants to merge 6 commits into
mainfrom
claude/lab

Conversation

@Astro-Han

Copy link
Copy Markdown
Contributor

Implements the MVP from #31. Independent of #32 / #33 — the lab runs on the FakeBackend with no credentials, and the real backend reads keys from an env var, not the credential store.

What & why

A headless lab for measuring an agent configuration: run a Config × Task grid in isolated sandboxes, capture each trajectory, score it with the task's own test command, and compare. Vary the model/connection (Config), hold the work (Task), read pass/fail off the table.

Config × Task → throwaway workspace → headless agent run → trajectory
↓
ResultRecord (JSONL) ← verification command

What's in this PR

New package @maka/lab (added to the workspaces list):

  • contractsTask (instruction + fixture workspace + verification command), Config (backend / connection / model), ResultRecord (one canonical row per run). Minimal on purpose; kept in @maka/lab for now, promotable to @maka/core once a second consumer exists.
  • sandbox — copies a task fixture into a throwaway dir; the source is never mutated and runs never bleed into each other. Isolation, not asking, is the safety model.
  • evaluator — runs the task's verification command (exit 0 = pass) in the throwaway workspace after the agent finishes, so a config can't grade itself; output-capped + timeout-killed.
  • runner — drives one headless turn through SessionManager, captures the trajectory via runtimeInvocationObserver, and auto-approves permission prompts (a benchmark has no human to confirm; the throwaway workspace is the boundary).
  • matrix — the full Config × Task cross product; a thrown run becomes a failed cell instead of aborting the grid.
  • results — ResultRecord JSONL is canonical truth; a git-diffable markdown comparison table is derived from it.
  • backends — the two concrete wirings, kept out of the engine: fake (deterministic) and ai-sdk (real model; key from a named env var, minimal AiSdkBackend = model + builtin tools + execute-mode permission).
  • CLImaka-lab run <spec.json> [--out <dir>] and maka-lab compare <results.jsonl>.
  • examplesdemo.spec.json (fake, no key) and fix-add.spec.json (a real coding task).

Live validation

The real ai-sdk backend was smoke-tested end-to-end on DeepSeek (deepseek-chat) with examples/fix-add (a buggy add() + a failing node:test):

  • completed, passed, exit 0, a 124-event trajectory using Edit / Write / Bash
  • the source fixture stayed buggy — the agent only edited the throwaway copy, proving sandbox isolation + headless permission auto-approve work against a real model

Testing

15 deterministic tests (FakeBackend), all green: CLI end-to-end smoke (spawns the built bin → results.jsonl + comparison.mdcompare), matrix cross-product + failed-cell, JSONL round-trip, table rendering, sandbox isolation, evaluator pass/fail/timeout, walking-skeleton e2e. Plus the live DeepSeek smoke above.

Out of scope (deliberate — pure additions later)

Parallel matrix execution, LLM/rule evaluators (today: command exit code), Docker / network isolation, SWE-bench pack ingestion, a richer report than the markdown grid, and promoting the contracts into @maka/core.

#31)
First end-to-end slice of the agent experiment lab: run one Config
against one Task in an isolated sandbox, capture the trajectory, score
it, and return a ResultRecord. De-risks the integration (headless
SessionManager run + sandbox + evaluator composing into one real run);
matrix/compare/CLI and real-model backends are later additions.
- contracts.ts: Task (instruction + fixture workspace + verification
command), Config (backend/connection/model), ResultRecord. Minimal;
systemPrompt/Execution/SWE-bench-pack ingestion deferred.
- sandbox.ts: prepareWorkspace copies the fixture to a throwaway dir so
the agent never mutates the source (isolation, not asking).
- evaluator.ts: runVerification runs the Task's test command in the
throwaway workspace after the agent finishes (config can't grade
itself), with output cap + timeout kill.
- runner.ts: runExperiment wires a pure-Node SessionManager (injected
registerBackends keeps model/credential wiring out of the lab core),
drives one turn, captures the InvocationResult trajectory via
runtimeInvocationObserver, scores it, returns a ResultRecord.
- tests: FakeBackend e2e (pass + fail fixtures, trajectory persisted as
runtime-events.jsonl), sandbox isolation, evaluator pass/fail/timeout.
Independent of the credential migration (#32/#33): the skeleton runs on
FakeBackend and needs no credentials.
Completes the MVP loop on top of the walking skeleton: run a grid of
Configs × Tasks, persist canonical results, and compare.
- matrix.ts: runMatrix runs the full cross product (sequential; a thrown
run becomes a failed cell instead of aborting the grid) with a
per-cell onResult callback.
- results.ts: ResultRecord JSONL is canonical truth; toComparisonTable
derives a git-diffable markdown grid (tasks × configs, ✅/❌/⚠️,
pass-rate footer). ResultRecord gains an optional `error`.
- backends.ts: the two concrete backend wirings, kept out of the engine.
registerFakeBackend (deterministic). registerAiSdkBackend resolves a
Config's slug against spec connections, reads the API key from a named
env var (no secrets at rest), and wires a minimal AiSdkBackend (model +
builtin tools + execute-mode permission). Telemetry/artifact/synthesis
hooks omitted — a benchmark scores via the verification command.
- runner.ts: the drain loop now auto-approves permission requests — a
headless benchmark has no human to confirm; throwaway-workspace
isolation is the safety net.
- cli.ts: `maka-lab run <spec.json> [--out <dir>]` and
`maka-lab compare <results.jsonl>`; task fixtures resolve relative to
the spec. Exposed as the `maka-lab` bin.
- tests (15 total): CLI end-to-end smoke (spawns the built bin on a fake
spec → results.jsonl + comparison.md → compare), matrix cross-product
+ failed-cell, JSONL round-trip, table rendering.
The real ai-sdk backend is typecheck-verified but not unit-tested (a live
model call is non-deterministic and costs money); it needs a live smoke
test with a real API key. Everything else is deterministic and green.
README documents the Config × Task model, the CLI (run/compare), the
spec shape (incl. a real ai-sdk connection with apiKeyEnv), the two
backends, and what's deliberately out of MVP scope. examples/demo.spec.json
+ examples/demo/marker.txt run green on the fake backend with no API key:
maka-lab run examples/demo.spec.json --out /tmp/maka-lab-demo
examples/fix-add — a buggy add() with a failing node:test. A real model
run must read the files, fix src.mjs, and turn `node --test` green.
Live smoke confirmed end-to-end on DeepSeek (deepseek-chat): completed,
passed, exit 0, 124-event trajectory using Edit/Write/Bash; the source
fixture stayed buggy (the agent only touched the throwaway copy), proving
sandbox isolation + headless permission auto-approve work against a real
backend. README points at it as the canonical real-run example.
…r, table (#36)
- runner: stop blanket-approving permission prompts. Allow ordinary tool
use, DENY dangerous categories (fs_destructive / git_destructive /
privileged / browser) by default — the workspace is a copy, not a jail,
so a tool can still escape via absolute paths/network. Opt in with
allowDangerousTools (real sandbox only). Corrects the misleading
"workspace is the safety net" comment. Also: a run that didn't complete
can no longer read as passed.
- evaluator: spawn detached + SIGKILL the process group on timeout, so
backgrounded grandchildren die too (a plain child.kill leaked them).
- sandbox: reject fixture symlinks (fs.cp preserved them verbatim → escape
to source/host) and clean up the temp dir if the copy fails (it was
leaked before the runner's finally registered).
- results: cellKey was separated by a NUL byte, making results.ts a binary
file to git — switch to a JSON key. Render a failed run as ⚠️ (distinct
from ❌ verification-failed) and exclude it from the pass count. Escape
`|`/newline in ids so they can't break the table.
- cli: reject unknown flags and require a value after --out.
- README: honest permission/safety wording.
+5 tests (symlink reject, failed-run render, id escaping, unknown flag,
missing flag value); 20 total green. Cross-process credential lock and a
real container sandbox remain deliberately out of scope.
The throwaway workspace stops a run from mutating the source fixture, but
it is NOT a security sandbox and a config could rewrite its own test to
pass. This addresses the two open review P1s without ripping out the real
backend (kept usable for your-own-models-on-trusted-tasks):
- Clean-room grading (verification integrity): Task.verification.protectedPaths
lists the grading assets; the runner restores them from the pristine
fixture AFTER the agent finishes and BEFORE the verification command runs,
so a model that rewrote its own test has that edit reverted. Each path is
removed first (drops an agent-planted symlink) and rejected if it escapes
the workspace. examples/fix-add now protects test.mjs.
- Honest docs (host isolation): the README states plainly that a real run
executes tool calls on your machine with your privileges — same exposure
as running Maka directly — so run only models/tasks you trust. Per-run
container isolation (mount workspace only, env allowlist, network policy)
is the named next hardening; allowDangerousTools is for inside it.
Tests (+4, 24 total green): restoreProtectedPaths unit (reverts protected,
keeps the rest, rejects escapes) + a malicious TamperBackend integration
proving protectedPaths reverts the cheat (passed=false) while the unguarded
task lets the cheat win (passed=true) — the guard is load-bearing. The
integration grades via `node check.mjs` (exit-code), not `node --test`, so
the verification child doesn't collide with the lab's own test runner.
@Astro-Han
Astro-Han deleted the claude/lab branch June 17, 2026 12:51
jackwener pushed a commit that referenced this pull request Jun 21, 2026
…r, table (#36)
- runner: stop blanket-approving permission prompts. Allow ordinary tool
use, DENY dangerous categories (fs_destructive / git_destructive /
privileged / browser) by default — the workspace is a copy, not a jail,
so a tool can still escape via absolute paths/network. Opt in with
allowDangerousTools (real sandbox only). Corrects the misleading
"workspace is the safety net" comment. Also: a run that didn't complete
can no longer read as passed.
- evaluator: spawn detached + SIGKILL the process group on timeout, so
backgrounded grandchildren die too (a plain child.kill leaked them).
- sandbox: reject fixture symlinks (fs.cp preserved them verbatim → escape
to source/host) and clean up the temp dir if the copy fails (it was
leaked before the runner's finally registered).
- results: cellKey was separated by a NUL byte, making results.ts a binary
file to git — switch to a JSON key. Render a failed run as ⚠️ (distinct
from ❌ verification-failed) and exclude it from the pass count. Escape
`|`/newline in ids so they can't break the table.
- cli: reject unknown flags and require a value after --out.
- README: honest permission/safety wording.
+5 tests (symlink reject, failed-run render, id escaping, unknown flag,
missing flag value); 20 total green. Cross-process credential lock and a
real container sandbox remain deliberately out of scope.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
The throwaway workspace stops a run from mutating the source fixture, but
it is NOT a security sandbox and a config could rewrite its own test to
pass. This addresses the two open review P1s without ripping out the real
backend (kept usable for your-own-models-on-trusted-tasks):
- Clean-room grading (verification integrity): Task.verification.protectedPaths
lists the grading assets; the runner restores them from the pristine
fixture AFTER the agent finishes and BEFORE the verification command runs,
so a model that rewrote its own test has that edit reverted. Each path is
removed first (drops an agent-planted symlink) and rejected if it escapes
the workspace. examples/fix-add now protects test.mjs.
- Honest docs (host isolation): the README states plainly that a real run
executes tool calls on your machine with your privileges — same exposure
as running Maka directly — so run only models/tasks you trust. Per-run
container isolation (mount workspace only, env allowlist, network policy)
is the named next hardening; allowDangerousTools is for inside it.
Tests (+4, 24 total green): restoreProtectedPaths unit (reverts protected,
keeps the rest, rejects escapes) + a malicious TamperBackend integration
proving protectedPaths reverts the cheat (passed=false) while the unguarded
task lets the cheat win (passed=true) — the guard is load-bearing. The
integration grades via `node check.mjs` (exit-code), not `node --test`, so
the verification child doesn't collide with the lab's own test runner.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
…r, table (#36)
- runner: stop blanket-approving permission prompts. Allow ordinary tool
use, DENY dangerous categories (fs_destructive / git_destructive /
privileged / browser) by default — the workspace is a copy, not a jail,
so a tool can still escape via absolute paths/network. Opt in with
allowDangerousTools (real sandbox only). Corrects the misleading
"workspace is the safety net" comment. Also: a run that didn't complete
can no longer read as passed.
- evaluator: spawn detached + SIGKILL the process group on timeout, so
backgrounded grandchildren die too (a plain child.kill leaked them).
- sandbox: reject fixture symlinks (fs.cp preserved them verbatim → escape
to source/host) and clean up the temp dir if the copy fails (it was
leaked before the runner's finally registered).
- results: cellKey was separated by a NUL byte, making results.ts a binary
file to git — switch to a JSON key. Render a failed run as ⚠️ (distinct
from ❌ verification-failed) and exclude it from the pass count. Escape
`|`/newline in ids so they can't break the table.
- cli: reject unknown flags and require a value after --out.
- README: honest permission/safety wording.
+5 tests (symlink reject, failed-run render, id escaping, unknown flag,
missing flag value); 20 total green. Cross-process credential lock and a
real container sandbox remain deliberately out of scope.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
The throwaway workspace stops a run from mutating the source fixture, but
it is NOT a security sandbox and a config could rewrite its own test to
pass. This addresses the two open review P1s without ripping out the real
backend (kept usable for your-own-models-on-trusted-tasks):
- Clean-room grading (verification integrity): Task.verification.protectedPaths
lists the grading assets; the runner restores them from the pristine
fixture AFTER the agent finishes and BEFORE the verification command runs,
so a model that rewrote its own test has that edit reverted. Each path is
removed first (drops an agent-planted symlink) and rejected if it escapes
the workspace. examples/fix-add now protects test.mjs.
- Honest docs (host isolation): the README states plainly that a real run
executes tool calls on your machine with your privileges — same exposure
as running Maka directly — so run only models/tasks you trust. Per-run
container isolation (mount workspace only, env allowlist, network policy)
is the named next hardening; allowDangerousTools is for inside it.
Tests (+4, 24 total green): restoreProtectedPaths unit (reverts protected,
keeps the rest, rejects escapes) + a malicious TamperBackend integration
proving protectedPaths reverts the cheat (passed=false) while the unguarded
task lets the cheat win (passed=true) — the guard is load-bearing. The
integration grades via `node check.mjs` (exit-code), not `node --test`, so
the verification child doesn't collide with the lab's own test runner.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
…r, table (#36)
- runner: stop blanket-approving permission prompts. Allow ordinary tool
use, DENY dangerous categories (fs_destructive / git_destructive /
privileged / browser) by default — the workspace is a copy, not a jail,
so a tool can still escape via absolute paths/network. Opt in with
allowDangerousTools (real sandbox only). Corrects the misleading
"workspace is the safety net" comment. Also: a run that didn't complete
can no longer read as passed.
- evaluator: spawn detached + SIGKILL the process group on timeout, so
backgrounded grandchildren die too (a plain child.kill leaked them).
- sandbox: reject fixture symlinks (fs.cp preserved them verbatim → escape
to source/host) and clean up the temp dir if the copy fails (it was
leaked before the runner's finally registered).
- results: cellKey was separated by a NUL byte, making results.ts a binary
file to git — switch to a JSON key. Render a failed run as ⚠️ (distinct
from ❌ verification-failed) and exclude it from the pass count. Escape
`|`/newline in ids so they can't break the table.
- cli: reject unknown flags and require a value after --out.
- README: honest permission/safety wording.
+5 tests (symlink reject, failed-run render, id escaping, unknown flag,
missing flag value); 20 total green. Cross-process credential lock and a
real container sandbox remain deliberately out of scope.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
The throwaway workspace stops a run from mutating the source fixture, but
it is NOT a security sandbox and a config could rewrite its own test to
pass. This addresses the two open review P1s without ripping out the real
backend (kept usable for your-own-models-on-trusted-tasks):
- Clean-room grading (verification integrity): Task.verification.protectedPaths
lists the grading assets; the runner restores them from the pristine
fixture AFTER the agent finishes and BEFORE the verification command runs,
so a model that rewrote its own test has that edit reverted. Each path is
removed first (drops an agent-planted symlink) and rejected if it escapes
the workspace. examples/fix-add now protects test.mjs.
- Honest docs (host isolation): the README states plainly that a real run
executes tool calls on your machine with your privileges — same exposure
as running Maka directly — so run only models/tasks you trust. Per-run
container isolation (mount workspace only, env allowlist, network policy)
is the named next hardening; allowDangerousTools is for inside it.
Tests (+4, 24 total green): restoreProtectedPaths unit (reverts protected,
keeps the rest, rejects escapes) + a malicious TamperBackend integration
proving protectedPaths reverts the cheat (passed=false) while the unguarded
task lets the cheat win (passed=true) — the guard is load-bearing. The
integration grades via `node check.mjs` (exit-code), not `node --test`, so
the verification child doesn't collide with the lab's own test runner.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' feat(lab): @maka/lab — headless agent experiment lab (MVP, #31) by Astro-Han · Pull Request #36 · apache/maka · GitHub
Skip to content

feat(lab): @maka/lab — headless agent experiment lab (MVP, #31) - #36

Closed
Astro-Han wants to merge 6 commits into
mainfrom
claude/lab
Closed

feat(lab): @maka/lab — headless agent experiment lab (MVP, #31)#36
Astro-Han wants to merge 6 commits into
mainfrom
claude/lab

Conversation

@Astro-Han

Copy link
Copy Markdown
Contributor

Implements the MVP from #31. Independent of #32 / #33 — the lab runs on the FakeBackend with no credentials, and the real backend reads keys from an env var, not the credential store.

What & why

A headless lab for measuring an agent configuration: run a Config × Task grid in isolated sandboxes, capture each trajectory, score it with the task's own test command, and compare. Vary the model/connection (Config), hold the work (Task), read pass/fail off the table.

Config × Task → throwaway workspace → headless agent run → trajectory
↓
ResultRecord (JSONL) ← verification command

What's in this PR

New package @maka/lab (added to the workspaces list):

  • contractsTask (instruction + fixture workspace + verification command), Config (backend / connection / model), ResultRecord (one canonical row per run). Minimal on purpose; kept in @maka/lab for now, promotable to @maka/core once a second consumer exists.
  • sandbox — copies a task fixture into a throwaway dir; the source is never mutated and runs never bleed into each other. Isolation, not asking, is the safety model.
  • evaluator — runs the task's verification command (exit 0 = pass) in the throwaway workspace after the agent finishes, so a config can't grade itself; output-capped + timeout-killed.
  • runner — drives one headless turn through SessionManager, captures the trajectory via runtimeInvocationObserver, and auto-approves permission prompts (a benchmark has no human to confirm; the throwaway workspace is the boundary).
  • matrix — the full Config × Task cross product; a thrown run becomes a failed cell instead of aborting the grid.
  • results — ResultRecord JSONL is canonical truth; a git-diffable markdown comparison table is derived from it.
  • backends — the two concrete wirings, kept out of the engine: fake (deterministic) and ai-sdk (real model; key from a named env var, minimal AiSdkBackend = model + builtin tools + execute-mode permission).
  • CLImaka-lab run <spec.json> [--out <dir>] and maka-lab compare <results.jsonl>.
  • examplesdemo.spec.json (fake, no key) and fix-add.spec.json (a real coding task).

Live validation

The real ai-sdk backend was smoke-tested end-to-end on DeepSeek (deepseek-chat) with examples/fix-add (a buggy add() + a failing node:test):

  • completed, passed, exit 0, a 124-event trajectory using Edit / Write / Bash
  • the source fixture stayed buggy — the agent only edited the throwaway copy, proving sandbox isolation + headless permission auto-approve work against a real model

Testing

15 deterministic tests (FakeBackend), all green: CLI end-to-end smoke (spawns the built bin → results.jsonl + comparison.mdcompare), matrix cross-product + failed-cell, JSONL round-trip, table rendering, sandbox isolation, evaluator pass/fail/timeout, walking-skeleton e2e. Plus the live DeepSeek smoke above.

Out of scope (deliberate — pure additions later)

Parallel matrix execution, LLM/rule evaluators (today: command exit code), Docker / network isolation, SWE-bench pack ingestion, a richer report than the markdown grid, and promoting the contracts into @maka/core.

#31)
First end-to-end slice of the agent experiment lab: run one Config
against one Task in an isolated sandbox, capture the trajectory, score
it, and return a ResultRecord. De-risks the integration (headless
SessionManager run + sandbox + evaluator composing into one real run);
matrix/compare/CLI and real-model backends are later additions.
- contracts.ts: Task (instruction + fixture workspace + verification
command), Config (backend/connection/model), ResultRecord. Minimal;
systemPrompt/Execution/SWE-bench-pack ingestion deferred.
- sandbox.ts: prepareWorkspace copies the fixture to a throwaway dir so
the agent never mutates the source (isolation, not asking).
- evaluator.ts: runVerification runs the Task's test command in the
throwaway workspace after the agent finishes (config can't grade
itself), with output cap + timeout kill.
- runner.ts: runExperiment wires a pure-Node SessionManager (injected
registerBackends keeps model/credential wiring out of the lab core),
drives one turn, captures the InvocationResult trajectory via
runtimeInvocationObserver, scores it, returns a ResultRecord.
- tests: FakeBackend e2e (pass + fail fixtures, trajectory persisted as
runtime-events.jsonl), sandbox isolation, evaluator pass/fail/timeout.
Independent of the credential migration (#32/#33): the skeleton runs on
FakeBackend and needs no credentials.
Completes the MVP loop on top of the walking skeleton: run a grid of
Configs × Tasks, persist canonical results, and compare.
- matrix.ts: runMatrix runs the full cross product (sequential; a thrown
run becomes a failed cell instead of aborting the grid) with a
per-cell onResult callback.
- results.ts: ResultRecord JSONL is canonical truth; toComparisonTable
derives a git-diffable markdown grid (tasks × configs, ✅/❌/⚠️,
pass-rate footer). ResultRecord gains an optional `error`.
- backends.ts: the two concrete backend wirings, kept out of the engine.
registerFakeBackend (deterministic). registerAiSdkBackend resolves a
Config's slug against spec connections, reads the API key from a named
env var (no secrets at rest), and wires a minimal AiSdkBackend (model +
builtin tools + execute-mode permission). Telemetry/artifact/synthesis
hooks omitted — a benchmark scores via the verification command.
- runner.ts: the drain loop now auto-approves permission requests — a
headless benchmark has no human to confirm; throwaway-workspace
isolation is the safety net.
- cli.ts: `maka-lab run <spec.json> [--out <dir>]` and
`maka-lab compare <results.jsonl>`; task fixtures resolve relative to
the spec. Exposed as the `maka-lab` bin.
- tests (15 total): CLI end-to-end smoke (spawns the built bin on a fake
spec → results.jsonl + comparison.md → compare), matrix cross-product
+ failed-cell, JSONL round-trip, table rendering.
The real ai-sdk backend is typecheck-verified but not unit-tested (a live
model call is non-deterministic and costs money); it needs a live smoke
test with a real API key. Everything else is deterministic and green.
README documents the Config × Task model, the CLI (run/compare), the
spec shape (incl. a real ai-sdk connection with apiKeyEnv), the two
backends, and what's deliberately out of MVP scope. examples/demo.spec.json
+ examples/demo/marker.txt run green on the fake backend with no API key:
maka-lab run examples/demo.spec.json --out /tmp/maka-lab-demo
examples/fix-add — a buggy add() with a failing node:test. A real model
run must read the files, fix src.mjs, and turn `node --test` green.
Live smoke confirmed end-to-end on DeepSeek (deepseek-chat): completed,
passed, exit 0, 124-event trajectory using Edit/Write/Bash; the source
fixture stayed buggy (the agent only touched the throwaway copy), proving
sandbox isolation + headless permission auto-approve work against a real
backend. README points at it as the canonical real-run example.
…r, table (#36)
- runner: stop blanket-approving permission prompts. Allow ordinary tool
use, DENY dangerous categories (fs_destructive / git_destructive /
privileged / browser) by default — the workspace is a copy, not a jail,
so a tool can still escape via absolute paths/network. Opt in with
allowDangerousTools (real sandbox only). Corrects the misleading
"workspace is the safety net" comment. Also: a run that didn't complete
can no longer read as passed.
- evaluator: spawn detached + SIGKILL the process group on timeout, so
backgrounded grandchildren die too (a plain child.kill leaked them).
- sandbox: reject fixture symlinks (fs.cp preserved them verbatim → escape
to source/host) and clean up the temp dir if the copy fails (it was
leaked before the runner's finally registered).
- results: cellKey was separated by a NUL byte, making results.ts a binary
file to git — switch to a JSON key. Render a failed run as ⚠️ (distinct
from ❌ verification-failed) and exclude it from the pass count. Escape
`|`/newline in ids so they can't break the table.
- cli: reject unknown flags and require a value after --out.
- README: honest permission/safety wording.
+5 tests (symlink reject, failed-run render, id escaping, unknown flag,
missing flag value); 20 total green. Cross-process credential lock and a
real container sandbox remain deliberately out of scope.
The throwaway workspace stops a run from mutating the source fixture, but
it is NOT a security sandbox and a config could rewrite its own test to
pass. This addresses the two open review P1s without ripping out the real
backend (kept usable for your-own-models-on-trusted-tasks):
- Clean-room grading (verification integrity): Task.verification.protectedPaths
lists the grading assets; the runner restores them from the pristine
fixture AFTER the agent finishes and BEFORE the verification command runs,
so a model that rewrote its own test has that edit reverted. Each path is
removed first (drops an agent-planted symlink) and rejected if it escapes
the workspace. examples/fix-add now protects test.mjs.
- Honest docs (host isolation): the README states plainly that a real run
executes tool calls on your machine with your privileges — same exposure
as running Maka directly — so run only models/tasks you trust. Per-run
container isolation (mount workspace only, env allowlist, network policy)
is the named next hardening; allowDangerousTools is for inside it.
Tests (+4, 24 total green): restoreProtectedPaths unit (reverts protected,
keeps the rest, rejects escapes) + a malicious TamperBackend integration
proving protectedPaths reverts the cheat (passed=false) while the unguarded
task lets the cheat win (passed=true) — the guard is load-bearing. The
integration grades via `node check.mjs` (exit-code), not `node --test`, so
the verification child doesn't collide with the lab's own test runner.
@Astro-Han
Astro-Han deleted the claude/lab branch June 17, 2026 12:51
jackwener pushed a commit that referenced this pull request Jun 21, 2026
…r, table (#36)
- runner: stop blanket-approving permission prompts. Allow ordinary tool
use, DENY dangerous categories (fs_destructive / git_destructive /
privileged / browser) by default — the workspace is a copy, not a jail,
so a tool can still escape via absolute paths/network. Opt in with
allowDangerousTools (real sandbox only). Corrects the misleading
"workspace is the safety net" comment. Also: a run that didn't complete
can no longer read as passed.
- evaluator: spawn detached + SIGKILL the process group on timeout, so
backgrounded grandchildren die too (a plain child.kill leaked them).
- sandbox: reject fixture symlinks (fs.cp preserved them verbatim → escape
to source/host) and clean up the temp dir if the copy fails (it was
leaked before the runner's finally registered).
- results: cellKey was separated by a NUL byte, making results.ts a binary
file to git — switch to a JSON key. Render a failed run as ⚠️ (distinct
from ❌ verification-failed) and exclude it from the pass count. Escape
`|`/newline in ids so they can't break the table.
- cli: reject unknown flags and require a value after --out.
- README: honest permission/safety wording.
+5 tests (symlink reject, failed-run render, id escaping, unknown flag,
missing flag value); 20 total green. Cross-process credential lock and a
real container sandbox remain deliberately out of scope.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
The throwaway workspace stops a run from mutating the source fixture, but
it is NOT a security sandbox and a config could rewrite its own test to
pass. This addresses the two open review P1s without ripping out the real
backend (kept usable for your-own-models-on-trusted-tasks):
- Clean-room grading (verification integrity): Task.verification.protectedPaths
lists the grading assets; the runner restores them from the pristine
fixture AFTER the agent finishes and BEFORE the verification command runs,
so a model that rewrote its own test has that edit reverted. Each path is
removed first (drops an agent-planted symlink) and rejected if it escapes
the workspace. examples/fix-add now protects test.mjs.
- Honest docs (host isolation): the README states plainly that a real run
executes tool calls on your machine with your privileges — same exposure
as running Maka directly — so run only models/tasks you trust. Per-run
container isolation (mount workspace only, env allowlist, network policy)
is the named next hardening; allowDangerousTools is for inside it.
Tests (+4, 24 total green): restoreProtectedPaths unit (reverts protected,
keeps the rest, rejects escapes) + a malicious TamperBackend integration
proving protectedPaths reverts the cheat (passed=false) while the unguarded
task lets the cheat win (passed=true) — the guard is load-bearing. The
integration grades via `node check.mjs` (exit-code), not `node --test`, so
the verification child doesn't collide with the lab's own test runner.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
…r, table (#36)
- runner: stop blanket-approving permission prompts. Allow ordinary tool
use, DENY dangerous categories (fs_destructive / git_destructive /
privileged / browser) by default — the workspace is a copy, not a jail,
so a tool can still escape via absolute paths/network. Opt in with
allowDangerousTools (real sandbox only). Corrects the misleading
"workspace is the safety net" comment. Also: a run that didn't complete
can no longer read as passed.
- evaluator: spawn detached + SIGKILL the process group on timeout, so
backgrounded grandchildren die too (a plain child.kill leaked them).
- sandbox: reject fixture symlinks (fs.cp preserved them verbatim → escape
to source/host) and clean up the temp dir if the copy fails (it was
leaked before the runner's finally registered).
- results: cellKey was separated by a NUL byte, making results.ts a binary
file to git — switch to a JSON key. Render a failed run as ⚠️ (distinct
from ❌ verification-failed) and exclude it from the pass count. Escape
`|`/newline in ids so they can't break the table.
- cli: reject unknown flags and require a value after --out.
- README: honest permission/safety wording.
+5 tests (symlink reject, failed-run render, id escaping, unknown flag,
missing flag value); 20 total green. Cross-process credential lock and a
real container sandbox remain deliberately out of scope.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
The throwaway workspace stops a run from mutating the source fixture, but
it is NOT a security sandbox and a config could rewrite its own test to
pass. This addresses the two open review P1s without ripping out the real
backend (kept usable for your-own-models-on-trusted-tasks):
- Clean-room grading (verification integrity): Task.verification.protectedPaths
lists the grading assets; the runner restores them from the pristine
fixture AFTER the agent finishes and BEFORE the verification command runs,
so a model that rewrote its own test has that edit reverted. Each path is
removed first (drops an agent-planted symlink) and rejected if it escapes
the workspace. examples/fix-add now protects test.mjs.
- Honest docs (host isolation): the README states plainly that a real run
executes tool calls on your machine with your privileges — same exposure
as running Maka directly — so run only models/tasks you trust. Per-run
container isolation (mount workspace only, env allowlist, network policy)
is the named next hardening; allowDangerousTools is for inside it.
Tests (+4, 24 total green): restoreProtectedPaths unit (reverts protected,
keeps the rest, rejects escapes) + a malicious TamperBackend integration
proving protectedPaths reverts the cheat (passed=false) while the unguarded
task lets the cheat win (passed=true) — the guard is load-bearing. The
integration grades via `node check.mjs` (exit-code), not `node --test`, so
the verification child doesn't collide with the lab's own test runner.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
…r, table (#36)
- runner: stop blanket-approving permission prompts. Allow ordinary tool
use, DENY dangerous categories (fs_destructive / git_destructive /
privileged / browser) by default — the workspace is a copy, not a jail,
so a tool can still escape via absolute paths/network. Opt in with
allowDangerousTools (real sandbox only). Corrects the misleading
"workspace is the safety net" comment. Also: a run that didn't complete
can no longer read as passed.
- evaluator: spawn detached + SIGKILL the process group on timeout, so
backgrounded grandchildren die too (a plain child.kill leaked them).
- sandbox: reject fixture symlinks (fs.cp preserved them verbatim → escape
to source/host) and clean up the temp dir if the copy fails (it was
leaked before the runner's finally registered).
- results: cellKey was separated by a NUL byte, making results.ts a binary
file to git — switch to a JSON key. Render a failed run as ⚠️ (distinct
from ❌ verification-failed) and exclude it from the pass count. Escape
`|`/newline in ids so they can't break the table.
- cli: reject unknown flags and require a value after --out.
- README: honest permission/safety wording.
+5 tests (symlink reject, failed-run render, id escaping, unknown flag,
missing flag value); 20 total green. Cross-process credential lock and a
real container sandbox remain deliberately out of scope.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
The throwaway workspace stops a run from mutating the source fixture, but
it is NOT a security sandbox and a config could rewrite its own test to
pass. This addresses the two open review P1s without ripping out the real
backend (kept usable for your-own-models-on-trusted-tasks):
- Clean-room grading (verification integrity): Task.verification.protectedPaths
lists the grading assets; the runner restores them from the pristine
fixture AFTER the agent finishes and BEFORE the verification command runs,
so a model that rewrote its own test has that edit reverted. Each path is
removed first (drops an agent-planted symlink) and rejected if it escapes
the workspace. examples/fix-add now protects test.mjs.
- Honest docs (host isolation): the README states plainly that a real run
executes tool calls on your machine with your privileges — same exposure
as running Maka directly — so run only models/tasks you trust. Per-run
container isolation (mount workspace only, env allowlist, network policy)
is the named next hardening; allowDangerousTools is for inside it.
Tests (+4, 24 total green): restoreProtectedPaths unit (reverts protected,
keeps the rest, rejects escapes) + a malicious TamperBackend integration
proving protectedPaths reverts the cheat (passed=false) while the unguarded
task lets the cheat win (passed=true) — the guard is load-bearing. The
integration grades via `node check.mjs` (exit-code), not `node --test`, so
the verification child doesn't collide with the lab's own test runner.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' feat(lab): @maka/lab — headless agent experiment lab (MVP, #31) by Astro-Han · Pull Request #36 · apache/maka · GitHub
Skip to content

feat(lab): @maka/lab — headless agent experiment lab (MVP, #31) - #36

Closed
Astro-Han wants to merge 6 commits into
mainfrom
claude/lab
Closed

feat(lab): @maka/lab — headless agent experiment lab (MVP, #31)#36
Astro-Han wants to merge 6 commits into
mainfrom
claude/lab

Conversation

@Astro-Han

Copy link
Copy Markdown
Contributor

Implements the MVP from #31. Independent of #32 / #33 — the lab runs on the FakeBackend with no credentials, and the real backend reads keys from an env var, not the credential store.

What & why

A headless lab for measuring an agent configuration: run a Config × Task grid in isolated sandboxes, capture each trajectory, score it with the task's own test command, and compare. Vary the model/connection (Config), hold the work (Task), read pass/fail off the table.

Config × Task → throwaway workspace → headless agent run → trajectory
↓
ResultRecord (JSONL) ← verification command

What's in this PR

New package @maka/lab (added to the workspaces list):

  • contractsTask (instruction + fixture workspace + verification command), Config (backend / connection / model), ResultRecord (one canonical row per run). Minimal on purpose; kept in @maka/lab for now, promotable to @maka/core once a second consumer exists.
  • sandbox — copies a task fixture into a throwaway dir; the source is never mutated and runs never bleed into each other. Isolation, not asking, is the safety model.
  • evaluator — runs the task's verification command (exit 0 = pass) in the throwaway workspace after the agent finishes, so a config can't grade itself; output-capped + timeout-killed.
  • runner — drives one headless turn through SessionManager, captures the trajectory via runtimeInvocationObserver, and auto-approves permission prompts (a benchmark has no human to confirm; the throwaway workspace is the boundary).
  • matrix — the full Config × Task cross product; a thrown run becomes a failed cell instead of aborting the grid.
  • results — ResultRecord JSONL is canonical truth; a git-diffable markdown comparison table is derived from it.
  • backends — the two concrete wirings, kept out of the engine: fake (deterministic) and ai-sdk (real model; key from a named env var, minimal AiSdkBackend = model + builtin tools + execute-mode permission).
  • CLImaka-lab run <spec.json> [--out <dir>] and maka-lab compare <results.jsonl>.
  • examplesdemo.spec.json (fake, no key) and fix-add.spec.json (a real coding task).

Live validation

The real ai-sdk backend was smoke-tested end-to-end on DeepSeek (deepseek-chat) with examples/fix-add (a buggy add() + a failing node:test):

  • completed, passed, exit 0, a 124-event trajectory using Edit / Write / Bash
  • the source fixture stayed buggy — the agent only edited the throwaway copy, proving sandbox isolation + headless permission auto-approve work against a real model

Testing

15 deterministic tests (FakeBackend), all green: CLI end-to-end smoke (spawns the built bin → results.jsonl + comparison.mdcompare), matrix cross-product + failed-cell, JSONL round-trip, table rendering, sandbox isolation, evaluator pass/fail/timeout, walking-skeleton e2e. Plus the live DeepSeek smoke above.

Out of scope (deliberate — pure additions later)

Parallel matrix execution, LLM/rule evaluators (today: command exit code), Docker / network isolation, SWE-bench pack ingestion, a richer report than the markdown grid, and promoting the contracts into @maka/core.

#31)
First end-to-end slice of the agent experiment lab: run one Config
against one Task in an isolated sandbox, capture the trajectory, score
it, and return a ResultRecord. De-risks the integration (headless
SessionManager run + sandbox + evaluator composing into one real run);
matrix/compare/CLI and real-model backends are later additions.
- contracts.ts: Task (instruction + fixture workspace + verification
command), Config (backend/connection/model), ResultRecord. Minimal;
systemPrompt/Execution/SWE-bench-pack ingestion deferred.
- sandbox.ts: prepareWorkspace copies the fixture to a throwaway dir so
the agent never mutates the source (isolation, not asking).
- evaluator.ts: runVerification runs the Task's test command in the
throwaway workspace after the agent finishes (config can't grade
itself), with output cap + timeout kill.
- runner.ts: runExperiment wires a pure-Node SessionManager (injected
registerBackends keeps model/credential wiring out of the lab core),
drives one turn, captures the InvocationResult trajectory via
runtimeInvocationObserver, scores it, returns a ResultRecord.
- tests: FakeBackend e2e (pass + fail fixtures, trajectory persisted as
runtime-events.jsonl), sandbox isolation, evaluator pass/fail/timeout.
Independent of the credential migration (#32/#33): the skeleton runs on
FakeBackend and needs no credentials.
Completes the MVP loop on top of the walking skeleton: run a grid of
Configs × Tasks, persist canonical results, and compare.
- matrix.ts: runMatrix runs the full cross product (sequential; a thrown
run becomes a failed cell instead of aborting the grid) with a
per-cell onResult callback.
- results.ts: ResultRecord JSONL is canonical truth; toComparisonTable
derives a git-diffable markdown grid (tasks × configs, ✅/❌/⚠️,
pass-rate footer). ResultRecord gains an optional `error`.
- backends.ts: the two concrete backend wirings, kept out of the engine.
registerFakeBackend (deterministic). registerAiSdkBackend resolves a
Config's slug against spec connections, reads the API key from a named
env var (no secrets at rest), and wires a minimal AiSdkBackend (model +
builtin tools + execute-mode permission). Telemetry/artifact/synthesis
hooks omitted — a benchmark scores via the verification command.
- runner.ts: the drain loop now auto-approves permission requests — a
headless benchmark has no human to confirm; throwaway-workspace
isolation is the safety net.
- cli.ts: `maka-lab run <spec.json> [--out <dir>]` and
`maka-lab compare <results.jsonl>`; task fixtures resolve relative to
the spec. Exposed as the `maka-lab` bin.
- tests (15 total): CLI end-to-end smoke (spawns the built bin on a fake
spec → results.jsonl + comparison.md → compare), matrix cross-product
+ failed-cell, JSONL round-trip, table rendering.
The real ai-sdk backend is typecheck-verified but not unit-tested (a live
model call is non-deterministic and costs money); it needs a live smoke
test with a real API key. Everything else is deterministic and green.
README documents the Config × Task model, the CLI (run/compare), the
spec shape (incl. a real ai-sdk connection with apiKeyEnv), the two
backends, and what's deliberately out of MVP scope. examples/demo.spec.json
+ examples/demo/marker.txt run green on the fake backend with no API key:
maka-lab run examples/demo.spec.json --out /tmp/maka-lab-demo
examples/fix-add — a buggy add() with a failing node:test. A real model
run must read the files, fix src.mjs, and turn `node --test` green.
Live smoke confirmed end-to-end on DeepSeek (deepseek-chat): completed,
passed, exit 0, 124-event trajectory using Edit/Write/Bash; the source
fixture stayed buggy (the agent only touched the throwaway copy), proving
sandbox isolation + headless permission auto-approve work against a real
backend. README points at it as the canonical real-run example.
…r, table (#36)
- runner: stop blanket-approving permission prompts. Allow ordinary tool
use, DENY dangerous categories (fs_destructive / git_destructive /
privileged / browser) by default — the workspace is a copy, not a jail,
so a tool can still escape via absolute paths/network. Opt in with
allowDangerousTools (real sandbox only). Corrects the misleading
"workspace is the safety net" comment. Also: a run that didn't complete
can no longer read as passed.
- evaluator: spawn detached + SIGKILL the process group on timeout, so
backgrounded grandchildren die too (a plain child.kill leaked them).
- sandbox: reject fixture symlinks (fs.cp preserved them verbatim → escape
to source/host) and clean up the temp dir if the copy fails (it was
leaked before the runner's finally registered).
- results: cellKey was separated by a NUL byte, making results.ts a binary
file to git — switch to a JSON key. Render a failed run as ⚠️ (distinct
from ❌ verification-failed) and exclude it from the pass count. Escape
`|`/newline in ids so they can't break the table.
- cli: reject unknown flags and require a value after --out.
- README: honest permission/safety wording.
+5 tests (symlink reject, failed-run render, id escaping, unknown flag,
missing flag value); 20 total green. Cross-process credential lock and a
real container sandbox remain deliberately out of scope.
The throwaway workspace stops a run from mutating the source fixture, but
it is NOT a security sandbox and a config could rewrite its own test to
pass. This addresses the two open review P1s without ripping out the real
backend (kept usable for your-own-models-on-trusted-tasks):
- Clean-room grading (verification integrity): Task.verification.protectedPaths
lists the grading assets; the runner restores them from the pristine
fixture AFTER the agent finishes and BEFORE the verification command runs,
so a model that rewrote its own test has that edit reverted. Each path is
removed first (drops an agent-planted symlink) and rejected if it escapes
the workspace. examples/fix-add now protects test.mjs.
- Honest docs (host isolation): the README states plainly that a real run
executes tool calls on your machine with your privileges — same exposure
as running Maka directly — so run only models/tasks you trust. Per-run
container isolation (mount workspace only, env allowlist, network policy)
is the named next hardening; allowDangerousTools is for inside it.
Tests (+4, 24 total green): restoreProtectedPaths unit (reverts protected,
keeps the rest, rejects escapes) + a malicious TamperBackend integration
proving protectedPaths reverts the cheat (passed=false) while the unguarded
task lets the cheat win (passed=true) — the guard is load-bearing. The
integration grades via `node check.mjs` (exit-code), not `node --test`, so
the verification child doesn't collide with the lab's own test runner.
@Astro-Han
Astro-Han deleted the claude/lab branch June 17, 2026 12:51
jackwener pushed a commit that referenced this pull request Jun 21, 2026
…r, table (#36)
- runner: stop blanket-approving permission prompts. Allow ordinary tool
use, DENY dangerous categories (fs_destructive / git_destructive /
privileged / browser) by default — the workspace is a copy, not a jail,
so a tool can still escape via absolute paths/network. Opt in with
allowDangerousTools (real sandbox only). Corrects the misleading
"workspace is the safety net" comment. Also: a run that didn't complete
can no longer read as passed.
- evaluator: spawn detached + SIGKILL the process group on timeout, so
backgrounded grandchildren die too (a plain child.kill leaked them).
- sandbox: reject fixture symlinks (fs.cp preserved them verbatim → escape
to source/host) and clean up the temp dir if the copy fails (it was
leaked before the runner's finally registered).
- results: cellKey was separated by a NUL byte, making results.ts a binary
file to git — switch to a JSON key. Render a failed run as ⚠️ (distinct
from ❌ verification-failed) and exclude it from the pass count. Escape
`|`/newline in ids so they can't break the table.
- cli: reject unknown flags and require a value after --out.
- README: honest permission/safety wording.
+5 tests (symlink reject, failed-run render, id escaping, unknown flag,
missing flag value); 20 total green. Cross-process credential lock and a
real container sandbox remain deliberately out of scope.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
The throwaway workspace stops a run from mutating the source fixture, but
it is NOT a security sandbox and a config could rewrite its own test to
pass. This addresses the two open review P1s without ripping out the real
backend (kept usable for your-own-models-on-trusted-tasks):
- Clean-room grading (verification integrity): Task.verification.protectedPaths
lists the grading assets; the runner restores them from the pristine
fixture AFTER the agent finishes and BEFORE the verification command runs,
so a model that rewrote its own test has that edit reverted. Each path is
removed first (drops an agent-planted symlink) and rejected if it escapes
the workspace. examples/fix-add now protects test.mjs.
- Honest docs (host isolation): the README states plainly that a real run
executes tool calls on your machine with your privileges — same exposure
as running Maka directly — so run only models/tasks you trust. Per-run
container isolation (mount workspace only, env allowlist, network policy)
is the named next hardening; allowDangerousTools is for inside it.
Tests (+4, 24 total green): restoreProtectedPaths unit (reverts protected,
keeps the rest, rejects escapes) + a malicious TamperBackend integration
proving protectedPaths reverts the cheat (passed=false) while the unguarded
task lets the cheat win (passed=true) — the guard is load-bearing. The
integration grades via `node check.mjs` (exit-code), not `node --test`, so
the verification child doesn't collide with the lab's own test runner.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
…r, table (#36)
- runner: stop blanket-approving permission prompts. Allow ordinary tool
use, DENY dangerous categories (fs_destructive / git_destructive /
privileged / browser) by default — the workspace is a copy, not a jail,
so a tool can still escape via absolute paths/network. Opt in with
allowDangerousTools (real sandbox only). Corrects the misleading
"workspace is the safety net" comment. Also: a run that didn't complete
can no longer read as passed.
- evaluator: spawn detached + SIGKILL the process group on timeout, so
backgrounded grandchildren die too (a plain child.kill leaked them).
- sandbox: reject fixture symlinks (fs.cp preserved them verbatim → escape
to source/host) and clean up the temp dir if the copy fails (it was
leaked before the runner's finally registered).
- results: cellKey was separated by a NUL byte, making results.ts a binary
file to git — switch to a JSON key. Render a failed run as ⚠️ (distinct
from ❌ verification-failed) and exclude it from the pass count. Escape
`|`/newline in ids so they can't break the table.
- cli: reject unknown flags and require a value after --out.
- README: honest permission/safety wording.
+5 tests (symlink reject, failed-run render, id escaping, unknown flag,
missing flag value); 20 total green. Cross-process credential lock and a
real container sandbox remain deliberately out of scope.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
The throwaway workspace stops a run from mutating the source fixture, but
it is NOT a security sandbox and a config could rewrite its own test to
pass. This addresses the two open review P1s without ripping out the real
backend (kept usable for your-own-models-on-trusted-tasks):
- Clean-room grading (verification integrity): Task.verification.protectedPaths
lists the grading assets; the runner restores them from the pristine
fixture AFTER the agent finishes and BEFORE the verification command runs,
so a model that rewrote its own test has that edit reverted. Each path is
removed first (drops an agent-planted symlink) and rejected if it escapes
the workspace. examples/fix-add now protects test.mjs.
- Honest docs (host isolation): the README states plainly that a real run
executes tool calls on your machine with your privileges — same exposure
as running Maka directly — so run only models/tasks you trust. Per-run
container isolation (mount workspace only, env allowlist, network policy)
is the named next hardening; allowDangerousTools is for inside it.
Tests (+4, 24 total green): restoreProtectedPaths unit (reverts protected,
keeps the rest, rejects escapes) + a malicious TamperBackend integration
proving protectedPaths reverts the cheat (passed=false) while the unguarded
task lets the cheat win (passed=true) — the guard is load-bearing. The
integration grades via `node check.mjs` (exit-code), not `node --test`, so
the verification child doesn't collide with the lab's own test runner.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
…r, table (#36)
- runner: stop blanket-approving permission prompts. Allow ordinary tool
use, DENY dangerous categories (fs_destructive / git_destructive /
privileged / browser) by default — the workspace is a copy, not a jail,
so a tool can still escape via absolute paths/network. Opt in with
allowDangerousTools (real sandbox only). Corrects the misleading
"workspace is the safety net" comment. Also: a run that didn't complete
can no longer read as passed.
- evaluator: spawn detached + SIGKILL the process group on timeout, so
backgrounded grandchildren die too (a plain child.kill leaked them).
- sandbox: reject fixture symlinks (fs.cp preserved them verbatim → escape
to source/host) and clean up the temp dir if the copy fails (it was
leaked before the runner's finally registered).
- results: cellKey was separated by a NUL byte, making results.ts a binary
file to git — switch to a JSON key. Render a failed run as ⚠️ (distinct
from ❌ verification-failed) and exclude it from the pass count. Escape
`|`/newline in ids so they can't break the table.
- cli: reject unknown flags and require a value after --out.
- README: honest permission/safety wording.
+5 tests (symlink reject, failed-run render, id escaping, unknown flag,
missing flag value); 20 total green. Cross-process credential lock and a
real container sandbox remain deliberately out of scope.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
The throwaway workspace stops a run from mutating the source fixture, but
it is NOT a security sandbox and a config could rewrite its own test to
pass. This addresses the two open review P1s without ripping out the real
backend (kept usable for your-own-models-on-trusted-tasks):
- Clean-room grading (verification integrity): Task.verification.protectedPaths
lists the grading assets; the runner restores them from the pristine
fixture AFTER the agent finishes and BEFORE the verification command runs,
so a model that rewrote its own test has that edit reverted. Each path is
removed first (drops an agent-planted symlink) and rejected if it escapes
the workspace. examples/fix-add now protects test.mjs.
- Honest docs (host isolation): the README states plainly that a real run
executes tool calls on your machine with your privileges — same exposure
as running Maka directly — so run only models/tasks you trust. Per-run
container isolation (mount workspace only, env allowlist, network policy)
is the named next hardening; allowDangerousTools is for inside it.
Tests (+4, 24 total green): restoreProtectedPaths unit (reverts protected,
keeps the rest, rejects escapes) + a malicious TamperBackend integration
proving protectedPaths reverts the cheat (passed=false) while the unguarded
task lets the cheat win (passed=true) — the guard is load-bearing. The
integration grades via `node check.mjs` (exit-code), not `node --test`, so
the verification child doesn't collide with the lab's own test runner.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' feat(lab): @maka/lab — headless agent experiment lab (MVP, #31) by Astro-Han · Pull Request #36 · apache/maka · GitHub
Skip to content

feat(lab): @maka/lab — headless agent experiment lab (MVP, #31) - #36

Closed
Astro-Han wants to merge 6 commits into
mainfrom
claude/lab
Closed

feat(lab): @maka/lab — headless agent experiment lab (MVP, #31)#36
Astro-Han wants to merge 6 commits into
mainfrom
claude/lab

Conversation

@Astro-Han

Copy link
Copy Markdown
Contributor

Implements the MVP from #31. Independent of #32 / #33 — the lab runs on the FakeBackend with no credentials, and the real backend reads keys from an env var, not the credential store.

What & why

A headless lab for measuring an agent configuration: run a Config × Task grid in isolated sandboxes, capture each trajectory, score it with the task's own test command, and compare. Vary the model/connection (Config), hold the work (Task), read pass/fail off the table.

Config × Task → throwaway workspace → headless agent run → trajectory
↓
ResultRecord (JSONL) ← verification command

What's in this PR

New package @maka/lab (added to the workspaces list):

  • contractsTask (instruction + fixture workspace + verification command), Config (backend / connection / model), ResultRecord (one canonical row per run). Minimal on purpose; kept in @maka/lab for now, promotable to @maka/core once a second consumer exists.
  • sandbox — copies a task fixture into a throwaway dir; the source is never mutated and runs never bleed into each other. Isolation, not asking, is the safety model.
  • evaluator — runs the task's verification command (exit 0 = pass) in the throwaway workspace after the agent finishes, so a config can't grade itself; output-capped + timeout-killed.
  • runner — drives one headless turn through SessionManager, captures the trajectory via runtimeInvocationObserver, and auto-approves permission prompts (a benchmark has no human to confirm; the throwaway workspace is the boundary).
  • matrix — the full Config × Task cross product; a thrown run becomes a failed cell instead of aborting the grid.
  • results — ResultRecord JSONL is canonical truth; a git-diffable markdown comparison table is derived from it.
  • backends — the two concrete wirings, kept out of the engine: fake (deterministic) and ai-sdk (real model; key from a named env var, minimal AiSdkBackend = model + builtin tools + execute-mode permission).
  • CLImaka-lab run <spec.json> [--out <dir>] and maka-lab compare <results.jsonl>.
  • examplesdemo.spec.json (fake, no key) and fix-add.spec.json (a real coding task).

Live validation

The real ai-sdk backend was smoke-tested end-to-end on DeepSeek (deepseek-chat) with examples/fix-add (a buggy add() + a failing node:test):

  • completed, passed, exit 0, a 124-event trajectory using Edit / Write / Bash
  • the source fixture stayed buggy — the agent only edited the throwaway copy, proving sandbox isolation + headless permission auto-approve work against a real model

Testing

15 deterministic tests (FakeBackend), all green: CLI end-to-end smoke (spawns the built bin → results.jsonl + comparison.mdcompare), matrix cross-product + failed-cell, JSONL round-trip, table rendering, sandbox isolation, evaluator pass/fail/timeout, walking-skeleton e2e. Plus the live DeepSeek smoke above.

Out of scope (deliberate — pure additions later)

Parallel matrix execution, LLM/rule evaluators (today: command exit code), Docker / network isolation, SWE-bench pack ingestion, a richer report than the markdown grid, and promoting the contracts into @maka/core.

#31)
First end-to-end slice of the agent experiment lab: run one Config
against one Task in an isolated sandbox, capture the trajectory, score
it, and return a ResultRecord. De-risks the integration (headless
SessionManager run + sandbox + evaluator composing into one real run);
matrix/compare/CLI and real-model backends are later additions.
- contracts.ts: Task (instruction + fixture workspace + verification
command), Config (backend/connection/model), ResultRecord. Minimal;
systemPrompt/Execution/SWE-bench-pack ingestion deferred.
- sandbox.ts: prepareWorkspace copies the fixture to a throwaway dir so
the agent never mutates the source (isolation, not asking).
- evaluator.ts: runVerification runs the Task's test command in the
throwaway workspace after the agent finishes (config can't grade
itself), with output cap + timeout kill.
- runner.ts: runExperiment wires a pure-Node SessionManager (injected
registerBackends keeps model/credential wiring out of the lab core),
drives one turn, captures the InvocationResult trajectory via
runtimeInvocationObserver, scores it, returns a ResultRecord.
- tests: FakeBackend e2e (pass + fail fixtures, trajectory persisted as
runtime-events.jsonl), sandbox isolation, evaluator pass/fail/timeout.
Independent of the credential migration (#32/#33): the skeleton runs on
FakeBackend and needs no credentials.
Completes the MVP loop on top of the walking skeleton: run a grid of
Configs × Tasks, persist canonical results, and compare.
- matrix.ts: runMatrix runs the full cross product (sequential; a thrown
run becomes a failed cell instead of aborting the grid) with a
per-cell onResult callback.
- results.ts: ResultRecord JSONL is canonical truth; toComparisonTable
derives a git-diffable markdown grid (tasks × configs, ✅/❌/⚠️,
pass-rate footer). ResultRecord gains an optional `error`.
- backends.ts: the two concrete backend wirings, kept out of the engine.
registerFakeBackend (deterministic). registerAiSdkBackend resolves a
Config's slug against spec connections, reads the API key from a named
env var (no secrets at rest), and wires a minimal AiSdkBackend (model +
builtin tools + execute-mode permission). Telemetry/artifact/synthesis
hooks omitted — a benchmark scores via the verification command.
- runner.ts: the drain loop now auto-approves permission requests — a
headless benchmark has no human to confirm; throwaway-workspace
isolation is the safety net.
- cli.ts: `maka-lab run <spec.json> [--out <dir>]` and
`maka-lab compare <results.jsonl>`; task fixtures resolve relative to
the spec. Exposed as the `maka-lab` bin.
- tests (15 total): CLI end-to-end smoke (spawns the built bin on a fake
spec → results.jsonl + comparison.md → compare), matrix cross-product
+ failed-cell, JSONL round-trip, table rendering.
The real ai-sdk backend is typecheck-verified but not unit-tested (a live
model call is non-deterministic and costs money); it needs a live smoke
test with a real API key. Everything else is deterministic and green.
README documents the Config × Task model, the CLI (run/compare), the
spec shape (incl. a real ai-sdk connection with apiKeyEnv), the two
backends, and what's deliberately out of MVP scope. examples/demo.spec.json
+ examples/demo/marker.txt run green on the fake backend with no API key:
maka-lab run examples/demo.spec.json --out /tmp/maka-lab-demo
examples/fix-add — a buggy add() with a failing node:test. A real model
run must read the files, fix src.mjs, and turn `node --test` green.
Live smoke confirmed end-to-end on DeepSeek (deepseek-chat): completed,
passed, exit 0, 124-event trajectory using Edit/Write/Bash; the source
fixture stayed buggy (the agent only touched the throwaway copy), proving
sandbox isolation + headless permission auto-approve work against a real
backend. README points at it as the canonical real-run example.
…r, table (#36)
- runner: stop blanket-approving permission prompts. Allow ordinary tool
use, DENY dangerous categories (fs_destructive / git_destructive /
privileged / browser) by default — the workspace is a copy, not a jail,
so a tool can still escape via absolute paths/network. Opt in with
allowDangerousTools (real sandbox only). Corrects the misleading
"workspace is the safety net" comment. Also: a run that didn't complete
can no longer read as passed.
- evaluator: spawn detached + SIGKILL the process group on timeout, so
backgrounded grandchildren die too (a plain child.kill leaked them).
- sandbox: reject fixture symlinks (fs.cp preserved them verbatim → escape
to source/host) and clean up the temp dir if the copy fails (it was
leaked before the runner's finally registered).
- results: cellKey was separated by a NUL byte, making results.ts a binary
file to git — switch to a JSON key. Render a failed run as ⚠️ (distinct
from ❌ verification-failed) and exclude it from the pass count. Escape
`|`/newline in ids so they can't break the table.
- cli: reject unknown flags and require a value after --out.
- README: honest permission/safety wording.
+5 tests (symlink reject, failed-run render, id escaping, unknown flag,
missing flag value); 20 total green. Cross-process credential lock and a
real container sandbox remain deliberately out of scope.
The throwaway workspace stops a run from mutating the source fixture, but
it is NOT a security sandbox and a config could rewrite its own test to
pass. This addresses the two open review P1s without ripping out the real
backend (kept usable for your-own-models-on-trusted-tasks):
- Clean-room grading (verification integrity): Task.verification.protectedPaths
lists the grading assets; the runner restores them from the pristine
fixture AFTER the agent finishes and BEFORE the verification command runs,
so a model that rewrote its own test has that edit reverted. Each path is
removed first (drops an agent-planted symlink) and rejected if it escapes
the workspace. examples/fix-add now protects test.mjs.
- Honest docs (host isolation): the README states plainly that a real run
executes tool calls on your machine with your privileges — same exposure
as running Maka directly — so run only models/tasks you trust. Per-run
container isolation (mount workspace only, env allowlist, network policy)
is the named next hardening; allowDangerousTools is for inside it.
Tests (+4, 24 total green): restoreProtectedPaths unit (reverts protected,
keeps the rest, rejects escapes) + a malicious TamperBackend integration
proving protectedPaths reverts the cheat (passed=false) while the unguarded
task lets the cheat win (passed=true) — the guard is load-bearing. The
integration grades via `node check.mjs` (exit-code), not `node --test`, so
the verification child doesn't collide with the lab's own test runner.
@Astro-Han
Astro-Han deleted the claude/lab branch June 17, 2026 12:51
jackwener pushed a commit that referenced this pull request Jun 21, 2026
…r, table (#36)
- runner: stop blanket-approving permission prompts. Allow ordinary tool
use, DENY dangerous categories (fs_destructive / git_destructive /
privileged / browser) by default — the workspace is a copy, not a jail,
so a tool can still escape via absolute paths/network. Opt in with
allowDangerousTools (real sandbox only). Corrects the misleading
"workspace is the safety net" comment. Also: a run that didn't complete
can no longer read as passed.
- evaluator: spawn detached + SIGKILL the process group on timeout, so
backgrounded grandchildren die too (a plain child.kill leaked them).
- sandbox: reject fixture symlinks (fs.cp preserved them verbatim → escape
to source/host) and clean up the temp dir if the copy fails (it was
leaked before the runner's finally registered).
- results: cellKey was separated by a NUL byte, making results.ts a binary
file to git — switch to a JSON key. Render a failed run as ⚠️ (distinct
from ❌ verification-failed) and exclude it from the pass count. Escape
`|`/newline in ids so they can't break the table.
- cli: reject unknown flags and require a value after --out.
- README: honest permission/safety wording.
+5 tests (symlink reject, failed-run render, id escaping, unknown flag,
missing flag value); 20 total green. Cross-process credential lock and a
real container sandbox remain deliberately out of scope.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
The throwaway workspace stops a run from mutating the source fixture, but
it is NOT a security sandbox and a config could rewrite its own test to
pass. This addresses the two open review P1s without ripping out the real
backend (kept usable for your-own-models-on-trusted-tasks):
- Clean-room grading (verification integrity): Task.verification.protectedPaths
lists the grading assets; the runner restores them from the pristine
fixture AFTER the agent finishes and BEFORE the verification command runs,
so a model that rewrote its own test has that edit reverted. Each path is
removed first (drops an agent-planted symlink) and rejected if it escapes
the workspace. examples/fix-add now protects test.mjs.
- Honest docs (host isolation): the README states plainly that a real run
executes tool calls on your machine with your privileges — same exposure
as running Maka directly — so run only models/tasks you trust. Per-run
container isolation (mount workspace only, env allowlist, network policy)
is the named next hardening; allowDangerousTools is for inside it.
Tests (+4, 24 total green): restoreProtectedPaths unit (reverts protected,
keeps the rest, rejects escapes) + a malicious TamperBackend integration
proving protectedPaths reverts the cheat (passed=false) while the unguarded
task lets the cheat win (passed=true) — the guard is load-bearing. The
integration grades via `node check.mjs` (exit-code), not `node --test`, so
the verification child doesn't collide with the lab's own test runner.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
…r, table (#36)
- runner: stop blanket-approving permission prompts. Allow ordinary tool
use, DENY dangerous categories (fs_destructive / git_destructive /
privileged / browser) by default — the workspace is a copy, not a jail,
so a tool can still escape via absolute paths/network. Opt in with
allowDangerousTools (real sandbox only). Corrects the misleading
"workspace is the safety net" comment. Also: a run that didn't complete
can no longer read as passed.
- evaluator: spawn detached + SIGKILL the process group on timeout, so
backgrounded grandchildren die too (a plain child.kill leaked them).
- sandbox: reject fixture symlinks (fs.cp preserved them verbatim → escape
to source/host) and clean up the temp dir if the copy fails (it was
leaked before the runner's finally registered).
- results: cellKey was separated by a NUL byte, making results.ts a binary
file to git — switch to a JSON key. Render a failed run as ⚠️ (distinct
from ❌ verification-failed) and exclude it from the pass count. Escape
`|`/newline in ids so they can't break the table.
- cli: reject unknown flags and require a value after --out.
- README: honest permission/safety wording.
+5 tests (symlink reject, failed-run render, id escaping, unknown flag,
missing flag value); 20 total green. Cross-process credential lock and a
real container sandbox remain deliberately out of scope.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
The throwaway workspace stops a run from mutating the source fixture, but
it is NOT a security sandbox and a config could rewrite its own test to
pass. This addresses the two open review P1s without ripping out the real
backend (kept usable for your-own-models-on-trusted-tasks):
- Clean-room grading (verification integrity): Task.verification.protectedPaths
lists the grading assets; the runner restores them from the pristine
fixture AFTER the agent finishes and BEFORE the verification command runs,
so a model that rewrote its own test has that edit reverted. Each path is
removed first (drops an agent-planted symlink) and rejected if it escapes
the workspace. examples/fix-add now protects test.mjs.
- Honest docs (host isolation): the README states plainly that a real run
executes tool calls on your machine with your privileges — same exposure
as running Maka directly — so run only models/tasks you trust. Per-run
container isolation (mount workspace only, env allowlist, network policy)
is the named next hardening; allowDangerousTools is for inside it.
Tests (+4, 24 total green): restoreProtectedPaths unit (reverts protected,
keeps the rest, rejects escapes) + a malicious TamperBackend integration
proving protectedPaths reverts the cheat (passed=false) while the unguarded
task lets the cheat win (passed=true) — the guard is load-bearing. The
integration grades via `node check.mjs` (exit-code), not `node --test`, so
the verification child doesn't collide with the lab's own test runner.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
…r, table (#36)
- runner: stop blanket-approving permission prompts. Allow ordinary tool
use, DENY dangerous categories (fs_destructive / git_destructive /
privileged / browser) by default — the workspace is a copy, not a jail,
so a tool can still escape via absolute paths/network. Opt in with
allowDangerousTools (real sandbox only). Corrects the misleading
"workspace is the safety net" comment. Also: a run that didn't complete
can no longer read as passed.
- evaluator: spawn detached + SIGKILL the process group on timeout, so
backgrounded grandchildren die too (a plain child.kill leaked them).
- sandbox: reject fixture symlinks (fs.cp preserved them verbatim → escape
to source/host) and clean up the temp dir if the copy fails (it was
leaked before the runner's finally registered).
- results: cellKey was separated by a NUL byte, making results.ts a binary
file to git — switch to a JSON key. Render a failed run as ⚠️ (distinct
from ❌ verification-failed) and exclude it from the pass count. Escape
`|`/newline in ids so they can't break the table.
- cli: reject unknown flags and require a value after --out.
- README: honest permission/safety wording.
+5 tests (symlink reject, failed-run render, id escaping, unknown flag,
missing flag value); 20 total green. Cross-process credential lock and a
real container sandbox remain deliberately out of scope.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
The throwaway workspace stops a run from mutating the source fixture, but
it is NOT a security sandbox and a config could rewrite its own test to
pass. This addresses the two open review P1s without ripping out the real
backend (kept usable for your-own-models-on-trusted-tasks):
- Clean-room grading (verification integrity): Task.verification.protectedPaths
lists the grading assets; the runner restores them from the pristine
fixture AFTER the agent finishes and BEFORE the verification command runs,
so a model that rewrote its own test has that edit reverted. Each path is
removed first (drops an agent-planted symlink) and rejected if it escapes
the workspace. examples/fix-add now protects test.mjs.
- Honest docs (host isolation): the README states plainly that a real run
executes tool calls on your machine with your privileges — same exposure
as running Maka directly — so run only models/tasks you trust. Per-run
container isolation (mount workspace only, env allowlist, network policy)
is the named next hardening; allowDangerousTools is for inside it.
Tests (+4, 24 total green): restoreProtectedPaths unit (reverts protected,
keeps the rest, rejects escapes) + a malicious TamperBackend integration
proving protectedPaths reverts the cheat (passed=false) while the unguarded
task lets the cheat win (passed=true) — the guard is load-bearing. The
integration grades via `node check.mjs` (exit-code), not `node --test`, so
the verification child doesn't collide with the lab's own test runner.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); feat(lab): @maka/lab — headless agent experiment lab (MVP, #31) by Astro-Han · Pull Request #36 · apache/maka · GitHub
Skip to content

feat(lab): @maka/lab — headless agent experiment lab (MVP, #31) - #36

Closed
Astro-Han wants to merge 6 commits into
mainfrom
claude/lab
Closed

feat(lab): @maka/lab — headless agent experiment lab (MVP, #31)#36
Astro-Han wants to merge 6 commits into
mainfrom
claude/lab

Conversation

@Astro-Han

Copy link
Copy Markdown
Contributor

Implements the MVP from #31. Independent of #32 / #33 — the lab runs on the FakeBackend with no credentials, and the real backend reads keys from an env var, not the credential store.

What & why

A headless lab for measuring an agent configuration: run a Config × Task grid in isolated sandboxes, capture each trajectory, score it with the task's own test command, and compare. Vary the model/connection (Config), hold the work (Task), read pass/fail off the table.

Config × Task → throwaway workspace → headless agent run → trajectory
↓
ResultRecord (JSONL) ← verification command

What's in this PR

New package @maka/lab (added to the workspaces list):

  • contractsTask (instruction + fixture workspace + verification command), Config (backend / connection / model), ResultRecord (one canonical row per run). Minimal on purpose; kept in @maka/lab for now, promotable to @maka/core once a second consumer exists.
  • sandbox — copies a task fixture into a throwaway dir; the source is never mutated and runs never bleed into each other. Isolation, not asking, is the safety model.
  • evaluator — runs the task's verification command (exit 0 = pass) in the throwaway workspace after the agent finishes, so a config can't grade itself; output-capped + timeout-killed.
  • runner — drives one headless turn through SessionManager, captures the trajectory via runtimeInvocationObserver, and auto-approves permission prompts (a benchmark has no human to confirm; the throwaway workspace is the boundary).
  • matrix — the full Config × Task cross product; a thrown run becomes a failed cell instead of aborting the grid.
  • results — ResultRecord JSONL is canonical truth; a git-diffable markdown comparison table is derived from it.
  • backends — the two concrete wirings, kept out of the engine: fake (deterministic) and ai-sdk (real model; key from a named env var, minimal AiSdkBackend = model + builtin tools + execute-mode permission).
  • CLImaka-lab run <spec.json> [--out <dir>] and maka-lab compare <results.jsonl>.
  • examplesdemo.spec.json (fake, no key) and fix-add.spec.json (a real coding task).

Live validation

The real ai-sdk backend was smoke-tested end-to-end on DeepSeek (deepseek-chat) with examples/fix-add (a buggy add() + a failing node:test):

  • completed, passed, exit 0, a 124-event trajectory using Edit / Write / Bash
  • the source fixture stayed buggy — the agent only edited the throwaway copy, proving sandbox isolation + headless permission auto-approve work against a real model

Testing

15 deterministic tests (FakeBackend), all green: CLI end-to-end smoke (spawns the built bin → results.jsonl + comparison.mdcompare), matrix cross-product + failed-cell, JSONL round-trip, table rendering, sandbox isolation, evaluator pass/fail/timeout, walking-skeleton e2e. Plus the live DeepSeek smoke above.

Out of scope (deliberate — pure additions later)

Parallel matrix execution, LLM/rule evaluators (today: command exit code), Docker / network isolation, SWE-bench pack ingestion, a richer report than the markdown grid, and promoting the contracts into @maka/core.

#31)
First end-to-end slice of the agent experiment lab: run one Config
against one Task in an isolated sandbox, capture the trajectory, score
it, and return a ResultRecord. De-risks the integration (headless
SessionManager run + sandbox + evaluator composing into one real run);
matrix/compare/CLI and real-model backends are later additions.
- contracts.ts: Task (instruction + fixture workspace + verification
command), Config (backend/connection/model), ResultRecord. Minimal;
systemPrompt/Execution/SWE-bench-pack ingestion deferred.
- sandbox.ts: prepareWorkspace copies the fixture to a throwaway dir so
the agent never mutates the source (isolation, not asking).
- evaluator.ts: runVerification runs the Task's test command in the
throwaway workspace after the agent finishes (config can't grade
itself), with output cap + timeout kill.
- runner.ts: runExperiment wires a pure-Node SessionManager (injected
registerBackends keeps model/credential wiring out of the lab core),
drives one turn, captures the InvocationResult trajectory via
runtimeInvocationObserver, scores it, returns a ResultRecord.
- tests: FakeBackend e2e (pass + fail fixtures, trajectory persisted as
runtime-events.jsonl), sandbox isolation, evaluator pass/fail/timeout.
Independent of the credential migration (#32/#33): the skeleton runs on
FakeBackend and needs no credentials.
Completes the MVP loop on top of the walking skeleton: run a grid of
Configs × Tasks, persist canonical results, and compare.
- matrix.ts: runMatrix runs the full cross product (sequential; a thrown
run becomes a failed cell instead of aborting the grid) with a
per-cell onResult callback.
- results.ts: ResultRecord JSONL is canonical truth; toComparisonTable
derives a git-diffable markdown grid (tasks × configs, ✅/❌/⚠️,
pass-rate footer). ResultRecord gains an optional `error`.
- backends.ts: the two concrete backend wirings, kept out of the engine.
registerFakeBackend (deterministic). registerAiSdkBackend resolves a
Config's slug against spec connections, reads the API key from a named
env var (no secrets at rest), and wires a minimal AiSdkBackend (model +
builtin tools + execute-mode permission). Telemetry/artifact/synthesis
hooks omitted — a benchmark scores via the verification command.
- runner.ts: the drain loop now auto-approves permission requests — a
headless benchmark has no human to confirm; throwaway-workspace
isolation is the safety net.
- cli.ts: `maka-lab run <spec.json> [--out <dir>]` and
`maka-lab compare <results.jsonl>`; task fixtures resolve relative to
the spec. Exposed as the `maka-lab` bin.
- tests (15 total): CLI end-to-end smoke (spawns the built bin on a fake
spec → results.jsonl + comparison.md → compare), matrix cross-product
+ failed-cell, JSONL round-trip, table rendering.
The real ai-sdk backend is typecheck-verified but not unit-tested (a live
model call is non-deterministic and costs money); it needs a live smoke
test with a real API key. Everything else is deterministic and green.
README documents the Config × Task model, the CLI (run/compare), the
spec shape (incl. a real ai-sdk connection with apiKeyEnv), the two
backends, and what's deliberately out of MVP scope. examples/demo.spec.json
+ examples/demo/marker.txt run green on the fake backend with no API key:
maka-lab run examples/demo.spec.json --out /tmp/maka-lab-demo
examples/fix-add — a buggy add() with a failing node:test. A real model
run must read the files, fix src.mjs, and turn `node --test` green.
Live smoke confirmed end-to-end on DeepSeek (deepseek-chat): completed,
passed, exit 0, 124-event trajectory using Edit/Write/Bash; the source
fixture stayed buggy (the agent only touched the throwaway copy), proving
sandbox isolation + headless permission auto-approve work against a real
backend. README points at it as the canonical real-run example.
…r, table (#36)
- runner: stop blanket-approving permission prompts. Allow ordinary tool
use, DENY dangerous categories (fs_destructive / git_destructive /
privileged / browser) by default — the workspace is a copy, not a jail,
so a tool can still escape via absolute paths/network. Opt in with
allowDangerousTools (real sandbox only). Corrects the misleading
"workspace is the safety net" comment. Also: a run that didn't complete
can no longer read as passed.
- evaluator: spawn detached + SIGKILL the process group on timeout, so
backgrounded grandchildren die too (a plain child.kill leaked them).
- sandbox: reject fixture symlinks (fs.cp preserved them verbatim → escape
to source/host) and clean up the temp dir if the copy fails (it was
leaked before the runner's finally registered).
- results: cellKey was separated by a NUL byte, making results.ts a binary
file to git — switch to a JSON key. Render a failed run as ⚠️ (distinct
from ❌ verification-failed) and exclude it from the pass count. Escape
`|`/newline in ids so they can't break the table.
- cli: reject unknown flags and require a value after --out.
- README: honest permission/safety wording.
+5 tests (symlink reject, failed-run render, id escaping, unknown flag,
missing flag value); 20 total green. Cross-process credential lock and a
real container sandbox remain deliberately out of scope.
The throwaway workspace stops a run from mutating the source fixture, but
it is NOT a security sandbox and a config could rewrite its own test to
pass. This addresses the two open review P1s without ripping out the real
backend (kept usable for your-own-models-on-trusted-tasks):
- Clean-room grading (verification integrity): Task.verification.protectedPaths
lists the grading assets; the runner restores them from the pristine
fixture AFTER the agent finishes and BEFORE the verification command runs,
so a model that rewrote its own test has that edit reverted. Each path is
removed first (drops an agent-planted symlink) and rejected if it escapes
the workspace. examples/fix-add now protects test.mjs.
- Honest docs (host isolation): the README states plainly that a real run
executes tool calls on your machine with your privileges — same exposure
as running Maka directly — so run only models/tasks you trust. Per-run
container isolation (mount workspace only, env allowlist, network policy)
is the named next hardening; allowDangerousTools is for inside it.
Tests (+4, 24 total green): restoreProtectedPaths unit (reverts protected,
keeps the rest, rejects escapes) + a malicious TamperBackend integration
proving protectedPaths reverts the cheat (passed=false) while the unguarded
task lets the cheat win (passed=true) — the guard is load-bearing. The
integration grades via `node check.mjs` (exit-code), not `node --test`, so
the verification child doesn't collide with the lab's own test runner.
@Astro-Han
Astro-Han deleted the claude/lab branch June 17, 2026 12:51
jackwener pushed a commit that referenced this pull request Jun 21, 2026
…r, table (#36)
- runner: stop blanket-approving permission prompts. Allow ordinary tool
use, DENY dangerous categories (fs_destructive / git_destructive /
privileged / browser) by default — the workspace is a copy, not a jail,
so a tool can still escape via absolute paths/network. Opt in with
allowDangerousTools (real sandbox only). Corrects the misleading
"workspace is the safety net" comment. Also: a run that didn't complete
can no longer read as passed.
- evaluator: spawn detached + SIGKILL the process group on timeout, so
backgrounded grandchildren die too (a plain child.kill leaked them).
- sandbox: reject fixture symlinks (fs.cp preserved them verbatim → escape
to source/host) and clean up the temp dir if the copy fails (it was
leaked before the runner's finally registered).
- results: cellKey was separated by a NUL byte, making results.ts a binary
file to git — switch to a JSON key. Render a failed run as ⚠️ (distinct
from ❌ verification-failed) and exclude it from the pass count. Escape
`|`/newline in ids so they can't break the table.
- cli: reject unknown flags and require a value after --out.
- README: honest permission/safety wording.
+5 tests (symlink reject, failed-run render, id escaping, unknown flag,
missing flag value); 20 total green. Cross-process credential lock and a
real container sandbox remain deliberately out of scope.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
The throwaway workspace stops a run from mutating the source fixture, but
it is NOT a security sandbox and a config could rewrite its own test to
pass. This addresses the two open review P1s without ripping out the real
backend (kept usable for your-own-models-on-trusted-tasks):
- Clean-room grading (verification integrity): Task.verification.protectedPaths
lists the grading assets; the runner restores them from the pristine
fixture AFTER the agent finishes and BEFORE the verification command runs,
so a model that rewrote its own test has that edit reverted. Each path is
removed first (drops an agent-planted symlink) and rejected if it escapes
the workspace. examples/fix-add now protects test.mjs.
- Honest docs (host isolation): the README states plainly that a real run
executes tool calls on your machine with your privileges — same exposure
as running Maka directly — so run only models/tasks you trust. Per-run
container isolation (mount workspace only, env allowlist, network policy)
is the named next hardening; allowDangerousTools is for inside it.
Tests (+4, 24 total green): restoreProtectedPaths unit (reverts protected,
keeps the rest, rejects escapes) + a malicious TamperBackend integration
proving protectedPaths reverts the cheat (passed=false) while the unguarded
task lets the cheat win (passed=true) — the guard is load-bearing. The
integration grades via `node check.mjs` (exit-code), not `node --test`, so
the verification child doesn't collide with the lab's own test runner.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
…r, table (#36)
- runner: stop blanket-approving permission prompts. Allow ordinary tool
use, DENY dangerous categories (fs_destructive / git_destructive /
privileged / browser) by default — the workspace is a copy, not a jail,
so a tool can still escape via absolute paths/network. Opt in with
allowDangerousTools (real sandbox only). Corrects the misleading
"workspace is the safety net" comment. Also: a run that didn't complete
can no longer read as passed.
- evaluator: spawn detached + SIGKILL the process group on timeout, so
backgrounded grandchildren die too (a plain child.kill leaked them).
- sandbox: reject fixture symlinks (fs.cp preserved them verbatim → escape
to source/host) and clean up the temp dir if the copy fails (it was
leaked before the runner's finally registered).
- results: cellKey was separated by a NUL byte, making results.ts a binary
file to git — switch to a JSON key. Render a failed run as ⚠️ (distinct
from ❌ verification-failed) and exclude it from the pass count. Escape
`|`/newline in ids so they can't break the table.
- cli: reject unknown flags and require a value after --out.
- README: honest permission/safety wording.
+5 tests (symlink reject, failed-run render, id escaping, unknown flag,
missing flag value); 20 total green. Cross-process credential lock and a
real container sandbox remain deliberately out of scope.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
The throwaway workspace stops a run from mutating the source fixture, but
it is NOT a security sandbox and a config could rewrite its own test to
pass. This addresses the two open review P1s without ripping out the real
backend (kept usable for your-own-models-on-trusted-tasks):
- Clean-room grading (verification integrity): Task.verification.protectedPaths
lists the grading assets; the runner restores them from the pristine
fixture AFTER the agent finishes and BEFORE the verification command runs,
so a model that rewrote its own test has that edit reverted. Each path is
removed first (drops an agent-planted symlink) and rejected if it escapes
the workspace. examples/fix-add now protects test.mjs.
- Honest docs (host isolation): the README states plainly that a real run
executes tool calls on your machine with your privileges — same exposure
as running Maka directly — so run only models/tasks you trust. Per-run
container isolation (mount workspace only, env allowlist, network policy)
is the named next hardening; allowDangerousTools is for inside it.
Tests (+4, 24 total green): restoreProtectedPaths unit (reverts protected,
keeps the rest, rejects escapes) + a malicious TamperBackend integration
proving protectedPaths reverts the cheat (passed=false) while the unguarded
task lets the cheat win (passed=true) — the guard is load-bearing. The
integration grades via `node check.mjs` (exit-code), not `node --test`, so
the verification child doesn't collide with the lab's own test runner.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
…r, table (#36)
- runner: stop blanket-approving permission prompts. Allow ordinary tool
use, DENY dangerous categories (fs_destructive / git_destructive /
privileged / browser) by default — the workspace is a copy, not a jail,
so a tool can still escape via absolute paths/network. Opt in with
allowDangerousTools (real sandbox only). Corrects the misleading
"workspace is the safety net" comment. Also: a run that didn't complete
can no longer read as passed.
- evaluator: spawn detached + SIGKILL the process group on timeout, so
backgrounded grandchildren die too (a plain child.kill leaked them).
- sandbox: reject fixture symlinks (fs.cp preserved them verbatim → escape
to source/host) and clean up the temp dir if the copy fails (it was
leaked before the runner's finally registered).
- results: cellKey was separated by a NUL byte, making results.ts a binary
file to git — switch to a JSON key. Render a failed run as ⚠️ (distinct
from ❌ verification-failed) and exclude it from the pass count. Escape
`|`/newline in ids so they can't break the table.
- cli: reject unknown flags and require a value after --out.
- README: honest permission/safety wording.
+5 tests (symlink reject, failed-run render, id escaping, unknown flag,
missing flag value); 20 total green. Cross-process credential lock and a
real container sandbox remain deliberately out of scope.
jackwener pushed a commit that referenced this pull request Jun 21, 2026
The throwaway workspace stops a run from mutating the source fixture, but
it is NOT a security sandbox and a config could rewrite its own test to
pass. This addresses the two open review P1s without ripping out the real
backend (kept usable for your-own-models-on-trusted-tasks):
- Clean-room grading (verification integrity): Task.verification.protectedPaths
lists the grading assets; the runner restores them from the pristine
fixture AFTER the agent finishes and BEFORE the verification command runs,
so a model that rewrote its own test has that edit reverted. Each path is
removed first (drops an agent-planted symlink) and rejected if it escapes
the workspace. examples/fix-add now protects test.mjs.
- Honest docs (host isolation): the README states plainly that a real run
executes tool calls on your machine with your privileges — same exposure
as running Maka directly — so run only models/tasks you trust. Per-run
container isolation (mount workspace only, env allowlist, network policy)
is the named next hardening; allowDangerousTools is for inside it.
Tests (+4, 24 total green): restoreProtectedPaths unit (reverts protected,
keeps the rest, rejects escapes) + a malicious TamperBackend integration
proving protectedPaths reverts the cheat (passed=false) while the unguarded
task lets the cheat win (passed=true) — the guard is load-bearing. The
integration grades via `node check.mjs` (exit-code), not `node --test`, so
the verification child doesn't collide with the lab's own test runner.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han