RFC: Harness RSI loop — autoresearch structure for system prompt optimization
Context
maka's eval framework is ready (PR #62/#63). The goal is closed-loop harness self-improvement: an agent autonomously optimizes its own system prompt against a benchmark task set, no human in the loop, runnable until stopped.
Prior art: Karpathy's autoresearch (minimal loop, single agent), HarnessX/AEGIS (heavy 4-stage engine, +14.5%), Self-Harness (regression-gated edit loop).
Architecture (verified)
RuntimeRunner in-container + Harbor verifier. Harbor provides the Docker container (from TB task Dockerfile). Inside that container, a run-cell.mjs entrypoint runs maka's RuntimeRunner directly — the agent works in-place at /app (the Dockerfile WORKDIR), using file tools + Bash all in-container (one filesystem). After the agent finishes, Harbor runs tests/test.sh which checks /app state.
Controller (host):
1. For each held-in task: harbor run --path <tb-task> --agent-import-path maka_agent:MakaAgent
--ae MAKA_SYSTEM_PROMPT=<prompt> --ae DEEPSEEK_API_KEY=<key>
2. Parse Harbor result.json → per-task reward (0/1)
3. Parse adapter's errorClass output → per-task errorClass
Harbor container (per task):
1. Build from task Dockerfile (WORKDIR /app, has tools, test env)
2. MakaAgent.run() → adapter starts run-cell.mjs in /app
→ RuntimeRunner runs agent (file tools + Bash, all in /app, one filesystem)
→ RuntimeRunner records trajectory to runtime-events.jsonl
→ adapter extracts errorClass from InvocationResult
3. Harbor runs tests/test.sh → checks /app state → reward 0 or 1
Why not runExperiment? Verified: runExperiment copies fixture to mkdtemp('/tmp/maka-headless-ws-XXXX'), runs verifier there, deletes it. But TB test scripts write absolute paths to /app (e.g. Path("/app/regex.txt") in test_outputs.py). 89/89 Dockerfiles use WORKDIR /app. runExperiment's temp workspace would never be checked by TB tests. RuntimeRunner works in-place at /app — no copy, no sync problem.
Why not Harbor verifier only? We need maka's errorClass (max_tokens / verification_failed / runtime_error) for signal enrichment. The adapter extracts this from RuntimeRunner's InvocationResult and writes it alongside Harbor's reward.
File/bash filesystem split (v5 P0-2, resolved): the entire run-cell.mjs process runs in-container. RuntimeRunner's file tools (Write/Edit/Read) and Bash tool (toolExecutor) all operate on the container's filesystem. No host-container split.
Prompt delivery: controller passes system_prompt.md content via MAKA_SYSTEM_PROMPT env var. run-cell.mjs reads it, passes to Config.systemPrompt. Runtime stamps a hash of the effectiveConfig.systemPrompt string (at backend constructor, after resolveSystemPrompt) into trajectory events. Controller verifies the hash round-trips — if trajectory hash ≠ hash of prompt sent, the round is invalid (plumbing failure).
Implementation prerequisite:systemPromptHash does not exist in runtime yet (verified: only requestShapeHash exists). Must add a small instrumentation: hash the resolved Config.systemPrompt and write to a runtime event field. Pure addition, no existing code changes.
Harbor/container failure ≠ benchmark failure. Container crash, image build failure, Harbor timeout → no result.json or invalid output. Controller treats this as infra error. Harbor reward + adapter errorClass only decide benchmark score.
Signal quality
Fix 1 — Coverage-aware gate.max_tokens/runtime_error tasks should count as fail (not drop from denominator). Gate uses pass / eligible (unscored = fail), NOT pass / scored.
keep = held_in_pass_rate (pass/eligible) improved beyond noise_band_held_in
AND held_in_coverage (scored/eligible) not degraded
AND held_out_pass_rate non-inferior vs original baseline (>= -noise_band_held_out)
AND held_out_coverage not degraded
AND no infra_errors (any infra error → no keep)
Controller computes pass/eligible itself. Harbor gives per-task reward (0/1); adapter gives per-task errorClass; controller combines: pass = reward==1, eligible = true (all tasks eligible unless infra error), scored = errorClass != 'max_tokens' && errorClass != 'runtime_error'.
Fix 2 — Dual noise band.noise_band = wilson_half_width(N, p_mean) × √(1 + 1/n_baseline) — CI of the difference, not single-proportion CI. Separate bands for held-in and held-out.
- Layer 0 (free): dual noise thresholds.
- Layer 1 (cheap): 60 held-in + 20 held-out from TB 2.0's 89 tasks.
- Layer 2 (deferred for v1): repeat sampling. v1 goal is structural validation.
Fix 3 — Agent reads failing trajectories. Controller extracts per-task digest from runtime-events.jsonl (raw events): last 2 tool calls (from function_call events with args), errorClass, steps/duration, verifier stdout/stderr first 3 lines.
Overfitting guard
Held-in vs current kept baseline (monotonic improvement). After KEEP: held_in_reference = max(previous_reference, new_run_pass_rate − noise_band_held_in). Banks real improvements without enshrining lucky peak, never falls below prior state.
Held-out vs original baseline (fixed floor). Floor = original_held_out_mean − noise_band_held_out. Prevents cumulative drift.
Held-out physical isolation. Held-out task content, trajectories, results.jsonl live outside agent cwd. Agent sees only program.md, system_prompt.md, results.tsv, held-in digests.
Reward-hack quarantine gate. Controller scans function_call.args in raw runtime-events.jsonl for verifier-specific patterns (expected-output strings, not filenames). Flagged → QUARANTINE (discard). Scanner fails closed if raw events missing.
Infra error handling
Any infra error → no KEEP. Infra = missing result.json, container crash, image build failure, Harbor timeout. Retry once, else discard+log.
>20% infra → stop. Stop condition, not keep-allowance.
Infra classification: model API failure (DeepSeek 429/500) → runtime_error (scored=false, counts in eligible). Container/Docker/Harbor failure → infra (excluded from scoring).
Budget
Controller sums costUsd from runtime-events.jsonl token-usage events. costUsd exists on RuntimeEvent. BUILTIN_PRICING keyed as deepseek:deepseek-chat (provider:model format) — controller must compose key as ${provider}:${config.model}.
Zero-cost guard: if summed costUsd = 0 but tokens > 0 → plumbing failure (pricing not wired). Do NOT count as free round.
Cost ceiling: $30 (default). 80 tasks × single-sample × ~10 rounds ≈ $15-25.
Controller / agent separation
Controller (host-side Node.js): Harbor invocation, result parsing, metric computation, budget, git state machine, keep/discard, crash/resume WAL (append jsonl LAST as source of truth; reconcile last_kept_commit from jsonl on resume). Diff guard AFTER agent edit, BEFORE commit — only system_prompt.md (tracked) may change; controller files gitignored.
Agent (meta-agent): scripted LLM call (deepseek-chat). Reads results.tsv + digests → outputs prompt edit + description. Log IS its memory. No interactive tools.
Implementation entrypoint
Not maka-headless eval CLI (only wires fake backend). run-cell.mjs in container directly imports RuntimeRunner from @maka/runtime:
import{RuntimeRunner,AiSdkBackend, ... }from'@maka/runtime';// agent runs in /app (Dockerfile WORKDIR), in-place// RuntimeRunner → InvocationResult → extract errorClass// write errorClass + trajectory to /output/maka-agent repo cloned + built at /opt/maka-agent (outside /app).
Verified facts
| Question | Answer | Source |
|---|
TB test.sh writes /app? | Yes — test_outputs.py uses Path("/app/regex.txt") etc. | grep 89 tasks |
| runExperiment compatible? | No — prepareWorkspace uses mkdtemp('/tmp/...'), can't be /app | sandbox.ts:20 |
| systemPromptHash exists? | No — only requestShapeHash | grep runtime/src |
| pricing key format? | provider:model (deepseek:deepseek-chat) | builtin-pricing.ts:14 |
| Harbor + Docker verified? | Yes — oracle reward 1.0, maka adapter runs, verifier works | smoke test 6/19 |
| file/bash split resolved? | Yes — whole process in-container | architecture decision |
Directory layout
~/.local/maka-eval/
├── harbor-adapter/
│ ├── maka_agent.py # Harbor agent adapter
│ └── run-cell.mjs # in-container entrypoint (RuntimeRunner)
├── harness-rsi/
│ ├── program.md # loop instructions (agent reads)
│ ├── system_prompt.md # optimization target (agent edits ONLY)
│ ├── controller.mjs # state machine (host-side)
│ ├── results.jsonl # canonical truth (OUTSIDE agent cwd)
│ ├── results.tsv # derived view for agent (IN agent cwd, gitignored)
│ └── baseline/ # 3× baseline calibration
├── ds-key.txt
└── fixtures/
The loop
INIT:
1. Run baseline 3× on held-in + held-out
2. Compute means + observed_spread
3. noise_band = max(wilson_half_width(N, p) × √(4/3), observed_spread)
4. held_in_reference = baseline_held_in_mean
5. Save original_held_out_mean (never updated)
LOOP:
1. Write pending round record
2. Agent: read results.tsv + digests → edit system_prompt.md
3. Controller: diff guard (after edit, before commit) — only system_prompt.md
4. Controller: git commit
5. Controller: run held-in via Harbor → parse reward + errorClass per task
6. Check infra — any infra → no keep, retry once, else discard
7. Compute held_in_delta vs held_in_reference
8. |delta| < noise_band → DISCARD (inconclusive)
9. delta <= 0 → DISCARD (regression)
10. delta > noise_band:
a. Run held-out via Harbor
b. Check held-out infra
c. held_out_delta vs original_held_out_mean
d. < -noise_band → DISCARD (regression)
e. coverage degraded → DISCARD
f. reward-hack scan → QUARANTINE if flagged
g. >= -noise_band AND coverage OK → KEEP
h. held_in_reference = max(held_in_reference, pass_rate − noise_band)
11. Append to results.jsonl (source of truth)
12. Advance last_kept_commit (if KEEP, AFTER jsonl append)
13. Update results.tsv
14. Check budget (sum costUsd, zero-cost guard)
Success criteria (structural validation)
- Loop runs unattended ≥10 iterations without human intervention
- Zero keeps is passing — v1 validates structure, not prompt quality
- If keeps occur: no held-out regression, no unreviewed reward-hack flags
- Cost ≤ $30
- Prompt round-trip hash verified every round
Open questions
- Task selection: which 60 of 89 for held-in? Filter by speed/stability.
- Parallelism:
--n-concurrent — determinism? Key results by task id. - Simplicity criterion: deletion + non-inferior = KEEP? Needs own branch.
- Pricing wiring: verify in-container path registers pricing lookup.
Non-goals (v1)
- No tool description optimization
- No control flow changes
- No AEGIS 4-stage pipeline
- No repeat sampling / pass@k
- No runExperiment (RuntimeRunner + Harbor verifier instead)
RFC: Harness RSI loop — autoresearch structure for system prompt optimization
Context
maka's eval framework is ready (PR #62/#63). The goal is closed-loop harness self-improvement: an agent autonomously optimizes its own system prompt against a benchmark task set, no human in the loop, runnable until stopped.
Prior art: Karpathy's autoresearch (minimal loop, single agent), HarnessX/AEGIS (heavy 4-stage engine, +14.5%), Self-Harness (regression-gated edit loop).
Architecture (verified)
RuntimeRunner in-container + Harbor verifier. Harbor provides the Docker container (from TB task Dockerfile). Inside that container, a
run-cell.mjsentrypoint runs maka'sRuntimeRunnerdirectly — the agent works in-place at/app(the Dockerfile WORKDIR), using file tools + Bash all in-container (one filesystem). After the agent finishes, Harbor runstests/test.shwhich checks/appstate.Why not runExperiment? Verified:
runExperimentcopies fixture tomkdtemp('/tmp/maka-headless-ws-XXXX'), runs verifier there, deletes it. But TB test scripts write absolute paths to/app(e.g.Path("/app/regex.txt")in test_outputs.py). 89/89 Dockerfiles useWORKDIR /app.runExperiment's temp workspace would never be checked by TB tests. RuntimeRunner works in-place at/app— no copy, no sync problem.Why not Harbor verifier only? We need maka's
errorClass(max_tokens / verification_failed / runtime_error) for signal enrichment. The adapter extracts this fromRuntimeRunner'sInvocationResultand writes it alongside Harbor's reward.File/bash filesystem split (v5 P0-2, resolved): the entire
run-cell.mjsprocess runs in-container.RuntimeRunner's file tools (Write/Edit/Read) and Bash tool (toolExecutor) all operate on the container's filesystem. No host-container split.Prompt delivery: controller passes
system_prompt.mdcontent viaMAKA_SYSTEM_PROMPTenv var.run-cell.mjsreads it, passes toConfig.systemPrompt. Runtime stamps a hash of the effectiveConfig.systemPromptstring (at backend constructor, afterresolveSystemPrompt) into trajectory events. Controller verifies the hash round-trips — if trajectory hash ≠ hash of prompt sent, the round is invalid (plumbing failure).Implementation prerequisite:
systemPromptHashdoes not exist in runtime yet (verified: onlyrequestShapeHashexists). Must add a small instrumentation: hash the resolvedConfig.systemPromptand write to a runtime event field. Pure addition, no existing code changes.Harbor/container failure ≠ benchmark failure. Container crash, image build failure, Harbor timeout → no
result.jsonor invalid output. Controller treats this as infra error. Harbor reward + adapter errorClass only decide benchmark score.Signal quality
Fix 1 — Coverage-aware gate.
max_tokens/runtime_errortasks should count as fail (not drop from denominator). Gate usespass / eligible(unscored = fail), NOTpass / scored.Controller computes
pass/eligibleitself. Harbor gives per-task reward (0/1); adapter gives per-task errorClass; controller combines:pass = reward==1,eligible = true(all tasks eligible unless infra error),scored = errorClass != 'max_tokens' && errorClass != 'runtime_error'.Fix 2 — Dual noise band.
noise_band = wilson_half_width(N, p_mean) × √(1 + 1/n_baseline)— CI of the difference, not single-proportion CI. Separate bands for held-in and held-out.Fix 3 — Agent reads failing trajectories. Controller extracts per-task digest from
runtime-events.jsonl(raw events): last 2 tool calls (fromfunction_callevents with args), errorClass, steps/duration, verifier stdout/stderr first 3 lines.Overfitting guard
Held-in vs current kept baseline (monotonic improvement). After KEEP:
held_in_reference = max(previous_reference, new_run_pass_rate − noise_band_held_in). Banks real improvements without enshrining lucky peak, never falls below prior state.Held-out vs original baseline (fixed floor). Floor =
original_held_out_mean − noise_band_held_out. Prevents cumulative drift.Held-out physical isolation. Held-out task content, trajectories,
results.jsonllive outside agent cwd. Agent sees onlyprogram.md,system_prompt.md,results.tsv, held-in digests.Reward-hack quarantine gate. Controller scans
function_call.argsin rawruntime-events.jsonlfor verifier-specific patterns (expected-output strings, not filenames). Flagged → QUARANTINE (discard). Scanner fails closed if raw events missing.Infra error handling
Any infra error → no KEEP. Infra = missing result.json, container crash, image build failure, Harbor timeout. Retry once, else discard+log.
>20%infra → stop. Stop condition, not keep-allowance.Infra classification: model API failure (DeepSeek 429/500) →
runtime_error(scored=false, counts in eligible). Container/Docker/Harbor failure → infra (excluded from scoring).Budget
Controller sums
costUsdfromruntime-events.jsonltoken-usage events.costUsdexists onRuntimeEvent.BUILTIN_PRICINGkeyed asdeepseek:deepseek-chat(provider:model format) — controller must compose key as${provider}:${config.model}.Zero-cost guard: if summed
costUsd= 0 but tokens > 0 → plumbing failure (pricing not wired). Do NOT count as free round.Cost ceiling: $30 (default). 80 tasks × single-sample × ~10 rounds ≈ $15-25.
Controller / agent separation
Controller (host-side Node.js): Harbor invocation, result parsing, metric computation, budget, git state machine, keep/discard, crash/resume WAL (append jsonl LAST as source of truth; reconcile
last_kept_commitfrom jsonl on resume). Diff guard AFTER agent edit, BEFORE commit — onlysystem_prompt.md(tracked) may change; controller files gitignored.Agent (meta-agent): scripted LLM call (deepseek-chat). Reads
results.tsv+ digests → outputs prompt edit + description. Log IS its memory. No interactive tools.Implementation entrypoint
Not
maka-headless evalCLI (only wires fake backend).run-cell.mjsin container directly importsRuntimeRunnerfrom@maka/runtime:maka-agent repo cloned + built at
/opt/maka-agent(outside/app).Verified facts
/app?Path("/app/regex.txt")etc.prepareWorkspaceusesmkdtemp('/tmp/...'), can't be/appsandbox.ts:20requestShapeHashprovider:model(deepseek:deepseek-chat)builtin-pricing.ts:14Directory layout
The loop
Success criteria (structural validation)
Open questions
--n-concurrent— determinism? Key results by task id.Non-goals (v1)