RFC: Harness RSI loop — autonomous system prompt optimization #64

Description

@Astro-Han

RFC: Harness RSI loop — autoresearch structure for system prompt optimization

Context

maka's eval framework is ready (PR #62/#63). The goal is closed-loop harness self-improvement: an agent autonomously optimizes its own system prompt against a benchmark task set, no human in the loop, runnable until stopped.

Prior art: Karpathy's autoresearch (minimal loop, single agent), HarnessX/AEGIS (heavy 4-stage engine, +14.5%), Self-Harness (regression-gated edit loop).

Architecture (verified)

RuntimeRunner in-container + Harbor verifier. Harbor provides the Docker container (from TB task Dockerfile). Inside that container, a run-cell.mjs entrypoint runs maka's RuntimeRunner directly — the agent works in-place at /app (the Dockerfile WORKDIR), using file tools + Bash all in-container (one filesystem). After the agent finishes, Harbor runs tests/test.sh which checks /app state.

Controller (host):
1. For each held-in task: harbor run --path <tb-task> --agent-import-path maka_agent:MakaAgent
--ae MAKA_SYSTEM_PROMPT=<prompt> --ae DEEPSEEK_API_KEY=<key>
2. Parse Harbor result.json → per-task reward (0/1)
3. Parse adapter's errorClass output → per-task errorClass
Harbor container (per task):
1. Build from task Dockerfile (WORKDIR /app, has tools, test env)
2. MakaAgent.run() → adapter starts run-cell.mjs in /app
→ RuntimeRunner runs agent (file tools + Bash, all in /app, one filesystem)
→ RuntimeRunner records trajectory to runtime-events.jsonl
→ adapter extracts errorClass from InvocationResult
3. Harbor runs tests/test.sh → checks /app state → reward 0 or 1

Why not runExperiment? Verified: runExperiment copies fixture to mkdtemp('/tmp/maka-headless-ws-XXXX'), runs verifier there, deletes it. But TB test scripts write absolute paths to /app (e.g. Path("/app/regex.txt") in test_outputs.py). 89/89 Dockerfiles use WORKDIR /app. runExperiment's temp workspace would never be checked by TB tests. RuntimeRunner works in-place at /app — no copy, no sync problem.

Why not Harbor verifier only? We need maka's errorClass (max_tokens / verification_failed / runtime_error) for signal enrichment. The adapter extracts this from RuntimeRunner's InvocationResult and writes it alongside Harbor's reward.

File/bash filesystem split (v5 P0-2, resolved): the entire run-cell.mjs process runs in-container. RuntimeRunner's file tools (Write/Edit/Read) and Bash tool (toolExecutor) all operate on the container's filesystem. No host-container split.

Prompt delivery: controller passes system_prompt.md content via MAKA_SYSTEM_PROMPT env var. run-cell.mjs reads it, passes to Config.systemPrompt. Runtime stamps a hash of the effectiveConfig.systemPrompt string (at backend constructor, after resolveSystemPrompt) into trajectory events. Controller verifies the hash round-trips — if trajectory hash ≠ hash of prompt sent, the round is invalid (plumbing failure).

Implementation prerequisite:systemPromptHash does not exist in runtime yet (verified: only requestShapeHash exists). Must add a small instrumentation: hash the resolved Config.systemPrompt and write to a runtime event field. Pure addition, no existing code changes.

Harbor/container failure ≠ benchmark failure. Container crash, image build failure, Harbor timeout → no result.json or invalid output. Controller treats this as infra error. Harbor reward + adapter errorClass only decide benchmark score.

Signal quality

Fix 1 — Coverage-aware gate.max_tokens/runtime_error tasks should count as fail (not drop from denominator). Gate uses pass / eligible (unscored = fail), NOT pass / scored.

keep = held_in_pass_rate (pass/eligible) improved beyond noise_band_held_in
AND held_in_coverage (scored/eligible) not degraded
AND held_out_pass_rate non-inferior vs original baseline (>= -noise_band_held_out)
AND held_out_coverage not degraded
AND no infra_errors (any infra error → no keep)

Controller computes pass/eligible itself. Harbor gives per-task reward (0/1); adapter gives per-task errorClass; controller combines: pass = reward==1, eligible = true (all tasks eligible unless infra error), scored = errorClass != 'max_tokens' && errorClass != 'runtime_error'.

Fix 2 — Dual noise band.noise_band = wilson_half_width(N, p_mean) × √(1 + 1/n_baseline) — CI of the difference, not single-proportion CI. Separate bands for held-in and held-out.

  • Layer 0 (free): dual noise thresholds.
  • Layer 1 (cheap): 60 held-in + 20 held-out from TB 2.0's 89 tasks.
  • Layer 2 (deferred for v1): repeat sampling. v1 goal is structural validation.

Fix 3 — Agent reads failing trajectories. Controller extracts per-task digest from runtime-events.jsonl (raw events): last 2 tool calls (from function_call events with args), errorClass, steps/duration, verifier stdout/stderr first 3 lines.

Overfitting guard

Held-in vs current kept baseline (monotonic improvement). After KEEP: held_in_reference = max(previous_reference, new_run_pass_rate − noise_band_held_in). Banks real improvements without enshrining lucky peak, never falls below prior state.

Held-out vs original baseline (fixed floor). Floor = original_held_out_mean − noise_band_held_out. Prevents cumulative drift.

Held-out physical isolation. Held-out task content, trajectories, results.jsonl live outside agent cwd. Agent sees only program.md, system_prompt.md, results.tsv, held-in digests.

Reward-hack quarantine gate. Controller scans function_call.args in raw runtime-events.jsonl for verifier-specific patterns (expected-output strings, not filenames). Flagged → QUARANTINE (discard). Scanner fails closed if raw events missing.

Infra error handling

Any infra error → no KEEP. Infra = missing result.json, container crash, image build failure, Harbor timeout. Retry once, else discard+log.

>20% infra → stop. Stop condition, not keep-allowance.

Infra classification: model API failure (DeepSeek 429/500) → runtime_error (scored=false, counts in eligible). Container/Docker/Harbor failure → infra (excluded from scoring).

Budget

Controller sums costUsd from runtime-events.jsonl token-usage events. costUsd exists on RuntimeEvent. BUILTIN_PRICING keyed as deepseek:deepseek-chat (provider:model format) — controller must compose key as ${provider}:${config.model}.

Zero-cost guard: if summed costUsd = 0 but tokens > 0 → plumbing failure (pricing not wired). Do NOT count as free round.

Cost ceiling: $30 (default). 80 tasks × single-sample × ~10 rounds ≈ $15-25.

Controller / agent separation

Controller (host-side Node.js): Harbor invocation, result parsing, metric computation, budget, git state machine, keep/discard, crash/resume WAL (append jsonl LAST as source of truth; reconcile last_kept_commit from jsonl on resume). Diff guard AFTER agent edit, BEFORE commit — only system_prompt.md (tracked) may change; controller files gitignored.

Agent (meta-agent): scripted LLM call (deepseek-chat). Reads results.tsv + digests → outputs prompt edit + description. Log IS its memory. No interactive tools.

Implementation entrypoint

Not maka-headless eval CLI (only wires fake backend). run-cell.mjs in container directly imports RuntimeRunner from @maka/runtime:

import{RuntimeRunner,AiSdkBackend, ... }from'@maka/runtime';// agent runs in /app (Dockerfile WORKDIR), in-place// RuntimeRunner → InvocationResult → extract errorClass// write errorClass + trajectory to /output/

maka-agent repo cloned + built at /opt/maka-agent (outside /app).

Verified facts

QuestionAnswerSource
TB test.sh writes /app?Yes — test_outputs.py uses Path("/app/regex.txt") etc.grep 89 tasks
runExperiment compatible?No — prepareWorkspace uses mkdtemp('/tmp/...'), can't be /appsandbox.ts:20
systemPromptHash exists?No — only requestShapeHashgrep runtime/src
pricing key format?provider:model (deepseek:deepseek-chat)builtin-pricing.ts:14
Harbor + Docker verified?Yes — oracle reward 1.0, maka adapter runs, verifier workssmoke test 6/19
file/bash split resolved?Yes — whole process in-containerarchitecture decision

Directory layout

~/.local/maka-eval/
├── harbor-adapter/
│ ├── maka_agent.py # Harbor agent adapter
│ └── run-cell.mjs # in-container entrypoint (RuntimeRunner)
├── harness-rsi/
│ ├── program.md # loop instructions (agent reads)
│ ├── system_prompt.md # optimization target (agent edits ONLY)
│ ├── controller.mjs # state machine (host-side)
│ ├── results.jsonl # canonical truth (OUTSIDE agent cwd)
│ ├── results.tsv # derived view for agent (IN agent cwd, gitignored)
│ └── baseline/ # 3× baseline calibration
├── ds-key.txt
└── fixtures/

The loop

INIT:
1. Run baseline 3× on held-in + held-out
2. Compute means + observed_spread
3. noise_band = max(wilson_half_width(N, p) × √(4/3), observed_spread)
4. held_in_reference = baseline_held_in_mean
5. Save original_held_out_mean (never updated)
LOOP:
1. Write pending round record
2. Agent: read results.tsv + digests → edit system_prompt.md
3. Controller: diff guard (after edit, before commit) — only system_prompt.md
4. Controller: git commit
5. Controller: run held-in via Harbor → parse reward + errorClass per task
6. Check infra — any infra → no keep, retry once, else discard
7. Compute held_in_delta vs held_in_reference
8. |delta| < noise_band → DISCARD (inconclusive)
9. delta <= 0 → DISCARD (regression)
10. delta > noise_band:
a. Run held-out via Harbor
b. Check held-out infra
c. held_out_delta vs original_held_out_mean
d. < -noise_band → DISCARD (regression)
e. coverage degraded → DISCARD
f. reward-hack scan → QUARANTINE if flagged
g. >= -noise_band AND coverage OK → KEEP
h. held_in_reference = max(held_in_reference, pass_rate − noise_band)
11. Append to results.jsonl (source of truth)
12. Advance last_kept_commit (if KEEP, AFTER jsonl append)
13. Update results.tsv
14. Check budget (sum costUsd, zero-cost guard)

Success criteria (structural validation)

  • Loop runs unattended ≥10 iterations without human intervention
  • Zero keeps is passing — v1 validates structure, not prompt quality
  • If keeps occur: no held-out regression, no unreviewed reward-hack flags
  • Cost ≤ $30
  • Prompt round-trip hash verified every round

Open questions

  1. Task selection: which 60 of 89 for held-in? Filter by speed/stability.
  2. Parallelism:--n-concurrent — determinism? Key results by task id.
  3. Simplicity criterion: deletion + non-inferior = KEEP? Needs own branch.
  4. Pricing wiring: verify in-container path registers pricing lookup.

Non-goals (v1)

  • No tool description optimization
  • No control flow changes
  • No AEGIS 4-stage pipeline
  • No repeat sampling / pass@k
  • No runExperiment (RuntimeRunner + Harbor verifier instead)

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions

      , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
       blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
      }
      } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
      })();
      (function(){
      try {
      var __m = "github.com";
      var __re = new RegExp('^' + "github\\.com" + '
      
      Skip to content

      RFC: Harness RSI loop — autonomous system prompt optimization #64

      Description

      @Astro-Han

      RFC: Harness RSI loop — autoresearch structure for system prompt optimization

      Context

      maka's eval framework is ready (PR #62/#63). The goal is closed-loop harness self-improvement: an agent autonomously optimizes its own system prompt against a benchmark task set, no human in the loop, runnable until stopped.

      Prior art: Karpathy's autoresearch (minimal loop, single agent), HarnessX/AEGIS (heavy 4-stage engine, +14.5%), Self-Harness (regression-gated edit loop).

      Architecture (verified)

      RuntimeRunner in-container + Harbor verifier. Harbor provides the Docker container (from TB task Dockerfile). Inside that container, a run-cell.mjs entrypoint runs maka's RuntimeRunner directly — the agent works in-place at /app (the Dockerfile WORKDIR), using file tools + Bash all in-container (one filesystem). After the agent finishes, Harbor runs tests/test.sh which checks /app state.

      Controller (host):
      1. For each held-in task: harbor run --path <tb-task> --agent-import-path maka_agent:MakaAgent
      --ae MAKA_SYSTEM_PROMPT=<prompt> --ae DEEPSEEK_API_KEY=<key>
      2. Parse Harbor result.json → per-task reward (0/1)
      3. Parse adapter's errorClass output → per-task errorClass
      Harbor container (per task):
      1. Build from task Dockerfile (WORKDIR /app, has tools, test env)
      2. MakaAgent.run() → adapter starts run-cell.mjs in /app
      → RuntimeRunner runs agent (file tools + Bash, all in /app, one filesystem)
      → RuntimeRunner records trajectory to runtime-events.jsonl
      → adapter extracts errorClass from InvocationResult
      3. Harbor runs tests/test.sh → checks /app state → reward 0 or 1
      

      Why not runExperiment? Verified: runExperiment copies fixture to mkdtemp('/tmp/maka-headless-ws-XXXX'), runs verifier there, deletes it. But TB test scripts write absolute paths to /app (e.g. Path("/app/regex.txt") in test_outputs.py). 89/89 Dockerfiles use WORKDIR /app. runExperiment's temp workspace would never be checked by TB tests. RuntimeRunner works in-place at /app — no copy, no sync problem.

      Why not Harbor verifier only? We need maka's errorClass (max_tokens / verification_failed / runtime_error) for signal enrichment. The adapter extracts this from RuntimeRunner's InvocationResult and writes it alongside Harbor's reward.

      File/bash filesystem split (v5 P0-2, resolved): the entire run-cell.mjs process runs in-container. RuntimeRunner's file tools (Write/Edit/Read) and Bash tool (toolExecutor) all operate on the container's filesystem. No host-container split.

      Prompt delivery: controller passes system_prompt.md content via MAKA_SYSTEM_PROMPT env var. run-cell.mjs reads it, passes to Config.systemPrompt. Runtime stamps a hash of the effectiveConfig.systemPrompt string (at backend constructor, after resolveSystemPrompt) into trajectory events. Controller verifies the hash round-trips — if trajectory hash ≠ hash of prompt sent, the round is invalid (plumbing failure).

      Implementation prerequisite:systemPromptHash does not exist in runtime yet (verified: only requestShapeHash exists). Must add a small instrumentation: hash the resolved Config.systemPrompt and write to a runtime event field. Pure addition, no existing code changes.

      Harbor/container failure ≠ benchmark failure. Container crash, image build failure, Harbor timeout → no result.json or invalid output. Controller treats this as infra error. Harbor reward + adapter errorClass only decide benchmark score.

      Signal quality

      Fix 1 — Coverage-aware gate.max_tokens/runtime_error tasks should count as fail (not drop from denominator). Gate uses pass / eligible (unscored = fail), NOT pass / scored.

      keep = held_in_pass_rate (pass/eligible) improved beyond noise_band_held_in
      AND held_in_coverage (scored/eligible) not degraded
      AND held_out_pass_rate non-inferior vs original baseline (>= -noise_band_held_out)
      AND held_out_coverage not degraded
      AND no infra_errors (any infra error → no keep)
      

      Controller computes pass/eligible itself. Harbor gives per-task reward (0/1); adapter gives per-task errorClass; controller combines: pass = reward==1, eligible = true (all tasks eligible unless infra error), scored = errorClass != 'max_tokens' && errorClass != 'runtime_error'.

      Fix 2 — Dual noise band.noise_band = wilson_half_width(N, p_mean) × √(1 + 1/n_baseline) — CI of the difference, not single-proportion CI. Separate bands for held-in and held-out.

      • Layer 0 (free): dual noise thresholds.
      • Layer 1 (cheap): 60 held-in + 20 held-out from TB 2.0's 89 tasks.
      • Layer 2 (deferred for v1): repeat sampling. v1 goal is structural validation.

      Fix 3 — Agent reads failing trajectories. Controller extracts per-task digest from runtime-events.jsonl (raw events): last 2 tool calls (from function_call events with args), errorClass, steps/duration, verifier stdout/stderr first 3 lines.

      Overfitting guard

      Held-in vs current kept baseline (monotonic improvement). After KEEP: held_in_reference = max(previous_reference, new_run_pass_rate − noise_band_held_in). Banks real improvements without enshrining lucky peak, never falls below prior state.

      Held-out vs original baseline (fixed floor). Floor = original_held_out_mean − noise_band_held_out. Prevents cumulative drift.

      Held-out physical isolation. Held-out task content, trajectories, results.jsonl live outside agent cwd. Agent sees only program.md, system_prompt.md, results.tsv, held-in digests.

      Reward-hack quarantine gate. Controller scans function_call.args in raw runtime-events.jsonl for verifier-specific patterns (expected-output strings, not filenames). Flagged → QUARANTINE (discard). Scanner fails closed if raw events missing.

      Infra error handling

      Any infra error → no KEEP. Infra = missing result.json, container crash, image build failure, Harbor timeout. Retry once, else discard+log.

      >20% infra → stop. Stop condition, not keep-allowance.

      Infra classification: model API failure (DeepSeek 429/500) → runtime_error (scored=false, counts in eligible). Container/Docker/Harbor failure → infra (excluded from scoring).

      Budget

      Controller sums costUsd from runtime-events.jsonl token-usage events. costUsd exists on RuntimeEvent. BUILTIN_PRICING keyed as deepseek:deepseek-chat (provider:model format) — controller must compose key as ${provider}:${config.model}.

      Zero-cost guard: if summed costUsd = 0 but tokens > 0 → plumbing failure (pricing not wired). Do NOT count as free round.

      Cost ceiling: $30 (default). 80 tasks × single-sample × ~10 rounds ≈ $15-25.

      Controller / agent separation

      Controller (host-side Node.js): Harbor invocation, result parsing, metric computation, budget, git state machine, keep/discard, crash/resume WAL (append jsonl LAST as source of truth; reconcile last_kept_commit from jsonl on resume). Diff guard AFTER agent edit, BEFORE commit — only system_prompt.md (tracked) may change; controller files gitignored.

      Agent (meta-agent): scripted LLM call (deepseek-chat). Reads results.tsv + digests → outputs prompt edit + description. Log IS its memory. No interactive tools.

      Implementation entrypoint

      Not maka-headless eval CLI (only wires fake backend). run-cell.mjs in container directly imports RuntimeRunner from @maka/runtime:

      import{RuntimeRunner,AiSdkBackend, ... }from'@maka/runtime';// agent runs in /app (Dockerfile WORKDIR), in-place// RuntimeRunner → InvocationResult → extract errorClass// write errorClass + trajectory to /output/

      maka-agent repo cloned + built at /opt/maka-agent (outside /app).

      Verified facts

      QuestionAnswerSource
      TB test.sh writes /app?Yes — test_outputs.py uses Path("/app/regex.txt") etc.grep 89 tasks
      runExperiment compatible?No — prepareWorkspace uses mkdtemp('/tmp/...'), can't be /appsandbox.ts:20
      systemPromptHash exists?No — only requestShapeHashgrep runtime/src
      pricing key format?provider:model (deepseek:deepseek-chat)builtin-pricing.ts:14
      Harbor + Docker verified?Yes — oracle reward 1.0, maka adapter runs, verifier workssmoke test 6/19
      file/bash split resolved?Yes — whole process in-containerarchitecture decision

      Directory layout

      ~/.local/maka-eval/
      ├── harbor-adapter/
      │ ├── maka_agent.py # Harbor agent adapter
      │ └── run-cell.mjs # in-container entrypoint (RuntimeRunner)
      ├── harness-rsi/
      │ ├── program.md # loop instructions (agent reads)
      │ ├── system_prompt.md # optimization target (agent edits ONLY)
      │ ├── controller.mjs # state machine (host-side)
      │ ├── results.jsonl # canonical truth (OUTSIDE agent cwd)
      │ ├── results.tsv # derived view for agent (IN agent cwd, gitignored)
      │ └── baseline/ # 3× baseline calibration
      ├── ds-key.txt
      └── fixtures/
      

      The loop

      INIT:
      1. Run baseline 3× on held-in + held-out
      2. Compute means + observed_spread
      3. noise_band = max(wilson_half_width(N, p) × √(4/3), observed_spread)
      4. held_in_reference = baseline_held_in_mean
      5. Save original_held_out_mean (never updated)
      LOOP:
      1. Write pending round record
      2. Agent: read results.tsv + digests → edit system_prompt.md
      3. Controller: diff guard (after edit, before commit) — only system_prompt.md
      4. Controller: git commit
      5. Controller: run held-in via Harbor → parse reward + errorClass per task
      6. Check infra — any infra → no keep, retry once, else discard
      7. Compute held_in_delta vs held_in_reference
      8. |delta| < noise_band → DISCARD (inconclusive)
      9. delta <= 0 → DISCARD (regression)
      10. delta > noise_band:
      a. Run held-out via Harbor
      b. Check held-out infra
      c. held_out_delta vs original_held_out_mean
      d. < -noise_band → DISCARD (regression)
      e. coverage degraded → DISCARD
      f. reward-hack scan → QUARANTINE if flagged
      g. >= -noise_band AND coverage OK → KEEP
      h. held_in_reference = max(held_in_reference, pass_rate − noise_band)
      11. Append to results.jsonl (source of truth)
      12. Advance last_kept_commit (if KEEP, AFTER jsonl append)
      13. Update results.tsv
      14. Check budget (sum costUsd, zero-cost guard)
      

      Success criteria (structural validation)

      • Loop runs unattended ≥10 iterations without human intervention
      • Zero keeps is passing — v1 validates structure, not prompt quality
      • If keeps occur: no held-out regression, no unreviewed reward-hack flags
      • Cost ≤ $30
      • Prompt round-trip hash verified every round

      Open questions

      1. Task selection: which 60 of 89 for held-in? Filter by speed/stability.
      2. Parallelism:--n-concurrent — determinism? Key results by task id.
      3. Simplicity criterion: deletion + non-inferior = KEEP? Needs own branch.
      4. Pricing wiring: verify in-container path registers pricing lookup.

      Non-goals (v1)

      • No tool description optimization
      • No control flow changes
      • No AEGIS 4-stage pipeline
      • No repeat sampling / pass@k
      • No runExperiment (RuntimeRunner + Harbor verifier instead)

      Metadata

      Metadata

      Assignees

      No one assigned

        Labels

        enhancementNew feature or request

        Type

        No type

        Projects

        No projects

          Milestone

          No milestone

          Relationships

          None yet

          Development

          No branches or pull requests

          Issue actions

          , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
          Skip to content

          RFC: Harness RSI loop — autonomous system prompt optimization #64

          Description

          @Astro-Han

          RFC: Harness RSI loop — autoresearch structure for system prompt optimization

          Context

          maka's eval framework is ready (PR #62/#63). The goal is closed-loop harness self-improvement: an agent autonomously optimizes its own system prompt against a benchmark task set, no human in the loop, runnable until stopped.

          Prior art: Karpathy's autoresearch (minimal loop, single agent), HarnessX/AEGIS (heavy 4-stage engine, +14.5%), Self-Harness (regression-gated edit loop).

          Architecture (verified)

          RuntimeRunner in-container + Harbor verifier. Harbor provides the Docker container (from TB task Dockerfile). Inside that container, a run-cell.mjs entrypoint runs maka's RuntimeRunner directly — the agent works in-place at /app (the Dockerfile WORKDIR), using file tools + Bash all in-container (one filesystem). After the agent finishes, Harbor runs tests/test.sh which checks /app state.

          Controller (host):
          1. For each held-in task: harbor run --path <tb-task> --agent-import-path maka_agent:MakaAgent
          --ae MAKA_SYSTEM_PROMPT=<prompt> --ae DEEPSEEK_API_KEY=<key>
          2. Parse Harbor result.json → per-task reward (0/1)
          3. Parse adapter's errorClass output → per-task errorClass
          Harbor container (per task):
          1. Build from task Dockerfile (WORKDIR /app, has tools, test env)
          2. MakaAgent.run() → adapter starts run-cell.mjs in /app
          → RuntimeRunner runs agent (file tools + Bash, all in /app, one filesystem)
          → RuntimeRunner records trajectory to runtime-events.jsonl
          → adapter extracts errorClass from InvocationResult
          3. Harbor runs tests/test.sh → checks /app state → reward 0 or 1
          

          Why not runExperiment? Verified: runExperiment copies fixture to mkdtemp('/tmp/maka-headless-ws-XXXX'), runs verifier there, deletes it. But TB test scripts write absolute paths to /app (e.g. Path("/app/regex.txt") in test_outputs.py). 89/89 Dockerfiles use WORKDIR /app. runExperiment's temp workspace would never be checked by TB tests. RuntimeRunner works in-place at /app — no copy, no sync problem.

          Why not Harbor verifier only? We need maka's errorClass (max_tokens / verification_failed / runtime_error) for signal enrichment. The adapter extracts this from RuntimeRunner's InvocationResult and writes it alongside Harbor's reward.

          File/bash filesystem split (v5 P0-2, resolved): the entire run-cell.mjs process runs in-container. RuntimeRunner's file tools (Write/Edit/Read) and Bash tool (toolExecutor) all operate on the container's filesystem. No host-container split.

          Prompt delivery: controller passes system_prompt.md content via MAKA_SYSTEM_PROMPT env var. run-cell.mjs reads it, passes to Config.systemPrompt. Runtime stamps a hash of the effectiveConfig.systemPrompt string (at backend constructor, after resolveSystemPrompt) into trajectory events. Controller verifies the hash round-trips — if trajectory hash ≠ hash of prompt sent, the round is invalid (plumbing failure).

          Implementation prerequisite:systemPromptHash does not exist in runtime yet (verified: only requestShapeHash exists). Must add a small instrumentation: hash the resolved Config.systemPrompt and write to a runtime event field. Pure addition, no existing code changes.

          Harbor/container failure ≠ benchmark failure. Container crash, image build failure, Harbor timeout → no result.json or invalid output. Controller treats this as infra error. Harbor reward + adapter errorClass only decide benchmark score.

          Signal quality

          Fix 1 — Coverage-aware gate.max_tokens/runtime_error tasks should count as fail (not drop from denominator). Gate uses pass / eligible (unscored = fail), NOT pass / scored.

          keep = held_in_pass_rate (pass/eligible) improved beyond noise_band_held_in
          AND held_in_coverage (scored/eligible) not degraded
          AND held_out_pass_rate non-inferior vs original baseline (>= -noise_band_held_out)
          AND held_out_coverage not degraded
          AND no infra_errors (any infra error → no keep)
          

          Controller computes pass/eligible itself. Harbor gives per-task reward (0/1); adapter gives per-task errorClass; controller combines: pass = reward==1, eligible = true (all tasks eligible unless infra error), scored = errorClass != 'max_tokens' && errorClass != 'runtime_error'.

          Fix 2 — Dual noise band.noise_band = wilson_half_width(N, p_mean) × √(1 + 1/n_baseline) — CI of the difference, not single-proportion CI. Separate bands for held-in and held-out.

          • Layer 0 (free): dual noise thresholds.
          • Layer 1 (cheap): 60 held-in + 20 held-out from TB 2.0's 89 tasks.
          • Layer 2 (deferred for v1): repeat sampling. v1 goal is structural validation.

          Fix 3 — Agent reads failing trajectories. Controller extracts per-task digest from runtime-events.jsonl (raw events): last 2 tool calls (from function_call events with args), errorClass, steps/duration, verifier stdout/stderr first 3 lines.

          Overfitting guard

          Held-in vs current kept baseline (monotonic improvement). After KEEP: held_in_reference = max(previous_reference, new_run_pass_rate − noise_band_held_in). Banks real improvements without enshrining lucky peak, never falls below prior state.

          Held-out vs original baseline (fixed floor). Floor = original_held_out_mean − noise_band_held_out. Prevents cumulative drift.

          Held-out physical isolation. Held-out task content, trajectories, results.jsonl live outside agent cwd. Agent sees only program.md, system_prompt.md, results.tsv, held-in digests.

          Reward-hack quarantine gate. Controller scans function_call.args in raw runtime-events.jsonl for verifier-specific patterns (expected-output strings, not filenames). Flagged → QUARANTINE (discard). Scanner fails closed if raw events missing.

          Infra error handling

          Any infra error → no KEEP. Infra = missing result.json, container crash, image build failure, Harbor timeout. Retry once, else discard+log.

          >20% infra → stop. Stop condition, not keep-allowance.

          Infra classification: model API failure (DeepSeek 429/500) → runtime_error (scored=false, counts in eligible). Container/Docker/Harbor failure → infra (excluded from scoring).

          Budget

          Controller sums costUsd from runtime-events.jsonl token-usage events. costUsd exists on RuntimeEvent. BUILTIN_PRICING keyed as deepseek:deepseek-chat (provider:model format) — controller must compose key as ${provider}:${config.model}.

          Zero-cost guard: if summed costUsd = 0 but tokens > 0 → plumbing failure (pricing not wired). Do NOT count as free round.

          Cost ceiling: $30 (default). 80 tasks × single-sample × ~10 rounds ≈ $15-25.

          Controller / agent separation

          Controller (host-side Node.js): Harbor invocation, result parsing, metric computation, budget, git state machine, keep/discard, crash/resume WAL (append jsonl LAST as source of truth; reconcile last_kept_commit from jsonl on resume). Diff guard AFTER agent edit, BEFORE commit — only system_prompt.md (tracked) may change; controller files gitignored.

          Agent (meta-agent): scripted LLM call (deepseek-chat). Reads results.tsv + digests → outputs prompt edit + description. Log IS its memory. No interactive tools.

          Implementation entrypoint

          Not maka-headless eval CLI (only wires fake backend). run-cell.mjs in container directly imports RuntimeRunner from @maka/runtime:

          import{RuntimeRunner,AiSdkBackend, ... }from'@maka/runtime';// agent runs in /app (Dockerfile WORKDIR), in-place// RuntimeRunner → InvocationResult → extract errorClass// write errorClass + trajectory to /output/

          maka-agent repo cloned + built at /opt/maka-agent (outside /app).

          Verified facts

          QuestionAnswerSource
          TB test.sh writes /app?Yes — test_outputs.py uses Path("/app/regex.txt") etc.grep 89 tasks
          runExperiment compatible?No — prepareWorkspace uses mkdtemp('/tmp/...'), can't be /appsandbox.ts:20
          systemPromptHash exists?No — only requestShapeHashgrep runtime/src
          pricing key format?provider:model (deepseek:deepseek-chat)builtin-pricing.ts:14
          Harbor + Docker verified?Yes — oracle reward 1.0, maka adapter runs, verifier workssmoke test 6/19
          file/bash split resolved?Yes — whole process in-containerarchitecture decision

          Directory layout

          ~/.local/maka-eval/
          ├── harbor-adapter/
          │ ├── maka_agent.py # Harbor agent adapter
          │ └── run-cell.mjs # in-container entrypoint (RuntimeRunner)
          ├── harness-rsi/
          │ ├── program.md # loop instructions (agent reads)
          │ ├── system_prompt.md # optimization target (agent edits ONLY)
          │ ├── controller.mjs # state machine (host-side)
          │ ├── results.jsonl # canonical truth (OUTSIDE agent cwd)
          │ ├── results.tsv # derived view for agent (IN agent cwd, gitignored)
          │ └── baseline/ # 3× baseline calibration
          ├── ds-key.txt
          └── fixtures/
          

          The loop

          INIT:
          1. Run baseline 3× on held-in + held-out
          2. Compute means + observed_spread
          3. noise_band = max(wilson_half_width(N, p) × √(4/3), observed_spread)
          4. held_in_reference = baseline_held_in_mean
          5. Save original_held_out_mean (never updated)
          LOOP:
          1. Write pending round record
          2. Agent: read results.tsv + digests → edit system_prompt.md
          3. Controller: diff guard (after edit, before commit) — only system_prompt.md
          4. Controller: git commit
          5. Controller: run held-in via Harbor → parse reward + errorClass per task
          6. Check infra — any infra → no keep, retry once, else discard
          7. Compute held_in_delta vs held_in_reference
          8. |delta| < noise_band → DISCARD (inconclusive)
          9. delta <= 0 → DISCARD (regression)
          10. delta > noise_band:
          a. Run held-out via Harbor
          b. Check held-out infra
          c. held_out_delta vs original_held_out_mean
          d. < -noise_band → DISCARD (regression)
          e. coverage degraded → DISCARD
          f. reward-hack scan → QUARANTINE if flagged
          g. >= -noise_band AND coverage OK → KEEP
          h. held_in_reference = max(held_in_reference, pass_rate − noise_band)
          11. Append to results.jsonl (source of truth)
          12. Advance last_kept_commit (if KEEP, AFTER jsonl append)
          13. Update results.tsv
          14. Check budget (sum costUsd, zero-cost guard)
          

          Success criteria (structural validation)

          • Loop runs unattended ≥10 iterations without human intervention
          • Zero keeps is passing — v1 validates structure, not prompt quality
          • If keeps occur: no held-out regression, no unreviewed reward-hack flags
          • Cost ≤ $30
          • Prompt round-trip hash verified every round

          Open questions

          1. Task selection: which 60 of 89 for held-in? Filter by speed/stability.
          2. Parallelism:--n-concurrent — determinism? Key results by task id.
          3. Simplicity criterion: deletion + non-inferior = KEEP? Needs own branch.
          4. Pricing wiring: verify in-container path registers pricing lookup.

          Non-goals (v1)

          • No tool description optimization
          • No control flow changes
          • No AEGIS 4-stage pipeline
          • No repeat sampling / pass@k
          • No runExperiment (RuntimeRunner + Harbor verifier instead)

          Metadata

          Metadata

          Assignees

          No one assigned

            Labels

            enhancementNew feature or request

            Type

            No type

            Projects

            No projects

              Milestone

              No milestone

              Relationships

              None yet

              Development

              No branches or pull requests

              Issue actions

              , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
              Skip to content

              RFC: Harness RSI loop — autonomous system prompt optimization #64

              Description

              @Astro-Han

              RFC: Harness RSI loop — autoresearch structure for system prompt optimization

              Context

              maka's eval framework is ready (PR #62/#63). The goal is closed-loop harness self-improvement: an agent autonomously optimizes its own system prompt against a benchmark task set, no human in the loop, runnable until stopped.

              Prior art: Karpathy's autoresearch (minimal loop, single agent), HarnessX/AEGIS (heavy 4-stage engine, +14.5%), Self-Harness (regression-gated edit loop).

              Architecture (verified)

              RuntimeRunner in-container + Harbor verifier. Harbor provides the Docker container (from TB task Dockerfile). Inside that container, a run-cell.mjs entrypoint runs maka's RuntimeRunner directly — the agent works in-place at /app (the Dockerfile WORKDIR), using file tools + Bash all in-container (one filesystem). After the agent finishes, Harbor runs tests/test.sh which checks /app state.

              Controller (host):
              1. For each held-in task: harbor run --path <tb-task> --agent-import-path maka_agent:MakaAgent
              --ae MAKA_SYSTEM_PROMPT=<prompt> --ae DEEPSEEK_API_KEY=<key>
              2. Parse Harbor result.json → per-task reward (0/1)
              3. Parse adapter's errorClass output → per-task errorClass
              Harbor container (per task):
              1. Build from task Dockerfile (WORKDIR /app, has tools, test env)
              2. MakaAgent.run() → adapter starts run-cell.mjs in /app
              → RuntimeRunner runs agent (file tools + Bash, all in /app, one filesystem)
              → RuntimeRunner records trajectory to runtime-events.jsonl
              → adapter extracts errorClass from InvocationResult
              3. Harbor runs tests/test.sh → checks /app state → reward 0 or 1
              

              Why not runExperiment? Verified: runExperiment copies fixture to mkdtemp('/tmp/maka-headless-ws-XXXX'), runs verifier there, deletes it. But TB test scripts write absolute paths to /app (e.g. Path("/app/regex.txt") in test_outputs.py). 89/89 Dockerfiles use WORKDIR /app. runExperiment's temp workspace would never be checked by TB tests. RuntimeRunner works in-place at /app — no copy, no sync problem.

              Why not Harbor verifier only? We need maka's errorClass (max_tokens / verification_failed / runtime_error) for signal enrichment. The adapter extracts this from RuntimeRunner's InvocationResult and writes it alongside Harbor's reward.

              File/bash filesystem split (v5 P0-2, resolved): the entire run-cell.mjs process runs in-container. RuntimeRunner's file tools (Write/Edit/Read) and Bash tool (toolExecutor) all operate on the container's filesystem. No host-container split.

              Prompt delivery: controller passes system_prompt.md content via MAKA_SYSTEM_PROMPT env var. run-cell.mjs reads it, passes to Config.systemPrompt. Runtime stamps a hash of the effectiveConfig.systemPrompt string (at backend constructor, after resolveSystemPrompt) into trajectory events. Controller verifies the hash round-trips — if trajectory hash ≠ hash of prompt sent, the round is invalid (plumbing failure).

              Implementation prerequisite:systemPromptHash does not exist in runtime yet (verified: only requestShapeHash exists). Must add a small instrumentation: hash the resolved Config.systemPrompt and write to a runtime event field. Pure addition, no existing code changes.

              Harbor/container failure ≠ benchmark failure. Container crash, image build failure, Harbor timeout → no result.json or invalid output. Controller treats this as infra error. Harbor reward + adapter errorClass only decide benchmark score.

              Signal quality

              Fix 1 — Coverage-aware gate.max_tokens/runtime_error tasks should count as fail (not drop from denominator). Gate uses pass / eligible (unscored = fail), NOT pass / scored.

              keep = held_in_pass_rate (pass/eligible) improved beyond noise_band_held_in
              AND held_in_coverage (scored/eligible) not degraded
              AND held_out_pass_rate non-inferior vs original baseline (>= -noise_band_held_out)
              AND held_out_coverage not degraded
              AND no infra_errors (any infra error → no keep)
              

              Controller computes pass/eligible itself. Harbor gives per-task reward (0/1); adapter gives per-task errorClass; controller combines: pass = reward==1, eligible = true (all tasks eligible unless infra error), scored = errorClass != 'max_tokens' && errorClass != 'runtime_error'.

              Fix 2 — Dual noise band.noise_band = wilson_half_width(N, p_mean) × √(1 + 1/n_baseline) — CI of the difference, not single-proportion CI. Separate bands for held-in and held-out.

              • Layer 0 (free): dual noise thresholds.
              • Layer 1 (cheap): 60 held-in + 20 held-out from TB 2.0's 89 tasks.
              • Layer 2 (deferred for v1): repeat sampling. v1 goal is structural validation.

              Fix 3 — Agent reads failing trajectories. Controller extracts per-task digest from runtime-events.jsonl (raw events): last 2 tool calls (from function_call events with args), errorClass, steps/duration, verifier stdout/stderr first 3 lines.

              Overfitting guard

              Held-in vs current kept baseline (monotonic improvement). After KEEP: held_in_reference = max(previous_reference, new_run_pass_rate − noise_band_held_in). Banks real improvements without enshrining lucky peak, never falls below prior state.

              Held-out vs original baseline (fixed floor). Floor = original_held_out_mean − noise_band_held_out. Prevents cumulative drift.

              Held-out physical isolation. Held-out task content, trajectories, results.jsonl live outside agent cwd. Agent sees only program.md, system_prompt.md, results.tsv, held-in digests.

              Reward-hack quarantine gate. Controller scans function_call.args in raw runtime-events.jsonl for verifier-specific patterns (expected-output strings, not filenames). Flagged → QUARANTINE (discard). Scanner fails closed if raw events missing.

              Infra error handling

              Any infra error → no KEEP. Infra = missing result.json, container crash, image build failure, Harbor timeout. Retry once, else discard+log.

              >20% infra → stop. Stop condition, not keep-allowance.

              Infra classification: model API failure (DeepSeek 429/500) → runtime_error (scored=false, counts in eligible). Container/Docker/Harbor failure → infra (excluded from scoring).

              Budget

              Controller sums costUsd from runtime-events.jsonl token-usage events. costUsd exists on RuntimeEvent. BUILTIN_PRICING keyed as deepseek:deepseek-chat (provider:model format) — controller must compose key as ${provider}:${config.model}.

              Zero-cost guard: if summed costUsd = 0 but tokens > 0 → plumbing failure (pricing not wired). Do NOT count as free round.

              Cost ceiling: $30 (default). 80 tasks × single-sample × ~10 rounds ≈ $15-25.

              Controller / agent separation

              Controller (host-side Node.js): Harbor invocation, result parsing, metric computation, budget, git state machine, keep/discard, crash/resume WAL (append jsonl LAST as source of truth; reconcile last_kept_commit from jsonl on resume). Diff guard AFTER agent edit, BEFORE commit — only system_prompt.md (tracked) may change; controller files gitignored.

              Agent (meta-agent): scripted LLM call (deepseek-chat). Reads results.tsv + digests → outputs prompt edit + description. Log IS its memory. No interactive tools.

              Implementation entrypoint

              Not maka-headless eval CLI (only wires fake backend). run-cell.mjs in container directly imports RuntimeRunner from @maka/runtime:

              import{RuntimeRunner,AiSdkBackend, ... }from'@maka/runtime';// agent runs in /app (Dockerfile WORKDIR), in-place// RuntimeRunner → InvocationResult → extract errorClass// write errorClass + trajectory to /output/

              maka-agent repo cloned + built at /opt/maka-agent (outside /app).

              Verified facts

              QuestionAnswerSource
              TB test.sh writes /app?Yes — test_outputs.py uses Path("/app/regex.txt") etc.grep 89 tasks
              runExperiment compatible?No — prepareWorkspace uses mkdtemp('/tmp/...'), can't be /appsandbox.ts:20
              systemPromptHash exists?No — only requestShapeHashgrep runtime/src
              pricing key format?provider:model (deepseek:deepseek-chat)builtin-pricing.ts:14
              Harbor + Docker verified?Yes — oracle reward 1.0, maka adapter runs, verifier workssmoke test 6/19
              file/bash split resolved?Yes — whole process in-containerarchitecture decision

              Directory layout

              ~/.local/maka-eval/
              ├── harbor-adapter/
              │ ├── maka_agent.py # Harbor agent adapter
              │ └── run-cell.mjs # in-container entrypoint (RuntimeRunner)
              ├── harness-rsi/
              │ ├── program.md # loop instructions (agent reads)
              │ ├── system_prompt.md # optimization target (agent edits ONLY)
              │ ├── controller.mjs # state machine (host-side)
              │ ├── results.jsonl # canonical truth (OUTSIDE agent cwd)
              │ ├── results.tsv # derived view for agent (IN agent cwd, gitignored)
              │ └── baseline/ # 3× baseline calibration
              ├── ds-key.txt
              └── fixtures/
              

              The loop

              INIT:
              1. Run baseline 3× on held-in + held-out
              2. Compute means + observed_spread
              3. noise_band = max(wilson_half_width(N, p) × √(4/3), observed_spread)
              4. held_in_reference = baseline_held_in_mean
              5. Save original_held_out_mean (never updated)
              LOOP:
              1. Write pending round record
              2. Agent: read results.tsv + digests → edit system_prompt.md
              3. Controller: diff guard (after edit, before commit) — only system_prompt.md
              4. Controller: git commit
              5. Controller: run held-in via Harbor → parse reward + errorClass per task
              6. Check infra — any infra → no keep, retry once, else discard
              7. Compute held_in_delta vs held_in_reference
              8. |delta| < noise_band → DISCARD (inconclusive)
              9. delta <= 0 → DISCARD (regression)
              10. delta > noise_band:
              a. Run held-out via Harbor
              b. Check held-out infra
              c. held_out_delta vs original_held_out_mean
              d. < -noise_band → DISCARD (regression)
              e. coverage degraded → DISCARD
              f. reward-hack scan → QUARANTINE if flagged
              g. >= -noise_band AND coverage OK → KEEP
              h. held_in_reference = max(held_in_reference, pass_rate − noise_band)
              11. Append to results.jsonl (source of truth)
              12. Advance last_kept_commit (if KEEP, AFTER jsonl append)
              13. Update results.tsv
              14. Check budget (sum costUsd, zero-cost guard)
              

              Success criteria (structural validation)

              • Loop runs unattended ≥10 iterations without human intervention
              • Zero keeps is passing — v1 validates structure, not prompt quality
              • If keeps occur: no held-out regression, no unreviewed reward-hack flags
              • Cost ≤ $30
              • Prompt round-trip hash verified every round

              Open questions

              1. Task selection: which 60 of 89 for held-in? Filter by speed/stability.
              2. Parallelism:--n-concurrent — determinism? Key results by task id.
              3. Simplicity criterion: deletion + non-inferior = KEEP? Needs own branch.
              4. Pricing wiring: verify in-container path registers pricing lookup.

              Non-goals (v1)

              • No tool description optimization
              • No control flow changes
              • No AEGIS 4-stage pipeline
              • No repeat sampling / pass@k
              • No runExperiment (RuntimeRunner + Harbor verifier instead)

              Metadata

              Metadata

              Assignees

              No one assigned

                Labels

                enhancementNew feature or request

                Type

                No type

                Projects

                No projects

                  Milestone

                  No milestone

                  Relationships

                  None yet

                  Development

                  No branches or pull requests

                  Issue actions

                  , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
                  Skip to content

                  RFC: Harness RSI loop — autonomous system prompt optimization #64

                  Description

                  @Astro-Han

                  RFC: Harness RSI loop — autoresearch structure for system prompt optimization

                  Context

                  maka's eval framework is ready (PR #62/#63). The goal is closed-loop harness self-improvement: an agent autonomously optimizes its own system prompt against a benchmark task set, no human in the loop, runnable until stopped.

                  Prior art: Karpathy's autoresearch (minimal loop, single agent), HarnessX/AEGIS (heavy 4-stage engine, +14.5%), Self-Harness (regression-gated edit loop).

                  Architecture (verified)

                  RuntimeRunner in-container + Harbor verifier. Harbor provides the Docker container (from TB task Dockerfile). Inside that container, a run-cell.mjs entrypoint runs maka's RuntimeRunner directly — the agent works in-place at /app (the Dockerfile WORKDIR), using file tools + Bash all in-container (one filesystem). After the agent finishes, Harbor runs tests/test.sh which checks /app state.

                  Controller (host):
                  1. For each held-in task: harbor run --path <tb-task> --agent-import-path maka_agent:MakaAgent
                  --ae MAKA_SYSTEM_PROMPT=<prompt> --ae DEEPSEEK_API_KEY=<key>
                  2. Parse Harbor result.json → per-task reward (0/1)
                  3. Parse adapter's errorClass output → per-task errorClass
                  Harbor container (per task):
                  1. Build from task Dockerfile (WORKDIR /app, has tools, test env)
                  2. MakaAgent.run() → adapter starts run-cell.mjs in /app
                  → RuntimeRunner runs agent (file tools + Bash, all in /app, one filesystem)
                  → RuntimeRunner records trajectory to runtime-events.jsonl
                  → adapter extracts errorClass from InvocationResult
                  3. Harbor runs tests/test.sh → checks /app state → reward 0 or 1
                  

                  Why not runExperiment? Verified: runExperiment copies fixture to mkdtemp('/tmp/maka-headless-ws-XXXX'), runs verifier there, deletes it. But TB test scripts write absolute paths to /app (e.g. Path("/app/regex.txt") in test_outputs.py). 89/89 Dockerfiles use WORKDIR /app. runExperiment's temp workspace would never be checked by TB tests. RuntimeRunner works in-place at /app — no copy, no sync problem.

                  Why not Harbor verifier only? We need maka's errorClass (max_tokens / verification_failed / runtime_error) for signal enrichment. The adapter extracts this from RuntimeRunner's InvocationResult and writes it alongside Harbor's reward.

                  File/bash filesystem split (v5 P0-2, resolved): the entire run-cell.mjs process runs in-container. RuntimeRunner's file tools (Write/Edit/Read) and Bash tool (toolExecutor) all operate on the container's filesystem. No host-container split.

                  Prompt delivery: controller passes system_prompt.md content via MAKA_SYSTEM_PROMPT env var. run-cell.mjs reads it, passes to Config.systemPrompt. Runtime stamps a hash of the effectiveConfig.systemPrompt string (at backend constructor, after resolveSystemPrompt) into trajectory events. Controller verifies the hash round-trips — if trajectory hash ≠ hash of prompt sent, the round is invalid (plumbing failure).

                  Implementation prerequisite:systemPromptHash does not exist in runtime yet (verified: only requestShapeHash exists). Must add a small instrumentation: hash the resolved Config.systemPrompt and write to a runtime event field. Pure addition, no existing code changes.

                  Harbor/container failure ≠ benchmark failure. Container crash, image build failure, Harbor timeout → no result.json or invalid output. Controller treats this as infra error. Harbor reward + adapter errorClass only decide benchmark score.

                  Signal quality

                  Fix 1 — Coverage-aware gate.max_tokens/runtime_error tasks should count as fail (not drop from denominator). Gate uses pass / eligible (unscored = fail), NOT pass / scored.

                  keep = held_in_pass_rate (pass/eligible) improved beyond noise_band_held_in
                  AND held_in_coverage (scored/eligible) not degraded
                  AND held_out_pass_rate non-inferior vs original baseline (>= -noise_band_held_out)
                  AND held_out_coverage not degraded
                  AND no infra_errors (any infra error → no keep)
                  

                  Controller computes pass/eligible itself. Harbor gives per-task reward (0/1); adapter gives per-task errorClass; controller combines: pass = reward==1, eligible = true (all tasks eligible unless infra error), scored = errorClass != 'max_tokens' && errorClass != 'runtime_error'.

                  Fix 2 — Dual noise band.noise_band = wilson_half_width(N, p_mean) × √(1 + 1/n_baseline) — CI of the difference, not single-proportion CI. Separate bands for held-in and held-out.

                  • Layer 0 (free): dual noise thresholds.
                  • Layer 1 (cheap): 60 held-in + 20 held-out from TB 2.0's 89 tasks.
                  • Layer 2 (deferred for v1): repeat sampling. v1 goal is structural validation.

                  Fix 3 — Agent reads failing trajectories. Controller extracts per-task digest from runtime-events.jsonl (raw events): last 2 tool calls (from function_call events with args), errorClass, steps/duration, verifier stdout/stderr first 3 lines.

                  Overfitting guard

                  Held-in vs current kept baseline (monotonic improvement). After KEEP: held_in_reference = max(previous_reference, new_run_pass_rate − noise_band_held_in). Banks real improvements without enshrining lucky peak, never falls below prior state.

                  Held-out vs original baseline (fixed floor). Floor = original_held_out_mean − noise_band_held_out. Prevents cumulative drift.

                  Held-out physical isolation. Held-out task content, trajectories, results.jsonl live outside agent cwd. Agent sees only program.md, system_prompt.md, results.tsv, held-in digests.

                  Reward-hack quarantine gate. Controller scans function_call.args in raw runtime-events.jsonl for verifier-specific patterns (expected-output strings, not filenames). Flagged → QUARANTINE (discard). Scanner fails closed if raw events missing.

                  Infra error handling

                  Any infra error → no KEEP. Infra = missing result.json, container crash, image build failure, Harbor timeout. Retry once, else discard+log.

                  >20% infra → stop. Stop condition, not keep-allowance.

                  Infra classification: model API failure (DeepSeek 429/500) → runtime_error (scored=false, counts in eligible). Container/Docker/Harbor failure → infra (excluded from scoring).

                  Budget

                  Controller sums costUsd from runtime-events.jsonl token-usage events. costUsd exists on RuntimeEvent. BUILTIN_PRICING keyed as deepseek:deepseek-chat (provider:model format) — controller must compose key as ${provider}:${config.model}.

                  Zero-cost guard: if summed costUsd = 0 but tokens > 0 → plumbing failure (pricing not wired). Do NOT count as free round.

                  Cost ceiling: $30 (default). 80 tasks × single-sample × ~10 rounds ≈ $15-25.

                  Controller / agent separation

                  Controller (host-side Node.js): Harbor invocation, result parsing, metric computation, budget, git state machine, keep/discard, crash/resume WAL (append jsonl LAST as source of truth; reconcile last_kept_commit from jsonl on resume). Diff guard AFTER agent edit, BEFORE commit — only system_prompt.md (tracked) may change; controller files gitignored.

                  Agent (meta-agent): scripted LLM call (deepseek-chat). Reads results.tsv + digests → outputs prompt edit + description. Log IS its memory. No interactive tools.

                  Implementation entrypoint

                  Not maka-headless eval CLI (only wires fake backend). run-cell.mjs in container directly imports RuntimeRunner from @maka/runtime:

                  import{RuntimeRunner,AiSdkBackend, ... }from'@maka/runtime';// agent runs in /app (Dockerfile WORKDIR), in-place// RuntimeRunner → InvocationResult → extract errorClass// write errorClass + trajectory to /output/

                  maka-agent repo cloned + built at /opt/maka-agent (outside /app).

                  Verified facts

                  QuestionAnswerSource
                  TB test.sh writes /app?Yes — test_outputs.py uses Path("/app/regex.txt") etc.grep 89 tasks
                  runExperiment compatible?No — prepareWorkspace uses mkdtemp('/tmp/...'), can't be /appsandbox.ts:20
                  systemPromptHash exists?No — only requestShapeHashgrep runtime/src
                  pricing key format?provider:model (deepseek:deepseek-chat)builtin-pricing.ts:14
                  Harbor + Docker verified?Yes — oracle reward 1.0, maka adapter runs, verifier workssmoke test 6/19
                  file/bash split resolved?Yes — whole process in-containerarchitecture decision

                  Directory layout

                  ~/.local/maka-eval/
                  ├── harbor-adapter/
                  │ ├── maka_agent.py # Harbor agent adapter
                  │ └── run-cell.mjs # in-container entrypoint (RuntimeRunner)
                  ├── harness-rsi/
                  │ ├── program.md # loop instructions (agent reads)
                  │ ├── system_prompt.md # optimization target (agent edits ONLY)
                  │ ├── controller.mjs # state machine (host-side)
                  │ ├── results.jsonl # canonical truth (OUTSIDE agent cwd)
                  │ ├── results.tsv # derived view for agent (IN agent cwd, gitignored)
                  │ └── baseline/ # 3× baseline calibration
                  ├── ds-key.txt
                  └── fixtures/
                  

                  The loop

                  INIT:
                  1. Run baseline 3× on held-in + held-out
                  2. Compute means + observed_spread
                  3. noise_band = max(wilson_half_width(N, p) × √(4/3), observed_spread)
                  4. held_in_reference = baseline_held_in_mean
                  5. Save original_held_out_mean (never updated)
                  LOOP:
                  1. Write pending round record
                  2. Agent: read results.tsv + digests → edit system_prompt.md
                  3. Controller: diff guard (after edit, before commit) — only system_prompt.md
                  4. Controller: git commit
                  5. Controller: run held-in via Harbor → parse reward + errorClass per task
                  6. Check infra — any infra → no keep, retry once, else discard
                  7. Compute held_in_delta vs held_in_reference
                  8. |delta| < noise_band → DISCARD (inconclusive)
                  9. delta <= 0 → DISCARD (regression)
                  10. delta > noise_band:
                  a. Run held-out via Harbor
                  b. Check held-out infra
                  c. held_out_delta vs original_held_out_mean
                  d. < -noise_band → DISCARD (regression)
                  e. coverage degraded → DISCARD
                  f. reward-hack scan → QUARANTINE if flagged
                  g. >= -noise_band AND coverage OK → KEEP
                  h. held_in_reference = max(held_in_reference, pass_rate − noise_band)
                  11. Append to results.jsonl (source of truth)
                  12. Advance last_kept_commit (if KEEP, AFTER jsonl append)
                  13. Update results.tsv
                  14. Check budget (sum costUsd, zero-cost guard)
                  

                  Success criteria (structural validation)

                  • Loop runs unattended ≥10 iterations without human intervention
                  • Zero keeps is passing — v1 validates structure, not prompt quality
                  • If keeps occur: no held-out regression, no unreviewed reward-hack flags
                  • Cost ≤ $30
                  • Prompt round-trip hash verified every round

                  Open questions

                  1. Task selection: which 60 of 89 for held-in? Filter by speed/stability.
                  2. Parallelism:--n-concurrent — determinism? Key results by task id.
                  3. Simplicity criterion: deletion + non-inferior = KEEP? Needs own branch.
                  4. Pricing wiring: verify in-container path registers pricing lookup.

                  Non-goals (v1)

                  • No tool description optimization
                  • No control flow changes
                  • No AEGIS 4-stage pipeline
                  • No repeat sampling / pass@k
                  • No runExperiment (RuntimeRunner + Harbor verifier instead)

                  Metadata

                  Metadata

                  Assignees

                  No one assigned

                    Labels

                    enhancementNew feature or request

                    Type

                    No type

                    Projects

                    No projects

                      Milestone

                      No milestone

                      Relationships

                      None yet

                      Development

                      No branches or pull requests

                      Issue actions

                      , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
                      Skip to content

                      RFC: Harness RSI loop — autonomous system prompt optimization #64

                      Description

                      @Astro-Han

                      RFC: Harness RSI loop — autoresearch structure for system prompt optimization

                      Context

                      maka's eval framework is ready (PR #62/#63). The goal is closed-loop harness self-improvement: an agent autonomously optimizes its own system prompt against a benchmark task set, no human in the loop, runnable until stopped.

                      Prior art: Karpathy's autoresearch (minimal loop, single agent), HarnessX/AEGIS (heavy 4-stage engine, +14.5%), Self-Harness (regression-gated edit loop).

                      Architecture (verified)

                      RuntimeRunner in-container + Harbor verifier. Harbor provides the Docker container (from TB task Dockerfile). Inside that container, a run-cell.mjs entrypoint runs maka's RuntimeRunner directly — the agent works in-place at /app (the Dockerfile WORKDIR), using file tools + Bash all in-container (one filesystem). After the agent finishes, Harbor runs tests/test.sh which checks /app state.

                      Controller (host):
                      1. For each held-in task: harbor run --path <tb-task> --agent-import-path maka_agent:MakaAgent
                      --ae MAKA_SYSTEM_PROMPT=<prompt> --ae DEEPSEEK_API_KEY=<key>
                      2. Parse Harbor result.json → per-task reward (0/1)
                      3. Parse adapter's errorClass output → per-task errorClass
                      Harbor container (per task):
                      1. Build from task Dockerfile (WORKDIR /app, has tools, test env)
                      2. MakaAgent.run() → adapter starts run-cell.mjs in /app
                      → RuntimeRunner runs agent (file tools + Bash, all in /app, one filesystem)
                      → RuntimeRunner records trajectory to runtime-events.jsonl
                      → adapter extracts errorClass from InvocationResult
                      3. Harbor runs tests/test.sh → checks /app state → reward 0 or 1
                      

                      Why not runExperiment? Verified: runExperiment copies fixture to mkdtemp('/tmp/maka-headless-ws-XXXX'), runs verifier there, deletes it. But TB test scripts write absolute paths to /app (e.g. Path("/app/regex.txt") in test_outputs.py). 89/89 Dockerfiles use WORKDIR /app. runExperiment's temp workspace would never be checked by TB tests. RuntimeRunner works in-place at /app — no copy, no sync problem.

                      Why not Harbor verifier only? We need maka's errorClass (max_tokens / verification_failed / runtime_error) for signal enrichment. The adapter extracts this from RuntimeRunner's InvocationResult and writes it alongside Harbor's reward.

                      File/bash filesystem split (v5 P0-2, resolved): the entire run-cell.mjs process runs in-container. RuntimeRunner's file tools (Write/Edit/Read) and Bash tool (toolExecutor) all operate on the container's filesystem. No host-container split.

                      Prompt delivery: controller passes system_prompt.md content via MAKA_SYSTEM_PROMPT env var. run-cell.mjs reads it, passes to Config.systemPrompt. Runtime stamps a hash of the effectiveConfig.systemPrompt string (at backend constructor, after resolveSystemPrompt) into trajectory events. Controller verifies the hash round-trips — if trajectory hash ≠ hash of prompt sent, the round is invalid (plumbing failure).

                      Implementation prerequisite:systemPromptHash does not exist in runtime yet (verified: only requestShapeHash exists). Must add a small instrumentation: hash the resolved Config.systemPrompt and write to a runtime event field. Pure addition, no existing code changes.

                      Harbor/container failure ≠ benchmark failure. Container crash, image build failure, Harbor timeout → no result.json or invalid output. Controller treats this as infra error. Harbor reward + adapter errorClass only decide benchmark score.

                      Signal quality

                      Fix 1 — Coverage-aware gate.max_tokens/runtime_error tasks should count as fail (not drop from denominator). Gate uses pass / eligible (unscored = fail), NOT pass / scored.

                      keep = held_in_pass_rate (pass/eligible) improved beyond noise_band_held_in
                      AND held_in_coverage (scored/eligible) not degraded
                      AND held_out_pass_rate non-inferior vs original baseline (>= -noise_band_held_out)
                      AND held_out_coverage not degraded
                      AND no infra_errors (any infra error → no keep)
                      

                      Controller computes pass/eligible itself. Harbor gives per-task reward (0/1); adapter gives per-task errorClass; controller combines: pass = reward==1, eligible = true (all tasks eligible unless infra error), scored = errorClass != 'max_tokens' && errorClass != 'runtime_error'.

                      Fix 2 — Dual noise band.noise_band = wilson_half_width(N, p_mean) × √(1 + 1/n_baseline) — CI of the difference, not single-proportion CI. Separate bands for held-in and held-out.

                      • Layer 0 (free): dual noise thresholds.
                      • Layer 1 (cheap): 60 held-in + 20 held-out from TB 2.0's 89 tasks.
                      • Layer 2 (deferred for v1): repeat sampling. v1 goal is structural validation.

                      Fix 3 — Agent reads failing trajectories. Controller extracts per-task digest from runtime-events.jsonl (raw events): last 2 tool calls (from function_call events with args), errorClass, steps/duration, verifier stdout/stderr first 3 lines.

                      Overfitting guard

                      Held-in vs current kept baseline (monotonic improvement). After KEEP: held_in_reference = max(previous_reference, new_run_pass_rate − noise_band_held_in). Banks real improvements without enshrining lucky peak, never falls below prior state.

                      Held-out vs original baseline (fixed floor). Floor = original_held_out_mean − noise_band_held_out. Prevents cumulative drift.

                      Held-out physical isolation. Held-out task content, trajectories, results.jsonl live outside agent cwd. Agent sees only program.md, system_prompt.md, results.tsv, held-in digests.

                      Reward-hack quarantine gate. Controller scans function_call.args in raw runtime-events.jsonl for verifier-specific patterns (expected-output strings, not filenames). Flagged → QUARANTINE (discard). Scanner fails closed if raw events missing.

                      Infra error handling

                      Any infra error → no KEEP. Infra = missing result.json, container crash, image build failure, Harbor timeout. Retry once, else discard+log.

                      >20% infra → stop. Stop condition, not keep-allowance.

                      Infra classification: model API failure (DeepSeek 429/500) → runtime_error (scored=false, counts in eligible). Container/Docker/Harbor failure → infra (excluded from scoring).

                      Budget

                      Controller sums costUsd from runtime-events.jsonl token-usage events. costUsd exists on RuntimeEvent. BUILTIN_PRICING keyed as deepseek:deepseek-chat (provider:model format) — controller must compose key as ${provider}:${config.model}.

                      Zero-cost guard: if summed costUsd = 0 but tokens > 0 → plumbing failure (pricing not wired). Do NOT count as free round.

                      Cost ceiling: $30 (default). 80 tasks × single-sample × ~10 rounds ≈ $15-25.

                      Controller / agent separation

                      Controller (host-side Node.js): Harbor invocation, result parsing, metric computation, budget, git state machine, keep/discard, crash/resume WAL (append jsonl LAST as source of truth; reconcile last_kept_commit from jsonl on resume). Diff guard AFTER agent edit, BEFORE commit — only system_prompt.md (tracked) may change; controller files gitignored.

                      Agent (meta-agent): scripted LLM call (deepseek-chat). Reads results.tsv + digests → outputs prompt edit + description. Log IS its memory. No interactive tools.

                      Implementation entrypoint

                      Not maka-headless eval CLI (only wires fake backend). run-cell.mjs in container directly imports RuntimeRunner from @maka/runtime:

                      import{RuntimeRunner,AiSdkBackend, ... }from'@maka/runtime';// agent runs in /app (Dockerfile WORKDIR), in-place// RuntimeRunner → InvocationResult → extract errorClass// write errorClass + trajectory to /output/

                      maka-agent repo cloned + built at /opt/maka-agent (outside /app).

                      Verified facts

                      QuestionAnswerSource
                      TB test.sh writes /app?Yes — test_outputs.py uses Path("/app/regex.txt") etc.grep 89 tasks
                      runExperiment compatible?No — prepareWorkspace uses mkdtemp('/tmp/...'), can't be /appsandbox.ts:20
                      systemPromptHash exists?No — only requestShapeHashgrep runtime/src
                      pricing key format?provider:model (deepseek:deepseek-chat)builtin-pricing.ts:14
                      Harbor + Docker verified?Yes — oracle reward 1.0, maka adapter runs, verifier workssmoke test 6/19
                      file/bash split resolved?Yes — whole process in-containerarchitecture decision

                      Directory layout

                      ~/.local/maka-eval/
                      ├── harbor-adapter/
                      │ ├── maka_agent.py # Harbor agent adapter
                      │ └── run-cell.mjs # in-container entrypoint (RuntimeRunner)
                      ├── harness-rsi/
                      │ ├── program.md # loop instructions (agent reads)
                      │ ├── system_prompt.md # optimization target (agent edits ONLY)
                      │ ├── controller.mjs # state machine (host-side)
                      │ ├── results.jsonl # canonical truth (OUTSIDE agent cwd)
                      │ ├── results.tsv # derived view for agent (IN agent cwd, gitignored)
                      │ └── baseline/ # 3× baseline calibration
                      ├── ds-key.txt
                      └── fixtures/
                      

                      The loop

                      INIT:
                      1. Run baseline 3× on held-in + held-out
                      2. Compute means + observed_spread
                      3. noise_band = max(wilson_half_width(N, p) × √(4/3), observed_spread)
                      4. held_in_reference = baseline_held_in_mean
                      5. Save original_held_out_mean (never updated)
                      LOOP:
                      1. Write pending round record
                      2. Agent: read results.tsv + digests → edit system_prompt.md
                      3. Controller: diff guard (after edit, before commit) — only system_prompt.md
                      4. Controller: git commit
                      5. Controller: run held-in via Harbor → parse reward + errorClass per task
                      6. Check infra — any infra → no keep, retry once, else discard
                      7. Compute held_in_delta vs held_in_reference
                      8. |delta| < noise_band → DISCARD (inconclusive)
                      9. delta <= 0 → DISCARD (regression)
                      10. delta > noise_band:
                      a. Run held-out via Harbor
                      b. Check held-out infra
                      c. held_out_delta vs original_held_out_mean
                      d. < -noise_band → DISCARD (regression)
                      e. coverage degraded → DISCARD
                      f. reward-hack scan → QUARANTINE if flagged
                      g. >= -noise_band AND coverage OK → KEEP
                      h. held_in_reference = max(held_in_reference, pass_rate − noise_band)
                      11. Append to results.jsonl (source of truth)
                      12. Advance last_kept_commit (if KEEP, AFTER jsonl append)
                      13. Update results.tsv
                      14. Check budget (sum costUsd, zero-cost guard)
                      

                      Success criteria (structural validation)

                      • Loop runs unattended ≥10 iterations without human intervention
                      • Zero keeps is passing — v1 validates structure, not prompt quality
                      • If keeps occur: no held-out regression, no unreviewed reward-hack flags
                      • Cost ≤ $30
                      • Prompt round-trip hash verified every round

                      Open questions

                      1. Task selection: which 60 of 89 for held-in? Filter by speed/stability.
                      2. Parallelism:--n-concurrent — determinism? Key results by task id.
                      3. Simplicity criterion: deletion + non-inferior = KEEP? Needs own branch.
                      4. Pricing wiring: verify in-container path registers pricing lookup.

                      Non-goals (v1)

                      • No tool description optimization
                      • No control flow changes
                      • No AEGIS 4-stage pipeline
                      • No repeat sampling / pass@k
                      • No runExperiment (RuntimeRunner + Harbor verifier instead)

                      Metadata

                      Metadata

                      Assignees

                      No one assigned

                        Labels

                        enhancementNew feature or request

                        Type

                        No type

                        Projects

                        No projects

                          Milestone

                          No milestone

                          Relationships

                          None yet

                          Development

                          No branches or pull requests

                          Issue actions

                          , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
                          Skip to content

                          RFC: Harness RSI loop — autonomous system prompt optimization #64

                          Description

                          @Astro-Han

                          RFC: Harness RSI loop — autoresearch structure for system prompt optimization

                          Context

                          maka's eval framework is ready (PR #62/#63). The goal is closed-loop harness self-improvement: an agent autonomously optimizes its own system prompt against a benchmark task set, no human in the loop, runnable until stopped.

                          Prior art: Karpathy's autoresearch (minimal loop, single agent), HarnessX/AEGIS (heavy 4-stage engine, +14.5%), Self-Harness (regression-gated edit loop).

                          Architecture (verified)

                          RuntimeRunner in-container + Harbor verifier. Harbor provides the Docker container (from TB task Dockerfile). Inside that container, a run-cell.mjs entrypoint runs maka's RuntimeRunner directly — the agent works in-place at /app (the Dockerfile WORKDIR), using file tools + Bash all in-container (one filesystem). After the agent finishes, Harbor runs tests/test.sh which checks /app state.

                          Controller (host):
                          1. For each held-in task: harbor run --path <tb-task> --agent-import-path maka_agent:MakaAgent
                          --ae MAKA_SYSTEM_PROMPT=<prompt> --ae DEEPSEEK_API_KEY=<key>
                          2. Parse Harbor result.json → per-task reward (0/1)
                          3. Parse adapter's errorClass output → per-task errorClass
                          Harbor container (per task):
                          1. Build from task Dockerfile (WORKDIR /app, has tools, test env)
                          2. MakaAgent.run() → adapter starts run-cell.mjs in /app
                          → RuntimeRunner runs agent (file tools + Bash, all in /app, one filesystem)
                          → RuntimeRunner records trajectory to runtime-events.jsonl
                          → adapter extracts errorClass from InvocationResult
                          3. Harbor runs tests/test.sh → checks /app state → reward 0 or 1
                          

                          Why not runExperiment? Verified: runExperiment copies fixture to mkdtemp('/tmp/maka-headless-ws-XXXX'), runs verifier there, deletes it. But TB test scripts write absolute paths to /app (e.g. Path("/app/regex.txt") in test_outputs.py). 89/89 Dockerfiles use WORKDIR /app. runExperiment's temp workspace would never be checked by TB tests. RuntimeRunner works in-place at /app — no copy, no sync problem.

                          Why not Harbor verifier only? We need maka's errorClass (max_tokens / verification_failed / runtime_error) for signal enrichment. The adapter extracts this from RuntimeRunner's InvocationResult and writes it alongside Harbor's reward.

                          File/bash filesystem split (v5 P0-2, resolved): the entire run-cell.mjs process runs in-container. RuntimeRunner's file tools (Write/Edit/Read) and Bash tool (toolExecutor) all operate on the container's filesystem. No host-container split.

                          Prompt delivery: controller passes system_prompt.md content via MAKA_SYSTEM_PROMPT env var. run-cell.mjs reads it, passes to Config.systemPrompt. Runtime stamps a hash of the effectiveConfig.systemPrompt string (at backend constructor, after resolveSystemPrompt) into trajectory events. Controller verifies the hash round-trips — if trajectory hash ≠ hash of prompt sent, the round is invalid (plumbing failure).

                          Implementation prerequisite:systemPromptHash does not exist in runtime yet (verified: only requestShapeHash exists). Must add a small instrumentation: hash the resolved Config.systemPrompt and write to a runtime event field. Pure addition, no existing code changes.

                          Harbor/container failure ≠ benchmark failure. Container crash, image build failure, Harbor timeout → no result.json or invalid output. Controller treats this as infra error. Harbor reward + adapter errorClass only decide benchmark score.

                          Signal quality

                          Fix 1 — Coverage-aware gate.max_tokens/runtime_error tasks should count as fail (not drop from denominator). Gate uses pass / eligible (unscored = fail), NOT pass / scored.

                          keep = held_in_pass_rate (pass/eligible) improved beyond noise_band_held_in
                          AND held_in_coverage (scored/eligible) not degraded
                          AND held_out_pass_rate non-inferior vs original baseline (>= -noise_band_held_out)
                          AND held_out_coverage not degraded
                          AND no infra_errors (any infra error → no keep)
                          

                          Controller computes pass/eligible itself. Harbor gives per-task reward (0/1); adapter gives per-task errorClass; controller combines: pass = reward==1, eligible = true (all tasks eligible unless infra error), scored = errorClass != 'max_tokens' && errorClass != 'runtime_error'.

                          Fix 2 — Dual noise band.noise_band = wilson_half_width(N, p_mean) × √(1 + 1/n_baseline) — CI of the difference, not single-proportion CI. Separate bands for held-in and held-out.

                          • Layer 0 (free): dual noise thresholds.
                          • Layer 1 (cheap): 60 held-in + 20 held-out from TB 2.0's 89 tasks.
                          • Layer 2 (deferred for v1): repeat sampling. v1 goal is structural validation.

                          Fix 3 — Agent reads failing trajectories. Controller extracts per-task digest from runtime-events.jsonl (raw events): last 2 tool calls (from function_call events with args), errorClass, steps/duration, verifier stdout/stderr first 3 lines.

                          Overfitting guard

                          Held-in vs current kept baseline (monotonic improvement). After KEEP: held_in_reference = max(previous_reference, new_run_pass_rate − noise_band_held_in). Banks real improvements without enshrining lucky peak, never falls below prior state.

                          Held-out vs original baseline (fixed floor). Floor = original_held_out_mean − noise_band_held_out. Prevents cumulative drift.

                          Held-out physical isolation. Held-out task content, trajectories, results.jsonl live outside agent cwd. Agent sees only program.md, system_prompt.md, results.tsv, held-in digests.

                          Reward-hack quarantine gate. Controller scans function_call.args in raw runtime-events.jsonl for verifier-specific patterns (expected-output strings, not filenames). Flagged → QUARANTINE (discard). Scanner fails closed if raw events missing.

                          Infra error handling

                          Any infra error → no KEEP. Infra = missing result.json, container crash, image build failure, Harbor timeout. Retry once, else discard+log.

                          >20% infra → stop. Stop condition, not keep-allowance.

                          Infra classification: model API failure (DeepSeek 429/500) → runtime_error (scored=false, counts in eligible). Container/Docker/Harbor failure → infra (excluded from scoring).

                          Budget

                          Controller sums costUsd from runtime-events.jsonl token-usage events. costUsd exists on RuntimeEvent. BUILTIN_PRICING keyed as deepseek:deepseek-chat (provider:model format) — controller must compose key as ${provider}:${config.model}.

                          Zero-cost guard: if summed costUsd = 0 but tokens > 0 → plumbing failure (pricing not wired). Do NOT count as free round.

                          Cost ceiling: $30 (default). 80 tasks × single-sample × ~10 rounds ≈ $15-25.

                          Controller / agent separation

                          Controller (host-side Node.js): Harbor invocation, result parsing, metric computation, budget, git state machine, keep/discard, crash/resume WAL (append jsonl LAST as source of truth; reconcile last_kept_commit from jsonl on resume). Diff guard AFTER agent edit, BEFORE commit — only system_prompt.md (tracked) may change; controller files gitignored.

                          Agent (meta-agent): scripted LLM call (deepseek-chat). Reads results.tsv + digests → outputs prompt edit + description. Log IS its memory. No interactive tools.

                          Implementation entrypoint

                          Not maka-headless eval CLI (only wires fake backend). run-cell.mjs in container directly imports RuntimeRunner from @maka/runtime:

                          import{RuntimeRunner,AiSdkBackend, ... }from'@maka/runtime';// agent runs in /app (Dockerfile WORKDIR), in-place// RuntimeRunner → InvocationResult → extract errorClass// write errorClass + trajectory to /output/

                          maka-agent repo cloned + built at /opt/maka-agent (outside /app).

                          Verified facts

                          QuestionAnswerSource
                          TB test.sh writes /app?Yes — test_outputs.py uses Path("/app/regex.txt") etc.grep 89 tasks
                          runExperiment compatible?No — prepareWorkspace uses mkdtemp('/tmp/...'), can't be /appsandbox.ts:20
                          systemPromptHash exists?No — only requestShapeHashgrep runtime/src
                          pricing key format?provider:model (deepseek:deepseek-chat)builtin-pricing.ts:14
                          Harbor + Docker verified?Yes — oracle reward 1.0, maka adapter runs, verifier workssmoke test 6/19
                          file/bash split resolved?Yes — whole process in-containerarchitecture decision

                          Directory layout

                          ~/.local/maka-eval/
                          ├── harbor-adapter/
                          │ ├── maka_agent.py # Harbor agent adapter
                          │ └── run-cell.mjs # in-container entrypoint (RuntimeRunner)
                          ├── harness-rsi/
                          │ ├── program.md # loop instructions (agent reads)
                          │ ├── system_prompt.md # optimization target (agent edits ONLY)
                          │ ├── controller.mjs # state machine (host-side)
                          │ ├── results.jsonl # canonical truth (OUTSIDE agent cwd)
                          │ ├── results.tsv # derived view for agent (IN agent cwd, gitignored)
                          │ └── baseline/ # 3× baseline calibration
                          ├── ds-key.txt
                          └── fixtures/
                          

                          The loop

                          INIT:
                          1. Run baseline 3× on held-in + held-out
                          2. Compute means + observed_spread
                          3. noise_band = max(wilson_half_width(N, p) × √(4/3), observed_spread)
                          4. held_in_reference = baseline_held_in_mean
                          5. Save original_held_out_mean (never updated)
                          LOOP:
                          1. Write pending round record
                          2. Agent: read results.tsv + digests → edit system_prompt.md
                          3. Controller: diff guard (after edit, before commit) — only system_prompt.md
                          4. Controller: git commit
                          5. Controller: run held-in via Harbor → parse reward + errorClass per task
                          6. Check infra — any infra → no keep, retry once, else discard
                          7. Compute held_in_delta vs held_in_reference
                          8. |delta| < noise_band → DISCARD (inconclusive)
                          9. delta <= 0 → DISCARD (regression)
                          10. delta > noise_band:
                          a. Run held-out via Harbor
                          b. Check held-out infra
                          c. held_out_delta vs original_held_out_mean
                          d. < -noise_band → DISCARD (regression)
                          e. coverage degraded → DISCARD
                          f. reward-hack scan → QUARANTINE if flagged
                          g. >= -noise_band AND coverage OK → KEEP
                          h. held_in_reference = max(held_in_reference, pass_rate − noise_band)
                          11. Append to results.jsonl (source of truth)
                          12. Advance last_kept_commit (if KEEP, AFTER jsonl append)
                          13. Update results.tsv
                          14. Check budget (sum costUsd, zero-cost guard)
                          

                          Success criteria (structural validation)

                          • Loop runs unattended ≥10 iterations without human intervention
                          • Zero keeps is passing — v1 validates structure, not prompt quality
                          • If keeps occur: no held-out regression, no unreviewed reward-hack flags
                          • Cost ≤ $30
                          • Prompt round-trip hash verified every round

                          Open questions

                          1. Task selection: which 60 of 89 for held-in? Filter by speed/stability.
                          2. Parallelism:--n-concurrent — determinism? Key results by task id.
                          3. Simplicity criterion: deletion + non-inferior = KEEP? Needs own branch.
                          4. Pricing wiring: verify in-container path registers pricing lookup.

                          Non-goals (v1)

                          • No tool description optimization
                          • No control flow changes
                          • No AEGIS 4-stage pipeline
                          • No repeat sampling / pass@k
                          • No runExperiment (RuntimeRunner + Harbor verifier instead)

                          Metadata

                          Metadata

                          Assignees

                          No one assigned

                            Labels

                            enhancementNew feature or request

                            Type

                            No type

                            Projects

                            No projects

                              Milestone

                              No milestone

                              Relationships

                              None yet

                              Development

                              No branches or pull requests

                              Issue actions

                              , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
                              Skip to content

                              RFC: Harness RSI loop — autonomous system prompt optimization #64

                              Description

                              @Astro-Han

                              RFC: Harness RSI loop — autoresearch structure for system prompt optimization

                              Context

                              maka's eval framework is ready (PR #62/#63). The goal is closed-loop harness self-improvement: an agent autonomously optimizes its own system prompt against a benchmark task set, no human in the loop, runnable until stopped.

                              Prior art: Karpathy's autoresearch (minimal loop, single agent), HarnessX/AEGIS (heavy 4-stage engine, +14.5%), Self-Harness (regression-gated edit loop).

                              Architecture (verified)

                              RuntimeRunner in-container + Harbor verifier. Harbor provides the Docker container (from TB task Dockerfile). Inside that container, a run-cell.mjs entrypoint runs maka's RuntimeRunner directly — the agent works in-place at /app (the Dockerfile WORKDIR), using file tools + Bash all in-container (one filesystem). After the agent finishes, Harbor runs tests/test.sh which checks /app state.

                              Controller (host):
                              1. For each held-in task: harbor run --path <tb-task> --agent-import-path maka_agent:MakaAgent
                              --ae MAKA_SYSTEM_PROMPT=<prompt> --ae DEEPSEEK_API_KEY=<key>
                              2. Parse Harbor result.json → per-task reward (0/1)
                              3. Parse adapter's errorClass output → per-task errorClass
                              Harbor container (per task):
                              1. Build from task Dockerfile (WORKDIR /app, has tools, test env)
                              2. MakaAgent.run() → adapter starts run-cell.mjs in /app
                              → RuntimeRunner runs agent (file tools + Bash, all in /app, one filesystem)
                              → RuntimeRunner records trajectory to runtime-events.jsonl
                              → adapter extracts errorClass from InvocationResult
                              3. Harbor runs tests/test.sh → checks /app state → reward 0 or 1
                              

                              Why not runExperiment? Verified: runExperiment copies fixture to mkdtemp('/tmp/maka-headless-ws-XXXX'), runs verifier there, deletes it. But TB test scripts write absolute paths to /app (e.g. Path("/app/regex.txt") in test_outputs.py). 89/89 Dockerfiles use WORKDIR /app. runExperiment's temp workspace would never be checked by TB tests. RuntimeRunner works in-place at /app — no copy, no sync problem.

                              Why not Harbor verifier only? We need maka's errorClass (max_tokens / verification_failed / runtime_error) for signal enrichment. The adapter extracts this from RuntimeRunner's InvocationResult and writes it alongside Harbor's reward.

                              File/bash filesystem split (v5 P0-2, resolved): the entire run-cell.mjs process runs in-container. RuntimeRunner's file tools (Write/Edit/Read) and Bash tool (toolExecutor) all operate on the container's filesystem. No host-container split.

                              Prompt delivery: controller passes system_prompt.md content via MAKA_SYSTEM_PROMPT env var. run-cell.mjs reads it, passes to Config.systemPrompt. Runtime stamps a hash of the effectiveConfig.systemPrompt string (at backend constructor, after resolveSystemPrompt) into trajectory events. Controller verifies the hash round-trips — if trajectory hash ≠ hash of prompt sent, the round is invalid (plumbing failure).

                              Implementation prerequisite:systemPromptHash does not exist in runtime yet (verified: only requestShapeHash exists). Must add a small instrumentation: hash the resolved Config.systemPrompt and write to a runtime event field. Pure addition, no existing code changes.

                              Harbor/container failure ≠ benchmark failure. Container crash, image build failure, Harbor timeout → no result.json or invalid output. Controller treats this as infra error. Harbor reward + adapter errorClass only decide benchmark score.

                              Signal quality

                              Fix 1 — Coverage-aware gate.max_tokens/runtime_error tasks should count as fail (not drop from denominator). Gate uses pass / eligible (unscored = fail), NOT pass / scored.

                              keep = held_in_pass_rate (pass/eligible) improved beyond noise_band_held_in
                              AND held_in_coverage (scored/eligible) not degraded
                              AND held_out_pass_rate non-inferior vs original baseline (>= -noise_band_held_out)
                              AND held_out_coverage not degraded
                              AND no infra_errors (any infra error → no keep)
                              

                              Controller computes pass/eligible itself. Harbor gives per-task reward (0/1); adapter gives per-task errorClass; controller combines: pass = reward==1, eligible = true (all tasks eligible unless infra error), scored = errorClass != 'max_tokens' && errorClass != 'runtime_error'.

                              Fix 2 — Dual noise band.noise_band = wilson_half_width(N, p_mean) × √(1 + 1/n_baseline) — CI of the difference, not single-proportion CI. Separate bands for held-in and held-out.

                              • Layer 0 (free): dual noise thresholds.
                              • Layer 1 (cheap): 60 held-in + 20 held-out from TB 2.0's 89 tasks.
                              • Layer 2 (deferred for v1): repeat sampling. v1 goal is structural validation.

                              Fix 3 — Agent reads failing trajectories. Controller extracts per-task digest from runtime-events.jsonl (raw events): last 2 tool calls (from function_call events with args), errorClass, steps/duration, verifier stdout/stderr first 3 lines.

                              Overfitting guard

                              Held-in vs current kept baseline (monotonic improvement). After KEEP: held_in_reference = max(previous_reference, new_run_pass_rate − noise_band_held_in). Banks real improvements without enshrining lucky peak, never falls below prior state.

                              Held-out vs original baseline (fixed floor). Floor = original_held_out_mean − noise_band_held_out. Prevents cumulative drift.

                              Held-out physical isolation. Held-out task content, trajectories, results.jsonl live outside agent cwd. Agent sees only program.md, system_prompt.md, results.tsv, held-in digests.

                              Reward-hack quarantine gate. Controller scans function_call.args in raw runtime-events.jsonl for verifier-specific patterns (expected-output strings, not filenames). Flagged → QUARANTINE (discard). Scanner fails closed if raw events missing.

                              Infra error handling

                              Any infra error → no KEEP. Infra = missing result.json, container crash, image build failure, Harbor timeout. Retry once, else discard+log.

                              >20% infra → stop. Stop condition, not keep-allowance.

                              Infra classification: model API failure (DeepSeek 429/500) → runtime_error (scored=false, counts in eligible). Container/Docker/Harbor failure → infra (excluded from scoring).

                              Budget

                              Controller sums costUsd from runtime-events.jsonl token-usage events. costUsd exists on RuntimeEvent. BUILTIN_PRICING keyed as deepseek:deepseek-chat (provider:model format) — controller must compose key as ${provider}:${config.model}.

                              Zero-cost guard: if summed costUsd = 0 but tokens > 0 → plumbing failure (pricing not wired). Do NOT count as free round.

                              Cost ceiling: $30 (default). 80 tasks × single-sample × ~10 rounds ≈ $15-25.

                              Controller / agent separation

                              Controller (host-side Node.js): Harbor invocation, result parsing, metric computation, budget, git state machine, keep/discard, crash/resume WAL (append jsonl LAST as source of truth; reconcile last_kept_commit from jsonl on resume). Diff guard AFTER agent edit, BEFORE commit — only system_prompt.md (tracked) may change; controller files gitignored.

                              Agent (meta-agent): scripted LLM call (deepseek-chat). Reads results.tsv + digests → outputs prompt edit + description. Log IS its memory. No interactive tools.

                              Implementation entrypoint

                              Not maka-headless eval CLI (only wires fake backend). run-cell.mjs in container directly imports RuntimeRunner from @maka/runtime:

                              import{RuntimeRunner,AiSdkBackend, ... }from'@maka/runtime';// agent runs in /app (Dockerfile WORKDIR), in-place// RuntimeRunner → InvocationResult → extract errorClass// write errorClass + trajectory to /output/

                              maka-agent repo cloned + built at /opt/maka-agent (outside /app).

                              Verified facts

                              QuestionAnswerSource
                              TB test.sh writes /app?Yes — test_outputs.py uses Path("/app/regex.txt") etc.grep 89 tasks
                              runExperiment compatible?No — prepareWorkspace uses mkdtemp('/tmp/...'), can't be /appsandbox.ts:20
                              systemPromptHash exists?No — only requestShapeHashgrep runtime/src
                              pricing key format?provider:model (deepseek:deepseek-chat)builtin-pricing.ts:14
                              Harbor + Docker verified?Yes — oracle reward 1.0, maka adapter runs, verifier workssmoke test 6/19
                              file/bash split resolved?Yes — whole process in-containerarchitecture decision

                              Directory layout

                              ~/.local/maka-eval/
                              ├── harbor-adapter/
                              │ ├── maka_agent.py # Harbor agent adapter
                              │ └── run-cell.mjs # in-container entrypoint (RuntimeRunner)
                              ├── harness-rsi/
                              │ ├── program.md # loop instructions (agent reads)
                              │ ├── system_prompt.md # optimization target (agent edits ONLY)
                              │ ├── controller.mjs # state machine (host-side)
                              │ ├── results.jsonl # canonical truth (OUTSIDE agent cwd)
                              │ ├── results.tsv # derived view for agent (IN agent cwd, gitignored)
                              │ └── baseline/ # 3× baseline calibration
                              ├── ds-key.txt
                              └── fixtures/
                              

                              The loop

                              INIT:
                              1. Run baseline 3× on held-in + held-out
                              2. Compute means + observed_spread
                              3. noise_band = max(wilson_half_width(N, p) × √(4/3), observed_spread)
                              4. held_in_reference = baseline_held_in_mean
                              5. Save original_held_out_mean (never updated)
                              LOOP:
                              1. Write pending round record
                              2. Agent: read results.tsv + digests → edit system_prompt.md
                              3. Controller: diff guard (after edit, before commit) — only system_prompt.md
                              4. Controller: git commit
                              5. Controller: run held-in via Harbor → parse reward + errorClass per task
                              6. Check infra — any infra → no keep, retry once, else discard
                              7. Compute held_in_delta vs held_in_reference
                              8. |delta| < noise_band → DISCARD (inconclusive)
                              9. delta <= 0 → DISCARD (regression)
                              10. delta > noise_band:
                              a. Run held-out via Harbor
                              b. Check held-out infra
                              c. held_out_delta vs original_held_out_mean
                              d. < -noise_band → DISCARD (regression)
                              e. coverage degraded → DISCARD
                              f. reward-hack scan → QUARANTINE if flagged
                              g. >= -noise_band AND coverage OK → KEEP
                              h. held_in_reference = max(held_in_reference, pass_rate − noise_band)
                              11. Append to results.jsonl (source of truth)
                              12. Advance last_kept_commit (if KEEP, AFTER jsonl append)
                              13. Update results.tsv
                              14. Check budget (sum costUsd, zero-cost guard)
                              

                              Success criteria (structural validation)

                              • Loop runs unattended ≥10 iterations without human intervention
                              • Zero keeps is passing — v1 validates structure, not prompt quality
                              • If keeps occur: no held-out regression, no unreviewed reward-hack flags
                              • Cost ≤ $30
                              • Prompt round-trip hash verified every round

                              Open questions

                              1. Task selection: which 60 of 89 for held-in? Filter by speed/stability.
                              2. Parallelism:--n-concurrent — determinism? Key results by task id.
                              3. Simplicity criterion: deletion + non-inferior = KEEP? Needs own branch.
                              4. Pricing wiring: verify in-container path registers pricing lookup.

                              Non-goals (v1)

                              • No tool description optimization
                              • No control flow changes
                              • No AEGIS 4-stage pipeline
                              • No repeat sampling / pass@k
                              • No runExperiment (RuntimeRunner + Harbor verifier instead)

                              Metadata

                              Metadata

                              Assignees

                              No one assigned

                                Labels

                                enhancementNew feature or request

                                Type

                                No type

                                Projects

                                No projects

                                  Milestone

                                  No milestone

                                  Relationships

                                  None yet

                                  Development

                                  No branches or pull requests

                                  Issue actions