feat(headless): Config.systemPrompt + runtime_error failure class - #62

Merged
Astro-Han merged 4 commits into
mainfrom
codex/eval-systemprompt-failureclass
Jun 19, 2026
Merged

feat(headless): Config.systemPrompt + runtime_error failure class#62
Astro-Han merged 4 commits into
mainfrom
codex/eval-systemprompt-failureclass

Conversation

@Astro-Han

Copy link
Copy Markdown
Contributor

Summary

Two fixes that unblock real-model benchmark smoke (Terminal-Bench on DeepSeek), exposed by the diagnosis in the prior session:

  • Config.systemPrompt — real models running benchmark tasks without a system prompt narrate their reasoning in text instead of calling tools, hit the 8192 output token limit, and never produce the required artifact. Adds systemPrompt?: string to Config; the registerBackends factory closure reads it and passes it to AiSdkBackend directly (same pattern as the desktop interactive path). Not routed through BackendFactoryContext.systemPrompt (that's the child-agent instruction channel) and not persisted into SessionHeader (benchmark sessions are throwaway). Also exports BENCHMARK_BASE_SYSTEM_PROMPT, a generic tool-first prefix verified against DeepSeek-chat on a Terminal-Bench regex task.
  • failureClass: runtime_error — a backend that ends with complete(stopReason='error') without a preceding error event left the run ledger's failureClass as unknown, so benchmark scoring could not distinguish runtime failures from max_tokens / incomplete_tool_calls. Classifies it as runtime_error instead.

Both reviewed by Codex and Claude (consult mode) — confirmed headless-only / runtime-pass / no-persistence is the right scope.

Test plan

  • @maka/headless 91/91 (2 new tests: factory closure reads config.systemPrompt; undefined when omitted)
  • @maka/runtime 586/586 (1 new test: complete(stopReason=error)runtime_error not unknown)
  • npm run typecheck clean

Real models running benchmark tasks without a system prompt narrate
their reasoning in text instead of calling tools, hit the output
token limit, and never produce the required artifact.
Add systemPrompt?: string to Config. The registerBackends factory
closure reads it and passes it to AiSdkBackend directly — same
pattern as the desktop interactive path. systemPrompt is NOT routed
through BackendFactoryContext.systemPrompt (that channel is the
child-agent instruction, not the main-session prompt) and is NOT
persisted into SessionHeader (benchmark sessions are throwaway).
Also exports BENCHMARK_BASE_SYSTEM_PROMPT, a generic tool-first
prefix verified against DeepSeek-chat on a Terminal-Bench regex task.
Reviewed by Codex and Claude (consult mode) — both confirmed
headless-only / runtime-pass / no-persistence is the right scope.
…t unknown
When a backend ends with stopReason='error' but never emits a
preceding error event (observed with DeepSeek-reasoner after it
tried to call an unavailable tool and entered a reasoning loop),
the run ledger's failureClass fell back to 'unknown' — making
benchmark failures unattributable.
Now turnStatusFromEvent returns errorClass 'runtime_error' for
stopReason='error', and recordSessionEvent calls markRunFailed so
finalize does not fall back to 'unknown'. This lets the benchmark
scorer distinguish runtime failures from max_tokens / verification_failed.
…mment
Codex [P2]: BENCHMARK_BASE_SYSTEM_PROMPT said 'You do not have a Bash
tool' but the standard isolated headless surface (buildIsolatedHeadlessTools)
includes an isolated Bash. Reworded to 'do not self-test or run
verification scripts' — acknowledges Bash exists for producing artifacts,
steers away from self-grading.
Claude [P3]: complete(stopReason=error) path now emits two run_failed
events (markRunFailed + finishRun), same as the error-event path already
did. Added a comment making the conscious choice explicit.
Claude review follow-up: the existing tests proved the factory closure
can READ config.systemPrompt, but not that it actually PASSES it to the
backend constructor's systemPrompt parameter — the exact seam a real
AiSdkBackend factory uses.
Adds an ai-sdk stub backend that records its constructor systemPrompt
input, and a test that runs runExperiment with registerBackends wiring
config.systemPrompt → backend ctor (mirroring the desktop pattern).
Verifies the prompt string arrives at the backend constructor intact.
@Astro-Han
Astro-Han merged commit ab61bc5 into mainJun 19, 2026
Astro-Han added a commit that referenced this pull request Jun 19, 2026
…ror (#63)
Fresh-eye review (Codex P1 / Claude P3) found that PR #62 only fixed
the AgentRunStore ledger failureClass, not the InvocationResult.failure.class
that benchmark ResultRecord reads. A complete(stopReason=error) with no
preceding error event still surfaced as errorClass='failed' in
ResultRecord — indistinguishable from other failures.
Root cause: failureFromTerminalEvent returned class=status (i.e. 'failed')
for any failed terminal, ignoring whether error content carried a precise
reason. The error-content branch (which extracts content.reason) was
unreachable for terminal events because the terminal check ran first.
Fix: in failureFromTerminalEvent, when status='failed':
- with error content reason/code → use that (e.g. 'tool_failed')
- without error content → 'runtime_error' (not 'failed')
This makes the invocation path consistent with the run ledger, so
ResultRecord.errorClass reads 'runtime_error' for the bare-error scenario.
Tests:
- runtime: updated locked test (failed+error content w/o reason → runtime_error),
added reason-code test, added bare-failed-terminal test
- headless: added end-to-end test verifying ResultRecord.errorClass='runtime_error'
for a backend that emits only complete(stopReason=error)
@Astro-Han
Astro-Han deleted the codex/eval-systemprompt-failureclass branch July 14, 2026 05:05
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

feat(headless): Config.systemPrompt + runtime_error failure class - #62

Merged
Astro-Han merged 4 commits into
mainfrom
codex/eval-systemprompt-failureclass
Jun 19, 2026
Merged

feat(headless): Config.systemPrompt + runtime_error failure class#62
Astro-Han merged 4 commits into
mainfrom
codex/eval-systemprompt-failureclass

Conversation

@Astro-Han

Copy link
Copy Markdown
Contributor

Summary

Two fixes that unblock real-model benchmark smoke (Terminal-Bench on DeepSeek), exposed by the diagnosis in the prior session:

  • Config.systemPrompt — real models running benchmark tasks without a system prompt narrate their reasoning in text instead of calling tools, hit the 8192 output token limit, and never produce the required artifact. Adds systemPrompt?: string to Config; the registerBackends factory closure reads it and passes it to AiSdkBackend directly (same pattern as the desktop interactive path). Not routed through BackendFactoryContext.systemPrompt (that's the child-agent instruction channel) and not persisted into SessionHeader (benchmark sessions are throwaway). Also exports BENCHMARK_BASE_SYSTEM_PROMPT, a generic tool-first prefix verified against DeepSeek-chat on a Terminal-Bench regex task.
  • failureClass: runtime_error — a backend that ends with complete(stopReason='error') without a preceding error event left the run ledger's failureClass as unknown, so benchmark scoring could not distinguish runtime failures from max_tokens / incomplete_tool_calls. Classifies it as runtime_error instead.

Both reviewed by Codex and Claude (consult mode) — confirmed headless-only / runtime-pass / no-persistence is the right scope.

Test plan

  • @maka/headless 91/91 (2 new tests: factory closure reads config.systemPrompt; undefined when omitted)
  • @maka/runtime 586/586 (1 new test: complete(stopReason=error)runtime_error not unknown)
  • npm run typecheck clean

Real models running benchmark tasks without a system prompt narrate
their reasoning in text instead of calling tools, hit the output
token limit, and never produce the required artifact.
Add systemPrompt?: string to Config. The registerBackends factory
closure reads it and passes it to AiSdkBackend directly — same
pattern as the desktop interactive path. systemPrompt is NOT routed
through BackendFactoryContext.systemPrompt (that channel is the
child-agent instruction, not the main-session prompt) and is NOT
persisted into SessionHeader (benchmark sessions are throwaway).
Also exports BENCHMARK_BASE_SYSTEM_PROMPT, a generic tool-first
prefix verified against DeepSeek-chat on a Terminal-Bench regex task.
Reviewed by Codex and Claude (consult mode) — both confirmed
headless-only / runtime-pass / no-persistence is the right scope.
…t unknown
When a backend ends with stopReason='error' but never emits a
preceding error event (observed with DeepSeek-reasoner after it
tried to call an unavailable tool and entered a reasoning loop),
the run ledger's failureClass fell back to 'unknown' — making
benchmark failures unattributable.
Now turnStatusFromEvent returns errorClass 'runtime_error' for
stopReason='error', and recordSessionEvent calls markRunFailed so
finalize does not fall back to 'unknown'. This lets the benchmark
scorer distinguish runtime failures from max_tokens / verification_failed.
…mment
Codex [P2]: BENCHMARK_BASE_SYSTEM_PROMPT said 'You do not have a Bash
tool' but the standard isolated headless surface (buildIsolatedHeadlessTools)
includes an isolated Bash. Reworded to 'do not self-test or run
verification scripts' — acknowledges Bash exists for producing artifacts,
steers away from self-grading.
Claude [P3]: complete(stopReason=error) path now emits two run_failed
events (markRunFailed + finishRun), same as the error-event path already
did. Added a comment making the conscious choice explicit.
Claude review follow-up: the existing tests proved the factory closure
can READ config.systemPrompt, but not that it actually PASSES it to the
backend constructor's systemPrompt parameter — the exact seam a real
AiSdkBackend factory uses.
Adds an ai-sdk stub backend that records its constructor systemPrompt
input, and a test that runs runExperiment with registerBackends wiring
config.systemPrompt → backend ctor (mirroring the desktop pattern).
Verifies the prompt string arrives at the backend constructor intact.
@Astro-Han
Astro-Han merged commit ab61bc5 into mainJun 19, 2026
Astro-Han added a commit that referenced this pull request Jun 19, 2026
…ror (#63)
Fresh-eye review (Codex P1 / Claude P3) found that PR #62 only fixed
the AgentRunStore ledger failureClass, not the InvocationResult.failure.class
that benchmark ResultRecord reads. A complete(stopReason=error) with no
preceding error event still surfaced as errorClass='failed' in
ResultRecord — indistinguishable from other failures.
Root cause: failureFromTerminalEvent returned class=status (i.e. 'failed')
for any failed terminal, ignoring whether error content carried a precise
reason. The error-content branch (which extracts content.reason) was
unreachable for terminal events because the terminal check ran first.
Fix: in failureFromTerminalEvent, when status='failed':
- with error content reason/code → use that (e.g. 'tool_failed')
- without error content → 'runtime_error' (not 'failed')
This makes the invocation path consistent with the run ledger, so
ResultRecord.errorClass reads 'runtime_error' for the bare-error scenario.
Tests:
- runtime: updated locked test (failed+error content w/o reason → runtime_error),
added reason-code test, added bare-failed-terminal test
- headless: added end-to-end test verifying ResultRecord.errorClass='runtime_error'
for a backend that emits only complete(stopReason=error)
@Astro-Han
Astro-Han deleted the codex/eval-systemprompt-failureclass branch July 14, 2026 05:05
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(headless): Config.systemPrompt + runtime_error failure class - #62

Merged
Astro-Han merged 4 commits into
mainfrom
codex/eval-systemprompt-failureclass
Jun 19, 2026
Merged

feat(headless): Config.systemPrompt + runtime_error failure class#62
Astro-Han merged 4 commits into
mainfrom
codex/eval-systemprompt-failureclass

Conversation

@Astro-Han

Copy link
Copy Markdown
Contributor

Summary

Two fixes that unblock real-model benchmark smoke (Terminal-Bench on DeepSeek), exposed by the diagnosis in the prior session:

  • Config.systemPrompt — real models running benchmark tasks without a system prompt narrate their reasoning in text instead of calling tools, hit the 8192 output token limit, and never produce the required artifact. Adds systemPrompt?: string to Config; the registerBackends factory closure reads it and passes it to AiSdkBackend directly (same pattern as the desktop interactive path). Not routed through BackendFactoryContext.systemPrompt (that's the child-agent instruction channel) and not persisted into SessionHeader (benchmark sessions are throwaway). Also exports BENCHMARK_BASE_SYSTEM_PROMPT, a generic tool-first prefix verified against DeepSeek-chat on a Terminal-Bench regex task.
  • failureClass: runtime_error — a backend that ends with complete(stopReason='error') without a preceding error event left the run ledger's failureClass as unknown, so benchmark scoring could not distinguish runtime failures from max_tokens / incomplete_tool_calls. Classifies it as runtime_error instead.

Both reviewed by Codex and Claude (consult mode) — confirmed headless-only / runtime-pass / no-persistence is the right scope.

Test plan

  • @maka/headless 91/91 (2 new tests: factory closure reads config.systemPrompt; undefined when omitted)
  • @maka/runtime 586/586 (1 new test: complete(stopReason=error)runtime_error not unknown)
  • npm run typecheck clean

Real models running benchmark tasks without a system prompt narrate
their reasoning in text instead of calling tools, hit the output
token limit, and never produce the required artifact.
Add systemPrompt?: string to Config. The registerBackends factory
closure reads it and passes it to AiSdkBackend directly — same
pattern as the desktop interactive path. systemPrompt is NOT routed
through BackendFactoryContext.systemPrompt (that channel is the
child-agent instruction, not the main-session prompt) and is NOT
persisted into SessionHeader (benchmark sessions are throwaway).
Also exports BENCHMARK_BASE_SYSTEM_PROMPT, a generic tool-first
prefix verified against DeepSeek-chat on a Terminal-Bench regex task.
Reviewed by Codex and Claude (consult mode) — both confirmed
headless-only / runtime-pass / no-persistence is the right scope.
…t unknown
When a backend ends with stopReason='error' but never emits a
preceding error event (observed with DeepSeek-reasoner after it
tried to call an unavailable tool and entered a reasoning loop),
the run ledger's failureClass fell back to 'unknown' — making
benchmark failures unattributable.
Now turnStatusFromEvent returns errorClass 'runtime_error' for
stopReason='error', and recordSessionEvent calls markRunFailed so
finalize does not fall back to 'unknown'. This lets the benchmark
scorer distinguish runtime failures from max_tokens / verification_failed.
…mment
Codex [P2]: BENCHMARK_BASE_SYSTEM_PROMPT said 'You do not have a Bash
tool' but the standard isolated headless surface (buildIsolatedHeadlessTools)
includes an isolated Bash. Reworded to 'do not self-test or run
verification scripts' — acknowledges Bash exists for producing artifacts,
steers away from self-grading.
Claude [P3]: complete(stopReason=error) path now emits two run_failed
events (markRunFailed + finishRun), same as the error-event path already
did. Added a comment making the conscious choice explicit.
Claude review follow-up: the existing tests proved the factory closure
can READ config.systemPrompt, but not that it actually PASSES it to the
backend constructor's systemPrompt parameter — the exact seam a real
AiSdkBackend factory uses.
Adds an ai-sdk stub backend that records its constructor systemPrompt
input, and a test that runs runExperiment with registerBackends wiring
config.systemPrompt → backend ctor (mirroring the desktop pattern).
Verifies the prompt string arrives at the backend constructor intact.
@Astro-Han
Astro-Han merged commit ab61bc5 into mainJun 19, 2026
Astro-Han added a commit that referenced this pull request Jun 19, 2026
…ror (#63)
Fresh-eye review (Codex P1 / Claude P3) found that PR #62 only fixed
the AgentRunStore ledger failureClass, not the InvocationResult.failure.class
that benchmark ResultRecord reads. A complete(stopReason=error) with no
preceding error event still surfaced as errorClass='failed' in
ResultRecord — indistinguishable from other failures.
Root cause: failureFromTerminalEvent returned class=status (i.e. 'failed')
for any failed terminal, ignoring whether error content carried a precise
reason. The error-content branch (which extracts content.reason) was
unreachable for terminal events because the terminal check ran first.
Fix: in failureFromTerminalEvent, when status='failed':
- with error content reason/code → use that (e.g. 'tool_failed')
- without error content → 'runtime_error' (not 'failed')
This makes the invocation path consistent with the run ledger, so
ResultRecord.errorClass reads 'runtime_error' for the bare-error scenario.
Tests:
- runtime: updated locked test (failed+error content w/o reason → runtime_error),
added reason-code test, added bare-failed-terminal test
- headless: added end-to-end test verifying ResultRecord.errorClass='runtime_error'
for a backend that emits only complete(stopReason=error)
@Astro-Han
Astro-Han deleted the codex/eval-systemprompt-failureclass branch July 14, 2026 05:05
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(headless): Config.systemPrompt + runtime_error failure class - #62

Merged
Astro-Han merged 4 commits into
mainfrom
codex/eval-systemprompt-failureclass
Jun 19, 2026
Merged

feat(headless): Config.systemPrompt + runtime_error failure class#62
Astro-Han merged 4 commits into
mainfrom
codex/eval-systemprompt-failureclass

Conversation

@Astro-Han

Copy link
Copy Markdown
Contributor

Summary

Two fixes that unblock real-model benchmark smoke (Terminal-Bench on DeepSeek), exposed by the diagnosis in the prior session:

  • Config.systemPrompt — real models running benchmark tasks without a system prompt narrate their reasoning in text instead of calling tools, hit the 8192 output token limit, and never produce the required artifact. Adds systemPrompt?: string to Config; the registerBackends factory closure reads it and passes it to AiSdkBackend directly (same pattern as the desktop interactive path). Not routed through BackendFactoryContext.systemPrompt (that's the child-agent instruction channel) and not persisted into SessionHeader (benchmark sessions are throwaway). Also exports BENCHMARK_BASE_SYSTEM_PROMPT, a generic tool-first prefix verified against DeepSeek-chat on a Terminal-Bench regex task.
  • failureClass: runtime_error — a backend that ends with complete(stopReason='error') without a preceding error event left the run ledger's failureClass as unknown, so benchmark scoring could not distinguish runtime failures from max_tokens / incomplete_tool_calls. Classifies it as runtime_error instead.

Both reviewed by Codex and Claude (consult mode) — confirmed headless-only / runtime-pass / no-persistence is the right scope.

Test plan

  • @maka/headless 91/91 (2 new tests: factory closure reads config.systemPrompt; undefined when omitted)
  • @maka/runtime 586/586 (1 new test: complete(stopReason=error)runtime_error not unknown)
  • npm run typecheck clean

Real models running benchmark tasks without a system prompt narrate
their reasoning in text instead of calling tools, hit the output
token limit, and never produce the required artifact.
Add systemPrompt?: string to Config. The registerBackends factory
closure reads it and passes it to AiSdkBackend directly — same
pattern as the desktop interactive path. systemPrompt is NOT routed
through BackendFactoryContext.systemPrompt (that channel is the
child-agent instruction, not the main-session prompt) and is NOT
persisted into SessionHeader (benchmark sessions are throwaway).
Also exports BENCHMARK_BASE_SYSTEM_PROMPT, a generic tool-first
prefix verified against DeepSeek-chat on a Terminal-Bench regex task.
Reviewed by Codex and Claude (consult mode) — both confirmed
headless-only / runtime-pass / no-persistence is the right scope.
…t unknown
When a backend ends with stopReason='error' but never emits a
preceding error event (observed with DeepSeek-reasoner after it
tried to call an unavailable tool and entered a reasoning loop),
the run ledger's failureClass fell back to 'unknown' — making
benchmark failures unattributable.
Now turnStatusFromEvent returns errorClass 'runtime_error' for
stopReason='error', and recordSessionEvent calls markRunFailed so
finalize does not fall back to 'unknown'. This lets the benchmark
scorer distinguish runtime failures from max_tokens / verification_failed.
…mment
Codex [P2]: BENCHMARK_BASE_SYSTEM_PROMPT said 'You do not have a Bash
tool' but the standard isolated headless surface (buildIsolatedHeadlessTools)
includes an isolated Bash. Reworded to 'do not self-test or run
verification scripts' — acknowledges Bash exists for producing artifacts,
steers away from self-grading.
Claude [P3]: complete(stopReason=error) path now emits two run_failed
events (markRunFailed + finishRun), same as the error-event path already
did. Added a comment making the conscious choice explicit.
Claude review follow-up: the existing tests proved the factory closure
can READ config.systemPrompt, but not that it actually PASSES it to the
backend constructor's systemPrompt parameter — the exact seam a real
AiSdkBackend factory uses.
Adds an ai-sdk stub backend that records its constructor systemPrompt
input, and a test that runs runExperiment with registerBackends wiring
config.systemPrompt → backend ctor (mirroring the desktop pattern).
Verifies the prompt string arrives at the backend constructor intact.
@Astro-Han
Astro-Han merged commit ab61bc5 into mainJun 19, 2026
Astro-Han added a commit that referenced this pull request Jun 19, 2026
…ror (#63)
Fresh-eye review (Codex P1 / Claude P3) found that PR #62 only fixed
the AgentRunStore ledger failureClass, not the InvocationResult.failure.class
that benchmark ResultRecord reads. A complete(stopReason=error) with no
preceding error event still surfaced as errorClass='failed' in
ResultRecord — indistinguishable from other failures.
Root cause: failureFromTerminalEvent returned class=status (i.e. 'failed')
for any failed terminal, ignoring whether error content carried a precise
reason. The error-content branch (which extracts content.reason) was
unreachable for terminal events because the terminal check ran first.
Fix: in failureFromTerminalEvent, when status='failed':
- with error content reason/code → use that (e.g. 'tool_failed')
- without error content → 'runtime_error' (not 'failed')
This makes the invocation path consistent with the run ledger, so
ResultRecord.errorClass reads 'runtime_error' for the bare-error scenario.
Tests:
- runtime: updated locked test (failed+error content w/o reason → runtime_error),
added reason-code test, added bare-failed-terminal test
- headless: added end-to-end test verifying ResultRecord.errorClass='runtime_error'
for a backend that emits only complete(stopReason=error)
@Astro-Han
Astro-Han deleted the codex/eval-systemprompt-failureclass branch July 14, 2026 05:05
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

feat(headless): Config.systemPrompt + runtime_error failure class - #62

Merged
Astro-Han merged 4 commits into
mainfrom
codex/eval-systemprompt-failureclass
Jun 19, 2026
Merged

feat(headless): Config.systemPrompt + runtime_error failure class#62
Astro-Han merged 4 commits into
mainfrom
codex/eval-systemprompt-failureclass

Conversation

@Astro-Han

Copy link
Copy Markdown
Contributor

Summary

Two fixes that unblock real-model benchmark smoke (Terminal-Bench on DeepSeek), exposed by the diagnosis in the prior session:

  • Config.systemPrompt — real models running benchmark tasks without a system prompt narrate their reasoning in text instead of calling tools, hit the 8192 output token limit, and never produce the required artifact. Adds systemPrompt?: string to Config; the registerBackends factory closure reads it and passes it to AiSdkBackend directly (same pattern as the desktop interactive path). Not routed through BackendFactoryContext.systemPrompt (that's the child-agent instruction channel) and not persisted into SessionHeader (benchmark sessions are throwaway). Also exports BENCHMARK_BASE_SYSTEM_PROMPT, a generic tool-first prefix verified against DeepSeek-chat on a Terminal-Bench regex task.
  • failureClass: runtime_error — a backend that ends with complete(stopReason='error') without a preceding error event left the run ledger's failureClass as unknown, so benchmark scoring could not distinguish runtime failures from max_tokens / incomplete_tool_calls. Classifies it as runtime_error instead.

Both reviewed by Codex and Claude (consult mode) — confirmed headless-only / runtime-pass / no-persistence is the right scope.

Test plan

  • @maka/headless 91/91 (2 new tests: factory closure reads config.systemPrompt; undefined when omitted)
  • @maka/runtime 586/586 (1 new test: complete(stopReason=error)runtime_error not unknown)
  • npm run typecheck clean

Real models running benchmark tasks without a system prompt narrate
their reasoning in text instead of calling tools, hit the output
token limit, and never produce the required artifact.
Add systemPrompt?: string to Config. The registerBackends factory
closure reads it and passes it to AiSdkBackend directly — same
pattern as the desktop interactive path. systemPrompt is NOT routed
through BackendFactoryContext.systemPrompt (that channel is the
child-agent instruction, not the main-session prompt) and is NOT
persisted into SessionHeader (benchmark sessions are throwaway).
Also exports BENCHMARK_BASE_SYSTEM_PROMPT, a generic tool-first
prefix verified against DeepSeek-chat on a Terminal-Bench regex task.
Reviewed by Codex and Claude (consult mode) — both confirmed
headless-only / runtime-pass / no-persistence is the right scope.
…t unknown
When a backend ends with stopReason='error' but never emits a
preceding error event (observed with DeepSeek-reasoner after it
tried to call an unavailable tool and entered a reasoning loop),
the run ledger's failureClass fell back to 'unknown' — making
benchmark failures unattributable.
Now turnStatusFromEvent returns errorClass 'runtime_error' for
stopReason='error', and recordSessionEvent calls markRunFailed so
finalize does not fall back to 'unknown'. This lets the benchmark
scorer distinguish runtime failures from max_tokens / verification_failed.
…mment
Codex [P2]: BENCHMARK_BASE_SYSTEM_PROMPT said 'You do not have a Bash
tool' but the standard isolated headless surface (buildIsolatedHeadlessTools)
includes an isolated Bash. Reworded to 'do not self-test or run
verification scripts' — acknowledges Bash exists for producing artifacts,
steers away from self-grading.
Claude [P3]: complete(stopReason=error) path now emits two run_failed
events (markRunFailed + finishRun), same as the error-event path already
did. Added a comment making the conscious choice explicit.
Claude review follow-up: the existing tests proved the factory closure
can READ config.systemPrompt, but not that it actually PASSES it to the
backend constructor's systemPrompt parameter — the exact seam a real
AiSdkBackend factory uses.
Adds an ai-sdk stub backend that records its constructor systemPrompt
input, and a test that runs runExperiment with registerBackends wiring
config.systemPrompt → backend ctor (mirroring the desktop pattern).
Verifies the prompt string arrives at the backend constructor intact.
@Astro-Han
Astro-Han merged commit ab61bc5 into mainJun 19, 2026
Astro-Han added a commit that referenced this pull request Jun 19, 2026
…ror (#63)
Fresh-eye review (Codex P1 / Claude P3) found that PR #62 only fixed
the AgentRunStore ledger failureClass, not the InvocationResult.failure.class
that benchmark ResultRecord reads. A complete(stopReason=error) with no
preceding error event still surfaced as errorClass='failed' in
ResultRecord — indistinguishable from other failures.
Root cause: failureFromTerminalEvent returned class=status (i.e. 'failed')
for any failed terminal, ignoring whether error content carried a precise
reason. The error-content branch (which extracts content.reason) was
unreachable for terminal events because the terminal check ran first.
Fix: in failureFromTerminalEvent, when status='failed':
- with error content reason/code → use that (e.g. 'tool_failed')
- without error content → 'runtime_error' (not 'failed')
This makes the invocation path consistent with the run ledger, so
ResultRecord.errorClass reads 'runtime_error' for the bare-error scenario.
Tests:
- runtime: updated locked test (failed+error content w/o reason → runtime_error),
added reason-code test, added bare-failed-terminal test
- headless: added end-to-end test verifying ResultRecord.errorClass='runtime_error'
for a backend that emits only complete(stopReason=error)
@Astro-Han
Astro-Han deleted the codex/eval-systemprompt-failureclass branch July 14, 2026 05:05
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(headless): Config.systemPrompt + runtime_error failure class - #62

Merged
Astro-Han merged 4 commits into
mainfrom
codex/eval-systemprompt-failureclass
Jun 19, 2026
Merged

feat(headless): Config.systemPrompt + runtime_error failure class#62
Astro-Han merged 4 commits into
mainfrom
codex/eval-systemprompt-failureclass

Conversation

@Astro-Han

Copy link
Copy Markdown
Contributor

Summary

Two fixes that unblock real-model benchmark smoke (Terminal-Bench on DeepSeek), exposed by the diagnosis in the prior session:

  • Config.systemPrompt — real models running benchmark tasks without a system prompt narrate their reasoning in text instead of calling tools, hit the 8192 output token limit, and never produce the required artifact. Adds systemPrompt?: string to Config; the registerBackends factory closure reads it and passes it to AiSdkBackend directly (same pattern as the desktop interactive path). Not routed through BackendFactoryContext.systemPrompt (that's the child-agent instruction channel) and not persisted into SessionHeader (benchmark sessions are throwaway). Also exports BENCHMARK_BASE_SYSTEM_PROMPT, a generic tool-first prefix verified against DeepSeek-chat on a Terminal-Bench regex task.
  • failureClass: runtime_error — a backend that ends with complete(stopReason='error') without a preceding error event left the run ledger's failureClass as unknown, so benchmark scoring could not distinguish runtime failures from max_tokens / incomplete_tool_calls. Classifies it as runtime_error instead.

Both reviewed by Codex and Claude (consult mode) — confirmed headless-only / runtime-pass / no-persistence is the right scope.

Test plan

  • @maka/headless 91/91 (2 new tests: factory closure reads config.systemPrompt; undefined when omitted)
  • @maka/runtime 586/586 (1 new test: complete(stopReason=error)runtime_error not unknown)
  • npm run typecheck clean

Real models running benchmark tasks without a system prompt narrate
their reasoning in text instead of calling tools, hit the output
token limit, and never produce the required artifact.
Add systemPrompt?: string to Config. The registerBackends factory
closure reads it and passes it to AiSdkBackend directly — same
pattern as the desktop interactive path. systemPrompt is NOT routed
through BackendFactoryContext.systemPrompt (that channel is the
child-agent instruction, not the main-session prompt) and is NOT
persisted into SessionHeader (benchmark sessions are throwaway).
Also exports BENCHMARK_BASE_SYSTEM_PROMPT, a generic tool-first
prefix verified against DeepSeek-chat on a Terminal-Bench regex task.
Reviewed by Codex and Claude (consult mode) — both confirmed
headless-only / runtime-pass / no-persistence is the right scope.
…t unknown
When a backend ends with stopReason='error' but never emits a
preceding error event (observed with DeepSeek-reasoner after it
tried to call an unavailable tool and entered a reasoning loop),
the run ledger's failureClass fell back to 'unknown' — making
benchmark failures unattributable.
Now turnStatusFromEvent returns errorClass 'runtime_error' for
stopReason='error', and recordSessionEvent calls markRunFailed so
finalize does not fall back to 'unknown'. This lets the benchmark
scorer distinguish runtime failures from max_tokens / verification_failed.
…mment
Codex [P2]: BENCHMARK_BASE_SYSTEM_PROMPT said 'You do not have a Bash
tool' but the standard isolated headless surface (buildIsolatedHeadlessTools)
includes an isolated Bash. Reworded to 'do not self-test or run
verification scripts' — acknowledges Bash exists for producing artifacts,
steers away from self-grading.
Claude [P3]: complete(stopReason=error) path now emits two run_failed
events (markRunFailed + finishRun), same as the error-event path already
did. Added a comment making the conscious choice explicit.
Claude review follow-up: the existing tests proved the factory closure
can READ config.systemPrompt, but not that it actually PASSES it to the
backend constructor's systemPrompt parameter — the exact seam a real
AiSdkBackend factory uses.
Adds an ai-sdk stub backend that records its constructor systemPrompt
input, and a test that runs runExperiment with registerBackends wiring
config.systemPrompt → backend ctor (mirroring the desktop pattern).
Verifies the prompt string arrives at the backend constructor intact.
@Astro-Han
Astro-Han merged commit ab61bc5 into mainJun 19, 2026
Astro-Han added a commit that referenced this pull request Jun 19, 2026
…ror (#63)
Fresh-eye review (Codex P1 / Claude P3) found that PR #62 only fixed
the AgentRunStore ledger failureClass, not the InvocationResult.failure.class
that benchmark ResultRecord reads. A complete(stopReason=error) with no
preceding error event still surfaced as errorClass='failed' in
ResultRecord — indistinguishable from other failures.
Root cause: failureFromTerminalEvent returned class=status (i.e. 'failed')
for any failed terminal, ignoring whether error content carried a precise
reason. The error-content branch (which extracts content.reason) was
unreachable for terminal events because the terminal check ran first.
Fix: in failureFromTerminalEvent, when status='failed':
- with error content reason/code → use that (e.g. 'tool_failed')
- without error content → 'runtime_error' (not 'failed')
This makes the invocation path consistent with the run ledger, so
ResultRecord.errorClass reads 'runtime_error' for the bare-error scenario.
Tests:
- runtime: updated locked test (failed+error content w/o reason → runtime_error),
added reason-code test, added bare-failed-terminal test
- headless: added end-to-end test verifying ResultRecord.errorClass='runtime_error'
for a backend that emits only complete(stopReason=error)
@Astro-Han
Astro-Han deleted the codex/eval-systemprompt-failureclass branch July 14, 2026 05:05
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(headless): Config.systemPrompt + runtime_error failure class - #62

Merged
Astro-Han merged 4 commits into
mainfrom
codex/eval-systemprompt-failureclass
Jun 19, 2026
Merged

feat(headless): Config.systemPrompt + runtime_error failure class#62
Astro-Han merged 4 commits into
mainfrom
codex/eval-systemprompt-failureclass

Conversation

@Astro-Han

Copy link
Copy Markdown
Contributor

Summary

Two fixes that unblock real-model benchmark smoke (Terminal-Bench on DeepSeek), exposed by the diagnosis in the prior session:

  • Config.systemPrompt — real models running benchmark tasks without a system prompt narrate their reasoning in text instead of calling tools, hit the 8192 output token limit, and never produce the required artifact. Adds systemPrompt?: string to Config; the registerBackends factory closure reads it and passes it to AiSdkBackend directly (same pattern as the desktop interactive path). Not routed through BackendFactoryContext.systemPrompt (that's the child-agent instruction channel) and not persisted into SessionHeader (benchmark sessions are throwaway). Also exports BENCHMARK_BASE_SYSTEM_PROMPT, a generic tool-first prefix verified against DeepSeek-chat on a Terminal-Bench regex task.
  • failureClass: runtime_error — a backend that ends with complete(stopReason='error') without a preceding error event left the run ledger's failureClass as unknown, so benchmark scoring could not distinguish runtime failures from max_tokens / incomplete_tool_calls. Classifies it as runtime_error instead.

Both reviewed by Codex and Claude (consult mode) — confirmed headless-only / runtime-pass / no-persistence is the right scope.

Test plan

  • @maka/headless 91/91 (2 new tests: factory closure reads config.systemPrompt; undefined when omitted)
  • @maka/runtime 586/586 (1 new test: complete(stopReason=error)runtime_error not unknown)
  • npm run typecheck clean

Real models running benchmark tasks without a system prompt narrate
their reasoning in text instead of calling tools, hit the output
token limit, and never produce the required artifact.
Add systemPrompt?: string to Config. The registerBackends factory
closure reads it and passes it to AiSdkBackend directly — same
pattern as the desktop interactive path. systemPrompt is NOT routed
through BackendFactoryContext.systemPrompt (that channel is the
child-agent instruction, not the main-session prompt) and is NOT
persisted into SessionHeader (benchmark sessions are throwaway).
Also exports BENCHMARK_BASE_SYSTEM_PROMPT, a generic tool-first
prefix verified against DeepSeek-chat on a Terminal-Bench regex task.
Reviewed by Codex and Claude (consult mode) — both confirmed
headless-only / runtime-pass / no-persistence is the right scope.
…t unknown
When a backend ends with stopReason='error' but never emits a
preceding error event (observed with DeepSeek-reasoner after it
tried to call an unavailable tool and entered a reasoning loop),
the run ledger's failureClass fell back to 'unknown' — making
benchmark failures unattributable.
Now turnStatusFromEvent returns errorClass 'runtime_error' for
stopReason='error', and recordSessionEvent calls markRunFailed so
finalize does not fall back to 'unknown'. This lets the benchmark
scorer distinguish runtime failures from max_tokens / verification_failed.
…mment
Codex [P2]: BENCHMARK_BASE_SYSTEM_PROMPT said 'You do not have a Bash
tool' but the standard isolated headless surface (buildIsolatedHeadlessTools)
includes an isolated Bash. Reworded to 'do not self-test or run
verification scripts' — acknowledges Bash exists for producing artifacts,
steers away from self-grading.
Claude [P3]: complete(stopReason=error) path now emits two run_failed
events (markRunFailed + finishRun), same as the error-event path already
did. Added a comment making the conscious choice explicit.
Claude review follow-up: the existing tests proved the factory closure
can READ config.systemPrompt, but not that it actually PASSES it to the
backend constructor's systemPrompt parameter — the exact seam a real
AiSdkBackend factory uses.
Adds an ai-sdk stub backend that records its constructor systemPrompt
input, and a test that runs runExperiment with registerBackends wiring
config.systemPrompt → backend ctor (mirroring the desktop pattern).
Verifies the prompt string arrives at the backend constructor intact.
@Astro-Han
Astro-Han merged commit ab61bc5 into mainJun 19, 2026
Astro-Han added a commit that referenced this pull request Jun 19, 2026
…ror (#63)
Fresh-eye review (Codex P1 / Claude P3) found that PR #62 only fixed
the AgentRunStore ledger failureClass, not the InvocationResult.failure.class
that benchmark ResultRecord reads. A complete(stopReason=error) with no
preceding error event still surfaced as errorClass='failed' in
ResultRecord — indistinguishable from other failures.
Root cause: failureFromTerminalEvent returned class=status (i.e. 'failed')
for any failed terminal, ignoring whether error content carried a precise
reason. The error-content branch (which extracts content.reason) was
unreachable for terminal events because the terminal check ran first.
Fix: in failureFromTerminalEvent, when status='failed':
- with error content reason/code → use that (e.g. 'tool_failed')
- without error content → 'runtime_error' (not 'failed')
This makes the invocation path consistent with the run ledger, so
ResultRecord.errorClass reads 'runtime_error' for the bare-error scenario.
Tests:
- runtime: updated locked test (failed+error content w/o reason → runtime_error),
added reason-code test, added bare-failed-terminal test
- headless: added end-to-end test verifying ResultRecord.errorClass='runtime_error'
for a backend that emits only complete(stopReason=error)
@Astro-Han
Astro-Han deleted the codex/eval-systemprompt-failureclass branch July 14, 2026 05:05
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

feat(headless): Config.systemPrompt + runtime_error failure class - #62

Merged
Astro-Han merged 4 commits into
mainfrom
codex/eval-systemprompt-failureclass
Jun 19, 2026
Merged

feat(headless): Config.systemPrompt + runtime_error failure class#62
Astro-Han merged 4 commits into
mainfrom
codex/eval-systemprompt-failureclass

Conversation

@Astro-Han

Copy link
Copy Markdown
Contributor

Summary

Two fixes that unblock real-model benchmark smoke (Terminal-Bench on DeepSeek), exposed by the diagnosis in the prior session:

  • Config.systemPrompt — real models running benchmark tasks without a system prompt narrate their reasoning in text instead of calling tools, hit the 8192 output token limit, and never produce the required artifact. Adds systemPrompt?: string to Config; the registerBackends factory closure reads it and passes it to AiSdkBackend directly (same pattern as the desktop interactive path). Not routed through BackendFactoryContext.systemPrompt (that's the child-agent instruction channel) and not persisted into SessionHeader (benchmark sessions are throwaway). Also exports BENCHMARK_BASE_SYSTEM_PROMPT, a generic tool-first prefix verified against DeepSeek-chat on a Terminal-Bench regex task.
  • failureClass: runtime_error — a backend that ends with complete(stopReason='error') without a preceding error event left the run ledger's failureClass as unknown, so benchmark scoring could not distinguish runtime failures from max_tokens / incomplete_tool_calls. Classifies it as runtime_error instead.

Both reviewed by Codex and Claude (consult mode) — confirmed headless-only / runtime-pass / no-persistence is the right scope.

Test plan

  • @maka/headless 91/91 (2 new tests: factory closure reads config.systemPrompt; undefined when omitted)
  • @maka/runtime 586/586 (1 new test: complete(stopReason=error)runtime_error not unknown)
  • npm run typecheck clean

Real models running benchmark tasks without a system prompt narrate
their reasoning in text instead of calling tools, hit the output
token limit, and never produce the required artifact.
Add systemPrompt?: string to Config. The registerBackends factory
closure reads it and passes it to AiSdkBackend directly — same
pattern as the desktop interactive path. systemPrompt is NOT routed
through BackendFactoryContext.systemPrompt (that channel is the
child-agent instruction, not the main-session prompt) and is NOT
persisted into SessionHeader (benchmark sessions are throwaway).
Also exports BENCHMARK_BASE_SYSTEM_PROMPT, a generic tool-first
prefix verified against DeepSeek-chat on a Terminal-Bench regex task.
Reviewed by Codex and Claude (consult mode) — both confirmed
headless-only / runtime-pass / no-persistence is the right scope.
…t unknown
When a backend ends with stopReason='error' but never emits a
preceding error event (observed with DeepSeek-reasoner after it
tried to call an unavailable tool and entered a reasoning loop),
the run ledger's failureClass fell back to 'unknown' — making
benchmark failures unattributable.
Now turnStatusFromEvent returns errorClass 'runtime_error' for
stopReason='error', and recordSessionEvent calls markRunFailed so
finalize does not fall back to 'unknown'. This lets the benchmark
scorer distinguish runtime failures from max_tokens / verification_failed.
…mment
Codex [P2]: BENCHMARK_BASE_SYSTEM_PROMPT said 'You do not have a Bash
tool' but the standard isolated headless surface (buildIsolatedHeadlessTools)
includes an isolated Bash. Reworded to 'do not self-test or run
verification scripts' — acknowledges Bash exists for producing artifacts,
steers away from self-grading.
Claude [P3]: complete(stopReason=error) path now emits two run_failed
events (markRunFailed + finishRun), same as the error-event path already
did. Added a comment making the conscious choice explicit.
Claude review follow-up: the existing tests proved the factory closure
can READ config.systemPrompt, but not that it actually PASSES it to the
backend constructor's systemPrompt parameter — the exact seam a real
AiSdkBackend factory uses.
Adds an ai-sdk stub backend that records its constructor systemPrompt
input, and a test that runs runExperiment with registerBackends wiring
config.systemPrompt → backend ctor (mirroring the desktop pattern).
Verifies the prompt string arrives at the backend constructor intact.
@Astro-Han
Astro-Han merged commit ab61bc5 into mainJun 19, 2026
Astro-Han added a commit that referenced this pull request Jun 19, 2026
…ror (#63)
Fresh-eye review (Codex P1 / Claude P3) found that PR #62 only fixed
the AgentRunStore ledger failureClass, not the InvocationResult.failure.class
that benchmark ResultRecord reads. A complete(stopReason=error) with no
preceding error event still surfaced as errorClass='failed' in
ResultRecord — indistinguishable from other failures.
Root cause: failureFromTerminalEvent returned class=status (i.e. 'failed')
for any failed terminal, ignoring whether error content carried a precise
reason. The error-content branch (which extracts content.reason) was
unreachable for terminal events because the terminal check ran first.
Fix: in failureFromTerminalEvent, when status='failed':
- with error content reason/code → use that (e.g. 'tool_failed')
- without error content → 'runtime_error' (not 'failed')
This makes the invocation path consistent with the run ledger, so
ResultRecord.errorClass reads 'runtime_error' for the bare-error scenario.
Tests:
- runtime: updated locked test (failed+error content w/o reason → runtime_error),
added reason-code test, added bare-failed-terminal test
- headless: added end-to-end test verifying ResultRecord.errorClass='runtime_error'
for a backend that emits only complete(stopReason=error)
@Astro-Han
Astro-Han deleted the codex/eval-systemprompt-failureclass branch July 14, 2026 05:05
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han