Skip to content

fix(headless): measure the wire the Maka runtime actually dials - #2286

Merged
Astro-Han merged 5 commits into
mainfrom
fix/headless-proxy-usage-protocol-wire
Aug 6, 2026
Merged

fix(headless): measure the wire the Maka runtime actually dials#2286
Astro-Han merged 5 commits into
mainfrom
fix/headless-proxy-usage-protocol-wire

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Aug 5, 2026

Copy link
Copy Markdown
Contributor

Summary

Two independent measurement bugs in the harness, both surfaced by the #2245 two-arm rerun.

1. The proxy measured a wire the Maka arm was not dialing. The provider proxy inferred its usage protocol from runtimeAdapter.kind, so every openai-compatible provider was parsed as Chat SSE. That inference went silently wrong when #2152 moved deepseek-v4-flash to the Responses API: the proxy never saw data: [DONE], recorded every request as interrupted with zero usage and no reasoning tokens, and incompleteTerminalProviderRequest then threw the whole graded cell away as an infra failure — cells that had in fact passed with reward 1.0.

The Maka arm runs the Maka runtime, so resolveModelRuntime is the authority on which wire it dials; ask it instead of guessing. Competitors run their own CLI and keep the adapter-kind path, which is what describes them. The A/B manifest's per-arm transport comes from the same call, so a run now records the protocol it measured rather than the one it assumed — the frozen three-way composition test asserted openai-chat for the Maka arm, describing a run that never happened.

packages/runtime gains a ./model-runtime export so the harness can reach the resolver.

2. The infra-retry allowlist could not name a dotted task.executeHarnessArmCohort spelled candidate round ids as ab-${arm}-r0-${task.id} by hand, while buildAbRoundId — the helper that actually writes them — normalizes . to -. A retry naming install-windows-3.11 was therefore rejected as unknown, and because the check throws rather than skips, that one task took the whole batch down with it: a rerun of three infra-failed cells exited in twelve seconds having run none of them, twice. Building the candidates with the same helper removes the second spelling.

Refs #2245.

Verification

npm run -w @maka/headless test — 1424 pass / 0 fail. npm run -w @maka/runtime test — 3270 pass / 0 fail. npm run lint and npm run format clean.

The dotted-retry regression test was verified red before the fix (adjudicated infra retry names unknown round ab-maka-r0-install-windows-3-11) and green after.

Reproduced (1) against the live 5-task Terminal-Bench 2.1 pilot on deepseek-v4-flash (the same seed order as the 2026-08-04 four-arm run, so count-dataset-tokens is the same cell in all three columns):

2026-08-04 (Chat wire)before this fixafter
path / parsed protocol/chat/completions · chat-sse/responses · chat-sse/responses · responses-sse
requests24, completed 2420, interrupted 2024, completed 24
reasoning tokens3358null (0 measured)5154 (24/24 measured)
cell verdictpassinfra_failedcompleted, pass

A second pilot cell (regex-log) is likewise 15/15 completed with 17058 reasoning tokens measured.

The full 89-task two-arm run then completed on this branch: 3479 Maka requests over 89 cells with 9 non-completed, versus every request non-completed before the fix.

Root cause

Not #2241/#2249. fc581d855 (#2152, provider-native web search) added deepseek-v4-flash → openai-responses to openAiAdapterApiProtocol. The 2026-08-04 benchmark ran on 9dae4bee, which predates it and therefore dialled Chat Completions; the wire changed under a harness that kept guessing the old one.

Review round 2

Review found the fix's own API could re-admit the bug it closes. providerProxyUsageProtocol took the model id as an optional, prefix-sensitive argument, and the two call sites applied the discipline while the manifest call did not. resolveModelRuntime does not recognize the catalog spelling deepseek/deepseek-v4-flash, so a caller forwarding it raw silently got the Chat guess back — the exact wrong number this PR exists to remove — and a caller passing nothing at all got it too, with TypeScript unable to say so.

The function now normalizes the id itself and refuses to guess for the maka arm. That deletes the caller-side discipline at both runners rather than testing it: there is no longer a spelling a call site can get wrong. Two assertions pin it, including the assert.throws for a missing id.

Unknown retry round ids are also now reported together, so an operator recovering a sweep with a batch of ids fixes every typo in one pass.

The provider proxy inferred its usage protocol from the adapter kind, so
every `openai-compatible` provider was parsed as Chat SSE. That inference
silently went wrong when deepseek-v4-flash moved to the Responses API:
the proxy never saw `data: [DONE]`, recorded all 20 requests of a cell as
`interrupted` with zero usage and no reasoning tokens, and the runner's
terminal-request check then threw the whole graded cell away as an infra
failure -- a cell that had in fact passed with reward 1.0.
The Maka arm runs the Maka runtime, so `resolveModelRuntime` is the
authority on which wire it dials; ask it instead of guessing. Competitors
run their own CLI and keep the adapter-kind path, which is what describes
them. The A/B manifest's per-arm transport comes from the same call, so a
run now records the protocol it measured rather than the one it assumed.
Verified on the 5-task pilot the guess had been failing.
…s them
The retry allowlist spelled `ab-${arm}-r0-${task.id}` by hand while
`buildAbRoundId` normalizes `.` to `-`, so a retry naming a dotted task
never matched. Because the check throws instead of skipping, one such
task took the whole batch down: a rerun of three infra-failed cells
exited in twelve seconds without running any of them.
…er spells the model
Review found the fix's own API could re-admit the bug it closes: the model id
was prefix-sensitive and optional, so a caller passing the catalog spelling
`deepseek/deepseek-v4-flash` — which `resolveModelRuntime` does not recognize
— silently got the Chat guess back, and a caller passing nothing at all got
it too. Normalize inside the function and refuse to guess for the maka arm,
which deletes the caller-side discipline at both runners rather than testing
it.
Also report every unknown retry round id at once.
… will build
The proxy asked the runtime which wire it would dial, but asked about a
connection the runtime never builds: an advertised protocol only reaches a
model entry for GitHub Copilot and the Kimi Coding Plan, and `connectionFromEnv`
drops it everywhere else. With MAKA_MODEL_API_PROTOCOL=openai-chat set against
DeepSeek, the proxy resolved Chat SSE while the runtime dialled Responses —
the wrong-wire failure this path exists to remove, readmitted through its own
input.
The gating rule is now one exported function that both sides read, and the
test that pinned the divergent answer asserts the runtime's.
The proxy path parsed the three protocol names with its own copy of the check
`provider-env` already owned, next to the gating rule it just started sharing
with it. Two copies of a value whitelist is how one of them ends up accepting
a name the other rejects.
@Astro-Han
Astro-Han marked this pull request as ready for review August 6, 2026 06:07
@Astro-Han
Astro-Han merged commit 5b7dbb4 into mainAug 6, 2026
12 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
fix(headless): measure the wire the Maka runtime actually dials by Astro-Han · Pull Request #2286 · apache/maka · GitHub
Skip to content

fix(headless): measure the wire the Maka runtime actually dials - #2286

Merged
Astro-Han merged 5 commits into
mainfrom
fix/headless-proxy-usage-protocol-wire
Aug 6, 2026
Merged

fix(headless): measure the wire the Maka runtime actually dials#2286
Astro-Han merged 5 commits into
mainfrom
fix/headless-proxy-usage-protocol-wire

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Aug 5, 2026

Copy link
Copy Markdown
Contributor

Summary

Two independent measurement bugs in the harness, both surfaced by the #2245 two-arm rerun.

1. The proxy measured a wire the Maka arm was not dialing. The provider proxy inferred its usage protocol from runtimeAdapter.kind, so every openai-compatible provider was parsed as Chat SSE. That inference went silently wrong when #2152 moved deepseek-v4-flash to the Responses API: the proxy never saw data: [DONE], recorded every request as interrupted with zero usage and no reasoning tokens, and incompleteTerminalProviderRequest then threw the whole graded cell away as an infra failure — cells that had in fact passed with reward 1.0.

The Maka arm runs the Maka runtime, so resolveModelRuntime is the authority on which wire it dials; ask it instead of guessing. Competitors run their own CLI and keep the adapter-kind path, which is what describes them. The A/B manifest's per-arm transport comes from the same call, so a run now records the protocol it measured rather than the one it assumed — the frozen three-way composition test asserted openai-chat for the Maka arm, describing a run that never happened.

packages/runtime gains a ./model-runtime export so the harness can reach the resolver.

2. The infra-retry allowlist could not name a dotted task.executeHarnessArmCohort spelled candidate round ids as ab-${arm}-r0-${task.id} by hand, while buildAbRoundId — the helper that actually writes them — normalizes . to -. A retry naming install-windows-3.11 was therefore rejected as unknown, and because the check throws rather than skips, that one task took the whole batch down with it: a rerun of three infra-failed cells exited in twelve seconds having run none of them, twice. Building the candidates with the same helper removes the second spelling.

Refs #2245.

Verification

npm run -w @maka/headless test — 1424 pass / 0 fail. npm run -w @maka/runtime test — 3270 pass / 0 fail. npm run lint and npm run format clean.

The dotted-retry regression test was verified red before the fix (adjudicated infra retry names unknown round ab-maka-r0-install-windows-3-11) and green after.

Reproduced (1) against the live 5-task Terminal-Bench 2.1 pilot on deepseek-v4-flash (the same seed order as the 2026-08-04 four-arm run, so count-dataset-tokens is the same cell in all three columns):

2026-08-04 (Chat wire)before this fixafter
path / parsed protocol/chat/completions · chat-sse/responses · chat-sse/responses · responses-sse
requests24, completed 2420, interrupted 2024, completed 24
reasoning tokens3358null (0 measured)5154 (24/24 measured)
cell verdictpassinfra_failedcompleted, pass

A second pilot cell (regex-log) is likewise 15/15 completed with 17058 reasoning tokens measured.

The full 89-task two-arm run then completed on this branch: 3479 Maka requests over 89 cells with 9 non-completed, versus every request non-completed before the fix.

Root cause

Not #2241/#2249. fc581d855 (#2152, provider-native web search) added deepseek-v4-flash → openai-responses to openAiAdapterApiProtocol. The 2026-08-04 benchmark ran on 9dae4bee, which predates it and therefore dialled Chat Completions; the wire changed under a harness that kept guessing the old one.

Review round 2

Review found the fix's own API could re-admit the bug it closes. providerProxyUsageProtocol took the model id as an optional, prefix-sensitive argument, and the two call sites applied the discipline while the manifest call did not. resolveModelRuntime does not recognize the catalog spelling deepseek/deepseek-v4-flash, so a caller forwarding it raw silently got the Chat guess back — the exact wrong number this PR exists to remove — and a caller passing nothing at all got it too, with TypeScript unable to say so.

The function now normalizes the id itself and refuses to guess for the maka arm. That deletes the caller-side discipline at both runners rather than testing it: there is no longer a spelling a call site can get wrong. Two assertions pin it, including the assert.throws for a missing id.

Unknown retry round ids are also now reported together, so an operator recovering a sweep with a batch of ids fixes every typo in one pass.

The provider proxy inferred its usage protocol from the adapter kind, so
every `openai-compatible` provider was parsed as Chat SSE. That inference
silently went wrong when deepseek-v4-flash moved to the Responses API:
the proxy never saw `data: [DONE]`, recorded all 20 requests of a cell as
`interrupted` with zero usage and no reasoning tokens, and the runner's
terminal-request check then threw the whole graded cell away as an infra
failure -- a cell that had in fact passed with reward 1.0.
The Maka arm runs the Maka runtime, so `resolveModelRuntime` is the
authority on which wire it dials; ask it instead of guessing. Competitors
run their own CLI and keep the adapter-kind path, which is what describes
them. The A/B manifest's per-arm transport comes from the same call, so a
run now records the protocol it measured rather than the one it assumed.
Verified on the 5-task pilot the guess had been failing.
…s them
The retry allowlist spelled `ab-${arm}-r0-${task.id}` by hand while
`buildAbRoundId` normalizes `.` to `-`, so a retry naming a dotted task
never matched. Because the check throws instead of skipping, one such
task took the whole batch down: a rerun of three infra-failed cells
exited in twelve seconds without running any of them.
…er spells the model
Review found the fix's own API could re-admit the bug it closes: the model id
was prefix-sensitive and optional, so a caller passing the catalog spelling
`deepseek/deepseek-v4-flash` — which `resolveModelRuntime` does not recognize
— silently got the Chat guess back, and a caller passing nothing at all got
it too. Normalize inside the function and refuse to guess for the maka arm,
which deletes the caller-side discipline at both runners rather than testing
it.
Also report every unknown retry round id at once.
… will build
The proxy asked the runtime which wire it would dial, but asked about a
connection the runtime never builds: an advertised protocol only reaches a
model entry for GitHub Copilot and the Kimi Coding Plan, and `connectionFromEnv`
drops it everywhere else. With MAKA_MODEL_API_PROTOCOL=openai-chat set against
DeepSeek, the proxy resolved Chat SSE while the runtime dialled Responses —
the wrong-wire failure this path exists to remove, readmitted through its own
input.
The gating rule is now one exported function that both sides read, and the
test that pinned the divergent answer asserts the runtime's.
The proxy path parsed the three protocol names with its own copy of the check
`provider-env` already owned, next to the gating rule it just started sharing
with it. Two copies of a value whitelist is how one of them ends up accepting
a name the other rejects.
@Astro-Han
Astro-Han marked this pull request as ready for review August 6, 2026 06:07
@Astro-Han
Astro-Han merged commit 5b7dbb4 into mainAug 6, 2026
12 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix(headless): measure the wire the Maka runtime actually dials by Astro-Han · Pull Request #2286 · apache/maka · GitHub
Skip to content

fix(headless): measure the wire the Maka runtime actually dials - #2286

Merged
Astro-Han merged 5 commits into
mainfrom
fix/headless-proxy-usage-protocol-wire
Aug 6, 2026
Merged

fix(headless): measure the wire the Maka runtime actually dials#2286
Astro-Han merged 5 commits into
mainfrom
fix/headless-proxy-usage-protocol-wire

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Aug 5, 2026

Copy link
Copy Markdown
Contributor

Summary

Two independent measurement bugs in the harness, both surfaced by the #2245 two-arm rerun.

1. The proxy measured a wire the Maka arm was not dialing. The provider proxy inferred its usage protocol from runtimeAdapter.kind, so every openai-compatible provider was parsed as Chat SSE. That inference went silently wrong when #2152 moved deepseek-v4-flash to the Responses API: the proxy never saw data: [DONE], recorded every request as interrupted with zero usage and no reasoning tokens, and incompleteTerminalProviderRequest then threw the whole graded cell away as an infra failure — cells that had in fact passed with reward 1.0.

The Maka arm runs the Maka runtime, so resolveModelRuntime is the authority on which wire it dials; ask it instead of guessing. Competitors run their own CLI and keep the adapter-kind path, which is what describes them. The A/B manifest's per-arm transport comes from the same call, so a run now records the protocol it measured rather than the one it assumed — the frozen three-way composition test asserted openai-chat for the Maka arm, describing a run that never happened.

packages/runtime gains a ./model-runtime export so the harness can reach the resolver.

2. The infra-retry allowlist could not name a dotted task.executeHarnessArmCohort spelled candidate round ids as ab-${arm}-r0-${task.id} by hand, while buildAbRoundId — the helper that actually writes them — normalizes . to -. A retry naming install-windows-3.11 was therefore rejected as unknown, and because the check throws rather than skips, that one task took the whole batch down with it: a rerun of three infra-failed cells exited in twelve seconds having run none of them, twice. Building the candidates with the same helper removes the second spelling.

Refs #2245.

Verification

npm run -w @maka/headless test — 1424 pass / 0 fail. npm run -w @maka/runtime test — 3270 pass / 0 fail. npm run lint and npm run format clean.

The dotted-retry regression test was verified red before the fix (adjudicated infra retry names unknown round ab-maka-r0-install-windows-3-11) and green after.

Reproduced (1) against the live 5-task Terminal-Bench 2.1 pilot on deepseek-v4-flash (the same seed order as the 2026-08-04 four-arm run, so count-dataset-tokens is the same cell in all three columns):

2026-08-04 (Chat wire)before this fixafter
path / parsed protocol/chat/completions · chat-sse/responses · chat-sse/responses · responses-sse
requests24, completed 2420, interrupted 2024, completed 24
reasoning tokens3358null (0 measured)5154 (24/24 measured)
cell verdictpassinfra_failedcompleted, pass

A second pilot cell (regex-log) is likewise 15/15 completed with 17058 reasoning tokens measured.

The full 89-task two-arm run then completed on this branch: 3479 Maka requests over 89 cells with 9 non-completed, versus every request non-completed before the fix.

Root cause

Not #2241/#2249. fc581d855 (#2152, provider-native web search) added deepseek-v4-flash → openai-responses to openAiAdapterApiProtocol. The 2026-08-04 benchmark ran on 9dae4bee, which predates it and therefore dialled Chat Completions; the wire changed under a harness that kept guessing the old one.

Review round 2

Review found the fix's own API could re-admit the bug it closes. providerProxyUsageProtocol took the model id as an optional, prefix-sensitive argument, and the two call sites applied the discipline while the manifest call did not. resolveModelRuntime does not recognize the catalog spelling deepseek/deepseek-v4-flash, so a caller forwarding it raw silently got the Chat guess back — the exact wrong number this PR exists to remove — and a caller passing nothing at all got it too, with TypeScript unable to say so.

The function now normalizes the id itself and refuses to guess for the maka arm. That deletes the caller-side discipline at both runners rather than testing it: there is no longer a spelling a call site can get wrong. Two assertions pin it, including the assert.throws for a missing id.

Unknown retry round ids are also now reported together, so an operator recovering a sweep with a batch of ids fixes every typo in one pass.

The provider proxy inferred its usage protocol from the adapter kind, so
every `openai-compatible` provider was parsed as Chat SSE. That inference
silently went wrong when deepseek-v4-flash moved to the Responses API:
the proxy never saw `data: [DONE]`, recorded all 20 requests of a cell as
`interrupted` with zero usage and no reasoning tokens, and the runner's
terminal-request check then threw the whole graded cell away as an infra
failure -- a cell that had in fact passed with reward 1.0.
The Maka arm runs the Maka runtime, so `resolveModelRuntime` is the
authority on which wire it dials; ask it instead of guessing. Competitors
run their own CLI and keep the adapter-kind path, which is what describes
them. The A/B manifest's per-arm transport comes from the same call, so a
run now records the protocol it measured rather than the one it assumed.
Verified on the 5-task pilot the guess had been failing.
…s them
The retry allowlist spelled `ab-${arm}-r0-${task.id}` by hand while
`buildAbRoundId` normalizes `.` to `-`, so a retry naming a dotted task
never matched. Because the check throws instead of skipping, one such
task took the whole batch down: a rerun of three infra-failed cells
exited in twelve seconds without running any of them.
…er spells the model
Review found the fix's own API could re-admit the bug it closes: the model id
was prefix-sensitive and optional, so a caller passing the catalog spelling
`deepseek/deepseek-v4-flash` — which `resolveModelRuntime` does not recognize
— silently got the Chat guess back, and a caller passing nothing at all got
it too. Normalize inside the function and refuse to guess for the maka arm,
which deletes the caller-side discipline at both runners rather than testing
it.
Also report every unknown retry round id at once.
… will build
The proxy asked the runtime which wire it would dial, but asked about a
connection the runtime never builds: an advertised protocol only reaches a
model entry for GitHub Copilot and the Kimi Coding Plan, and `connectionFromEnv`
drops it everywhere else. With MAKA_MODEL_API_PROTOCOL=openai-chat set against
DeepSeek, the proxy resolved Chat SSE while the runtime dialled Responses —
the wrong-wire failure this path exists to remove, readmitted through its own
input.
The gating rule is now one exported function that both sides read, and the
test that pinned the divergent answer asserts the runtime's.
The proxy path parsed the three protocol names with its own copy of the check
`provider-env` already owned, next to the gating rule it just started sharing
with it. Two copies of a value whitelist is how one of them ends up accepting
a name the other rejects.
@Astro-Han
Astro-Han marked this pull request as ready for review August 6, 2026 06:07
@Astro-Han
Astro-Han merged commit 5b7dbb4 into mainAug 6, 2026
12 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix(headless): measure the wire the Maka runtime actually dials by Astro-Han · Pull Request #2286 · apache/maka · GitHub
Skip to content

fix(headless): measure the wire the Maka runtime actually dials - #2286

Merged
Astro-Han merged 5 commits into
mainfrom
fix/headless-proxy-usage-protocol-wire
Aug 6, 2026
Merged

fix(headless): measure the wire the Maka runtime actually dials#2286
Astro-Han merged 5 commits into
mainfrom
fix/headless-proxy-usage-protocol-wire

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Aug 5, 2026

Copy link
Copy Markdown
Contributor

Summary

Two independent measurement bugs in the harness, both surfaced by the #2245 two-arm rerun.

1. The proxy measured a wire the Maka arm was not dialing. The provider proxy inferred its usage protocol from runtimeAdapter.kind, so every openai-compatible provider was parsed as Chat SSE. That inference went silently wrong when #2152 moved deepseek-v4-flash to the Responses API: the proxy never saw data: [DONE], recorded every request as interrupted with zero usage and no reasoning tokens, and incompleteTerminalProviderRequest then threw the whole graded cell away as an infra failure — cells that had in fact passed with reward 1.0.

The Maka arm runs the Maka runtime, so resolveModelRuntime is the authority on which wire it dials; ask it instead of guessing. Competitors run their own CLI and keep the adapter-kind path, which is what describes them. The A/B manifest's per-arm transport comes from the same call, so a run now records the protocol it measured rather than the one it assumed — the frozen three-way composition test asserted openai-chat for the Maka arm, describing a run that never happened.

packages/runtime gains a ./model-runtime export so the harness can reach the resolver.

2. The infra-retry allowlist could not name a dotted task.executeHarnessArmCohort spelled candidate round ids as ab-${arm}-r0-${task.id} by hand, while buildAbRoundId — the helper that actually writes them — normalizes . to -. A retry naming install-windows-3.11 was therefore rejected as unknown, and because the check throws rather than skips, that one task took the whole batch down with it: a rerun of three infra-failed cells exited in twelve seconds having run none of them, twice. Building the candidates with the same helper removes the second spelling.

Refs #2245.

Verification

npm run -w @maka/headless test — 1424 pass / 0 fail. npm run -w @maka/runtime test — 3270 pass / 0 fail. npm run lint and npm run format clean.

The dotted-retry regression test was verified red before the fix (adjudicated infra retry names unknown round ab-maka-r0-install-windows-3-11) and green after.

Reproduced (1) against the live 5-task Terminal-Bench 2.1 pilot on deepseek-v4-flash (the same seed order as the 2026-08-04 four-arm run, so count-dataset-tokens is the same cell in all three columns):

2026-08-04 (Chat wire)before this fixafter
path / parsed protocol/chat/completions · chat-sse/responses · chat-sse/responses · responses-sse
requests24, completed 2420, interrupted 2024, completed 24
reasoning tokens3358null (0 measured)5154 (24/24 measured)
cell verdictpassinfra_failedcompleted, pass

A second pilot cell (regex-log) is likewise 15/15 completed with 17058 reasoning tokens measured.

The full 89-task two-arm run then completed on this branch: 3479 Maka requests over 89 cells with 9 non-completed, versus every request non-completed before the fix.

Root cause

Not #2241/#2249. fc581d855 (#2152, provider-native web search) added deepseek-v4-flash → openai-responses to openAiAdapterApiProtocol. The 2026-08-04 benchmark ran on 9dae4bee, which predates it and therefore dialled Chat Completions; the wire changed under a harness that kept guessing the old one.

Review round 2

Review found the fix's own API could re-admit the bug it closes. providerProxyUsageProtocol took the model id as an optional, prefix-sensitive argument, and the two call sites applied the discipline while the manifest call did not. resolveModelRuntime does not recognize the catalog spelling deepseek/deepseek-v4-flash, so a caller forwarding it raw silently got the Chat guess back — the exact wrong number this PR exists to remove — and a caller passing nothing at all got it too, with TypeScript unable to say so.

The function now normalizes the id itself and refuses to guess for the maka arm. That deletes the caller-side discipline at both runners rather than testing it: there is no longer a spelling a call site can get wrong. Two assertions pin it, including the assert.throws for a missing id.

Unknown retry round ids are also now reported together, so an operator recovering a sweep with a batch of ids fixes every typo in one pass.

The provider proxy inferred its usage protocol from the adapter kind, so
every `openai-compatible` provider was parsed as Chat SSE. That inference
silently went wrong when deepseek-v4-flash moved to the Responses API:
the proxy never saw `data: [DONE]`, recorded all 20 requests of a cell as
`interrupted` with zero usage and no reasoning tokens, and the runner's
terminal-request check then threw the whole graded cell away as an infra
failure -- a cell that had in fact passed with reward 1.0.
The Maka arm runs the Maka runtime, so `resolveModelRuntime` is the
authority on which wire it dials; ask it instead of guessing. Competitors
run their own CLI and keep the adapter-kind path, which is what describes
them. The A/B manifest's per-arm transport comes from the same call, so a
run now records the protocol it measured rather than the one it assumed.
Verified on the 5-task pilot the guess had been failing.
…s them
The retry allowlist spelled `ab-${arm}-r0-${task.id}` by hand while
`buildAbRoundId` normalizes `.` to `-`, so a retry naming a dotted task
never matched. Because the check throws instead of skipping, one such
task took the whole batch down: a rerun of three infra-failed cells
exited in twelve seconds without running any of them.
…er spells the model
Review found the fix's own API could re-admit the bug it closes: the model id
was prefix-sensitive and optional, so a caller passing the catalog spelling
`deepseek/deepseek-v4-flash` — which `resolveModelRuntime` does not recognize
— silently got the Chat guess back, and a caller passing nothing at all got
it too. Normalize inside the function and refuse to guess for the maka arm,
which deletes the caller-side discipline at both runners rather than testing
it.
Also report every unknown retry round id at once.
… will build
The proxy asked the runtime which wire it would dial, but asked about a
connection the runtime never builds: an advertised protocol only reaches a
model entry for GitHub Copilot and the Kimi Coding Plan, and `connectionFromEnv`
drops it everywhere else. With MAKA_MODEL_API_PROTOCOL=openai-chat set against
DeepSeek, the proxy resolved Chat SSE while the runtime dialled Responses —
the wrong-wire failure this path exists to remove, readmitted through its own
input.
The gating rule is now one exported function that both sides read, and the
test that pinned the divergent answer asserts the runtime's.
The proxy path parsed the three protocol names with its own copy of the check
`provider-env` already owned, next to the gating rule it just started sharing
with it. Two copies of a value whitelist is how one of them ends up accepting
a name the other rejects.
@Astro-Han
Astro-Han marked this pull request as ready for review August 6, 2026 06:07
@Astro-Han
Astro-Han merged commit 5b7dbb4 into mainAug 6, 2026
12 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' fix(headless): measure the wire the Maka runtime actually dials by Astro-Han · Pull Request #2286 · apache/maka · GitHub
Skip to content

fix(headless): measure the wire the Maka runtime actually dials - #2286

Merged
Astro-Han merged 5 commits into
mainfrom
fix/headless-proxy-usage-protocol-wire
Aug 6, 2026
Merged

fix(headless): measure the wire the Maka runtime actually dials#2286
Astro-Han merged 5 commits into
mainfrom
fix/headless-proxy-usage-protocol-wire

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Aug 5, 2026

Copy link
Copy Markdown
Contributor

Summary

Two independent measurement bugs in the harness, both surfaced by the #2245 two-arm rerun.

1. The proxy measured a wire the Maka arm was not dialing. The provider proxy inferred its usage protocol from runtimeAdapter.kind, so every openai-compatible provider was parsed as Chat SSE. That inference went silently wrong when #2152 moved deepseek-v4-flash to the Responses API: the proxy never saw data: [DONE], recorded every request as interrupted with zero usage and no reasoning tokens, and incompleteTerminalProviderRequest then threw the whole graded cell away as an infra failure — cells that had in fact passed with reward 1.0.

The Maka arm runs the Maka runtime, so resolveModelRuntime is the authority on which wire it dials; ask it instead of guessing. Competitors run their own CLI and keep the adapter-kind path, which is what describes them. The A/B manifest's per-arm transport comes from the same call, so a run now records the protocol it measured rather than the one it assumed — the frozen three-way composition test asserted openai-chat for the Maka arm, describing a run that never happened.

packages/runtime gains a ./model-runtime export so the harness can reach the resolver.

2. The infra-retry allowlist could not name a dotted task.executeHarnessArmCohort spelled candidate round ids as ab-${arm}-r0-${task.id} by hand, while buildAbRoundId — the helper that actually writes them — normalizes . to -. A retry naming install-windows-3.11 was therefore rejected as unknown, and because the check throws rather than skips, that one task took the whole batch down with it: a rerun of three infra-failed cells exited in twelve seconds having run none of them, twice. Building the candidates with the same helper removes the second spelling.

Refs #2245.

Verification

npm run -w @maka/headless test — 1424 pass / 0 fail. npm run -w @maka/runtime test — 3270 pass / 0 fail. npm run lint and npm run format clean.

The dotted-retry regression test was verified red before the fix (adjudicated infra retry names unknown round ab-maka-r0-install-windows-3-11) and green after.

Reproduced (1) against the live 5-task Terminal-Bench 2.1 pilot on deepseek-v4-flash (the same seed order as the 2026-08-04 four-arm run, so count-dataset-tokens is the same cell in all three columns):

2026-08-04 (Chat wire)before this fixafter
path / parsed protocol/chat/completions · chat-sse/responses · chat-sse/responses · responses-sse
requests24, completed 2420, interrupted 2024, completed 24
reasoning tokens3358null (0 measured)5154 (24/24 measured)
cell verdictpassinfra_failedcompleted, pass

A second pilot cell (regex-log) is likewise 15/15 completed with 17058 reasoning tokens measured.

The full 89-task two-arm run then completed on this branch: 3479 Maka requests over 89 cells with 9 non-completed, versus every request non-completed before the fix.

Root cause

Not #2241/#2249. fc581d855 (#2152, provider-native web search) added deepseek-v4-flash → openai-responses to openAiAdapterApiProtocol. The 2026-08-04 benchmark ran on 9dae4bee, which predates it and therefore dialled Chat Completions; the wire changed under a harness that kept guessing the old one.

Review round 2

Review found the fix's own API could re-admit the bug it closes. providerProxyUsageProtocol took the model id as an optional, prefix-sensitive argument, and the two call sites applied the discipline while the manifest call did not. resolveModelRuntime does not recognize the catalog spelling deepseek/deepseek-v4-flash, so a caller forwarding it raw silently got the Chat guess back — the exact wrong number this PR exists to remove — and a caller passing nothing at all got it too, with TypeScript unable to say so.

The function now normalizes the id itself and refuses to guess for the maka arm. That deletes the caller-side discipline at both runners rather than testing it: there is no longer a spelling a call site can get wrong. Two assertions pin it, including the assert.throws for a missing id.

Unknown retry round ids are also now reported together, so an operator recovering a sweep with a batch of ids fixes every typo in one pass.

The provider proxy inferred its usage protocol from the adapter kind, so
every `openai-compatible` provider was parsed as Chat SSE. That inference
silently went wrong when deepseek-v4-flash moved to the Responses API:
the proxy never saw `data: [DONE]`, recorded all 20 requests of a cell as
`interrupted` with zero usage and no reasoning tokens, and the runner's
terminal-request check then threw the whole graded cell away as an infra
failure -- a cell that had in fact passed with reward 1.0.
The Maka arm runs the Maka runtime, so `resolveModelRuntime` is the
authority on which wire it dials; ask it instead of guessing. Competitors
run their own CLI and keep the adapter-kind path, which is what describes
them. The A/B manifest's per-arm transport comes from the same call, so a
run now records the protocol it measured rather than the one it assumed.
Verified on the 5-task pilot the guess had been failing.
…s them
The retry allowlist spelled `ab-${arm}-r0-${task.id}` by hand while
`buildAbRoundId` normalizes `.` to `-`, so a retry naming a dotted task
never matched. Because the check throws instead of skipping, one such
task took the whole batch down: a rerun of three infra-failed cells
exited in twelve seconds without running any of them.
…er spells the model
Review found the fix's own API could re-admit the bug it closes: the model id
was prefix-sensitive and optional, so a caller passing the catalog spelling
`deepseek/deepseek-v4-flash` — which `resolveModelRuntime` does not recognize
— silently got the Chat guess back, and a caller passing nothing at all got
it too. Normalize inside the function and refuse to guess for the maka arm,
which deletes the caller-side discipline at both runners rather than testing
it.
Also report every unknown retry round id at once.
… will build
The proxy asked the runtime which wire it would dial, but asked about a
connection the runtime never builds: an advertised protocol only reaches a
model entry for GitHub Copilot and the Kimi Coding Plan, and `connectionFromEnv`
drops it everywhere else. With MAKA_MODEL_API_PROTOCOL=openai-chat set against
DeepSeek, the proxy resolved Chat SSE while the runtime dialled Responses —
the wrong-wire failure this path exists to remove, readmitted through its own
input.
The gating rule is now one exported function that both sides read, and the
test that pinned the divergent answer asserts the runtime's.
The proxy path parsed the three protocol names with its own copy of the check
`provider-env` already owned, next to the gating rule it just started sharing
with it. Two copies of a value whitelist is how one of them ends up accepting
a name the other rejects.
@Astro-Han
Astro-Han marked this pull request as ready for review August 6, 2026 06:07
@Astro-Han
Astro-Han merged commit 5b7dbb4 into mainAug 6, 2026
12 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix(headless): measure the wire the Maka runtime actually dials by Astro-Han · Pull Request #2286 · apache/maka · GitHub
Skip to content

fix(headless): measure the wire the Maka runtime actually dials - #2286

Merged
Astro-Han merged 5 commits into
mainfrom
fix/headless-proxy-usage-protocol-wire
Aug 6, 2026
Merged

fix(headless): measure the wire the Maka runtime actually dials#2286
Astro-Han merged 5 commits into
mainfrom
fix/headless-proxy-usage-protocol-wire

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Aug 5, 2026

Copy link
Copy Markdown
Contributor

Summary

Two independent measurement bugs in the harness, both surfaced by the #2245 two-arm rerun.

1. The proxy measured a wire the Maka arm was not dialing. The provider proxy inferred its usage protocol from runtimeAdapter.kind, so every openai-compatible provider was parsed as Chat SSE. That inference went silently wrong when #2152 moved deepseek-v4-flash to the Responses API: the proxy never saw data: [DONE], recorded every request as interrupted with zero usage and no reasoning tokens, and incompleteTerminalProviderRequest then threw the whole graded cell away as an infra failure — cells that had in fact passed with reward 1.0.

The Maka arm runs the Maka runtime, so resolveModelRuntime is the authority on which wire it dials; ask it instead of guessing. Competitors run their own CLI and keep the adapter-kind path, which is what describes them. The A/B manifest's per-arm transport comes from the same call, so a run now records the protocol it measured rather than the one it assumed — the frozen three-way composition test asserted openai-chat for the Maka arm, describing a run that never happened.

packages/runtime gains a ./model-runtime export so the harness can reach the resolver.

2. The infra-retry allowlist could not name a dotted task.executeHarnessArmCohort spelled candidate round ids as ab-${arm}-r0-${task.id} by hand, while buildAbRoundId — the helper that actually writes them — normalizes . to -. A retry naming install-windows-3.11 was therefore rejected as unknown, and because the check throws rather than skips, that one task took the whole batch down with it: a rerun of three infra-failed cells exited in twelve seconds having run none of them, twice. Building the candidates with the same helper removes the second spelling.

Refs #2245.

Verification

npm run -w @maka/headless test — 1424 pass / 0 fail. npm run -w @maka/runtime test — 3270 pass / 0 fail. npm run lint and npm run format clean.

The dotted-retry regression test was verified red before the fix (adjudicated infra retry names unknown round ab-maka-r0-install-windows-3-11) and green after.

Reproduced (1) against the live 5-task Terminal-Bench 2.1 pilot on deepseek-v4-flash (the same seed order as the 2026-08-04 four-arm run, so count-dataset-tokens is the same cell in all three columns):

2026-08-04 (Chat wire)before this fixafter
path / parsed protocol/chat/completions · chat-sse/responses · chat-sse/responses · responses-sse
requests24, completed 2420, interrupted 2024, completed 24
reasoning tokens3358null (0 measured)5154 (24/24 measured)
cell verdictpassinfra_failedcompleted, pass

A second pilot cell (regex-log) is likewise 15/15 completed with 17058 reasoning tokens measured.

The full 89-task two-arm run then completed on this branch: 3479 Maka requests over 89 cells with 9 non-completed, versus every request non-completed before the fix.

Root cause

Not #2241/#2249. fc581d855 (#2152, provider-native web search) added deepseek-v4-flash → openai-responses to openAiAdapterApiProtocol. The 2026-08-04 benchmark ran on 9dae4bee, which predates it and therefore dialled Chat Completions; the wire changed under a harness that kept guessing the old one.

Review round 2

Review found the fix's own API could re-admit the bug it closes. providerProxyUsageProtocol took the model id as an optional, prefix-sensitive argument, and the two call sites applied the discipline while the manifest call did not. resolveModelRuntime does not recognize the catalog spelling deepseek/deepseek-v4-flash, so a caller forwarding it raw silently got the Chat guess back — the exact wrong number this PR exists to remove — and a caller passing nothing at all got it too, with TypeScript unable to say so.

The function now normalizes the id itself and refuses to guess for the maka arm. That deletes the caller-side discipline at both runners rather than testing it: there is no longer a spelling a call site can get wrong. Two assertions pin it, including the assert.throws for a missing id.

Unknown retry round ids are also now reported together, so an operator recovering a sweep with a batch of ids fixes every typo in one pass.

The provider proxy inferred its usage protocol from the adapter kind, so
every `openai-compatible` provider was parsed as Chat SSE. That inference
silently went wrong when deepseek-v4-flash moved to the Responses API:
the proxy never saw `data: [DONE]`, recorded all 20 requests of a cell as
`interrupted` with zero usage and no reasoning tokens, and the runner's
terminal-request check then threw the whole graded cell away as an infra
failure -- a cell that had in fact passed with reward 1.0.
The Maka arm runs the Maka runtime, so `resolveModelRuntime` is the
authority on which wire it dials; ask it instead of guessing. Competitors
run their own CLI and keep the adapter-kind path, which is what describes
them. The A/B manifest's per-arm transport comes from the same call, so a
run now records the protocol it measured rather than the one it assumed.
Verified on the 5-task pilot the guess had been failing.
…s them
The retry allowlist spelled `ab-${arm}-r0-${task.id}` by hand while
`buildAbRoundId` normalizes `.` to `-`, so a retry naming a dotted task
never matched. Because the check throws instead of skipping, one such
task took the whole batch down: a rerun of three infra-failed cells
exited in twelve seconds without running any of them.
…er spells the model
Review found the fix's own API could re-admit the bug it closes: the model id
was prefix-sensitive and optional, so a caller passing the catalog spelling
`deepseek/deepseek-v4-flash` — which `resolveModelRuntime` does not recognize
— silently got the Chat guess back, and a caller passing nothing at all got
it too. Normalize inside the function and refuse to guess for the maka arm,
which deletes the caller-side discipline at both runners rather than testing
it.
Also report every unknown retry round id at once.
… will build
The proxy asked the runtime which wire it would dial, but asked about a
connection the runtime never builds: an advertised protocol only reaches a
model entry for GitHub Copilot and the Kimi Coding Plan, and `connectionFromEnv`
drops it everywhere else. With MAKA_MODEL_API_PROTOCOL=openai-chat set against
DeepSeek, the proxy resolved Chat SSE while the runtime dialled Responses —
the wrong-wire failure this path exists to remove, readmitted through its own
input.
The gating rule is now one exported function that both sides read, and the
test that pinned the divergent answer asserts the runtime's.
The proxy path parsed the three protocol names with its own copy of the check
`provider-env` already owned, next to the gating rule it just started sharing
with it. Two copies of a value whitelist is how one of them ends up accepting
a name the other rejects.
@Astro-Han
Astro-Han marked this pull request as ready for review August 6, 2026 06:07
@Astro-Han
Astro-Han merged commit 5b7dbb4 into mainAug 6, 2026
12 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix(headless): measure the wire the Maka runtime actually dials by Astro-Han · Pull Request #2286 · apache/maka · GitHub
Skip to content

fix(headless): measure the wire the Maka runtime actually dials - #2286

Merged
Astro-Han merged 5 commits into
mainfrom
fix/headless-proxy-usage-protocol-wire
Aug 6, 2026
Merged

fix(headless): measure the wire the Maka runtime actually dials#2286
Astro-Han merged 5 commits into
mainfrom
fix/headless-proxy-usage-protocol-wire

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Aug 5, 2026

Copy link
Copy Markdown
Contributor

Summary

Two independent measurement bugs in the harness, both surfaced by the #2245 two-arm rerun.

1. The proxy measured a wire the Maka arm was not dialing. The provider proxy inferred its usage protocol from runtimeAdapter.kind, so every openai-compatible provider was parsed as Chat SSE. That inference went silently wrong when #2152 moved deepseek-v4-flash to the Responses API: the proxy never saw data: [DONE], recorded every request as interrupted with zero usage and no reasoning tokens, and incompleteTerminalProviderRequest then threw the whole graded cell away as an infra failure — cells that had in fact passed with reward 1.0.

The Maka arm runs the Maka runtime, so resolveModelRuntime is the authority on which wire it dials; ask it instead of guessing. Competitors run their own CLI and keep the adapter-kind path, which is what describes them. The A/B manifest's per-arm transport comes from the same call, so a run now records the protocol it measured rather than the one it assumed — the frozen three-way composition test asserted openai-chat for the Maka arm, describing a run that never happened.

packages/runtime gains a ./model-runtime export so the harness can reach the resolver.

2. The infra-retry allowlist could not name a dotted task.executeHarnessArmCohort spelled candidate round ids as ab-${arm}-r0-${task.id} by hand, while buildAbRoundId — the helper that actually writes them — normalizes . to -. A retry naming install-windows-3.11 was therefore rejected as unknown, and because the check throws rather than skips, that one task took the whole batch down with it: a rerun of three infra-failed cells exited in twelve seconds having run none of them, twice. Building the candidates with the same helper removes the second spelling.

Refs #2245.

Verification

npm run -w @maka/headless test — 1424 pass / 0 fail. npm run -w @maka/runtime test — 3270 pass / 0 fail. npm run lint and npm run format clean.

The dotted-retry regression test was verified red before the fix (adjudicated infra retry names unknown round ab-maka-r0-install-windows-3-11) and green after.

Reproduced (1) against the live 5-task Terminal-Bench 2.1 pilot on deepseek-v4-flash (the same seed order as the 2026-08-04 four-arm run, so count-dataset-tokens is the same cell in all three columns):

2026-08-04 (Chat wire)before this fixafter
path / parsed protocol/chat/completions · chat-sse/responses · chat-sse/responses · responses-sse
requests24, completed 2420, interrupted 2024, completed 24
reasoning tokens3358null (0 measured)5154 (24/24 measured)
cell verdictpassinfra_failedcompleted, pass

A second pilot cell (regex-log) is likewise 15/15 completed with 17058 reasoning tokens measured.

The full 89-task two-arm run then completed on this branch: 3479 Maka requests over 89 cells with 9 non-completed, versus every request non-completed before the fix.

Root cause

Not #2241/#2249. fc581d855 (#2152, provider-native web search) added deepseek-v4-flash → openai-responses to openAiAdapterApiProtocol. The 2026-08-04 benchmark ran on 9dae4bee, which predates it and therefore dialled Chat Completions; the wire changed under a harness that kept guessing the old one.

Review round 2

Review found the fix's own API could re-admit the bug it closes. providerProxyUsageProtocol took the model id as an optional, prefix-sensitive argument, and the two call sites applied the discipline while the manifest call did not. resolveModelRuntime does not recognize the catalog spelling deepseek/deepseek-v4-flash, so a caller forwarding it raw silently got the Chat guess back — the exact wrong number this PR exists to remove — and a caller passing nothing at all got it too, with TypeScript unable to say so.

The function now normalizes the id itself and refuses to guess for the maka arm. That deletes the caller-side discipline at both runners rather than testing it: there is no longer a spelling a call site can get wrong. Two assertions pin it, including the assert.throws for a missing id.

Unknown retry round ids are also now reported together, so an operator recovering a sweep with a batch of ids fixes every typo in one pass.

The provider proxy inferred its usage protocol from the adapter kind, so
every `openai-compatible` provider was parsed as Chat SSE. That inference
silently went wrong when deepseek-v4-flash moved to the Responses API:
the proxy never saw `data: [DONE]`, recorded all 20 requests of a cell as
`interrupted` with zero usage and no reasoning tokens, and the runner's
terminal-request check then threw the whole graded cell away as an infra
failure -- a cell that had in fact passed with reward 1.0.
The Maka arm runs the Maka runtime, so `resolveModelRuntime` is the
authority on which wire it dials; ask it instead of guessing. Competitors
run their own CLI and keep the adapter-kind path, which is what describes
them. The A/B manifest's per-arm transport comes from the same call, so a
run now records the protocol it measured rather than the one it assumed.
Verified on the 5-task pilot the guess had been failing.
…s them
The retry allowlist spelled `ab-${arm}-r0-${task.id}` by hand while
`buildAbRoundId` normalizes `.` to `-`, so a retry naming a dotted task
never matched. Because the check throws instead of skipping, one such
task took the whole batch down: a rerun of three infra-failed cells
exited in twelve seconds without running any of them.
…er spells the model
Review found the fix's own API could re-admit the bug it closes: the model id
was prefix-sensitive and optional, so a caller passing the catalog spelling
`deepseek/deepseek-v4-flash` — which `resolveModelRuntime` does not recognize
— silently got the Chat guess back, and a caller passing nothing at all got
it too. Normalize inside the function and refuse to guess for the maka arm,
which deletes the caller-side discipline at both runners rather than testing
it.
Also report every unknown retry round id at once.
… will build
The proxy asked the runtime which wire it would dial, but asked about a
connection the runtime never builds: an advertised protocol only reaches a
model entry for GitHub Copilot and the Kimi Coding Plan, and `connectionFromEnv`
drops it everywhere else. With MAKA_MODEL_API_PROTOCOL=openai-chat set against
DeepSeek, the proxy resolved Chat SSE while the runtime dialled Responses —
the wrong-wire failure this path exists to remove, readmitted through its own
input.
The gating rule is now one exported function that both sides read, and the
test that pinned the divergent answer asserts the runtime's.
The proxy path parsed the three protocol names with its own copy of the check
`provider-env` already owned, next to the gating rule it just started sharing
with it. Two copies of a value whitelist is how one of them ends up accepting
a name the other rejects.
@Astro-Han
Astro-Han marked this pull request as ready for review August 6, 2026 06:07
@Astro-Han
Astro-Han merged commit 5b7dbb4 into mainAug 6, 2026
12 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); fix(headless): measure the wire the Maka runtime actually dials by Astro-Han · Pull Request #2286 · apache/maka · GitHub
Skip to content

fix(headless): measure the wire the Maka runtime actually dials - #2286

Merged
Astro-Han merged 5 commits into
mainfrom
fix/headless-proxy-usage-protocol-wire
Aug 6, 2026
Merged

fix(headless): measure the wire the Maka runtime actually dials#2286
Astro-Han merged 5 commits into
mainfrom
fix/headless-proxy-usage-protocol-wire

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Aug 5, 2026

Copy link
Copy Markdown
Contributor

Summary

Two independent measurement bugs in the harness, both surfaced by the #2245 two-arm rerun.

1. The proxy measured a wire the Maka arm was not dialing. The provider proxy inferred its usage protocol from runtimeAdapter.kind, so every openai-compatible provider was parsed as Chat SSE. That inference went silently wrong when #2152 moved deepseek-v4-flash to the Responses API: the proxy never saw data: [DONE], recorded every request as interrupted with zero usage and no reasoning tokens, and incompleteTerminalProviderRequest then threw the whole graded cell away as an infra failure — cells that had in fact passed with reward 1.0.

The Maka arm runs the Maka runtime, so resolveModelRuntime is the authority on which wire it dials; ask it instead of guessing. Competitors run their own CLI and keep the adapter-kind path, which is what describes them. The A/B manifest's per-arm transport comes from the same call, so a run now records the protocol it measured rather than the one it assumed — the frozen three-way composition test asserted openai-chat for the Maka arm, describing a run that never happened.

packages/runtime gains a ./model-runtime export so the harness can reach the resolver.

2. The infra-retry allowlist could not name a dotted task.executeHarnessArmCohort spelled candidate round ids as ab-${arm}-r0-${task.id} by hand, while buildAbRoundId — the helper that actually writes them — normalizes . to -. A retry naming install-windows-3.11 was therefore rejected as unknown, and because the check throws rather than skips, that one task took the whole batch down with it: a rerun of three infra-failed cells exited in twelve seconds having run none of them, twice. Building the candidates with the same helper removes the second spelling.

Refs #2245.

Verification

npm run -w @maka/headless test — 1424 pass / 0 fail. npm run -w @maka/runtime test — 3270 pass / 0 fail. npm run lint and npm run format clean.

The dotted-retry regression test was verified red before the fix (adjudicated infra retry names unknown round ab-maka-r0-install-windows-3-11) and green after.

Reproduced (1) against the live 5-task Terminal-Bench 2.1 pilot on deepseek-v4-flash (the same seed order as the 2026-08-04 four-arm run, so count-dataset-tokens is the same cell in all three columns):

2026-08-04 (Chat wire)before this fixafter
path / parsed protocol/chat/completions · chat-sse/responses · chat-sse/responses · responses-sse
requests24, completed 2420, interrupted 2024, completed 24
reasoning tokens3358null (0 measured)5154 (24/24 measured)
cell verdictpassinfra_failedcompleted, pass

A second pilot cell (regex-log) is likewise 15/15 completed with 17058 reasoning tokens measured.

The full 89-task two-arm run then completed on this branch: 3479 Maka requests over 89 cells with 9 non-completed, versus every request non-completed before the fix.

Root cause

Not #2241/#2249. fc581d855 (#2152, provider-native web search) added deepseek-v4-flash → openai-responses to openAiAdapterApiProtocol. The 2026-08-04 benchmark ran on 9dae4bee, which predates it and therefore dialled Chat Completions; the wire changed under a harness that kept guessing the old one.

Review round 2

Review found the fix's own API could re-admit the bug it closes. providerProxyUsageProtocol took the model id as an optional, prefix-sensitive argument, and the two call sites applied the discipline while the manifest call did not. resolveModelRuntime does not recognize the catalog spelling deepseek/deepseek-v4-flash, so a caller forwarding it raw silently got the Chat guess back — the exact wrong number this PR exists to remove — and a caller passing nothing at all got it too, with TypeScript unable to say so.

The function now normalizes the id itself and refuses to guess for the maka arm. That deletes the caller-side discipline at both runners rather than testing it: there is no longer a spelling a call site can get wrong. Two assertions pin it, including the assert.throws for a missing id.

Unknown retry round ids are also now reported together, so an operator recovering a sweep with a batch of ids fixes every typo in one pass.

The provider proxy inferred its usage protocol from the adapter kind, so
every `openai-compatible` provider was parsed as Chat SSE. That inference
silently went wrong when deepseek-v4-flash moved to the Responses API:
the proxy never saw `data: [DONE]`, recorded all 20 requests of a cell as
`interrupted` with zero usage and no reasoning tokens, and the runner's
terminal-request check then threw the whole graded cell away as an infra
failure -- a cell that had in fact passed with reward 1.0.
The Maka arm runs the Maka runtime, so `resolveModelRuntime` is the
authority on which wire it dials; ask it instead of guessing. Competitors
run their own CLI and keep the adapter-kind path, which is what describes
them. The A/B manifest's per-arm transport comes from the same call, so a
run now records the protocol it measured rather than the one it assumed.
Verified on the 5-task pilot the guess had been failing.
…s them
The retry allowlist spelled `ab-${arm}-r0-${task.id}` by hand while
`buildAbRoundId` normalizes `.` to `-`, so a retry naming a dotted task
never matched. Because the check throws instead of skipping, one such
task took the whole batch down: a rerun of three infra-failed cells
exited in twelve seconds without running any of them.
…er spells the model
Review found the fix's own API could re-admit the bug it closes: the model id
was prefix-sensitive and optional, so a caller passing the catalog spelling
`deepseek/deepseek-v4-flash` — which `resolveModelRuntime` does not recognize
— silently got the Chat guess back, and a caller passing nothing at all got
it too. Normalize inside the function and refuse to guess for the maka arm,
which deletes the caller-side discipline at both runners rather than testing
it.
Also report every unknown retry round id at once.
… will build
The proxy asked the runtime which wire it would dial, but asked about a
connection the runtime never builds: an advertised protocol only reaches a
model entry for GitHub Copilot and the Kimi Coding Plan, and `connectionFromEnv`
drops it everywhere else. With MAKA_MODEL_API_PROTOCOL=openai-chat set against
DeepSeek, the proxy resolved Chat SSE while the runtime dialled Responses —
the wrong-wire failure this path exists to remove, readmitted through its own
input.
The gating rule is now one exported function that both sides read, and the
test that pinned the divergent answer asserts the runtime's.
The proxy path parsed the three protocol names with its own copy of the check
`provider-env` already owned, next to the gating rule it just started sharing
with it. Two copies of a value whitelist is how one of them ends up accepting
a name the other rejects.
@Astro-Han
Astro-Han marked this pull request as ready for review August 6, 2026 06:07
@Astro-Han
Astro-Han merged commit 5b7dbb4 into mainAug 6, 2026
12 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han