Fix the Gemini tool-replay 400, and stop it from costing 12 minutes of silence - #8

Merged
timothyerwin merged 6 commits into
mainfrom
fix-gemini-tool-replay
Sep 3, 2026
Merged

Fix the Gemini tool-replay 400, and stop it from costing 12 minutes of silence#8
timothyerwin merged 6 commits into
mainfrom
fix-gemini-tool-replay

Conversation

@timothyerwin

@timothyerwintimothyerwin commented Sep 3, 2026

Copy link
Copy Markdown
Contributor

Code review was stuck on "running" with nothing posted. Six bugs compounded. Four are fixed here; two were gateway-side and are already deployed (memcode-api-00048-sph).

The actual root cause chain

A Gemini tool call carries an opaque thoughtSignature that must be echoed back when the call is replayed. It was being dropped in two independent places, and three more bugs turned that one 400 into twelve minutes of silence.

1. The Gemini adapter dropped the signature

The adapter decoded functionCall into a tool_use block but discarded the signature, so any multi-turn tool conversation 400'd on its second turn. The reviewer is a tool-using agent, so it hit this on its first replay, every time.

Verified against live Vertex: 3.6 fails identically. This is a long-standing adapter bug, not a Gemini 3.8 regression. My earlier 3.8 checks were text, vision, PDF and thinking — all single-turn, which is exactly the path that works.

2. The hosted wire had nowhere to put it

Fixing the adapter was necessary but not sufficient, and this is why the first gateway deploy did not fix it. On the hosted path the CLI never touches that adapter: the call crosses the wire in OpenAI-compat shape, which has no field for opaque per-call state. The signature was dropped in translation regardless.

It rides a namespaced memcode_signature extension now — the same channel pattern as memcode_opaque for reasoning blocks — through all four hops: streaming delta, accumulator, block assembly, and encode-back on replay. omitempty keeps the standard shape standard for providers that issue no signature.

3. A 400 walked the fallback chain

A malformed request, bad credentials, or an unknown model does not become valid because a different model receives it. Walking the chain on those turned one bug into N identical failures — 298 failed calls in the broken run. Now classified by category: request-shape 4xx is terminal; 408/429 and 5xx still walk.

4. Fallbacks pointed at the same vendor

gemini-flash fell back to gemini-pro — same adapter, same API, same bug. The chain existed only to fail twice as expensively. All 16 chains repointed so the first hop always changes vendor, with a catalog guard test against reintroduction.

Fallbacks are infrastructure resilience only. A capability gap still refuses rather than substituting, so the pin stays authoritative.

5. Nothing was fatal, so it ran for 12 minutes

A worker whose model call died returned errResult("agent failed: …") — an ordinary retryable tool error. For a terminal cause that is a lie, and the model believed it: ~150 identical failed calls, one every two seconds, ending only on the pipeline's own timeout. The cause never surfaced anywhere a user could see it.

Terminal worker errors now arm a turn-fatal error. The loop ends the turn and the existing path carries the cause the rest of the way: printed to the session, stored as LastError, returned as a non-zero exit and a failed gateway job. Retryable causes are untouched.

Fixed gateway-side (already deployed, not in this PR)

  • The gateway's Gemini adapter registered no error extractor, so every Gemini error there was opaque to status classification — no retry policy, no terminal detection. All of them fell through to a generic 502.
  • compatTurnError mapped everything unrecognized to 502, documented as "fallback-chain territory for the CLI". For a provider 4xx we caused that is wrong. Request-shape 4xx now returns a terminal 400 with code upstream_request_invalid — deliberately 400 and not the upstream status, because an upstream 401 is our credential problem and the CLI reads a 401 as the user being signed out.

Verification

gofmt, go vet, go test -race ./..., staticcheck@2026.1 — all clean.

Live end-to-end against the deployed gateway: three sequential tool rounds on gemini-flash, all answers correct against ground truth, no retries and no fallback. The same prompt before these changes failed three times and fell over to haiku.

New regression tests: the hosted signature round-trip (and the no-signature case, so the standard shape stays clean); multi-turn adapter replay; the terminal/retryable status table; the same-vendor fallback guard; terminal-vs-transient worker failure; the gateway's upstream-4xx mapping including the 401 sign-out trap.

What this does not fix yet

The reviewer's CLI is built from a pinned tag (repoagent/Dockerfile, CLI_REF=v0.28.0), so the memcode_signature round-trip reaches production only after this merges, a release is tagged, and the repoagent image is rebuilt and redeployed at that tag.

Until then review still hits the Gemini 400 — but thanks to the deployed gateway half it now fails in seconds with the real cause, instead of hanging for twenty minutes.

Notes

Conflicts with #7 (retire-kimi-k2); both edit catalog/models.json fallbacks. Merge order matters.

Gemini issues an opaque thoughtSignature with every functionCall and REQUIRES
it echoed back when that call is replayed in a later turn. The adapter decoded
functionCall into a tool_use block but dropped the signature, so any multi-turn
tool conversation failed with a 400 on the second turn.
This is why code review broke: the reviewer is a tool-using agent, so it hit the
400 on its first replay, every time. It reproduces identically on Gemini 3.6, so
it is an adapter bug, not a 3.8 regression.
Adds toolUseFromPart as a testable decode seam, echoes the signature back in
blockToPart, and pins both directions with a multi-turn replay test plus a
no-signature case (an older transcript must still send none, not an empty one).
A 400 means WE malformed the request; a 401/403 means credentials; a 404 means
the model does not exist. None of those change because a different model
receives them, so walking the fallback chain just multiplies one bug into N
identical failures — 298 failed calls in the broken review run came from exactly
this. 408/429 and 5xx are timing, and still walk.
Classifies by category rather than by individual status, with a table test
pinning both halves.
A same-vendor fallback is not a fallback. gemini-flash fell back to gemini-pro,
which is the same adapter over the same API — so the thought-signature 400 hit
the backup identically, and the chain existed only to fail twice as expensively.
Repoints all 16 chains so the first hop always changes vendor (gemini-flash →
haiku, luna; gemini-pro → sonnet, terra; glm-5p2 → terra, luna), and adds a
catalog guard test so a same-vendor first hop cannot be reintroduced.
Fallbacks exist to survive infrastructure failure, never to reinterpret the
user's pin: a capability gap still refuses rather than substituting.
A delegated worker whose model call failed came back to the model as
errResult("agent failed: …"), an ordinary retryable tool error. For a terminal
cause that is a lie: it will fail identically next time. The model read it as
"try again" and did — the broken review run made ~150 identical failed calls,
one every two seconds for 12 minutes, and only ended on the pipeline's own
timeout. The real cause never appeared anywhere the user could see it.
Terminal worker errors now arm turn.fatalErr, the loop ends the turn on it, and
the existing path carries it the rest of the way: printed to the session, stored
as LastError, returned as a non-zero exit and a failed gateway job. Retryable
causes (429, 5xx, transport) are untouched and still just a tool error.
Exports llm.IsTerminal for the classification, and pins both sides.
Fixing the Gemini adapter was necessary but not sufficient. On the hosted path
the CLI never touches that adapter: the tool call crosses the wire in
OpenAI-compat shape, which has nowhere to put opaque per-call provider state, so
the signature was dropped in translation and every replay 400'd regardless.
It rides a namespaced memcode_signature extension now — the same channel pattern
as memcode_opaque for reasoning blocks — through all four hops: streaming delta,
accumulator, block assembly, and the encode back on replay. omitempty keeps the
standard shape standard for providers that issue no signature.
Verified live end to end against the deployed gateway: three sequential tool
rounds on gemini-flash, correct answers, no retries and no fallback. The same
prompt before this change failed three times and fell over to haiku.
catalog/models.json and models.json: resolved onto main's post-#7 catalog as the
base, so both retired Kimi rows stay gone, then re-applied the 14 cross-vendor
chains on top. Resolved structurally against the parsed catalog rather than by
editing conflict text — the file round-trips byte-identically through json, so
the rewrite touches only the fallback arrays.
policy_test.go: a comment-wording conflict only; took main's phrasing.
No chain references a retired model, and TestFallbackFirstHopLeavesTheVendor
still passes: every first hop changes vendor.
@timothyerwin
timothyerwin merged commit 29f67f9 into mainSep 3, 2026
2 checks passed
@timothyerwin
timothyerwin deleted the fix-gemini-tool-replay branch September 3, 2026 08:13
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@timothyerwin
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Fix the Gemini tool-replay 400, and stop it from costing 12 minutes of silence - #8

Merged
timothyerwin merged 6 commits into
mainfrom
fix-gemini-tool-replay
Sep 3, 2026
Merged

Fix the Gemini tool-replay 400, and stop it from costing 12 minutes of silence#8
timothyerwin merged 6 commits into
mainfrom
fix-gemini-tool-replay

Conversation

@timothyerwin

@timothyerwintimothyerwin commented Sep 3, 2026

Copy link
Copy Markdown
Contributor

Code review was stuck on "running" with nothing posted. Six bugs compounded. Four are fixed here; two were gateway-side and are already deployed (memcode-api-00048-sph).

The actual root cause chain

A Gemini tool call carries an opaque thoughtSignature that must be echoed back when the call is replayed. It was being dropped in two independent places, and three more bugs turned that one 400 into twelve minutes of silence.

1. The Gemini adapter dropped the signature

The adapter decoded functionCall into a tool_use block but discarded the signature, so any multi-turn tool conversation 400'd on its second turn. The reviewer is a tool-using agent, so it hit this on its first replay, every time.

Verified against live Vertex: 3.6 fails identically. This is a long-standing adapter bug, not a Gemini 3.8 regression. My earlier 3.8 checks were text, vision, PDF and thinking — all single-turn, which is exactly the path that works.

2. The hosted wire had nowhere to put it

Fixing the adapter was necessary but not sufficient, and this is why the first gateway deploy did not fix it. On the hosted path the CLI never touches that adapter: the call crosses the wire in OpenAI-compat shape, which has no field for opaque per-call state. The signature was dropped in translation regardless.

It rides a namespaced memcode_signature extension now — the same channel pattern as memcode_opaque for reasoning blocks — through all four hops: streaming delta, accumulator, block assembly, and encode-back on replay. omitempty keeps the standard shape standard for providers that issue no signature.

3. A 400 walked the fallback chain

A malformed request, bad credentials, or an unknown model does not become valid because a different model receives it. Walking the chain on those turned one bug into N identical failures — 298 failed calls in the broken run. Now classified by category: request-shape 4xx is terminal; 408/429 and 5xx still walk.

4. Fallbacks pointed at the same vendor

gemini-flash fell back to gemini-pro — same adapter, same API, same bug. The chain existed only to fail twice as expensively. All 16 chains repointed so the first hop always changes vendor, with a catalog guard test against reintroduction.

Fallbacks are infrastructure resilience only. A capability gap still refuses rather than substituting, so the pin stays authoritative.

5. Nothing was fatal, so it ran for 12 minutes

A worker whose model call died returned errResult("agent failed: …") — an ordinary retryable tool error. For a terminal cause that is a lie, and the model believed it: ~150 identical failed calls, one every two seconds, ending only on the pipeline's own timeout. The cause never surfaced anywhere a user could see it.

Terminal worker errors now arm a turn-fatal error. The loop ends the turn and the existing path carries the cause the rest of the way: printed to the session, stored as LastError, returned as a non-zero exit and a failed gateway job. Retryable causes are untouched.

Fixed gateway-side (already deployed, not in this PR)

  • The gateway's Gemini adapter registered no error extractor, so every Gemini error there was opaque to status classification — no retry policy, no terminal detection. All of them fell through to a generic 502.
  • compatTurnError mapped everything unrecognized to 502, documented as "fallback-chain territory for the CLI". For a provider 4xx we caused that is wrong. Request-shape 4xx now returns a terminal 400 with code upstream_request_invalid — deliberately 400 and not the upstream status, because an upstream 401 is our credential problem and the CLI reads a 401 as the user being signed out.

Verification

gofmt, go vet, go test -race ./..., staticcheck@2026.1 — all clean.

Live end-to-end against the deployed gateway: three sequential tool rounds on gemini-flash, all answers correct against ground truth, no retries and no fallback. The same prompt before these changes failed three times and fell over to haiku.

New regression tests: the hosted signature round-trip (and the no-signature case, so the standard shape stays clean); multi-turn adapter replay; the terminal/retryable status table; the same-vendor fallback guard; terminal-vs-transient worker failure; the gateway's upstream-4xx mapping including the 401 sign-out trap.

What this does not fix yet

The reviewer's CLI is built from a pinned tag (repoagent/Dockerfile, CLI_REF=v0.28.0), so the memcode_signature round-trip reaches production only after this merges, a release is tagged, and the repoagent image is rebuilt and redeployed at that tag.

Until then review still hits the Gemini 400 — but thanks to the deployed gateway half it now fails in seconds with the real cause, instead of hanging for twenty minutes.

Notes

Conflicts with #7 (retire-kimi-k2); both edit catalog/models.json fallbacks. Merge order matters.

Gemini issues an opaque thoughtSignature with every functionCall and REQUIRES
it echoed back when that call is replayed in a later turn. The adapter decoded
functionCall into a tool_use block but dropped the signature, so any multi-turn
tool conversation failed with a 400 on the second turn.
This is why code review broke: the reviewer is a tool-using agent, so it hit the
400 on its first replay, every time. It reproduces identically on Gemini 3.6, so
it is an adapter bug, not a 3.8 regression.
Adds toolUseFromPart as a testable decode seam, echoes the signature back in
blockToPart, and pins both directions with a multi-turn replay test plus a
no-signature case (an older transcript must still send none, not an empty one).
A 400 means WE malformed the request; a 401/403 means credentials; a 404 means
the model does not exist. None of those change because a different model
receives them, so walking the fallback chain just multiplies one bug into N
identical failures — 298 failed calls in the broken review run came from exactly
this. 408/429 and 5xx are timing, and still walk.
Classifies by category rather than by individual status, with a table test
pinning both halves.
A same-vendor fallback is not a fallback. gemini-flash fell back to gemini-pro,
which is the same adapter over the same API — so the thought-signature 400 hit
the backup identically, and the chain existed only to fail twice as expensively.
Repoints all 16 chains so the first hop always changes vendor (gemini-flash →
haiku, luna; gemini-pro → sonnet, terra; glm-5p2 → terra, luna), and adds a
catalog guard test so a same-vendor first hop cannot be reintroduced.
Fallbacks exist to survive infrastructure failure, never to reinterpret the
user's pin: a capability gap still refuses rather than substituting.
A delegated worker whose model call failed came back to the model as
errResult("agent failed: …"), an ordinary retryable tool error. For a terminal
cause that is a lie: it will fail identically next time. The model read it as
"try again" and did — the broken review run made ~150 identical failed calls,
one every two seconds for 12 minutes, and only ended on the pipeline's own
timeout. The real cause never appeared anywhere the user could see it.
Terminal worker errors now arm turn.fatalErr, the loop ends the turn on it, and
the existing path carries it the rest of the way: printed to the session, stored
as LastError, returned as a non-zero exit and a failed gateway job. Retryable
causes (429, 5xx, transport) are untouched and still just a tool error.
Exports llm.IsTerminal for the classification, and pins both sides.
Fixing the Gemini adapter was necessary but not sufficient. On the hosted path
the CLI never touches that adapter: the tool call crosses the wire in
OpenAI-compat shape, which has nowhere to put opaque per-call provider state, so
the signature was dropped in translation and every replay 400'd regardless.
It rides a namespaced memcode_signature extension now — the same channel pattern
as memcode_opaque for reasoning blocks — through all four hops: streaming delta,
accumulator, block assembly, and the encode back on replay. omitempty keeps the
standard shape standard for providers that issue no signature.
Verified live end to end against the deployed gateway: three sequential tool
rounds on gemini-flash, correct answers, no retries and no fallback. The same
prompt before this change failed three times and fell over to haiku.
catalog/models.json and models.json: resolved onto main's post-#7 catalog as the
base, so both retired Kimi rows stay gone, then re-applied the 14 cross-vendor
chains on top. Resolved structurally against the parsed catalog rather than by
editing conflict text — the file round-trips byte-identically through json, so
the rewrite touches only the fallback arrays.
policy_test.go: a comment-wording conflict only; took main's phrasing.
No chain references a retired model, and TestFallbackFirstHopLeavesTheVendor
still passes: every first hop changes vendor.
@timothyerwin
timothyerwin merged commit 29f67f9 into mainSep 3, 2026
2 checks passed
@timothyerwin
timothyerwin deleted the fix-gemini-tool-replay branch September 3, 2026 08:13
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@timothyerwin
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Fix the Gemini tool-replay 400, and stop it from costing 12 minutes of silence - #8

Merged
timothyerwin merged 6 commits into
mainfrom
fix-gemini-tool-replay
Sep 3, 2026
Merged

Fix the Gemini tool-replay 400, and stop it from costing 12 minutes of silence#8
timothyerwin merged 6 commits into
mainfrom
fix-gemini-tool-replay

Conversation

@timothyerwin

@timothyerwintimothyerwin commented Sep 3, 2026

Copy link
Copy Markdown
Contributor

Code review was stuck on "running" with nothing posted. Six bugs compounded. Four are fixed here; two were gateway-side and are already deployed (memcode-api-00048-sph).

The actual root cause chain

A Gemini tool call carries an opaque thoughtSignature that must be echoed back when the call is replayed. It was being dropped in two independent places, and three more bugs turned that one 400 into twelve minutes of silence.

1. The Gemini adapter dropped the signature

The adapter decoded functionCall into a tool_use block but discarded the signature, so any multi-turn tool conversation 400'd on its second turn. The reviewer is a tool-using agent, so it hit this on its first replay, every time.

Verified against live Vertex: 3.6 fails identically. This is a long-standing adapter bug, not a Gemini 3.8 regression. My earlier 3.8 checks were text, vision, PDF and thinking — all single-turn, which is exactly the path that works.

2. The hosted wire had nowhere to put it

Fixing the adapter was necessary but not sufficient, and this is why the first gateway deploy did not fix it. On the hosted path the CLI never touches that adapter: the call crosses the wire in OpenAI-compat shape, which has no field for opaque per-call state. The signature was dropped in translation regardless.

It rides a namespaced memcode_signature extension now — the same channel pattern as memcode_opaque for reasoning blocks — through all four hops: streaming delta, accumulator, block assembly, and encode-back on replay. omitempty keeps the standard shape standard for providers that issue no signature.

3. A 400 walked the fallback chain

A malformed request, bad credentials, or an unknown model does not become valid because a different model receives it. Walking the chain on those turned one bug into N identical failures — 298 failed calls in the broken run. Now classified by category: request-shape 4xx is terminal; 408/429 and 5xx still walk.

4. Fallbacks pointed at the same vendor

gemini-flash fell back to gemini-pro — same adapter, same API, same bug. The chain existed only to fail twice as expensively. All 16 chains repointed so the first hop always changes vendor, with a catalog guard test against reintroduction.

Fallbacks are infrastructure resilience only. A capability gap still refuses rather than substituting, so the pin stays authoritative.

5. Nothing was fatal, so it ran for 12 minutes

A worker whose model call died returned errResult("agent failed: …") — an ordinary retryable tool error. For a terminal cause that is a lie, and the model believed it: ~150 identical failed calls, one every two seconds, ending only on the pipeline's own timeout. The cause never surfaced anywhere a user could see it.

Terminal worker errors now arm a turn-fatal error. The loop ends the turn and the existing path carries the cause the rest of the way: printed to the session, stored as LastError, returned as a non-zero exit and a failed gateway job. Retryable causes are untouched.

Fixed gateway-side (already deployed, not in this PR)

  • The gateway's Gemini adapter registered no error extractor, so every Gemini error there was opaque to status classification — no retry policy, no terminal detection. All of them fell through to a generic 502.
  • compatTurnError mapped everything unrecognized to 502, documented as "fallback-chain territory for the CLI". For a provider 4xx we caused that is wrong. Request-shape 4xx now returns a terminal 400 with code upstream_request_invalid — deliberately 400 and not the upstream status, because an upstream 401 is our credential problem and the CLI reads a 401 as the user being signed out.

Verification

gofmt, go vet, go test -race ./..., staticcheck@2026.1 — all clean.

Live end-to-end against the deployed gateway: three sequential tool rounds on gemini-flash, all answers correct against ground truth, no retries and no fallback. The same prompt before these changes failed three times and fell over to haiku.

New regression tests: the hosted signature round-trip (and the no-signature case, so the standard shape stays clean); multi-turn adapter replay; the terminal/retryable status table; the same-vendor fallback guard; terminal-vs-transient worker failure; the gateway's upstream-4xx mapping including the 401 sign-out trap.

What this does not fix yet

The reviewer's CLI is built from a pinned tag (repoagent/Dockerfile, CLI_REF=v0.28.0), so the memcode_signature round-trip reaches production only after this merges, a release is tagged, and the repoagent image is rebuilt and redeployed at that tag.

Until then review still hits the Gemini 400 — but thanks to the deployed gateway half it now fails in seconds with the real cause, instead of hanging for twenty minutes.

Notes

Conflicts with #7 (retire-kimi-k2); both edit catalog/models.json fallbacks. Merge order matters.

Gemini issues an opaque thoughtSignature with every functionCall and REQUIRES
it echoed back when that call is replayed in a later turn. The adapter decoded
functionCall into a tool_use block but dropped the signature, so any multi-turn
tool conversation failed with a 400 on the second turn.
This is why code review broke: the reviewer is a tool-using agent, so it hit the
400 on its first replay, every time. It reproduces identically on Gemini 3.6, so
it is an adapter bug, not a 3.8 regression.
Adds toolUseFromPart as a testable decode seam, echoes the signature back in
blockToPart, and pins both directions with a multi-turn replay test plus a
no-signature case (an older transcript must still send none, not an empty one).
A 400 means WE malformed the request; a 401/403 means credentials; a 404 means
the model does not exist. None of those change because a different model
receives them, so walking the fallback chain just multiplies one bug into N
identical failures — 298 failed calls in the broken review run came from exactly
this. 408/429 and 5xx are timing, and still walk.
Classifies by category rather than by individual status, with a table test
pinning both halves.
A same-vendor fallback is not a fallback. gemini-flash fell back to gemini-pro,
which is the same adapter over the same API — so the thought-signature 400 hit
the backup identically, and the chain existed only to fail twice as expensively.
Repoints all 16 chains so the first hop always changes vendor (gemini-flash →
haiku, luna; gemini-pro → sonnet, terra; glm-5p2 → terra, luna), and adds a
catalog guard test so a same-vendor first hop cannot be reintroduced.
Fallbacks exist to survive infrastructure failure, never to reinterpret the
user's pin: a capability gap still refuses rather than substituting.
A delegated worker whose model call failed came back to the model as
errResult("agent failed: …"), an ordinary retryable tool error. For a terminal
cause that is a lie: it will fail identically next time. The model read it as
"try again" and did — the broken review run made ~150 identical failed calls,
one every two seconds for 12 minutes, and only ended on the pipeline's own
timeout. The real cause never appeared anywhere the user could see it.
Terminal worker errors now arm turn.fatalErr, the loop ends the turn on it, and
the existing path carries it the rest of the way: printed to the session, stored
as LastError, returned as a non-zero exit and a failed gateway job. Retryable
causes (429, 5xx, transport) are untouched and still just a tool error.
Exports llm.IsTerminal for the classification, and pins both sides.
Fixing the Gemini adapter was necessary but not sufficient. On the hosted path
the CLI never touches that adapter: the tool call crosses the wire in
OpenAI-compat shape, which has nowhere to put opaque per-call provider state, so
the signature was dropped in translation and every replay 400'd regardless.
It rides a namespaced memcode_signature extension now — the same channel pattern
as memcode_opaque for reasoning blocks — through all four hops: streaming delta,
accumulator, block assembly, and the encode back on replay. omitempty keeps the
standard shape standard for providers that issue no signature.
Verified live end to end against the deployed gateway: three sequential tool
rounds on gemini-flash, correct answers, no retries and no fallback. The same
prompt before this change failed three times and fell over to haiku.
catalog/models.json and models.json: resolved onto main's post-#7 catalog as the
base, so both retired Kimi rows stay gone, then re-applied the 14 cross-vendor
chains on top. Resolved structurally against the parsed catalog rather than by
editing conflict text — the file round-trips byte-identically through json, so
the rewrite touches only the fallback arrays.
policy_test.go: a comment-wording conflict only; took main's phrasing.
No chain references a retired model, and TestFallbackFirstHopLeavesTheVendor
still passes: every first hop changes vendor.
@timothyerwin
timothyerwin merged commit 29f67f9 into mainSep 3, 2026
2 checks passed
@timothyerwin
timothyerwin deleted the fix-gemini-tool-replay branch September 3, 2026 08:13
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@timothyerwin
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Fix the Gemini tool-replay 400, and stop it from costing 12 minutes of silence - #8

Merged
timothyerwin merged 6 commits into
mainfrom
fix-gemini-tool-replay
Sep 3, 2026
Merged

Fix the Gemini tool-replay 400, and stop it from costing 12 minutes of silence#8
timothyerwin merged 6 commits into
mainfrom
fix-gemini-tool-replay

Conversation

@timothyerwin

@timothyerwintimothyerwin commented Sep 3, 2026

Copy link
Copy Markdown
Contributor

Code review was stuck on "running" with nothing posted. Six bugs compounded. Four are fixed here; two were gateway-side and are already deployed (memcode-api-00048-sph).

The actual root cause chain

A Gemini tool call carries an opaque thoughtSignature that must be echoed back when the call is replayed. It was being dropped in two independent places, and three more bugs turned that one 400 into twelve minutes of silence.

1. The Gemini adapter dropped the signature

The adapter decoded functionCall into a tool_use block but discarded the signature, so any multi-turn tool conversation 400'd on its second turn. The reviewer is a tool-using agent, so it hit this on its first replay, every time.

Verified against live Vertex: 3.6 fails identically. This is a long-standing adapter bug, not a Gemini 3.8 regression. My earlier 3.8 checks were text, vision, PDF and thinking — all single-turn, which is exactly the path that works.

2. The hosted wire had nowhere to put it

Fixing the adapter was necessary but not sufficient, and this is why the first gateway deploy did not fix it. On the hosted path the CLI never touches that adapter: the call crosses the wire in OpenAI-compat shape, which has no field for opaque per-call state. The signature was dropped in translation regardless.

It rides a namespaced memcode_signature extension now — the same channel pattern as memcode_opaque for reasoning blocks — through all four hops: streaming delta, accumulator, block assembly, and encode-back on replay. omitempty keeps the standard shape standard for providers that issue no signature.

3. A 400 walked the fallback chain

A malformed request, bad credentials, or an unknown model does not become valid because a different model receives it. Walking the chain on those turned one bug into N identical failures — 298 failed calls in the broken run. Now classified by category: request-shape 4xx is terminal; 408/429 and 5xx still walk.

4. Fallbacks pointed at the same vendor

gemini-flash fell back to gemini-pro — same adapter, same API, same bug. The chain existed only to fail twice as expensively. All 16 chains repointed so the first hop always changes vendor, with a catalog guard test against reintroduction.

Fallbacks are infrastructure resilience only. A capability gap still refuses rather than substituting, so the pin stays authoritative.

5. Nothing was fatal, so it ran for 12 minutes

A worker whose model call died returned errResult("agent failed: …") — an ordinary retryable tool error. For a terminal cause that is a lie, and the model believed it: ~150 identical failed calls, one every two seconds, ending only on the pipeline's own timeout. The cause never surfaced anywhere a user could see it.

Terminal worker errors now arm a turn-fatal error. The loop ends the turn and the existing path carries the cause the rest of the way: printed to the session, stored as LastError, returned as a non-zero exit and a failed gateway job. Retryable causes are untouched.

Fixed gateway-side (already deployed, not in this PR)

  • The gateway's Gemini adapter registered no error extractor, so every Gemini error there was opaque to status classification — no retry policy, no terminal detection. All of them fell through to a generic 502.
  • compatTurnError mapped everything unrecognized to 502, documented as "fallback-chain territory for the CLI". For a provider 4xx we caused that is wrong. Request-shape 4xx now returns a terminal 400 with code upstream_request_invalid — deliberately 400 and not the upstream status, because an upstream 401 is our credential problem and the CLI reads a 401 as the user being signed out.

Verification

gofmt, go vet, go test -race ./..., staticcheck@2026.1 — all clean.

Live end-to-end against the deployed gateway: three sequential tool rounds on gemini-flash, all answers correct against ground truth, no retries and no fallback. The same prompt before these changes failed three times and fell over to haiku.

New regression tests: the hosted signature round-trip (and the no-signature case, so the standard shape stays clean); multi-turn adapter replay; the terminal/retryable status table; the same-vendor fallback guard; terminal-vs-transient worker failure; the gateway's upstream-4xx mapping including the 401 sign-out trap.

What this does not fix yet

The reviewer's CLI is built from a pinned tag (repoagent/Dockerfile, CLI_REF=v0.28.0), so the memcode_signature round-trip reaches production only after this merges, a release is tagged, and the repoagent image is rebuilt and redeployed at that tag.

Until then review still hits the Gemini 400 — but thanks to the deployed gateway half it now fails in seconds with the real cause, instead of hanging for twenty minutes.

Notes

Conflicts with #7 (retire-kimi-k2); both edit catalog/models.json fallbacks. Merge order matters.

Gemini issues an opaque thoughtSignature with every functionCall and REQUIRES
it echoed back when that call is replayed in a later turn. The adapter decoded
functionCall into a tool_use block but dropped the signature, so any multi-turn
tool conversation failed with a 400 on the second turn.
This is why code review broke: the reviewer is a tool-using agent, so it hit the
400 on its first replay, every time. It reproduces identically on Gemini 3.6, so
it is an adapter bug, not a 3.8 regression.
Adds toolUseFromPart as a testable decode seam, echoes the signature back in
blockToPart, and pins both directions with a multi-turn replay test plus a
no-signature case (an older transcript must still send none, not an empty one).
A 400 means WE malformed the request; a 401/403 means credentials; a 404 means
the model does not exist. None of those change because a different model
receives them, so walking the fallback chain just multiplies one bug into N
identical failures — 298 failed calls in the broken review run came from exactly
this. 408/429 and 5xx are timing, and still walk.
Classifies by category rather than by individual status, with a table test
pinning both halves.
A same-vendor fallback is not a fallback. gemini-flash fell back to gemini-pro,
which is the same adapter over the same API — so the thought-signature 400 hit
the backup identically, and the chain existed only to fail twice as expensively.
Repoints all 16 chains so the first hop always changes vendor (gemini-flash →
haiku, luna; gemini-pro → sonnet, terra; glm-5p2 → terra, luna), and adds a
catalog guard test so a same-vendor first hop cannot be reintroduced.
Fallbacks exist to survive infrastructure failure, never to reinterpret the
user's pin: a capability gap still refuses rather than substituting.
A delegated worker whose model call failed came back to the model as
errResult("agent failed: …"), an ordinary retryable tool error. For a terminal
cause that is a lie: it will fail identically next time. The model read it as
"try again" and did — the broken review run made ~150 identical failed calls,
one every two seconds for 12 minutes, and only ended on the pipeline's own
timeout. The real cause never appeared anywhere the user could see it.
Terminal worker errors now arm turn.fatalErr, the loop ends the turn on it, and
the existing path carries it the rest of the way: printed to the session, stored
as LastError, returned as a non-zero exit and a failed gateway job. Retryable
causes (429, 5xx, transport) are untouched and still just a tool error.
Exports llm.IsTerminal for the classification, and pins both sides.
Fixing the Gemini adapter was necessary but not sufficient. On the hosted path
the CLI never touches that adapter: the tool call crosses the wire in
OpenAI-compat shape, which has nowhere to put opaque per-call provider state, so
the signature was dropped in translation and every replay 400'd regardless.
It rides a namespaced memcode_signature extension now — the same channel pattern
as memcode_opaque for reasoning blocks — through all four hops: streaming delta,
accumulator, block assembly, and the encode back on replay. omitempty keeps the
standard shape standard for providers that issue no signature.
Verified live end to end against the deployed gateway: three sequential tool
rounds on gemini-flash, correct answers, no retries and no fallback. The same
prompt before this change failed three times and fell over to haiku.
catalog/models.json and models.json: resolved onto main's post-#7 catalog as the
base, so both retired Kimi rows stay gone, then re-applied the 14 cross-vendor
chains on top. Resolved structurally against the parsed catalog rather than by
editing conflict text — the file round-trips byte-identically through json, so
the rewrite touches only the fallback arrays.
policy_test.go: a comment-wording conflict only; took main's phrasing.
No chain references a retired model, and TestFallbackFirstHopLeavesTheVendor
still passes: every first hop changes vendor.
@timothyerwin
timothyerwin merged commit 29f67f9 into mainSep 3, 2026
2 checks passed
@timothyerwin
timothyerwin deleted the fix-gemini-tool-replay branch September 3, 2026 08:13
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@timothyerwin
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Fix the Gemini tool-replay 400, and stop it from costing 12 minutes of silence - #8

Merged
timothyerwin merged 6 commits into
mainfrom
fix-gemini-tool-replay
Sep 3, 2026
Merged

Fix the Gemini tool-replay 400, and stop it from costing 12 minutes of silence#8
timothyerwin merged 6 commits into
mainfrom
fix-gemini-tool-replay

Conversation

@timothyerwin

@timothyerwintimothyerwin commented Sep 3, 2026

Copy link
Copy Markdown
Contributor

Code review was stuck on "running" with nothing posted. Six bugs compounded. Four are fixed here; two were gateway-side and are already deployed (memcode-api-00048-sph).

The actual root cause chain

A Gemini tool call carries an opaque thoughtSignature that must be echoed back when the call is replayed. It was being dropped in two independent places, and three more bugs turned that one 400 into twelve minutes of silence.

1. The Gemini adapter dropped the signature

The adapter decoded functionCall into a tool_use block but discarded the signature, so any multi-turn tool conversation 400'd on its second turn. The reviewer is a tool-using agent, so it hit this on its first replay, every time.

Verified against live Vertex: 3.6 fails identically. This is a long-standing adapter bug, not a Gemini 3.8 regression. My earlier 3.8 checks were text, vision, PDF and thinking — all single-turn, which is exactly the path that works.

2. The hosted wire had nowhere to put it

Fixing the adapter was necessary but not sufficient, and this is why the first gateway deploy did not fix it. On the hosted path the CLI never touches that adapter: the call crosses the wire in OpenAI-compat shape, which has no field for opaque per-call state. The signature was dropped in translation regardless.

It rides a namespaced memcode_signature extension now — the same channel pattern as memcode_opaque for reasoning blocks — through all four hops: streaming delta, accumulator, block assembly, and encode-back on replay. omitempty keeps the standard shape standard for providers that issue no signature.

3. A 400 walked the fallback chain

A malformed request, bad credentials, or an unknown model does not become valid because a different model receives it. Walking the chain on those turned one bug into N identical failures — 298 failed calls in the broken run. Now classified by category: request-shape 4xx is terminal; 408/429 and 5xx still walk.

4. Fallbacks pointed at the same vendor

gemini-flash fell back to gemini-pro — same adapter, same API, same bug. The chain existed only to fail twice as expensively. All 16 chains repointed so the first hop always changes vendor, with a catalog guard test against reintroduction.

Fallbacks are infrastructure resilience only. A capability gap still refuses rather than substituting, so the pin stays authoritative.

5. Nothing was fatal, so it ran for 12 minutes

A worker whose model call died returned errResult("agent failed: …") — an ordinary retryable tool error. For a terminal cause that is a lie, and the model believed it: ~150 identical failed calls, one every two seconds, ending only on the pipeline's own timeout. The cause never surfaced anywhere a user could see it.

Terminal worker errors now arm a turn-fatal error. The loop ends the turn and the existing path carries the cause the rest of the way: printed to the session, stored as LastError, returned as a non-zero exit and a failed gateway job. Retryable causes are untouched.

Fixed gateway-side (already deployed, not in this PR)

  • The gateway's Gemini adapter registered no error extractor, so every Gemini error there was opaque to status classification — no retry policy, no terminal detection. All of them fell through to a generic 502.
  • compatTurnError mapped everything unrecognized to 502, documented as "fallback-chain territory for the CLI". For a provider 4xx we caused that is wrong. Request-shape 4xx now returns a terminal 400 with code upstream_request_invalid — deliberately 400 and not the upstream status, because an upstream 401 is our credential problem and the CLI reads a 401 as the user being signed out.

Verification

gofmt, go vet, go test -race ./..., staticcheck@2026.1 — all clean.

Live end-to-end against the deployed gateway: three sequential tool rounds on gemini-flash, all answers correct against ground truth, no retries and no fallback. The same prompt before these changes failed three times and fell over to haiku.

New regression tests: the hosted signature round-trip (and the no-signature case, so the standard shape stays clean); multi-turn adapter replay; the terminal/retryable status table; the same-vendor fallback guard; terminal-vs-transient worker failure; the gateway's upstream-4xx mapping including the 401 sign-out trap.

What this does not fix yet

The reviewer's CLI is built from a pinned tag (repoagent/Dockerfile, CLI_REF=v0.28.0), so the memcode_signature round-trip reaches production only after this merges, a release is tagged, and the repoagent image is rebuilt and redeployed at that tag.

Until then review still hits the Gemini 400 — but thanks to the deployed gateway half it now fails in seconds with the real cause, instead of hanging for twenty minutes.

Notes

Conflicts with #7 (retire-kimi-k2); both edit catalog/models.json fallbacks. Merge order matters.

Gemini issues an opaque thoughtSignature with every functionCall and REQUIRES
it echoed back when that call is replayed in a later turn. The adapter decoded
functionCall into a tool_use block but dropped the signature, so any multi-turn
tool conversation failed with a 400 on the second turn.
This is why code review broke: the reviewer is a tool-using agent, so it hit the
400 on its first replay, every time. It reproduces identically on Gemini 3.6, so
it is an adapter bug, not a 3.8 regression.
Adds toolUseFromPart as a testable decode seam, echoes the signature back in
blockToPart, and pins both directions with a multi-turn replay test plus a
no-signature case (an older transcript must still send none, not an empty one).
A 400 means WE malformed the request; a 401/403 means credentials; a 404 means
the model does not exist. None of those change because a different model
receives them, so walking the fallback chain just multiplies one bug into N
identical failures — 298 failed calls in the broken review run came from exactly
this. 408/429 and 5xx are timing, and still walk.
Classifies by category rather than by individual status, with a table test
pinning both halves.
A same-vendor fallback is not a fallback. gemini-flash fell back to gemini-pro,
which is the same adapter over the same API — so the thought-signature 400 hit
the backup identically, and the chain existed only to fail twice as expensively.
Repoints all 16 chains so the first hop always changes vendor (gemini-flash →
haiku, luna; gemini-pro → sonnet, terra; glm-5p2 → terra, luna), and adds a
catalog guard test so a same-vendor first hop cannot be reintroduced.
Fallbacks exist to survive infrastructure failure, never to reinterpret the
user's pin: a capability gap still refuses rather than substituting.
A delegated worker whose model call failed came back to the model as
errResult("agent failed: …"), an ordinary retryable tool error. For a terminal
cause that is a lie: it will fail identically next time. The model read it as
"try again" and did — the broken review run made ~150 identical failed calls,
one every two seconds for 12 minutes, and only ended on the pipeline's own
timeout. The real cause never appeared anywhere the user could see it.
Terminal worker errors now arm turn.fatalErr, the loop ends the turn on it, and
the existing path carries it the rest of the way: printed to the session, stored
as LastError, returned as a non-zero exit and a failed gateway job. Retryable
causes (429, 5xx, transport) are untouched and still just a tool error.
Exports llm.IsTerminal for the classification, and pins both sides.
Fixing the Gemini adapter was necessary but not sufficient. On the hosted path
the CLI never touches that adapter: the tool call crosses the wire in
OpenAI-compat shape, which has nowhere to put opaque per-call provider state, so
the signature was dropped in translation and every replay 400'd regardless.
It rides a namespaced memcode_signature extension now — the same channel pattern
as memcode_opaque for reasoning blocks — through all four hops: streaming delta,
accumulator, block assembly, and the encode back on replay. omitempty keeps the
standard shape standard for providers that issue no signature.
Verified live end to end against the deployed gateway: three sequential tool
rounds on gemini-flash, correct answers, no retries and no fallback. The same
prompt before this change failed three times and fell over to haiku.
catalog/models.json and models.json: resolved onto main's post-#7 catalog as the
base, so both retired Kimi rows stay gone, then re-applied the 14 cross-vendor
chains on top. Resolved structurally against the parsed catalog rather than by
editing conflict text — the file round-trips byte-identically through json, so
the rewrite touches only the fallback arrays.
policy_test.go: a comment-wording conflict only; took main's phrasing.
No chain references a retired model, and TestFallbackFirstHopLeavesTheVendor
still passes: every first hop changes vendor.
@timothyerwin
timothyerwin merged commit 29f67f9 into mainSep 3, 2026
2 checks passed
@timothyerwin
timothyerwin deleted the fix-gemini-tool-replay branch September 3, 2026 08:13
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@timothyerwin
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Fix the Gemini tool-replay 400, and stop it from costing 12 minutes of silence - #8

Merged
timothyerwin merged 6 commits into
mainfrom
fix-gemini-tool-replay
Sep 3, 2026
Merged

Fix the Gemini tool-replay 400, and stop it from costing 12 minutes of silence#8
timothyerwin merged 6 commits into
mainfrom
fix-gemini-tool-replay

Conversation

@timothyerwin

@timothyerwintimothyerwin commented Sep 3, 2026

Copy link
Copy Markdown
Contributor

Code review was stuck on "running" with nothing posted. Six bugs compounded. Four are fixed here; two were gateway-side and are already deployed (memcode-api-00048-sph).

The actual root cause chain

A Gemini tool call carries an opaque thoughtSignature that must be echoed back when the call is replayed. It was being dropped in two independent places, and three more bugs turned that one 400 into twelve minutes of silence.

1. The Gemini adapter dropped the signature

The adapter decoded functionCall into a tool_use block but discarded the signature, so any multi-turn tool conversation 400'd on its second turn. The reviewer is a tool-using agent, so it hit this on its first replay, every time.

Verified against live Vertex: 3.6 fails identically. This is a long-standing adapter bug, not a Gemini 3.8 regression. My earlier 3.8 checks were text, vision, PDF and thinking — all single-turn, which is exactly the path that works.

2. The hosted wire had nowhere to put it

Fixing the adapter was necessary but not sufficient, and this is why the first gateway deploy did not fix it. On the hosted path the CLI never touches that adapter: the call crosses the wire in OpenAI-compat shape, which has no field for opaque per-call state. The signature was dropped in translation regardless.

It rides a namespaced memcode_signature extension now — the same channel pattern as memcode_opaque for reasoning blocks — through all four hops: streaming delta, accumulator, block assembly, and encode-back on replay. omitempty keeps the standard shape standard for providers that issue no signature.

3. A 400 walked the fallback chain

A malformed request, bad credentials, or an unknown model does not become valid because a different model receives it. Walking the chain on those turned one bug into N identical failures — 298 failed calls in the broken run. Now classified by category: request-shape 4xx is terminal; 408/429 and 5xx still walk.

4. Fallbacks pointed at the same vendor

gemini-flash fell back to gemini-pro — same adapter, same API, same bug. The chain existed only to fail twice as expensively. All 16 chains repointed so the first hop always changes vendor, with a catalog guard test against reintroduction.

Fallbacks are infrastructure resilience only. A capability gap still refuses rather than substituting, so the pin stays authoritative.

5. Nothing was fatal, so it ran for 12 minutes

A worker whose model call died returned errResult("agent failed: …") — an ordinary retryable tool error. For a terminal cause that is a lie, and the model believed it: ~150 identical failed calls, one every two seconds, ending only on the pipeline's own timeout. The cause never surfaced anywhere a user could see it.

Terminal worker errors now arm a turn-fatal error. The loop ends the turn and the existing path carries the cause the rest of the way: printed to the session, stored as LastError, returned as a non-zero exit and a failed gateway job. Retryable causes are untouched.

Fixed gateway-side (already deployed, not in this PR)

  • The gateway's Gemini adapter registered no error extractor, so every Gemini error there was opaque to status classification — no retry policy, no terminal detection. All of them fell through to a generic 502.
  • compatTurnError mapped everything unrecognized to 502, documented as "fallback-chain territory for the CLI". For a provider 4xx we caused that is wrong. Request-shape 4xx now returns a terminal 400 with code upstream_request_invalid — deliberately 400 and not the upstream status, because an upstream 401 is our credential problem and the CLI reads a 401 as the user being signed out.

Verification

gofmt, go vet, go test -race ./..., staticcheck@2026.1 — all clean.

Live end-to-end against the deployed gateway: three sequential tool rounds on gemini-flash, all answers correct against ground truth, no retries and no fallback. The same prompt before these changes failed three times and fell over to haiku.

New regression tests: the hosted signature round-trip (and the no-signature case, so the standard shape stays clean); multi-turn adapter replay; the terminal/retryable status table; the same-vendor fallback guard; terminal-vs-transient worker failure; the gateway's upstream-4xx mapping including the 401 sign-out trap.

What this does not fix yet

The reviewer's CLI is built from a pinned tag (repoagent/Dockerfile, CLI_REF=v0.28.0), so the memcode_signature round-trip reaches production only after this merges, a release is tagged, and the repoagent image is rebuilt and redeployed at that tag.

Until then review still hits the Gemini 400 — but thanks to the deployed gateway half it now fails in seconds with the real cause, instead of hanging for twenty minutes.

Notes

Conflicts with #7 (retire-kimi-k2); both edit catalog/models.json fallbacks. Merge order matters.

Gemini issues an opaque thoughtSignature with every functionCall and REQUIRES
it echoed back when that call is replayed in a later turn. The adapter decoded
functionCall into a tool_use block but dropped the signature, so any multi-turn
tool conversation failed with a 400 on the second turn.
This is why code review broke: the reviewer is a tool-using agent, so it hit the
400 on its first replay, every time. It reproduces identically on Gemini 3.6, so
it is an adapter bug, not a 3.8 regression.
Adds toolUseFromPart as a testable decode seam, echoes the signature back in
blockToPart, and pins both directions with a multi-turn replay test plus a
no-signature case (an older transcript must still send none, not an empty one).
A 400 means WE malformed the request; a 401/403 means credentials; a 404 means
the model does not exist. None of those change because a different model
receives them, so walking the fallback chain just multiplies one bug into N
identical failures — 298 failed calls in the broken review run came from exactly
this. 408/429 and 5xx are timing, and still walk.
Classifies by category rather than by individual status, with a table test
pinning both halves.
A same-vendor fallback is not a fallback. gemini-flash fell back to gemini-pro,
which is the same adapter over the same API — so the thought-signature 400 hit
the backup identically, and the chain existed only to fail twice as expensively.
Repoints all 16 chains so the first hop always changes vendor (gemini-flash →
haiku, luna; gemini-pro → sonnet, terra; glm-5p2 → terra, luna), and adds a
catalog guard test so a same-vendor first hop cannot be reintroduced.
Fallbacks exist to survive infrastructure failure, never to reinterpret the
user's pin: a capability gap still refuses rather than substituting.
A delegated worker whose model call failed came back to the model as
errResult("agent failed: …"), an ordinary retryable tool error. For a terminal
cause that is a lie: it will fail identically next time. The model read it as
"try again" and did — the broken review run made ~150 identical failed calls,
one every two seconds for 12 minutes, and only ended on the pipeline's own
timeout. The real cause never appeared anywhere the user could see it.
Terminal worker errors now arm turn.fatalErr, the loop ends the turn on it, and
the existing path carries it the rest of the way: printed to the session, stored
as LastError, returned as a non-zero exit and a failed gateway job. Retryable
causes (429, 5xx, transport) are untouched and still just a tool error.
Exports llm.IsTerminal for the classification, and pins both sides.
Fixing the Gemini adapter was necessary but not sufficient. On the hosted path
the CLI never touches that adapter: the tool call crosses the wire in
OpenAI-compat shape, which has nowhere to put opaque per-call provider state, so
the signature was dropped in translation and every replay 400'd regardless.
It rides a namespaced memcode_signature extension now — the same channel pattern
as memcode_opaque for reasoning blocks — through all four hops: streaming delta,
accumulator, block assembly, and the encode back on replay. omitempty keeps the
standard shape standard for providers that issue no signature.
Verified live end to end against the deployed gateway: three sequential tool
rounds on gemini-flash, correct answers, no retries and no fallback. The same
prompt before this change failed three times and fell over to haiku.
catalog/models.json and models.json: resolved onto main's post-#7 catalog as the
base, so both retired Kimi rows stay gone, then re-applied the 14 cross-vendor
chains on top. Resolved structurally against the parsed catalog rather than by
editing conflict text — the file round-trips byte-identically through json, so
the rewrite touches only the fallback arrays.
policy_test.go: a comment-wording conflict only; took main's phrasing.
No chain references a retired model, and TestFallbackFirstHopLeavesTheVendor
still passes: every first hop changes vendor.
@timothyerwin
timothyerwin merged commit 29f67f9 into mainSep 3, 2026
2 checks passed
@timothyerwin
timothyerwin deleted the fix-gemini-tool-replay branch September 3, 2026 08:13
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@timothyerwin
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Fix the Gemini tool-replay 400, and stop it from costing 12 minutes of silence - #8

Merged
timothyerwin merged 6 commits into
mainfrom
fix-gemini-tool-replay
Sep 3, 2026
Merged

Fix the Gemini tool-replay 400, and stop it from costing 12 minutes of silence#8
timothyerwin merged 6 commits into
mainfrom
fix-gemini-tool-replay

Conversation

@timothyerwin

@timothyerwintimothyerwin commented Sep 3, 2026

Copy link
Copy Markdown
Contributor

Code review was stuck on "running" with nothing posted. Six bugs compounded. Four are fixed here; two were gateway-side and are already deployed (memcode-api-00048-sph).

The actual root cause chain

A Gemini tool call carries an opaque thoughtSignature that must be echoed back when the call is replayed. It was being dropped in two independent places, and three more bugs turned that one 400 into twelve minutes of silence.

1. The Gemini adapter dropped the signature

The adapter decoded functionCall into a tool_use block but discarded the signature, so any multi-turn tool conversation 400'd on its second turn. The reviewer is a tool-using agent, so it hit this on its first replay, every time.

Verified against live Vertex: 3.6 fails identically. This is a long-standing adapter bug, not a Gemini 3.8 regression. My earlier 3.8 checks were text, vision, PDF and thinking — all single-turn, which is exactly the path that works.

2. The hosted wire had nowhere to put it

Fixing the adapter was necessary but not sufficient, and this is why the first gateway deploy did not fix it. On the hosted path the CLI never touches that adapter: the call crosses the wire in OpenAI-compat shape, which has no field for opaque per-call state. The signature was dropped in translation regardless.

It rides a namespaced memcode_signature extension now — the same channel pattern as memcode_opaque for reasoning blocks — through all four hops: streaming delta, accumulator, block assembly, and encode-back on replay. omitempty keeps the standard shape standard for providers that issue no signature.

3. A 400 walked the fallback chain

A malformed request, bad credentials, or an unknown model does not become valid because a different model receives it. Walking the chain on those turned one bug into N identical failures — 298 failed calls in the broken run. Now classified by category: request-shape 4xx is terminal; 408/429 and 5xx still walk.

4. Fallbacks pointed at the same vendor

gemini-flash fell back to gemini-pro — same adapter, same API, same bug. The chain existed only to fail twice as expensively. All 16 chains repointed so the first hop always changes vendor, with a catalog guard test against reintroduction.

Fallbacks are infrastructure resilience only. A capability gap still refuses rather than substituting, so the pin stays authoritative.

5. Nothing was fatal, so it ran for 12 minutes

A worker whose model call died returned errResult("agent failed: …") — an ordinary retryable tool error. For a terminal cause that is a lie, and the model believed it: ~150 identical failed calls, one every two seconds, ending only on the pipeline's own timeout. The cause never surfaced anywhere a user could see it.

Terminal worker errors now arm a turn-fatal error. The loop ends the turn and the existing path carries the cause the rest of the way: printed to the session, stored as LastError, returned as a non-zero exit and a failed gateway job. Retryable causes are untouched.

Fixed gateway-side (already deployed, not in this PR)

  • The gateway's Gemini adapter registered no error extractor, so every Gemini error there was opaque to status classification — no retry policy, no terminal detection. All of them fell through to a generic 502.
  • compatTurnError mapped everything unrecognized to 502, documented as "fallback-chain territory for the CLI". For a provider 4xx we caused that is wrong. Request-shape 4xx now returns a terminal 400 with code upstream_request_invalid — deliberately 400 and not the upstream status, because an upstream 401 is our credential problem and the CLI reads a 401 as the user being signed out.

Verification

gofmt, go vet, go test -race ./..., staticcheck@2026.1 — all clean.

Live end-to-end against the deployed gateway: three sequential tool rounds on gemini-flash, all answers correct against ground truth, no retries and no fallback. The same prompt before these changes failed three times and fell over to haiku.

New regression tests: the hosted signature round-trip (and the no-signature case, so the standard shape stays clean); multi-turn adapter replay; the terminal/retryable status table; the same-vendor fallback guard; terminal-vs-transient worker failure; the gateway's upstream-4xx mapping including the 401 sign-out trap.

What this does not fix yet

The reviewer's CLI is built from a pinned tag (repoagent/Dockerfile, CLI_REF=v0.28.0), so the memcode_signature round-trip reaches production only after this merges, a release is tagged, and the repoagent image is rebuilt and redeployed at that tag.

Until then review still hits the Gemini 400 — but thanks to the deployed gateway half it now fails in seconds with the real cause, instead of hanging for twenty minutes.

Notes

Conflicts with #7 (retire-kimi-k2); both edit catalog/models.json fallbacks. Merge order matters.

Gemini issues an opaque thoughtSignature with every functionCall and REQUIRES
it echoed back when that call is replayed in a later turn. The adapter decoded
functionCall into a tool_use block but dropped the signature, so any multi-turn
tool conversation failed with a 400 on the second turn.
This is why code review broke: the reviewer is a tool-using agent, so it hit the
400 on its first replay, every time. It reproduces identically on Gemini 3.6, so
it is an adapter bug, not a 3.8 regression.
Adds toolUseFromPart as a testable decode seam, echoes the signature back in
blockToPart, and pins both directions with a multi-turn replay test plus a
no-signature case (an older transcript must still send none, not an empty one).
A 400 means WE malformed the request; a 401/403 means credentials; a 404 means
the model does not exist. None of those change because a different model
receives them, so walking the fallback chain just multiplies one bug into N
identical failures — 298 failed calls in the broken review run came from exactly
this. 408/429 and 5xx are timing, and still walk.
Classifies by category rather than by individual status, with a table test
pinning both halves.
A same-vendor fallback is not a fallback. gemini-flash fell back to gemini-pro,
which is the same adapter over the same API — so the thought-signature 400 hit
the backup identically, and the chain existed only to fail twice as expensively.
Repoints all 16 chains so the first hop always changes vendor (gemini-flash →
haiku, luna; gemini-pro → sonnet, terra; glm-5p2 → terra, luna), and adds a
catalog guard test so a same-vendor first hop cannot be reintroduced.
Fallbacks exist to survive infrastructure failure, never to reinterpret the
user's pin: a capability gap still refuses rather than substituting.
A delegated worker whose model call failed came back to the model as
errResult("agent failed: …"), an ordinary retryable tool error. For a terminal
cause that is a lie: it will fail identically next time. The model read it as
"try again" and did — the broken review run made ~150 identical failed calls,
one every two seconds for 12 minutes, and only ended on the pipeline's own
timeout. The real cause never appeared anywhere the user could see it.
Terminal worker errors now arm turn.fatalErr, the loop ends the turn on it, and
the existing path carries it the rest of the way: printed to the session, stored
as LastError, returned as a non-zero exit and a failed gateway job. Retryable
causes (429, 5xx, transport) are untouched and still just a tool error.
Exports llm.IsTerminal for the classification, and pins both sides.
Fixing the Gemini adapter was necessary but not sufficient. On the hosted path
the CLI never touches that adapter: the tool call crosses the wire in
OpenAI-compat shape, which has nowhere to put opaque per-call provider state, so
the signature was dropped in translation and every replay 400'd regardless.
It rides a namespaced memcode_signature extension now — the same channel pattern
as memcode_opaque for reasoning blocks — through all four hops: streaming delta,
accumulator, block assembly, and the encode back on replay. omitempty keeps the
standard shape standard for providers that issue no signature.
Verified live end to end against the deployed gateway: three sequential tool
rounds on gemini-flash, correct answers, no retries and no fallback. The same
prompt before this change failed three times and fell over to haiku.
catalog/models.json and models.json: resolved onto main's post-#7 catalog as the
base, so both retired Kimi rows stay gone, then re-applied the 14 cross-vendor
chains on top. Resolved structurally against the parsed catalog rather than by
editing conflict text — the file round-trips byte-identically through json, so
the rewrite touches only the fallback arrays.
policy_test.go: a comment-wording conflict only; took main's phrasing.
No chain references a retired model, and TestFallbackFirstHopLeavesTheVendor
still passes: every first hop changes vendor.
@timothyerwin
timothyerwin merged commit 29f67f9 into mainSep 3, 2026
2 checks passed
@timothyerwin
timothyerwin deleted the fix-gemini-tool-replay branch September 3, 2026 08:13
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@timothyerwin
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Fix the Gemini tool-replay 400, and stop it from costing 12 minutes of silence - #8

Merged
timothyerwin merged 6 commits into
mainfrom
fix-gemini-tool-replay
Sep 3, 2026
Merged

Fix the Gemini tool-replay 400, and stop it from costing 12 minutes of silence#8
timothyerwin merged 6 commits into
mainfrom
fix-gemini-tool-replay

Conversation

@timothyerwin

@timothyerwintimothyerwin commented Sep 3, 2026

Copy link
Copy Markdown
Contributor

Code review was stuck on "running" with nothing posted. Six bugs compounded. Four are fixed here; two were gateway-side and are already deployed (memcode-api-00048-sph).

The actual root cause chain

A Gemini tool call carries an opaque thoughtSignature that must be echoed back when the call is replayed. It was being dropped in two independent places, and three more bugs turned that one 400 into twelve minutes of silence.

1. The Gemini adapter dropped the signature

The adapter decoded functionCall into a tool_use block but discarded the signature, so any multi-turn tool conversation 400'd on its second turn. The reviewer is a tool-using agent, so it hit this on its first replay, every time.

Verified against live Vertex: 3.6 fails identically. This is a long-standing adapter bug, not a Gemini 3.8 regression. My earlier 3.8 checks were text, vision, PDF and thinking — all single-turn, which is exactly the path that works.

2. The hosted wire had nowhere to put it

Fixing the adapter was necessary but not sufficient, and this is why the first gateway deploy did not fix it. On the hosted path the CLI never touches that adapter: the call crosses the wire in OpenAI-compat shape, which has no field for opaque per-call state. The signature was dropped in translation regardless.

It rides a namespaced memcode_signature extension now — the same channel pattern as memcode_opaque for reasoning blocks — through all four hops: streaming delta, accumulator, block assembly, and encode-back on replay. omitempty keeps the standard shape standard for providers that issue no signature.

3. A 400 walked the fallback chain

A malformed request, bad credentials, or an unknown model does not become valid because a different model receives it. Walking the chain on those turned one bug into N identical failures — 298 failed calls in the broken run. Now classified by category: request-shape 4xx is terminal; 408/429 and 5xx still walk.

4. Fallbacks pointed at the same vendor

gemini-flash fell back to gemini-pro — same adapter, same API, same bug. The chain existed only to fail twice as expensively. All 16 chains repointed so the first hop always changes vendor, with a catalog guard test against reintroduction.

Fallbacks are infrastructure resilience only. A capability gap still refuses rather than substituting, so the pin stays authoritative.

5. Nothing was fatal, so it ran for 12 minutes

A worker whose model call died returned errResult("agent failed: …") — an ordinary retryable tool error. For a terminal cause that is a lie, and the model believed it: ~150 identical failed calls, one every two seconds, ending only on the pipeline's own timeout. The cause never surfaced anywhere a user could see it.

Terminal worker errors now arm a turn-fatal error. The loop ends the turn and the existing path carries the cause the rest of the way: printed to the session, stored as LastError, returned as a non-zero exit and a failed gateway job. Retryable causes are untouched.

Fixed gateway-side (already deployed, not in this PR)

  • The gateway's Gemini adapter registered no error extractor, so every Gemini error there was opaque to status classification — no retry policy, no terminal detection. All of them fell through to a generic 502.
  • compatTurnError mapped everything unrecognized to 502, documented as "fallback-chain territory for the CLI". For a provider 4xx we caused that is wrong. Request-shape 4xx now returns a terminal 400 with code upstream_request_invalid — deliberately 400 and not the upstream status, because an upstream 401 is our credential problem and the CLI reads a 401 as the user being signed out.

Verification

gofmt, go vet, go test -race ./..., staticcheck@2026.1 — all clean.

Live end-to-end against the deployed gateway: three sequential tool rounds on gemini-flash, all answers correct against ground truth, no retries and no fallback. The same prompt before these changes failed three times and fell over to haiku.

New regression tests: the hosted signature round-trip (and the no-signature case, so the standard shape stays clean); multi-turn adapter replay; the terminal/retryable status table; the same-vendor fallback guard; terminal-vs-transient worker failure; the gateway's upstream-4xx mapping including the 401 sign-out trap.

What this does not fix yet

The reviewer's CLI is built from a pinned tag (repoagent/Dockerfile, CLI_REF=v0.28.0), so the memcode_signature round-trip reaches production only after this merges, a release is tagged, and the repoagent image is rebuilt and redeployed at that tag.

Until then review still hits the Gemini 400 — but thanks to the deployed gateway half it now fails in seconds with the real cause, instead of hanging for twenty minutes.

Notes

Conflicts with #7 (retire-kimi-k2); both edit catalog/models.json fallbacks. Merge order matters.

Gemini issues an opaque thoughtSignature with every functionCall and REQUIRES
it echoed back when that call is replayed in a later turn. The adapter decoded
functionCall into a tool_use block but dropped the signature, so any multi-turn
tool conversation failed with a 400 on the second turn.
This is why code review broke: the reviewer is a tool-using agent, so it hit the
400 on its first replay, every time. It reproduces identically on Gemini 3.6, so
it is an adapter bug, not a 3.8 regression.
Adds toolUseFromPart as a testable decode seam, echoes the signature back in
blockToPart, and pins both directions with a multi-turn replay test plus a
no-signature case (an older transcript must still send none, not an empty one).
A 400 means WE malformed the request; a 401/403 means credentials; a 404 means
the model does not exist. None of those change because a different model
receives them, so walking the fallback chain just multiplies one bug into N
identical failures — 298 failed calls in the broken review run came from exactly
this. 408/429 and 5xx are timing, and still walk.
Classifies by category rather than by individual status, with a table test
pinning both halves.
A same-vendor fallback is not a fallback. gemini-flash fell back to gemini-pro,
which is the same adapter over the same API — so the thought-signature 400 hit
the backup identically, and the chain existed only to fail twice as expensively.
Repoints all 16 chains so the first hop always changes vendor (gemini-flash →
haiku, luna; gemini-pro → sonnet, terra; glm-5p2 → terra, luna), and adds a
catalog guard test so a same-vendor first hop cannot be reintroduced.
Fallbacks exist to survive infrastructure failure, never to reinterpret the
user's pin: a capability gap still refuses rather than substituting.
A delegated worker whose model call failed came back to the model as
errResult("agent failed: …"), an ordinary retryable tool error. For a terminal
cause that is a lie: it will fail identically next time. The model read it as
"try again" and did — the broken review run made ~150 identical failed calls,
one every two seconds for 12 minutes, and only ended on the pipeline's own
timeout. The real cause never appeared anywhere the user could see it.
Terminal worker errors now arm turn.fatalErr, the loop ends the turn on it, and
the existing path carries it the rest of the way: printed to the session, stored
as LastError, returned as a non-zero exit and a failed gateway job. Retryable
causes (429, 5xx, transport) are untouched and still just a tool error.
Exports llm.IsTerminal for the classification, and pins both sides.
Fixing the Gemini adapter was necessary but not sufficient. On the hosted path
the CLI never touches that adapter: the tool call crosses the wire in
OpenAI-compat shape, which has nowhere to put opaque per-call provider state, so
the signature was dropped in translation and every replay 400'd regardless.
It rides a namespaced memcode_signature extension now — the same channel pattern
as memcode_opaque for reasoning blocks — through all four hops: streaming delta,
accumulator, block assembly, and the encode back on replay. omitempty keeps the
standard shape standard for providers that issue no signature.
Verified live end to end against the deployed gateway: three sequential tool
rounds on gemini-flash, correct answers, no retries and no fallback. The same
prompt before this change failed three times and fell over to haiku.
catalog/models.json and models.json: resolved onto main's post-#7 catalog as the
base, so both retired Kimi rows stay gone, then re-applied the 14 cross-vendor
chains on top. Resolved structurally against the parsed catalog rather than by
editing conflict text — the file round-trips byte-identically through json, so
the rewrite touches only the fallback arrays.
policy_test.go: a comment-wording conflict only; took main's phrasing.
No chain references a retired model, and TestFallbackFirstHopLeavesTheVendor
still passes: every first hop changes vendor.
@timothyerwin
timothyerwin merged commit 29f67f9 into mainSep 3, 2026
2 checks passed
@timothyerwin
timothyerwin deleted the fix-gemini-tool-replay branch September 3, 2026 08:13
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@timothyerwin