Uh oh!
There was an error while loading. Please reload this page.
Fix the Gemini tool-replay 400, and stop it from costing 12 minutes of silence - #8
Merged
Conversation
Gemini issues an opaque thoughtSignature with every functionCall and REQUIRES it echoed back when that call is replayed in a later turn. The adapter decoded functionCall into a tool_use block but dropped the signature, so any multi-turn tool conversation failed with a 400 on the second turn. This is why code review broke: the reviewer is a tool-using agent, so it hit the 400 on its first replay, every time. It reproduces identically on Gemini 3.6, so it is an adapter bug, not a 3.8 regression. Adds toolUseFromPart as a testable decode seam, echoes the signature back in blockToPart, and pins both directions with a multi-turn replay test plus a no-signature case (an older transcript must still send none, not an empty one).
A 400 means WE malformed the request; a 401/403 means credentials; a 404 means the model does not exist. None of those change because a different model receives them, so walking the fallback chain just multiplies one bug into N identical failures — 298 failed calls in the broken review run came from exactly this. 408/429 and 5xx are timing, and still walk. Classifies by category rather than by individual status, with a table test pinning both halves.
A same-vendor fallback is not a fallback. gemini-flash fell back to gemini-pro, which is the same adapter over the same API — so the thought-signature 400 hit the backup identically, and the chain existed only to fail twice as expensively. Repoints all 16 chains so the first hop always changes vendor (gemini-flash → haiku, luna; gemini-pro → sonnet, terra; glm-5p2 → terra, luna), and adds a catalog guard test so a same-vendor first hop cannot be reintroduced. Fallbacks exist to survive infrastructure failure, never to reinterpret the user's pin: a capability gap still refuses rather than substituting.
A delegated worker whose model call failed came back to the model as
errResult("agent failed: …"), an ordinary retryable tool error. For a terminal
cause that is a lie: it will fail identically next time. The model read it as
"try again" and did — the broken review run made ~150 identical failed calls,
one every two seconds for 12 minutes, and only ended on the pipeline's own
timeout. The real cause never appeared anywhere the user could see it.
Terminal worker errors now arm turn.fatalErr, the loop ends the turn on it, and
the existing path carries it the rest of the way: printed to the session, stored
as LastError, returned as a non-zero exit and a failed gateway job. Retryable
causes (429, 5xx, transport) are untouched and still just a tool error.
Exports llm.IsTerminal for the classification, and pins both sides.Fixing the Gemini adapter was necessary but not sufficient. On the hosted path the CLI never touches that adapter: the tool call crosses the wire in OpenAI-compat shape, which has nowhere to put opaque per-call provider state, so the signature was dropped in translation and every replay 400'd regardless. It rides a namespaced memcode_signature extension now — the same channel pattern as memcode_opaque for reasoning blocks — through all four hops: streaming delta, accumulator, block assembly, and the encode back on replay. omitempty keeps the standard shape standard for providers that issue no signature. Verified live end to end against the deployed gateway: three sequential tool rounds on gemini-flash, correct answers, no retries and no fallback. The same prompt before this change failed three times and fell over to haiku.
catalog/models.json and models.json: resolved onto main's post-#7 catalog as the base, so both retired Kimi rows stay gone, then re-applied the 14 cross-vendor chains on top. Resolved structurally against the parsed catalog rather than by editing conflict text — the file round-trips byte-identically through json, so the rewrite touches only the fallback arrays. policy_test.go: a comment-wording conflict only; took main's phrasing. No chain references a retired model, and TestFallbackFirstHopLeavesTheVendor still passes: every first hop changes vendor.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for freeto join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Code review was stuck on "running" with nothing posted. Six bugs compounded. Four are fixed here; two were gateway-side and are already deployed (
memcode-api-00048-sph).The actual root cause chain
A Gemini tool call carries an opaque
thoughtSignaturethat must be echoed back when the call is replayed. It was being dropped in two independent places, and three more bugs turned that one 400 into twelve minutes of silence.1. The Gemini adapter dropped the signature
The adapter decoded
functionCallinto atool_useblock but discarded the signature, so any multi-turn tool conversation 400'd on its second turn. The reviewer is a tool-using agent, so it hit this on its first replay, every time.Verified against live Vertex: 3.6 fails identically. This is a long-standing adapter bug, not a Gemini 3.8 regression. My earlier 3.8 checks were text, vision, PDF and thinking — all single-turn, which is exactly the path that works.
2. The hosted wire had nowhere to put it
Fixing the adapter was necessary but not sufficient, and this is why the first gateway deploy did not fix it. On the hosted path the CLI never touches that adapter: the call crosses the wire in OpenAI-compat shape, which has no field for opaque per-call state. The signature was dropped in translation regardless.
It rides a namespaced
memcode_signatureextension now — the same channel pattern asmemcode_opaquefor reasoning blocks — through all four hops: streaming delta, accumulator, block assembly, and encode-back on replay.omitemptykeeps the standard shape standard for providers that issue no signature.3. A 400 walked the fallback chain
A malformed request, bad credentials, or an unknown model does not become valid because a different model receives it. Walking the chain on those turned one bug into N identical failures — 298 failed calls in the broken run. Now classified by category: request-shape 4xx is terminal; 408/429 and 5xx still walk.
4. Fallbacks pointed at the same vendor
gemini-flashfell back togemini-pro— same adapter, same API, same bug. The chain existed only to fail twice as expensively. All 16 chains repointed so the first hop always changes vendor, with a catalog guard test against reintroduction.Fallbacks are infrastructure resilience only. A capability gap still refuses rather than substituting, so the pin stays authoritative.
5. Nothing was fatal, so it ran for 12 minutes
A worker whose model call died returned
errResult("agent failed: …")— an ordinary retryable tool error. For a terminal cause that is a lie, and the model believed it: ~150 identical failed calls, one every two seconds, ending only on the pipeline's own timeout. The cause never surfaced anywhere a user could see it.Terminal worker errors now arm a turn-fatal error. The loop ends the turn and the existing path carries the cause the rest of the way: printed to the session, stored as
LastError, returned as a non-zero exit and a failed gateway job. Retryable causes are untouched.Fixed gateway-side (already deployed, not in this PR)
compatTurnErrormapped everything unrecognized to 502, documented as "fallback-chain territory for the CLI". For a provider 4xx we caused that is wrong. Request-shape 4xx now returns a terminal 400 with codeupstream_request_invalid— deliberately 400 and not the upstream status, because an upstream 401 is our credential problem and the CLI reads a 401 as the user being signed out.Verification
gofmt,go vet,go test -race ./...,staticcheck@2026.1— all clean.Live end-to-end against the deployed gateway: three sequential tool rounds on
gemini-flash, all answers correct against ground truth, no retries and no fallback. The same prompt before these changes failed three times and fell over to haiku.New regression tests: the hosted signature round-trip (and the no-signature case, so the standard shape stays clean); multi-turn adapter replay; the terminal/retryable status table; the same-vendor fallback guard; terminal-vs-transient worker failure; the gateway's upstream-4xx mapping including the 401 sign-out trap.
What this does not fix yet
The reviewer's CLI is built from a pinned tag (
repoagent/Dockerfile,CLI_REF=v0.28.0), so thememcode_signatureround-trip reaches production only after this merges, a release is tagged, and the repoagent image is rebuilt and redeployed at that tag.Until then review still hits the Gemini 400 — but thanks to the deployed gateway half it now fails in seconds with the real cause, instead of hanging for twenty minutes.
Notes
Conflicts with #7 (
retire-kimi-k2); both editcatalog/models.jsonfallbacks. Merge order matters.