fix(sync): stop the shell thread list from wedging on a stale cache - #40

Merged
DanielGGordon merged 1 commit into
mainfrom
t3code/mobile-chat-sync-fix
Jul 23, 2026
Merged

fix(sync): stop the shell thread list from wedging on a stale cache#40
DanielGGordon merged 1 commit into
mainfrom
t3code/mobile-chat-sync-fix

Conversation

@DanielGGordon

Copy link
Copy Markdown
Owner

The bug

On mobile, the sidebar thread list showed conversations days out of date, and refreshing the page did not help.

The list is not fetched on load. It is a snapshot cached in IndexedDB (t3code:connection-runtime, store shell) that the client renders immediately, skipping the HTTP snapshot entirely, then resumes over a WebSocket using a persisted event cursor. If that resume never completes, the stale cache is what the next load resumes from too — so the wedge is self-reinforcing and survives reloads.

Diagnosed against a live server holding 193 threads with a healthy event log (1→59,935, all projectors current). The data was fine; the client could not get to it.

Three gaps

  1. No retry.subscribeShell passed no retryExpectedFailureAfter, so an expected (non-transport) failure returned a drained stream and the subscription ended for the life of the page. The UI kept rendering the cached snapshot and silently never updated. The thread-detail path already got this right.

  2. Boot-time resume cursor.subscribeInput was computed once, outside the retry/reconnect loop. Every reconnect replayed the same catch-up window rather than resuming from where the client actually got to.

  3. Server always honored afterSequence. Two ways that fails:

    • Cursor far behind: every thread-aggregate event in the window costs a getThreadShellById query and emits a full thread shell. On the server that prompted this, a cursor ~6 days old meant ~11k events / ~17MB to rebuild a list a single snapshot expresses in a fraction of that. A client that cannot finish the replay never advances its cursor, so its next attempt is identical — stale indefinitely.
    • Cursor ahead of head: a rebuilt/restored event log (or a different server) replays nothing, and every live event is then dropped by the client's sequence <= snapshotSequence check. Stale forever, with no path to recovery.

The fix

  • Add retryExpectedFailureAfter: "250 millis" to the shell subscription, mirroring the thread path.
  • Resolve the resume cursor per subscribe. subscribe() now accepts a thunk, so cursor-carrying inputs are re-read on every attempt instead of captured once.
  • Server falls back to a full snapshot when the client's cursor is ahead of the projected sequence or more than SHELL_CATCH_UP_REPLAY_LIMIT (500) behind. The client applies snapshots wholesale, so this heals a wedged cache. If the server cannot read its own cursor it prefers the previous replay behavior.

Tests

Four new tests, each verified to fail without the corresponding fix (not vacuous):

  • shell-sync.test.ts — a stream that advances to sequence 9 then fails must resubscribe and ask for afterSequence: 9. Without the retry the recorded attempts are [5]; with the retry but a boot-time cursor they are [5, 5]; fixed, [5, 9].
  • server.test.ts — replay is still used for a recent cursor (control), and a snapshot is sent when the client is far behind or ahead of head.

packages/client-runtime 255/255 pass; apps/server typechecks clean.

Notes

  • One pre-existing apps/server test failure (preserves structured workspace rpc failures) and the scripts/test-deploy*.ts typecheck/lint errors are present on the base commit and unrelated to this change — confirmed by stashing.
  • Not addressed here: static assets ship with no Cache-Control/ETag/Last-Modified (apps/server/src/http.ts), so a phone can also pin a stale JS bundle after a redeploy. The thread-detail cache has the same class of stale-cursor-on-reconnect issue; this PR only makes subscribe() capable of fixing it. Worth follow-ups.
  • Existing wedged clients still need their site data cleared once; this prevents recurrence and heals clients whose cursor is unusable.

🤖 Generated with Claude Code

The sidebar thread list is a snapshot cached in IndexedDB that resumes
over a persisted event cursor. Two gaps let a client pin itself to a
stale list that survived reloads, because a reload re-reads the same
cache and re-attempts the same resume:
- The shell subscription passed no `retryExpectedFailureAfter`, so an
expected (non-transport) failure ended the stream for the life of the
page. The cached snapshot kept rendering and never updated again.
- The resume cursor was captured once at boot, so every reconnect
replayed the same catch-up window instead of resuming from where the
client actually got to.
- The server always honored `afterSequence`. A cursor far behind our
head asks for a replay costing far more than the snapshot it stands
in for (every thread-aggregate event is a projection lookup plus a
full thread shell on the socket), and a client that cannot afford
that replay never advances its cursor, so its next attempt is just as
expensive. A cursor *ahead* of our head — a rebuilt or restored event
log — replayed nothing while every live event was dropped by the
client's sequence check.
Adds the retry, resolves the resume cursor per subscribe, and falls
back to a snapshot when the client's cursor is unusable or too far
behind. `subscribe` now accepts a thunk so cursor-carrying inputs are
re-read on every attempt.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:L labels Jul 23, 2026
@DanielGGordon
DanielGGordon merged commit 6a4cfdd into mainJul 23, 2026
6 of 10 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:Lvouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@DanielGGordon
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

fix(sync): stop the shell thread list from wedging on a stale cache - #40

Merged
DanielGGordon merged 1 commit into
mainfrom
t3code/mobile-chat-sync-fix
Jul 23, 2026
Merged

fix(sync): stop the shell thread list from wedging on a stale cache#40
DanielGGordon merged 1 commit into
mainfrom
t3code/mobile-chat-sync-fix

Conversation

@DanielGGordon

Copy link
Copy Markdown
Owner

The bug

On mobile, the sidebar thread list showed conversations days out of date, and refreshing the page did not help.

The list is not fetched on load. It is a snapshot cached in IndexedDB (t3code:connection-runtime, store shell) that the client renders immediately, skipping the HTTP snapshot entirely, then resumes over a WebSocket using a persisted event cursor. If that resume never completes, the stale cache is what the next load resumes from too — so the wedge is self-reinforcing and survives reloads.

Diagnosed against a live server holding 193 threads with a healthy event log (1→59,935, all projectors current). The data was fine; the client could not get to it.

Three gaps

  1. No retry.subscribeShell passed no retryExpectedFailureAfter, so an expected (non-transport) failure returned a drained stream and the subscription ended for the life of the page. The UI kept rendering the cached snapshot and silently never updated. The thread-detail path already got this right.

  2. Boot-time resume cursor.subscribeInput was computed once, outside the retry/reconnect loop. Every reconnect replayed the same catch-up window rather than resuming from where the client actually got to.

  3. Server always honored afterSequence. Two ways that fails:

    • Cursor far behind: every thread-aggregate event in the window costs a getThreadShellById query and emits a full thread shell. On the server that prompted this, a cursor ~6 days old meant ~11k events / ~17MB to rebuild a list a single snapshot expresses in a fraction of that. A client that cannot finish the replay never advances its cursor, so its next attempt is identical — stale indefinitely.
    • Cursor ahead of head: a rebuilt/restored event log (or a different server) replays nothing, and every live event is then dropped by the client's sequence <= snapshotSequence check. Stale forever, with no path to recovery.

The fix

  • Add retryExpectedFailureAfter: "250 millis" to the shell subscription, mirroring the thread path.
  • Resolve the resume cursor per subscribe. subscribe() now accepts a thunk, so cursor-carrying inputs are re-read on every attempt instead of captured once.
  • Server falls back to a full snapshot when the client's cursor is ahead of the projected sequence or more than SHELL_CATCH_UP_REPLAY_LIMIT (500) behind. The client applies snapshots wholesale, so this heals a wedged cache. If the server cannot read its own cursor it prefers the previous replay behavior.

Tests

Four new tests, each verified to fail without the corresponding fix (not vacuous):

  • shell-sync.test.ts — a stream that advances to sequence 9 then fails must resubscribe and ask for afterSequence: 9. Without the retry the recorded attempts are [5]; with the retry but a boot-time cursor they are [5, 5]; fixed, [5, 9].
  • server.test.ts — replay is still used for a recent cursor (control), and a snapshot is sent when the client is far behind or ahead of head.

packages/client-runtime 255/255 pass; apps/server typechecks clean.

Notes

  • One pre-existing apps/server test failure (preserves structured workspace rpc failures) and the scripts/test-deploy*.ts typecheck/lint errors are present on the base commit and unrelated to this change — confirmed by stashing.
  • Not addressed here: static assets ship with no Cache-Control/ETag/Last-Modified (apps/server/src/http.ts), so a phone can also pin a stale JS bundle after a redeploy. The thread-detail cache has the same class of stale-cursor-on-reconnect issue; this PR only makes subscribe() capable of fixing it. Worth follow-ups.
  • Existing wedged clients still need their site data cleared once; this prevents recurrence and heals clients whose cursor is unusable.

🤖 Generated with Claude Code

The sidebar thread list is a snapshot cached in IndexedDB that resumes
over a persisted event cursor. Two gaps let a client pin itself to a
stale list that survived reloads, because a reload re-reads the same
cache and re-attempts the same resume:
- The shell subscription passed no `retryExpectedFailureAfter`, so an
expected (non-transport) failure ended the stream for the life of the
page. The cached snapshot kept rendering and never updated again.
- The resume cursor was captured once at boot, so every reconnect
replayed the same catch-up window instead of resuming from where the
client actually got to.
- The server always honored `afterSequence`. A cursor far behind our
head asks for a replay costing far more than the snapshot it stands
in for (every thread-aggregate event is a projection lookup plus a
full thread shell on the socket), and a client that cannot afford
that replay never advances its cursor, so its next attempt is just as
expensive. A cursor *ahead* of our head — a rebuilt or restored event
log — replayed nothing while every live event was dropped by the
client's sequence check.
Adds the retry, resolves the resume cursor per subscribe, and falls
back to a snapshot when the client's cursor is unusable or too far
behind. `subscribe` now accepts a thunk so cursor-carrying inputs are
re-read on every attempt.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:L labels Jul 23, 2026
@DanielGGordon
DanielGGordon merged commit 6a4cfdd into mainJul 23, 2026
6 of 10 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:Lvouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@DanielGGordon
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(sync): stop the shell thread list from wedging on a stale cache - #40

Merged
DanielGGordon merged 1 commit into
mainfrom
t3code/mobile-chat-sync-fix
Jul 23, 2026
Merged

fix(sync): stop the shell thread list from wedging on a stale cache#40
DanielGGordon merged 1 commit into
mainfrom
t3code/mobile-chat-sync-fix

Conversation

@DanielGGordon

Copy link
Copy Markdown
Owner

The bug

On mobile, the sidebar thread list showed conversations days out of date, and refreshing the page did not help.

The list is not fetched on load. It is a snapshot cached in IndexedDB (t3code:connection-runtime, store shell) that the client renders immediately, skipping the HTTP snapshot entirely, then resumes over a WebSocket using a persisted event cursor. If that resume never completes, the stale cache is what the next load resumes from too — so the wedge is self-reinforcing and survives reloads.

Diagnosed against a live server holding 193 threads with a healthy event log (1→59,935, all projectors current). The data was fine; the client could not get to it.

Three gaps

  1. No retry.subscribeShell passed no retryExpectedFailureAfter, so an expected (non-transport) failure returned a drained stream and the subscription ended for the life of the page. The UI kept rendering the cached snapshot and silently never updated. The thread-detail path already got this right.

  2. Boot-time resume cursor.subscribeInput was computed once, outside the retry/reconnect loop. Every reconnect replayed the same catch-up window rather than resuming from where the client actually got to.

  3. Server always honored afterSequence. Two ways that fails:

    • Cursor far behind: every thread-aggregate event in the window costs a getThreadShellById query and emits a full thread shell. On the server that prompted this, a cursor ~6 days old meant ~11k events / ~17MB to rebuild a list a single snapshot expresses in a fraction of that. A client that cannot finish the replay never advances its cursor, so its next attempt is identical — stale indefinitely.
    • Cursor ahead of head: a rebuilt/restored event log (or a different server) replays nothing, and every live event is then dropped by the client's sequence <= snapshotSequence check. Stale forever, with no path to recovery.

The fix

  • Add retryExpectedFailureAfter: "250 millis" to the shell subscription, mirroring the thread path.
  • Resolve the resume cursor per subscribe. subscribe() now accepts a thunk, so cursor-carrying inputs are re-read on every attempt instead of captured once.
  • Server falls back to a full snapshot when the client's cursor is ahead of the projected sequence or more than SHELL_CATCH_UP_REPLAY_LIMIT (500) behind. The client applies snapshots wholesale, so this heals a wedged cache. If the server cannot read its own cursor it prefers the previous replay behavior.

Tests

Four new tests, each verified to fail without the corresponding fix (not vacuous):

  • shell-sync.test.ts — a stream that advances to sequence 9 then fails must resubscribe and ask for afterSequence: 9. Without the retry the recorded attempts are [5]; with the retry but a boot-time cursor they are [5, 5]; fixed, [5, 9].
  • server.test.ts — replay is still used for a recent cursor (control), and a snapshot is sent when the client is far behind or ahead of head.

packages/client-runtime 255/255 pass; apps/server typechecks clean.

Notes

  • One pre-existing apps/server test failure (preserves structured workspace rpc failures) and the scripts/test-deploy*.ts typecheck/lint errors are present on the base commit and unrelated to this change — confirmed by stashing.
  • Not addressed here: static assets ship with no Cache-Control/ETag/Last-Modified (apps/server/src/http.ts), so a phone can also pin a stale JS bundle after a redeploy. The thread-detail cache has the same class of stale-cursor-on-reconnect issue; this PR only makes subscribe() capable of fixing it. Worth follow-ups.
  • Existing wedged clients still need their site data cleared once; this prevents recurrence and heals clients whose cursor is unusable.

🤖 Generated with Claude Code

The sidebar thread list is a snapshot cached in IndexedDB that resumes
over a persisted event cursor. Two gaps let a client pin itself to a
stale list that survived reloads, because a reload re-reads the same
cache and re-attempts the same resume:
- The shell subscription passed no `retryExpectedFailureAfter`, so an
expected (non-transport) failure ended the stream for the life of the
page. The cached snapshot kept rendering and never updated again.
- The resume cursor was captured once at boot, so every reconnect
replayed the same catch-up window instead of resuming from where the
client actually got to.
- The server always honored `afterSequence`. A cursor far behind our
head asks for a replay costing far more than the snapshot it stands
in for (every thread-aggregate event is a projection lookup plus a
full thread shell on the socket), and a client that cannot afford
that replay never advances its cursor, so its next attempt is just as
expensive. A cursor *ahead* of our head — a rebuilt or restored event
log — replayed nothing while every live event was dropped by the
client's sequence check.
Adds the retry, resolves the resume cursor per subscribe, and falls
back to a snapshot when the client's cursor is unusable or too far
behind. `subscribe` now accepts a thunk so cursor-carrying inputs are
re-read on every attempt.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:L labels Jul 23, 2026
@DanielGGordon
DanielGGordon merged commit 6a4cfdd into mainJul 23, 2026
6 of 10 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:Lvouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@DanielGGordon
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(sync): stop the shell thread list from wedging on a stale cache - #40

Merged
DanielGGordon merged 1 commit into
mainfrom
t3code/mobile-chat-sync-fix
Jul 23, 2026
Merged

fix(sync): stop the shell thread list from wedging on a stale cache#40
DanielGGordon merged 1 commit into
mainfrom
t3code/mobile-chat-sync-fix

Conversation

@DanielGGordon

Copy link
Copy Markdown
Owner

The bug

On mobile, the sidebar thread list showed conversations days out of date, and refreshing the page did not help.

The list is not fetched on load. It is a snapshot cached in IndexedDB (t3code:connection-runtime, store shell) that the client renders immediately, skipping the HTTP snapshot entirely, then resumes over a WebSocket using a persisted event cursor. If that resume never completes, the stale cache is what the next load resumes from too — so the wedge is self-reinforcing and survives reloads.

Diagnosed against a live server holding 193 threads with a healthy event log (1→59,935, all projectors current). The data was fine; the client could not get to it.

Three gaps

  1. No retry.subscribeShell passed no retryExpectedFailureAfter, so an expected (non-transport) failure returned a drained stream and the subscription ended for the life of the page. The UI kept rendering the cached snapshot and silently never updated. The thread-detail path already got this right.

  2. Boot-time resume cursor.subscribeInput was computed once, outside the retry/reconnect loop. Every reconnect replayed the same catch-up window rather than resuming from where the client actually got to.

  3. Server always honored afterSequence. Two ways that fails:

    • Cursor far behind: every thread-aggregate event in the window costs a getThreadShellById query and emits a full thread shell. On the server that prompted this, a cursor ~6 days old meant ~11k events / ~17MB to rebuild a list a single snapshot expresses in a fraction of that. A client that cannot finish the replay never advances its cursor, so its next attempt is identical — stale indefinitely.
    • Cursor ahead of head: a rebuilt/restored event log (or a different server) replays nothing, and every live event is then dropped by the client's sequence <= snapshotSequence check. Stale forever, with no path to recovery.

The fix

  • Add retryExpectedFailureAfter: "250 millis" to the shell subscription, mirroring the thread path.
  • Resolve the resume cursor per subscribe. subscribe() now accepts a thunk, so cursor-carrying inputs are re-read on every attempt instead of captured once.
  • Server falls back to a full snapshot when the client's cursor is ahead of the projected sequence or more than SHELL_CATCH_UP_REPLAY_LIMIT (500) behind. The client applies snapshots wholesale, so this heals a wedged cache. If the server cannot read its own cursor it prefers the previous replay behavior.

Tests

Four new tests, each verified to fail without the corresponding fix (not vacuous):

  • shell-sync.test.ts — a stream that advances to sequence 9 then fails must resubscribe and ask for afterSequence: 9. Without the retry the recorded attempts are [5]; with the retry but a boot-time cursor they are [5, 5]; fixed, [5, 9].
  • server.test.ts — replay is still used for a recent cursor (control), and a snapshot is sent when the client is far behind or ahead of head.

packages/client-runtime 255/255 pass; apps/server typechecks clean.

Notes

  • One pre-existing apps/server test failure (preserves structured workspace rpc failures) and the scripts/test-deploy*.ts typecheck/lint errors are present on the base commit and unrelated to this change — confirmed by stashing.
  • Not addressed here: static assets ship with no Cache-Control/ETag/Last-Modified (apps/server/src/http.ts), so a phone can also pin a stale JS bundle after a redeploy. The thread-detail cache has the same class of stale-cursor-on-reconnect issue; this PR only makes subscribe() capable of fixing it. Worth follow-ups.
  • Existing wedged clients still need their site data cleared once; this prevents recurrence and heals clients whose cursor is unusable.

🤖 Generated with Claude Code

The sidebar thread list is a snapshot cached in IndexedDB that resumes
over a persisted event cursor. Two gaps let a client pin itself to a
stale list that survived reloads, because a reload re-reads the same
cache and re-attempts the same resume:
- The shell subscription passed no `retryExpectedFailureAfter`, so an
expected (non-transport) failure ended the stream for the life of the
page. The cached snapshot kept rendering and never updated again.
- The resume cursor was captured once at boot, so every reconnect
replayed the same catch-up window instead of resuming from where the
client actually got to.
- The server always honored `afterSequence`. A cursor far behind our
head asks for a replay costing far more than the snapshot it stands
in for (every thread-aggregate event is a projection lookup plus a
full thread shell on the socket), and a client that cannot afford
that replay never advances its cursor, so its next attempt is just as
expensive. A cursor *ahead* of our head — a rebuilt or restored event
log — replayed nothing while every live event was dropped by the
client's sequence check.
Adds the retry, resolves the resume cursor per subscribe, and falls
back to a snapshot when the client's cursor is unusable or too far
behind. `subscribe` now accepts a thunk so cursor-carrying inputs are
re-read on every attempt.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:L labels Jul 23, 2026
@DanielGGordon
DanielGGordon merged commit 6a4cfdd into mainJul 23, 2026
6 of 10 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:Lvouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@DanielGGordon
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

fix(sync): stop the shell thread list from wedging on a stale cache - #40

Merged
DanielGGordon merged 1 commit into
mainfrom
t3code/mobile-chat-sync-fix
Jul 23, 2026
Merged

fix(sync): stop the shell thread list from wedging on a stale cache#40
DanielGGordon merged 1 commit into
mainfrom
t3code/mobile-chat-sync-fix

Conversation

@DanielGGordon

Copy link
Copy Markdown
Owner

The bug

On mobile, the sidebar thread list showed conversations days out of date, and refreshing the page did not help.

The list is not fetched on load. It is a snapshot cached in IndexedDB (t3code:connection-runtime, store shell) that the client renders immediately, skipping the HTTP snapshot entirely, then resumes over a WebSocket using a persisted event cursor. If that resume never completes, the stale cache is what the next load resumes from too — so the wedge is self-reinforcing and survives reloads.

Diagnosed against a live server holding 193 threads with a healthy event log (1→59,935, all projectors current). The data was fine; the client could not get to it.

Three gaps

  1. No retry.subscribeShell passed no retryExpectedFailureAfter, so an expected (non-transport) failure returned a drained stream and the subscription ended for the life of the page. The UI kept rendering the cached snapshot and silently never updated. The thread-detail path already got this right.

  2. Boot-time resume cursor.subscribeInput was computed once, outside the retry/reconnect loop. Every reconnect replayed the same catch-up window rather than resuming from where the client actually got to.

  3. Server always honored afterSequence. Two ways that fails:

    • Cursor far behind: every thread-aggregate event in the window costs a getThreadShellById query and emits a full thread shell. On the server that prompted this, a cursor ~6 days old meant ~11k events / ~17MB to rebuild a list a single snapshot expresses in a fraction of that. A client that cannot finish the replay never advances its cursor, so its next attempt is identical — stale indefinitely.
    • Cursor ahead of head: a rebuilt/restored event log (or a different server) replays nothing, and every live event is then dropped by the client's sequence <= snapshotSequence check. Stale forever, with no path to recovery.

The fix

  • Add retryExpectedFailureAfter: "250 millis" to the shell subscription, mirroring the thread path.
  • Resolve the resume cursor per subscribe. subscribe() now accepts a thunk, so cursor-carrying inputs are re-read on every attempt instead of captured once.
  • Server falls back to a full snapshot when the client's cursor is ahead of the projected sequence or more than SHELL_CATCH_UP_REPLAY_LIMIT (500) behind. The client applies snapshots wholesale, so this heals a wedged cache. If the server cannot read its own cursor it prefers the previous replay behavior.

Tests

Four new tests, each verified to fail without the corresponding fix (not vacuous):

  • shell-sync.test.ts — a stream that advances to sequence 9 then fails must resubscribe and ask for afterSequence: 9. Without the retry the recorded attempts are [5]; with the retry but a boot-time cursor they are [5, 5]; fixed, [5, 9].
  • server.test.ts — replay is still used for a recent cursor (control), and a snapshot is sent when the client is far behind or ahead of head.

packages/client-runtime 255/255 pass; apps/server typechecks clean.

Notes

  • One pre-existing apps/server test failure (preserves structured workspace rpc failures) and the scripts/test-deploy*.ts typecheck/lint errors are present on the base commit and unrelated to this change — confirmed by stashing.
  • Not addressed here: static assets ship with no Cache-Control/ETag/Last-Modified (apps/server/src/http.ts), so a phone can also pin a stale JS bundle after a redeploy. The thread-detail cache has the same class of stale-cursor-on-reconnect issue; this PR only makes subscribe() capable of fixing it. Worth follow-ups.
  • Existing wedged clients still need their site data cleared once; this prevents recurrence and heals clients whose cursor is unusable.

🤖 Generated with Claude Code

The sidebar thread list is a snapshot cached in IndexedDB that resumes
over a persisted event cursor. Two gaps let a client pin itself to a
stale list that survived reloads, because a reload re-reads the same
cache and re-attempts the same resume:
- The shell subscription passed no `retryExpectedFailureAfter`, so an
expected (non-transport) failure ended the stream for the life of the
page. The cached snapshot kept rendering and never updated again.
- The resume cursor was captured once at boot, so every reconnect
replayed the same catch-up window instead of resuming from where the
client actually got to.
- The server always honored `afterSequence`. A cursor far behind our
head asks for a replay costing far more than the snapshot it stands
in for (every thread-aggregate event is a projection lookup plus a
full thread shell on the socket), and a client that cannot afford
that replay never advances its cursor, so its next attempt is just as
expensive. A cursor *ahead* of our head — a rebuilt or restored event
log — replayed nothing while every live event was dropped by the
client's sequence check.
Adds the retry, resolves the resume cursor per subscribe, and falls
back to a snapshot when the client's cursor is unusable or too far
behind. `subscribe` now accepts a thunk so cursor-carrying inputs are
re-read on every attempt.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:L labels Jul 23, 2026
@DanielGGordon
DanielGGordon merged commit 6a4cfdd into mainJul 23, 2026
6 of 10 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:Lvouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@DanielGGordon
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(sync): stop the shell thread list from wedging on a stale cache - #40

Merged
DanielGGordon merged 1 commit into
mainfrom
t3code/mobile-chat-sync-fix
Jul 23, 2026
Merged

fix(sync): stop the shell thread list from wedging on a stale cache#40
DanielGGordon merged 1 commit into
mainfrom
t3code/mobile-chat-sync-fix

Conversation

@DanielGGordon

Copy link
Copy Markdown
Owner

The bug

On mobile, the sidebar thread list showed conversations days out of date, and refreshing the page did not help.

The list is not fetched on load. It is a snapshot cached in IndexedDB (t3code:connection-runtime, store shell) that the client renders immediately, skipping the HTTP snapshot entirely, then resumes over a WebSocket using a persisted event cursor. If that resume never completes, the stale cache is what the next load resumes from too — so the wedge is self-reinforcing and survives reloads.

Diagnosed against a live server holding 193 threads with a healthy event log (1→59,935, all projectors current). The data was fine; the client could not get to it.

Three gaps

  1. No retry.subscribeShell passed no retryExpectedFailureAfter, so an expected (non-transport) failure returned a drained stream and the subscription ended for the life of the page. The UI kept rendering the cached snapshot and silently never updated. The thread-detail path already got this right.

  2. Boot-time resume cursor.subscribeInput was computed once, outside the retry/reconnect loop. Every reconnect replayed the same catch-up window rather than resuming from where the client actually got to.

  3. Server always honored afterSequence. Two ways that fails:

    • Cursor far behind: every thread-aggregate event in the window costs a getThreadShellById query and emits a full thread shell. On the server that prompted this, a cursor ~6 days old meant ~11k events / ~17MB to rebuild a list a single snapshot expresses in a fraction of that. A client that cannot finish the replay never advances its cursor, so its next attempt is identical — stale indefinitely.
    • Cursor ahead of head: a rebuilt/restored event log (or a different server) replays nothing, and every live event is then dropped by the client's sequence <= snapshotSequence check. Stale forever, with no path to recovery.

The fix

  • Add retryExpectedFailureAfter: "250 millis" to the shell subscription, mirroring the thread path.
  • Resolve the resume cursor per subscribe. subscribe() now accepts a thunk, so cursor-carrying inputs are re-read on every attempt instead of captured once.
  • Server falls back to a full snapshot when the client's cursor is ahead of the projected sequence or more than SHELL_CATCH_UP_REPLAY_LIMIT (500) behind. The client applies snapshots wholesale, so this heals a wedged cache. If the server cannot read its own cursor it prefers the previous replay behavior.

Tests

Four new tests, each verified to fail without the corresponding fix (not vacuous):

  • shell-sync.test.ts — a stream that advances to sequence 9 then fails must resubscribe and ask for afterSequence: 9. Without the retry the recorded attempts are [5]; with the retry but a boot-time cursor they are [5, 5]; fixed, [5, 9].
  • server.test.ts — replay is still used for a recent cursor (control), and a snapshot is sent when the client is far behind or ahead of head.

packages/client-runtime 255/255 pass; apps/server typechecks clean.

Notes

  • One pre-existing apps/server test failure (preserves structured workspace rpc failures) and the scripts/test-deploy*.ts typecheck/lint errors are present on the base commit and unrelated to this change — confirmed by stashing.
  • Not addressed here: static assets ship with no Cache-Control/ETag/Last-Modified (apps/server/src/http.ts), so a phone can also pin a stale JS bundle after a redeploy. The thread-detail cache has the same class of stale-cursor-on-reconnect issue; this PR only makes subscribe() capable of fixing it. Worth follow-ups.
  • Existing wedged clients still need their site data cleared once; this prevents recurrence and heals clients whose cursor is unusable.

🤖 Generated with Claude Code

The sidebar thread list is a snapshot cached in IndexedDB that resumes
over a persisted event cursor. Two gaps let a client pin itself to a
stale list that survived reloads, because a reload re-reads the same
cache and re-attempts the same resume:
- The shell subscription passed no `retryExpectedFailureAfter`, so an
expected (non-transport) failure ended the stream for the life of the
page. The cached snapshot kept rendering and never updated again.
- The resume cursor was captured once at boot, so every reconnect
replayed the same catch-up window instead of resuming from where the
client actually got to.
- The server always honored `afterSequence`. A cursor far behind our
head asks for a replay costing far more than the snapshot it stands
in for (every thread-aggregate event is a projection lookup plus a
full thread shell on the socket), and a client that cannot afford
that replay never advances its cursor, so its next attempt is just as
expensive. A cursor *ahead* of our head — a rebuilt or restored event
log — replayed nothing while every live event was dropped by the
client's sequence check.
Adds the retry, resolves the resume cursor per subscribe, and falls
back to a snapshot when the client's cursor is unusable or too far
behind. `subscribe` now accepts a thunk so cursor-carrying inputs are
re-read on every attempt.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:L labels Jul 23, 2026
@DanielGGordon
DanielGGordon merged commit 6a4cfdd into mainJul 23, 2026
6 of 10 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:Lvouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@DanielGGordon
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(sync): stop the shell thread list from wedging on a stale cache - #40

Merged
DanielGGordon merged 1 commit into
mainfrom
t3code/mobile-chat-sync-fix
Jul 23, 2026
Merged

fix(sync): stop the shell thread list from wedging on a stale cache#40
DanielGGordon merged 1 commit into
mainfrom
t3code/mobile-chat-sync-fix

Conversation

@DanielGGordon

Copy link
Copy Markdown
Owner

The bug

On mobile, the sidebar thread list showed conversations days out of date, and refreshing the page did not help.

The list is not fetched on load. It is a snapshot cached in IndexedDB (t3code:connection-runtime, store shell) that the client renders immediately, skipping the HTTP snapshot entirely, then resumes over a WebSocket using a persisted event cursor. If that resume never completes, the stale cache is what the next load resumes from too — so the wedge is self-reinforcing and survives reloads.

Diagnosed against a live server holding 193 threads with a healthy event log (1→59,935, all projectors current). The data was fine; the client could not get to it.

Three gaps

  1. No retry.subscribeShell passed no retryExpectedFailureAfter, so an expected (non-transport) failure returned a drained stream and the subscription ended for the life of the page. The UI kept rendering the cached snapshot and silently never updated. The thread-detail path already got this right.

  2. Boot-time resume cursor.subscribeInput was computed once, outside the retry/reconnect loop. Every reconnect replayed the same catch-up window rather than resuming from where the client actually got to.

  3. Server always honored afterSequence. Two ways that fails:

    • Cursor far behind: every thread-aggregate event in the window costs a getThreadShellById query and emits a full thread shell. On the server that prompted this, a cursor ~6 days old meant ~11k events / ~17MB to rebuild a list a single snapshot expresses in a fraction of that. A client that cannot finish the replay never advances its cursor, so its next attempt is identical — stale indefinitely.
    • Cursor ahead of head: a rebuilt/restored event log (or a different server) replays nothing, and every live event is then dropped by the client's sequence <= snapshotSequence check. Stale forever, with no path to recovery.

The fix

  • Add retryExpectedFailureAfter: "250 millis" to the shell subscription, mirroring the thread path.
  • Resolve the resume cursor per subscribe. subscribe() now accepts a thunk, so cursor-carrying inputs are re-read on every attempt instead of captured once.
  • Server falls back to a full snapshot when the client's cursor is ahead of the projected sequence or more than SHELL_CATCH_UP_REPLAY_LIMIT (500) behind. The client applies snapshots wholesale, so this heals a wedged cache. If the server cannot read its own cursor it prefers the previous replay behavior.

Tests

Four new tests, each verified to fail without the corresponding fix (not vacuous):

  • shell-sync.test.ts — a stream that advances to sequence 9 then fails must resubscribe and ask for afterSequence: 9. Without the retry the recorded attempts are [5]; with the retry but a boot-time cursor they are [5, 5]; fixed, [5, 9].
  • server.test.ts — replay is still used for a recent cursor (control), and a snapshot is sent when the client is far behind or ahead of head.

packages/client-runtime 255/255 pass; apps/server typechecks clean.

Notes

  • One pre-existing apps/server test failure (preserves structured workspace rpc failures) and the scripts/test-deploy*.ts typecheck/lint errors are present on the base commit and unrelated to this change — confirmed by stashing.
  • Not addressed here: static assets ship with no Cache-Control/ETag/Last-Modified (apps/server/src/http.ts), so a phone can also pin a stale JS bundle after a redeploy. The thread-detail cache has the same class of stale-cursor-on-reconnect issue; this PR only makes subscribe() capable of fixing it. Worth follow-ups.
  • Existing wedged clients still need their site data cleared once; this prevents recurrence and heals clients whose cursor is unusable.

🤖 Generated with Claude Code

The sidebar thread list is a snapshot cached in IndexedDB that resumes
over a persisted event cursor. Two gaps let a client pin itself to a
stale list that survived reloads, because a reload re-reads the same
cache and re-attempts the same resume:
- The shell subscription passed no `retryExpectedFailureAfter`, so an
expected (non-transport) failure ended the stream for the life of the
page. The cached snapshot kept rendering and never updated again.
- The resume cursor was captured once at boot, so every reconnect
replayed the same catch-up window instead of resuming from where the
client actually got to.
- The server always honored `afterSequence`. A cursor far behind our
head asks for a replay costing far more than the snapshot it stands
in for (every thread-aggregate event is a projection lookup plus a
full thread shell on the socket), and a client that cannot afford
that replay never advances its cursor, so its next attempt is just as
expensive. A cursor *ahead* of our head — a rebuilt or restored event
log — replayed nothing while every live event was dropped by the
client's sequence check.
Adds the retry, resolves the resume cursor per subscribe, and falls
back to a snapshot when the client's cursor is unusable or too far
behind. `subscribe` now accepts a thunk so cursor-carrying inputs are
re-read on every attempt.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:L labels Jul 23, 2026
@DanielGGordon
DanielGGordon merged commit 6a4cfdd into mainJul 23, 2026
6 of 10 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:Lvouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@DanielGGordon
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

fix(sync): stop the shell thread list from wedging on a stale cache - #40

Merged
DanielGGordon merged 1 commit into
mainfrom
t3code/mobile-chat-sync-fix
Jul 23, 2026
Merged

fix(sync): stop the shell thread list from wedging on a stale cache#40
DanielGGordon merged 1 commit into
mainfrom
t3code/mobile-chat-sync-fix

Conversation

@DanielGGordon

Copy link
Copy Markdown
Owner

The bug

On mobile, the sidebar thread list showed conversations days out of date, and refreshing the page did not help.

The list is not fetched on load. It is a snapshot cached in IndexedDB (t3code:connection-runtime, store shell) that the client renders immediately, skipping the HTTP snapshot entirely, then resumes over a WebSocket using a persisted event cursor. If that resume never completes, the stale cache is what the next load resumes from too — so the wedge is self-reinforcing and survives reloads.

Diagnosed against a live server holding 193 threads with a healthy event log (1→59,935, all projectors current). The data was fine; the client could not get to it.

Three gaps

  1. No retry.subscribeShell passed no retryExpectedFailureAfter, so an expected (non-transport) failure returned a drained stream and the subscription ended for the life of the page. The UI kept rendering the cached snapshot and silently never updated. The thread-detail path already got this right.

  2. Boot-time resume cursor.subscribeInput was computed once, outside the retry/reconnect loop. Every reconnect replayed the same catch-up window rather than resuming from where the client actually got to.

  3. Server always honored afterSequence. Two ways that fails:

    • Cursor far behind: every thread-aggregate event in the window costs a getThreadShellById query and emits a full thread shell. On the server that prompted this, a cursor ~6 days old meant ~11k events / ~17MB to rebuild a list a single snapshot expresses in a fraction of that. A client that cannot finish the replay never advances its cursor, so its next attempt is identical — stale indefinitely.
    • Cursor ahead of head: a rebuilt/restored event log (or a different server) replays nothing, and every live event is then dropped by the client's sequence <= snapshotSequence check. Stale forever, with no path to recovery.

The fix

  • Add retryExpectedFailureAfter: "250 millis" to the shell subscription, mirroring the thread path.
  • Resolve the resume cursor per subscribe. subscribe() now accepts a thunk, so cursor-carrying inputs are re-read on every attempt instead of captured once.
  • Server falls back to a full snapshot when the client's cursor is ahead of the projected sequence or more than SHELL_CATCH_UP_REPLAY_LIMIT (500) behind. The client applies snapshots wholesale, so this heals a wedged cache. If the server cannot read its own cursor it prefers the previous replay behavior.

Tests

Four new tests, each verified to fail without the corresponding fix (not vacuous):

  • shell-sync.test.ts — a stream that advances to sequence 9 then fails must resubscribe and ask for afterSequence: 9. Without the retry the recorded attempts are [5]; with the retry but a boot-time cursor they are [5, 5]; fixed, [5, 9].
  • server.test.ts — replay is still used for a recent cursor (control), and a snapshot is sent when the client is far behind or ahead of head.

packages/client-runtime 255/255 pass; apps/server typechecks clean.

Notes

  • One pre-existing apps/server test failure (preserves structured workspace rpc failures) and the scripts/test-deploy*.ts typecheck/lint errors are present on the base commit and unrelated to this change — confirmed by stashing.
  • Not addressed here: static assets ship with no Cache-Control/ETag/Last-Modified (apps/server/src/http.ts), so a phone can also pin a stale JS bundle after a redeploy. The thread-detail cache has the same class of stale-cursor-on-reconnect issue; this PR only makes subscribe() capable of fixing it. Worth follow-ups.
  • Existing wedged clients still need their site data cleared once; this prevents recurrence and heals clients whose cursor is unusable.

🤖 Generated with Claude Code

The sidebar thread list is a snapshot cached in IndexedDB that resumes
over a persisted event cursor. Two gaps let a client pin itself to a
stale list that survived reloads, because a reload re-reads the same
cache and re-attempts the same resume:
- The shell subscription passed no `retryExpectedFailureAfter`, so an
expected (non-transport) failure ended the stream for the life of the
page. The cached snapshot kept rendering and never updated again.
- The resume cursor was captured once at boot, so every reconnect
replayed the same catch-up window instead of resuming from where the
client actually got to.
- The server always honored `afterSequence`. A cursor far behind our
head asks for a replay costing far more than the snapshot it stands
in for (every thread-aggregate event is a projection lookup plus a
full thread shell on the socket), and a client that cannot afford
that replay never advances its cursor, so its next attempt is just as
expensive. A cursor *ahead* of our head — a rebuilt or restored event
log — replayed nothing while every live event was dropped by the
client's sequence check.
Adds the retry, resolves the resume cursor per subscribe, and falls
back to a snapshot when the client's cursor is unusable or too far
behind. `subscribe` now accepts a thunk so cursor-carrying inputs are
re-read on every attempt.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:L labels Jul 23, 2026
@DanielGGordon
DanielGGordon merged commit 6a4cfdd into mainJul 23, 2026
6 of 10 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:Lvouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@DanielGGordon