fix(t3x): auto-resume never fired — 'thread-advanced' false cancellation on settled turns - #9

Merged
radroid merged 1 commit into
mainfrom
t3x/fix-thread-advanced-false-cancel
Jul 25, 2026
Merged

fix(t3x): auto-resume never fired — 'thread-advanced' false cancellation on settled turns#9
radroid merged 1 commit into
mainfrom
t3x/fix-thread-advanced-false-cancel

Conversation

@radroid

Copy link
Copy Markdown
Owner

Fixes#6.

What was broken

Every real-world auto-resume was cancelled at wake time with "Auto-resume cancelled: thread-advanced." — observed 3/3 on 2026-07-25 (threads 33c06a5a, 8f31df69, bb66acc2, all cancelled within 2s of their shared window reset), and the durable state showed firedAtMs: [] everywhere: the feature had never successfully fired.

Root cause

ProjectionSnapshotQuery.getSnapshot() reads threads from the SQLite projection, where latestTurn is joined on projection_threads.latest_turn_id — a column populated only while a turn is active. A usage limit normally lands mid-turn, so the guard baseline captures the running turn's id; by wake time the turn has settled and the snapshot reports latestTurn: null. The guard

if((thread.latestTurn?.turnId??null)!==baseline.latestTurnId)return"thread-advanced";

read null !== "<turnId>" as advancement and cancelled — then nothing ever re-armed, because an idle thread emits no further rejection events.

The fix

thread-advanced now requires positive evidence — a different, non-null turn id:

constcurrentTurnId=thread.latestTurn?.turnId??null;if(currentTurnId!==null&&currentTurnId!==baseline.latestTurnId)return"thread-advanced";

null at fire time means "no active turn", the expected state after a settled limit. Genuine user takeovers remain covered by user-took-over (checked first, keyed on the newest user message); active work remains covered by progressing. Known residual: an auto-started turn that both runs and settles during the wait with no user message is no longer detected via turn id — acceptable against a guard that previously cancelled 100% of legitimate resumes.

Testing

  • New unit case (guards.test.ts): fire-time latestTurn: null vs non-null baseline → no cancellation; plus the converse (baseline null, turn present → still cancels).
  • New integration case (Reactor.test.ts): replays the production incident — rejection arrives mid-running-turn, thread settles to latestTurn: null + session stopped during the wait, wake must dispatch exactly one resume turn and post no thread-advanced note.
  • Both verified to fail against the unfixed guard (stash-run-restore), then pass with the fix.
  • Full src/t3x/autoResume suite: 69/69 across 4 consecutive runs. pnpm typecheck output identical to main (two pre-existing non-error diagnostics in untouched files).
  • Diagnosis evidence (timeline, event-store replay, live DB columns) recorded in Auto-resume never fires in practice: 'thread-advanced' false cancellation when the limited turn settles during the wait #6.

🤖 Generated with Claude Code

…d turn settles
The projection populates projection_threads.latest_turn_id only while a turn
is active, so a usage limit that lands mid-turn captures the running turn's id
in the guard baseline — and by wake time (turn settled, session idle) the
snapshot reports latestTurn: null. cancelReason treated that null as evidence
the thread had advanced and cancelled the pending resume, which killed every
real-world resume (observed 3/3 on 2026-07-25; firedAtMs was empty across all
threads — the feature had never fired).
Advancement now requires positive evidence: a different, NON-NULL turn id.
Null means "no active turn" — the expected state after a settled limit. User
takeovers stay covered by the user-took-over guard (checked first) and active
work by the progressing guard.
Regression coverage: a guards.test.ts unit case (null-at-fire vs non-null
baseline must not cancel) plus a Reactor.test.ts integration case replaying
the production shape (schedule mid-running-turn, settle to latestTurn: null,
assert the resume fires with no thread-advanced note). Both verified to fail
against the previous guard.
Fixes#6
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 9a36a907-aa9f-4ca4-bf0e-1e8885209e18

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch t3x/fix-thread-advanced-false-cancel

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@radroid
radroid merged commit 237ca66 into mainJul 25, 2026
1 check passed
@radroid
radroid deleted the t3x/fix-thread-advanced-false-cancel branch July 25, 2026 12:56
radroid added a commit that referenced this pull request Aug 17, 2026
Closes#39.
Two defects compounded into a permanently stranded thread: a usage limit
armed a resume, the user typed "keep going" at the banner, that message
was itself rejected so it started nothing, and the wake tick then threw
the arm away as `user-took-over`. Net pending afterwards: zero. Measured
on the reporting install at 4 of 17 armed resumes (~24%) since #6.
1. `guards.ts` — drop the `user-took-over` branch. It is the same
negative-evidence mistake #6 fixed on the line below it: "a message
exists that wasn't there when we armed" is not evidence the human took
the wheel, and in practice it is the opposite signal. Everything the
branch reached for is still covered — `progressing` (actively driving),
`awaiting-input` (blocked on a prompt), `thread-advanced` (a different
live turn), and the per-thread switch for "stop entirely".
`baseline.newestUserMessageId` stays: it is part of the persisted
record, it is re-captured on every (re)schedule, and it is what makes a
stranded arm diagnosable from the state file.
2. `decide.ts` — narrow `already-pending`. A rejection naming a CONCRETE
reset time later than the pending one now supersedes it instead of
being dropped, so a `seven_day` limit landing on an armed `five_hour`
no longer leaves the arm firing into a window that is still shut.
Restricted to `windowOpensInFuture` on purpose: a ladder-derived time
is `nowMs + delay`, so it is later on every telemetry re-emit, and
letting those supersede would push the arm out forever and flood the
timeline. An earlier reset never supersedes either — the existing arm
is already the conservative choice.
A supersede re-writes the record with a fresh `captureBaseline`, so it
is also the re-arm path, and posts a distinct
`t3x.auto-resume.rescheduled` note rather than a second "scheduled".
The issue's third item — detecting the org monthly spend limit — is a
genuinely separate feature (no `account.rate-limits.updated` path reaches
it at all) and is deliberately not in scope here, per the triage.
Tests: 4 new behavioural cases, each verified to fail against the old
code — guards (a newer user message does not cancel; it still cancels
when that message is actually being worked on), decide (supersede,
earlier-window no-op, ladder-churn guard), and two reactor integration
tests covering the reported shape end to end. Reactor suite run 12x for
flake. `docs/t3x/loop/DESIGN.md` §6 gets a dated note: guard #9 still
stands, but the hazard it was defending against is gone.
radroid added a commit that referenced this pull request Aug 18, 2026
Closes#39.
Two defects compounded into a permanently stranded thread: a usage limit
armed a resume, the user typed "keep going" at the banner, that message
was itself rejected so it started nothing, and the wake tick then threw
the arm away as `user-took-over`. Net pending afterwards: zero. Measured
on the reporting install at 4 of 17 armed resumes (~24%) since #6.
1. `guards.ts` — drop the `user-took-over` branch. It is the same
negative-evidence mistake #6 fixed on the line below it: "a message
exists that wasn't there when we armed" is not evidence the human took
the wheel, and in practice it is the opposite signal. Everything the
branch reached for is still covered — `progressing` (actively driving),
`awaiting-input` (blocked on a prompt), `thread-advanced` (a different
live turn), and the per-thread switch for "stop entirely".
`baseline.newestUserMessageId` stays: it is part of the persisted
record, it is re-captured on every (re)schedule, and it is what makes a
stranded arm diagnosable from the state file.
2. `decide.ts` — narrow `already-pending`. A rejection naming a CONCRETE
reset time later than the pending one now supersedes it instead of
being dropped, so a `seven_day` limit landing on an armed `five_hour`
no longer leaves the arm firing into a window that is still shut.
Restricted to `windowOpensInFuture` on purpose: a ladder-derived time
is `nowMs + delay`, so it is later on every telemetry re-emit, and
letting those supersede would push the arm out forever and flood the
timeline. An earlier reset never supersedes either — the existing arm
is already the conservative choice.
A supersede re-writes the record with a fresh `captureBaseline`, so it
is also the re-arm path, and posts a distinct
`t3x.auto-resume.rescheduled` note rather than a second "scheduled".
The issue's third item — detecting the org monthly spend limit — is a
genuinely separate feature (no `account.rate-limits.updated` path reaches
it at all) and is deliberately not in scope here, per the triage.
Tests: 4 new behavioural cases, each verified to fail against the old
code — guards (a newer user message does not cancel; it still cancels
when that message is actually being worked on), decide (supersede,
earlier-window no-op, ladder-churn guard), and two reactor integration
tests covering the reported shape end to end. Reactor suite run 12x for
flake. `docs/t3x/loop/DESIGN.md` §6 gets a dated note: guard #9 still
stands, but the hazard it was defending against is gone.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Auto-resume never fires in practice: 'thread-advanced' false cancellation when the limited turn settles during the wait

1 participant

@radroid
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

fix(t3x): auto-resume never fired — 'thread-advanced' false cancellation on settled turns - #9

Merged
radroid merged 1 commit into
mainfrom
t3x/fix-thread-advanced-false-cancel
Jul 25, 2026
Merged

fix(t3x): auto-resume never fired — 'thread-advanced' false cancellation on settled turns#9
radroid merged 1 commit into
mainfrom
t3x/fix-thread-advanced-false-cancel

Conversation

@radroid

Copy link
Copy Markdown
Owner

Fixes#6.

What was broken

Every real-world auto-resume was cancelled at wake time with "Auto-resume cancelled: thread-advanced." — observed 3/3 on 2026-07-25 (threads 33c06a5a, 8f31df69, bb66acc2, all cancelled within 2s of their shared window reset), and the durable state showed firedAtMs: [] everywhere: the feature had never successfully fired.

Root cause

ProjectionSnapshotQuery.getSnapshot() reads threads from the SQLite projection, where latestTurn is joined on projection_threads.latest_turn_id — a column populated only while a turn is active. A usage limit normally lands mid-turn, so the guard baseline captures the running turn's id; by wake time the turn has settled and the snapshot reports latestTurn: null. The guard

if((thread.latestTurn?.turnId??null)!==baseline.latestTurnId)return"thread-advanced";

read null !== "<turnId>" as advancement and cancelled — then nothing ever re-armed, because an idle thread emits no further rejection events.

The fix

thread-advanced now requires positive evidence — a different, non-null turn id:

constcurrentTurnId=thread.latestTurn?.turnId??null;if(currentTurnId!==null&&currentTurnId!==baseline.latestTurnId)return"thread-advanced";

null at fire time means "no active turn", the expected state after a settled limit. Genuine user takeovers remain covered by user-took-over (checked first, keyed on the newest user message); active work remains covered by progressing. Known residual: an auto-started turn that both runs and settles during the wait with no user message is no longer detected via turn id — acceptable against a guard that previously cancelled 100% of legitimate resumes.

Testing

  • New unit case (guards.test.ts): fire-time latestTurn: null vs non-null baseline → no cancellation; plus the converse (baseline null, turn present → still cancels).
  • New integration case (Reactor.test.ts): replays the production incident — rejection arrives mid-running-turn, thread settles to latestTurn: null + session stopped during the wait, wake must dispatch exactly one resume turn and post no thread-advanced note.
  • Both verified to fail against the unfixed guard (stash-run-restore), then pass with the fix.
  • Full src/t3x/autoResume suite: 69/69 across 4 consecutive runs. pnpm typecheck output identical to main (two pre-existing non-error diagnostics in untouched files).
  • Diagnosis evidence (timeline, event-store replay, live DB columns) recorded in Auto-resume never fires in practice: 'thread-advanced' false cancellation when the limited turn settles during the wait #6.

🤖 Generated with Claude Code

…d turn settles
The projection populates projection_threads.latest_turn_id only while a turn
is active, so a usage limit that lands mid-turn captures the running turn's id
in the guard baseline — and by wake time (turn settled, session idle) the
snapshot reports latestTurn: null. cancelReason treated that null as evidence
the thread had advanced and cancelled the pending resume, which killed every
real-world resume (observed 3/3 on 2026-07-25; firedAtMs was empty across all
threads — the feature had never fired).
Advancement now requires positive evidence: a different, NON-NULL turn id.
Null means "no active turn" — the expected state after a settled limit. User
takeovers stay covered by the user-took-over guard (checked first) and active
work by the progressing guard.
Regression coverage: a guards.test.ts unit case (null-at-fire vs non-null
baseline must not cancel) plus a Reactor.test.ts integration case replaying
the production shape (schedule mid-running-turn, settle to latestTurn: null,
assert the resume fires with no thread-advanced note). Both verified to fail
against the previous guard.
Fixes#6
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 9a36a907-aa9f-4ca4-bf0e-1e8885209e18

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch t3x/fix-thread-advanced-false-cancel

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@radroid
radroid merged commit 237ca66 into mainJul 25, 2026
1 check passed
@radroid
radroid deleted the t3x/fix-thread-advanced-false-cancel branch July 25, 2026 12:56
radroid added a commit that referenced this pull request Aug 17, 2026
Closes#39.
Two defects compounded into a permanently stranded thread: a usage limit
armed a resume, the user typed "keep going" at the banner, that message
was itself rejected so it started nothing, and the wake tick then threw
the arm away as `user-took-over`. Net pending afterwards: zero. Measured
on the reporting install at 4 of 17 armed resumes (~24%) since #6.
1. `guards.ts` — drop the `user-took-over` branch. It is the same
negative-evidence mistake #6 fixed on the line below it: "a message
exists that wasn't there when we armed" is not evidence the human took
the wheel, and in practice it is the opposite signal. Everything the
branch reached for is still covered — `progressing` (actively driving),
`awaiting-input` (blocked on a prompt), `thread-advanced` (a different
live turn), and the per-thread switch for "stop entirely".
`baseline.newestUserMessageId` stays: it is part of the persisted
record, it is re-captured on every (re)schedule, and it is what makes a
stranded arm diagnosable from the state file.
2. `decide.ts` — narrow `already-pending`. A rejection naming a CONCRETE
reset time later than the pending one now supersedes it instead of
being dropped, so a `seven_day` limit landing on an armed `five_hour`
no longer leaves the arm firing into a window that is still shut.
Restricted to `windowOpensInFuture` on purpose: a ladder-derived time
is `nowMs + delay`, so it is later on every telemetry re-emit, and
letting those supersede would push the arm out forever and flood the
timeline. An earlier reset never supersedes either — the existing arm
is already the conservative choice.
A supersede re-writes the record with a fresh `captureBaseline`, so it
is also the re-arm path, and posts a distinct
`t3x.auto-resume.rescheduled` note rather than a second "scheduled".
The issue's third item — detecting the org monthly spend limit — is a
genuinely separate feature (no `account.rate-limits.updated` path reaches
it at all) and is deliberately not in scope here, per the triage.
Tests: 4 new behavioural cases, each verified to fail against the old
code — guards (a newer user message does not cancel; it still cancels
when that message is actually being worked on), decide (supersede,
earlier-window no-op, ladder-churn guard), and two reactor integration
tests covering the reported shape end to end. Reactor suite run 12x for
flake. `docs/t3x/loop/DESIGN.md` §6 gets a dated note: guard #9 still
stands, but the hazard it was defending against is gone.
radroid added a commit that referenced this pull request Aug 18, 2026
Closes#39.
Two defects compounded into a permanently stranded thread: a usage limit
armed a resume, the user typed "keep going" at the banner, that message
was itself rejected so it started nothing, and the wake tick then threw
the arm away as `user-took-over`. Net pending afterwards: zero. Measured
on the reporting install at 4 of 17 armed resumes (~24%) since #6.
1. `guards.ts` — drop the `user-took-over` branch. It is the same
negative-evidence mistake #6 fixed on the line below it: "a message
exists that wasn't there when we armed" is not evidence the human took
the wheel, and in practice it is the opposite signal. Everything the
branch reached for is still covered — `progressing` (actively driving),
`awaiting-input` (blocked on a prompt), `thread-advanced` (a different
live turn), and the per-thread switch for "stop entirely".
`baseline.newestUserMessageId` stays: it is part of the persisted
record, it is re-captured on every (re)schedule, and it is what makes a
stranded arm diagnosable from the state file.
2. `decide.ts` — narrow `already-pending`. A rejection naming a CONCRETE
reset time later than the pending one now supersedes it instead of
being dropped, so a `seven_day` limit landing on an armed `five_hour`
no longer leaves the arm firing into a window that is still shut.
Restricted to `windowOpensInFuture` on purpose: a ladder-derived time
is `nowMs + delay`, so it is later on every telemetry re-emit, and
letting those supersede would push the arm out forever and flood the
timeline. An earlier reset never supersedes either — the existing arm
is already the conservative choice.
A supersede re-writes the record with a fresh `captureBaseline`, so it
is also the re-arm path, and posts a distinct
`t3x.auto-resume.rescheduled` note rather than a second "scheduled".
The issue's third item — detecting the org monthly spend limit — is a
genuinely separate feature (no `account.rate-limits.updated` path reaches
it at all) and is deliberately not in scope here, per the triage.
Tests: 4 new behavioural cases, each verified to fail against the old
code — guards (a newer user message does not cancel; it still cancels
when that message is actually being worked on), decide (supersede,
earlier-window no-op, ladder-churn guard), and two reactor integration
tests covering the reported shape end to end. Reactor suite run 12x for
flake. `docs/t3x/loop/DESIGN.md` §6 gets a dated note: guard #9 still
stands, but the hazard it was defending against is gone.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Auto-resume never fires in practice: 'thread-advanced' false cancellation when the limited turn settles during the wait

1 participant

@radroid
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(t3x): auto-resume never fired — 'thread-advanced' false cancellation on settled turns - #9

Merged
radroid merged 1 commit into
mainfrom
t3x/fix-thread-advanced-false-cancel
Jul 25, 2026
Merged

fix(t3x): auto-resume never fired — 'thread-advanced' false cancellation on settled turns#9
radroid merged 1 commit into
mainfrom
t3x/fix-thread-advanced-false-cancel

Conversation

@radroid

Copy link
Copy Markdown
Owner

Fixes#6.

What was broken

Every real-world auto-resume was cancelled at wake time with "Auto-resume cancelled: thread-advanced." — observed 3/3 on 2026-07-25 (threads 33c06a5a, 8f31df69, bb66acc2, all cancelled within 2s of their shared window reset), and the durable state showed firedAtMs: [] everywhere: the feature had never successfully fired.

Root cause

ProjectionSnapshotQuery.getSnapshot() reads threads from the SQLite projection, where latestTurn is joined on projection_threads.latest_turn_id — a column populated only while a turn is active. A usage limit normally lands mid-turn, so the guard baseline captures the running turn's id; by wake time the turn has settled and the snapshot reports latestTurn: null. The guard

if((thread.latestTurn?.turnId??null)!==baseline.latestTurnId)return"thread-advanced";

read null !== "<turnId>" as advancement and cancelled — then nothing ever re-armed, because an idle thread emits no further rejection events.

The fix

thread-advanced now requires positive evidence — a different, non-null turn id:

constcurrentTurnId=thread.latestTurn?.turnId??null;if(currentTurnId!==null&&currentTurnId!==baseline.latestTurnId)return"thread-advanced";

null at fire time means "no active turn", the expected state after a settled limit. Genuine user takeovers remain covered by user-took-over (checked first, keyed on the newest user message); active work remains covered by progressing. Known residual: an auto-started turn that both runs and settles during the wait with no user message is no longer detected via turn id — acceptable against a guard that previously cancelled 100% of legitimate resumes.

Testing

  • New unit case (guards.test.ts): fire-time latestTurn: null vs non-null baseline → no cancellation; plus the converse (baseline null, turn present → still cancels).
  • New integration case (Reactor.test.ts): replays the production incident — rejection arrives mid-running-turn, thread settles to latestTurn: null + session stopped during the wait, wake must dispatch exactly one resume turn and post no thread-advanced note.
  • Both verified to fail against the unfixed guard (stash-run-restore), then pass with the fix.
  • Full src/t3x/autoResume suite: 69/69 across 4 consecutive runs. pnpm typecheck output identical to main (two pre-existing non-error diagnostics in untouched files).
  • Diagnosis evidence (timeline, event-store replay, live DB columns) recorded in Auto-resume never fires in practice: 'thread-advanced' false cancellation when the limited turn settles during the wait #6.

🤖 Generated with Claude Code

…d turn settles
The projection populates projection_threads.latest_turn_id only while a turn
is active, so a usage limit that lands mid-turn captures the running turn's id
in the guard baseline — and by wake time (turn settled, session idle) the
snapshot reports latestTurn: null. cancelReason treated that null as evidence
the thread had advanced and cancelled the pending resume, which killed every
real-world resume (observed 3/3 on 2026-07-25; firedAtMs was empty across all
threads — the feature had never fired).
Advancement now requires positive evidence: a different, NON-NULL turn id.
Null means "no active turn" — the expected state after a settled limit. User
takeovers stay covered by the user-took-over guard (checked first) and active
work by the progressing guard.
Regression coverage: a guards.test.ts unit case (null-at-fire vs non-null
baseline must not cancel) plus a Reactor.test.ts integration case replaying
the production shape (schedule mid-running-turn, settle to latestTurn: null,
assert the resume fires with no thread-advanced note). Both verified to fail
against the previous guard.
Fixes#6
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 9a36a907-aa9f-4ca4-bf0e-1e8885209e18

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch t3x/fix-thread-advanced-false-cancel

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@radroid
radroid merged commit 237ca66 into mainJul 25, 2026
1 check passed
@radroid
radroid deleted the t3x/fix-thread-advanced-false-cancel branch July 25, 2026 12:56
radroid added a commit that referenced this pull request Aug 17, 2026
Closes#39.
Two defects compounded into a permanently stranded thread: a usage limit
armed a resume, the user typed "keep going" at the banner, that message
was itself rejected so it started nothing, and the wake tick then threw
the arm away as `user-took-over`. Net pending afterwards: zero. Measured
on the reporting install at 4 of 17 armed resumes (~24%) since #6.
1. `guards.ts` — drop the `user-took-over` branch. It is the same
negative-evidence mistake #6 fixed on the line below it: "a message
exists that wasn't there when we armed" is not evidence the human took
the wheel, and in practice it is the opposite signal. Everything the
branch reached for is still covered — `progressing` (actively driving),
`awaiting-input` (blocked on a prompt), `thread-advanced` (a different
live turn), and the per-thread switch for "stop entirely".
`baseline.newestUserMessageId` stays: it is part of the persisted
record, it is re-captured on every (re)schedule, and it is what makes a
stranded arm diagnosable from the state file.
2. `decide.ts` — narrow `already-pending`. A rejection naming a CONCRETE
reset time later than the pending one now supersedes it instead of
being dropped, so a `seven_day` limit landing on an armed `five_hour`
no longer leaves the arm firing into a window that is still shut.
Restricted to `windowOpensInFuture` on purpose: a ladder-derived time
is `nowMs + delay`, so it is later on every telemetry re-emit, and
letting those supersede would push the arm out forever and flood the
timeline. An earlier reset never supersedes either — the existing arm
is already the conservative choice.
A supersede re-writes the record with a fresh `captureBaseline`, so it
is also the re-arm path, and posts a distinct
`t3x.auto-resume.rescheduled` note rather than a second "scheduled".
The issue's third item — detecting the org monthly spend limit — is a
genuinely separate feature (no `account.rate-limits.updated` path reaches
it at all) and is deliberately not in scope here, per the triage.
Tests: 4 new behavioural cases, each verified to fail against the old
code — guards (a newer user message does not cancel; it still cancels
when that message is actually being worked on), decide (supersede,
earlier-window no-op, ladder-churn guard), and two reactor integration
tests covering the reported shape end to end. Reactor suite run 12x for
flake. `docs/t3x/loop/DESIGN.md` §6 gets a dated note: guard #9 still
stands, but the hazard it was defending against is gone.
radroid added a commit that referenced this pull request Aug 18, 2026
Closes#39.
Two defects compounded into a permanently stranded thread: a usage limit
armed a resume, the user typed "keep going" at the banner, that message
was itself rejected so it started nothing, and the wake tick then threw
the arm away as `user-took-over`. Net pending afterwards: zero. Measured
on the reporting install at 4 of 17 armed resumes (~24%) since #6.
1. `guards.ts` — drop the `user-took-over` branch. It is the same
negative-evidence mistake #6 fixed on the line below it: "a message
exists that wasn't there when we armed" is not evidence the human took
the wheel, and in practice it is the opposite signal. Everything the
branch reached for is still covered — `progressing` (actively driving),
`awaiting-input` (blocked on a prompt), `thread-advanced` (a different
live turn), and the per-thread switch for "stop entirely".
`baseline.newestUserMessageId` stays: it is part of the persisted
record, it is re-captured on every (re)schedule, and it is what makes a
stranded arm diagnosable from the state file.
2. `decide.ts` — narrow `already-pending`. A rejection naming a CONCRETE
reset time later than the pending one now supersedes it instead of
being dropped, so a `seven_day` limit landing on an armed `five_hour`
no longer leaves the arm firing into a window that is still shut.
Restricted to `windowOpensInFuture` on purpose: a ladder-derived time
is `nowMs + delay`, so it is later on every telemetry re-emit, and
letting those supersede would push the arm out forever and flood the
timeline. An earlier reset never supersedes either — the existing arm
is already the conservative choice.
A supersede re-writes the record with a fresh `captureBaseline`, so it
is also the re-arm path, and posts a distinct
`t3x.auto-resume.rescheduled` note rather than a second "scheduled".
The issue's third item — detecting the org monthly spend limit — is a
genuinely separate feature (no `account.rate-limits.updated` path reaches
it at all) and is deliberately not in scope here, per the triage.
Tests: 4 new behavioural cases, each verified to fail against the old
code — guards (a newer user message does not cancel; it still cancels
when that message is actually being worked on), decide (supersede,
earlier-window no-op, ladder-churn guard), and two reactor integration
tests covering the reported shape end to end. Reactor suite run 12x for
flake. `docs/t3x/loop/DESIGN.md` §6 gets a dated note: guard #9 still
stands, but the hazard it was defending against is gone.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Auto-resume never fires in practice: 'thread-advanced' false cancellation when the limited turn settles during the wait

1 participant

@radroid
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(t3x): auto-resume never fired — 'thread-advanced' false cancellation on settled turns - #9

Merged
radroid merged 1 commit into
mainfrom
t3x/fix-thread-advanced-false-cancel
Jul 25, 2026
Merged

fix(t3x): auto-resume never fired — 'thread-advanced' false cancellation on settled turns#9
radroid merged 1 commit into
mainfrom
t3x/fix-thread-advanced-false-cancel

Conversation

@radroid

Copy link
Copy Markdown
Owner

Fixes#6.

What was broken

Every real-world auto-resume was cancelled at wake time with "Auto-resume cancelled: thread-advanced." — observed 3/3 on 2026-07-25 (threads 33c06a5a, 8f31df69, bb66acc2, all cancelled within 2s of their shared window reset), and the durable state showed firedAtMs: [] everywhere: the feature had never successfully fired.

Root cause

ProjectionSnapshotQuery.getSnapshot() reads threads from the SQLite projection, where latestTurn is joined on projection_threads.latest_turn_id — a column populated only while a turn is active. A usage limit normally lands mid-turn, so the guard baseline captures the running turn's id; by wake time the turn has settled and the snapshot reports latestTurn: null. The guard

if((thread.latestTurn?.turnId??null)!==baseline.latestTurnId)return"thread-advanced";

read null !== "<turnId>" as advancement and cancelled — then nothing ever re-armed, because an idle thread emits no further rejection events.

The fix

thread-advanced now requires positive evidence — a different, non-null turn id:

constcurrentTurnId=thread.latestTurn?.turnId??null;if(currentTurnId!==null&&currentTurnId!==baseline.latestTurnId)return"thread-advanced";

null at fire time means "no active turn", the expected state after a settled limit. Genuine user takeovers remain covered by user-took-over (checked first, keyed on the newest user message); active work remains covered by progressing. Known residual: an auto-started turn that both runs and settles during the wait with no user message is no longer detected via turn id — acceptable against a guard that previously cancelled 100% of legitimate resumes.

Testing

  • New unit case (guards.test.ts): fire-time latestTurn: null vs non-null baseline → no cancellation; plus the converse (baseline null, turn present → still cancels).
  • New integration case (Reactor.test.ts): replays the production incident — rejection arrives mid-running-turn, thread settles to latestTurn: null + session stopped during the wait, wake must dispatch exactly one resume turn and post no thread-advanced note.
  • Both verified to fail against the unfixed guard (stash-run-restore), then pass with the fix.
  • Full src/t3x/autoResume suite: 69/69 across 4 consecutive runs. pnpm typecheck output identical to main (two pre-existing non-error diagnostics in untouched files).
  • Diagnosis evidence (timeline, event-store replay, live DB columns) recorded in Auto-resume never fires in practice: 'thread-advanced' false cancellation when the limited turn settles during the wait #6.

🤖 Generated with Claude Code

…d turn settles
The projection populates projection_threads.latest_turn_id only while a turn
is active, so a usage limit that lands mid-turn captures the running turn's id
in the guard baseline — and by wake time (turn settled, session idle) the
snapshot reports latestTurn: null. cancelReason treated that null as evidence
the thread had advanced and cancelled the pending resume, which killed every
real-world resume (observed 3/3 on 2026-07-25; firedAtMs was empty across all
threads — the feature had never fired).
Advancement now requires positive evidence: a different, NON-NULL turn id.
Null means "no active turn" — the expected state after a settled limit. User
takeovers stay covered by the user-took-over guard (checked first) and active
work by the progressing guard.
Regression coverage: a guards.test.ts unit case (null-at-fire vs non-null
baseline must not cancel) plus a Reactor.test.ts integration case replaying
the production shape (schedule mid-running-turn, settle to latestTurn: null,
assert the resume fires with no thread-advanced note). Both verified to fail
against the previous guard.
Fixes#6
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 9a36a907-aa9f-4ca4-bf0e-1e8885209e18

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch t3x/fix-thread-advanced-false-cancel

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@radroid
radroid merged commit 237ca66 into mainJul 25, 2026
1 check passed
@radroid
radroid deleted the t3x/fix-thread-advanced-false-cancel branch July 25, 2026 12:56
radroid added a commit that referenced this pull request Aug 17, 2026
Closes#39.
Two defects compounded into a permanently stranded thread: a usage limit
armed a resume, the user typed "keep going" at the banner, that message
was itself rejected so it started nothing, and the wake tick then threw
the arm away as `user-took-over`. Net pending afterwards: zero. Measured
on the reporting install at 4 of 17 armed resumes (~24%) since #6.
1. `guards.ts` — drop the `user-took-over` branch. It is the same
negative-evidence mistake #6 fixed on the line below it: "a message
exists that wasn't there when we armed" is not evidence the human took
the wheel, and in practice it is the opposite signal. Everything the
branch reached for is still covered — `progressing` (actively driving),
`awaiting-input` (blocked on a prompt), `thread-advanced` (a different
live turn), and the per-thread switch for "stop entirely".
`baseline.newestUserMessageId` stays: it is part of the persisted
record, it is re-captured on every (re)schedule, and it is what makes a
stranded arm diagnosable from the state file.
2. `decide.ts` — narrow `already-pending`. A rejection naming a CONCRETE
reset time later than the pending one now supersedes it instead of
being dropped, so a `seven_day` limit landing on an armed `five_hour`
no longer leaves the arm firing into a window that is still shut.
Restricted to `windowOpensInFuture` on purpose: a ladder-derived time
is `nowMs + delay`, so it is later on every telemetry re-emit, and
letting those supersede would push the arm out forever and flood the
timeline. An earlier reset never supersedes either — the existing arm
is already the conservative choice.
A supersede re-writes the record with a fresh `captureBaseline`, so it
is also the re-arm path, and posts a distinct
`t3x.auto-resume.rescheduled` note rather than a second "scheduled".
The issue's third item — detecting the org monthly spend limit — is a
genuinely separate feature (no `account.rate-limits.updated` path reaches
it at all) and is deliberately not in scope here, per the triage.
Tests: 4 new behavioural cases, each verified to fail against the old
code — guards (a newer user message does not cancel; it still cancels
when that message is actually being worked on), decide (supersede,
earlier-window no-op, ladder-churn guard), and two reactor integration
tests covering the reported shape end to end. Reactor suite run 12x for
flake. `docs/t3x/loop/DESIGN.md` §6 gets a dated note: guard #9 still
stands, but the hazard it was defending against is gone.
radroid added a commit that referenced this pull request Aug 18, 2026
Closes#39.
Two defects compounded into a permanently stranded thread: a usage limit
armed a resume, the user typed "keep going" at the banner, that message
was itself rejected so it started nothing, and the wake tick then threw
the arm away as `user-took-over`. Net pending afterwards: zero. Measured
on the reporting install at 4 of 17 armed resumes (~24%) since #6.
1. `guards.ts` — drop the `user-took-over` branch. It is the same
negative-evidence mistake #6 fixed on the line below it: "a message
exists that wasn't there when we armed" is not evidence the human took
the wheel, and in practice it is the opposite signal. Everything the
branch reached for is still covered — `progressing` (actively driving),
`awaiting-input` (blocked on a prompt), `thread-advanced` (a different
live turn), and the per-thread switch for "stop entirely".
`baseline.newestUserMessageId` stays: it is part of the persisted
record, it is re-captured on every (re)schedule, and it is what makes a
stranded arm diagnosable from the state file.
2. `decide.ts` — narrow `already-pending`. A rejection naming a CONCRETE
reset time later than the pending one now supersedes it instead of
being dropped, so a `seven_day` limit landing on an armed `five_hour`
no longer leaves the arm firing into a window that is still shut.
Restricted to `windowOpensInFuture` on purpose: a ladder-derived time
is `nowMs + delay`, so it is later on every telemetry re-emit, and
letting those supersede would push the arm out forever and flood the
timeline. An earlier reset never supersedes either — the existing arm
is already the conservative choice.
A supersede re-writes the record with a fresh `captureBaseline`, so it
is also the re-arm path, and posts a distinct
`t3x.auto-resume.rescheduled` note rather than a second "scheduled".
The issue's third item — detecting the org monthly spend limit — is a
genuinely separate feature (no `account.rate-limits.updated` path reaches
it at all) and is deliberately not in scope here, per the triage.
Tests: 4 new behavioural cases, each verified to fail against the old
code — guards (a newer user message does not cancel; it still cancels
when that message is actually being worked on), decide (supersede,
earlier-window no-op, ladder-churn guard), and two reactor integration
tests covering the reported shape end to end. Reactor suite run 12x for
flake. `docs/t3x/loop/DESIGN.md` §6 gets a dated note: guard #9 still
stands, but the hazard it was defending against is gone.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Auto-resume never fires in practice: 'thread-advanced' false cancellation when the limited turn settles during the wait

1 participant

@radroid
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

fix(t3x): auto-resume never fired — 'thread-advanced' false cancellation on settled turns - #9

Merged
radroid merged 1 commit into
mainfrom
t3x/fix-thread-advanced-false-cancel
Jul 25, 2026
Merged

fix(t3x): auto-resume never fired — 'thread-advanced' false cancellation on settled turns#9
radroid merged 1 commit into
mainfrom
t3x/fix-thread-advanced-false-cancel

Conversation

@radroid

Copy link
Copy Markdown
Owner

Fixes#6.

What was broken

Every real-world auto-resume was cancelled at wake time with "Auto-resume cancelled: thread-advanced." — observed 3/3 on 2026-07-25 (threads 33c06a5a, 8f31df69, bb66acc2, all cancelled within 2s of their shared window reset), and the durable state showed firedAtMs: [] everywhere: the feature had never successfully fired.

Root cause

ProjectionSnapshotQuery.getSnapshot() reads threads from the SQLite projection, where latestTurn is joined on projection_threads.latest_turn_id — a column populated only while a turn is active. A usage limit normally lands mid-turn, so the guard baseline captures the running turn's id; by wake time the turn has settled and the snapshot reports latestTurn: null. The guard

if((thread.latestTurn?.turnId??null)!==baseline.latestTurnId)return"thread-advanced";

read null !== "<turnId>" as advancement and cancelled — then nothing ever re-armed, because an idle thread emits no further rejection events.

The fix

thread-advanced now requires positive evidence — a different, non-null turn id:

constcurrentTurnId=thread.latestTurn?.turnId??null;if(currentTurnId!==null&&currentTurnId!==baseline.latestTurnId)return"thread-advanced";

null at fire time means "no active turn", the expected state after a settled limit. Genuine user takeovers remain covered by user-took-over (checked first, keyed on the newest user message); active work remains covered by progressing. Known residual: an auto-started turn that both runs and settles during the wait with no user message is no longer detected via turn id — acceptable against a guard that previously cancelled 100% of legitimate resumes.

Testing

  • New unit case (guards.test.ts): fire-time latestTurn: null vs non-null baseline → no cancellation; plus the converse (baseline null, turn present → still cancels).
  • New integration case (Reactor.test.ts): replays the production incident — rejection arrives mid-running-turn, thread settles to latestTurn: null + session stopped during the wait, wake must dispatch exactly one resume turn and post no thread-advanced note.
  • Both verified to fail against the unfixed guard (stash-run-restore), then pass with the fix.
  • Full src/t3x/autoResume suite: 69/69 across 4 consecutive runs. pnpm typecheck output identical to main (two pre-existing non-error diagnostics in untouched files).
  • Diagnosis evidence (timeline, event-store replay, live DB columns) recorded in Auto-resume never fires in practice: 'thread-advanced' false cancellation when the limited turn settles during the wait #6.

🤖 Generated with Claude Code

…d turn settles
The projection populates projection_threads.latest_turn_id only while a turn
is active, so a usage limit that lands mid-turn captures the running turn's id
in the guard baseline — and by wake time (turn settled, session idle) the
snapshot reports latestTurn: null. cancelReason treated that null as evidence
the thread had advanced and cancelled the pending resume, which killed every
real-world resume (observed 3/3 on 2026-07-25; firedAtMs was empty across all
threads — the feature had never fired).
Advancement now requires positive evidence: a different, NON-NULL turn id.
Null means "no active turn" — the expected state after a settled limit. User
takeovers stay covered by the user-took-over guard (checked first) and active
work by the progressing guard.
Regression coverage: a guards.test.ts unit case (null-at-fire vs non-null
baseline must not cancel) plus a Reactor.test.ts integration case replaying
the production shape (schedule mid-running-turn, settle to latestTurn: null,
assert the resume fires with no thread-advanced note). Both verified to fail
against the previous guard.
Fixes#6
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 9a36a907-aa9f-4ca4-bf0e-1e8885209e18

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch t3x/fix-thread-advanced-false-cancel

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@radroid
radroid merged commit 237ca66 into mainJul 25, 2026
1 check passed
@radroid
radroid deleted the t3x/fix-thread-advanced-false-cancel branch July 25, 2026 12:56
radroid added a commit that referenced this pull request Aug 17, 2026
Closes#39.
Two defects compounded into a permanently stranded thread: a usage limit
armed a resume, the user typed "keep going" at the banner, that message
was itself rejected so it started nothing, and the wake tick then threw
the arm away as `user-took-over`. Net pending afterwards: zero. Measured
on the reporting install at 4 of 17 armed resumes (~24%) since #6.
1. `guards.ts` — drop the `user-took-over` branch. It is the same
negative-evidence mistake #6 fixed on the line below it: "a message
exists that wasn't there when we armed" is not evidence the human took
the wheel, and in practice it is the opposite signal. Everything the
branch reached for is still covered — `progressing` (actively driving),
`awaiting-input` (blocked on a prompt), `thread-advanced` (a different
live turn), and the per-thread switch for "stop entirely".
`baseline.newestUserMessageId` stays: it is part of the persisted
record, it is re-captured on every (re)schedule, and it is what makes a
stranded arm diagnosable from the state file.
2. `decide.ts` — narrow `already-pending`. A rejection naming a CONCRETE
reset time later than the pending one now supersedes it instead of
being dropped, so a `seven_day` limit landing on an armed `five_hour`
no longer leaves the arm firing into a window that is still shut.
Restricted to `windowOpensInFuture` on purpose: a ladder-derived time
is `nowMs + delay`, so it is later on every telemetry re-emit, and
letting those supersede would push the arm out forever and flood the
timeline. An earlier reset never supersedes either — the existing arm
is already the conservative choice.
A supersede re-writes the record with a fresh `captureBaseline`, so it
is also the re-arm path, and posts a distinct
`t3x.auto-resume.rescheduled` note rather than a second "scheduled".
The issue's third item — detecting the org monthly spend limit — is a
genuinely separate feature (no `account.rate-limits.updated` path reaches
it at all) and is deliberately not in scope here, per the triage.
Tests: 4 new behavioural cases, each verified to fail against the old
code — guards (a newer user message does not cancel; it still cancels
when that message is actually being worked on), decide (supersede,
earlier-window no-op, ladder-churn guard), and two reactor integration
tests covering the reported shape end to end. Reactor suite run 12x for
flake. `docs/t3x/loop/DESIGN.md` §6 gets a dated note: guard #9 still
stands, but the hazard it was defending against is gone.
radroid added a commit that referenced this pull request Aug 18, 2026
Closes#39.
Two defects compounded into a permanently stranded thread: a usage limit
armed a resume, the user typed "keep going" at the banner, that message
was itself rejected so it started nothing, and the wake tick then threw
the arm away as `user-took-over`. Net pending afterwards: zero. Measured
on the reporting install at 4 of 17 armed resumes (~24%) since #6.
1. `guards.ts` — drop the `user-took-over` branch. It is the same
negative-evidence mistake #6 fixed on the line below it: "a message
exists that wasn't there when we armed" is not evidence the human took
the wheel, and in practice it is the opposite signal. Everything the
branch reached for is still covered — `progressing` (actively driving),
`awaiting-input` (blocked on a prompt), `thread-advanced` (a different
live turn), and the per-thread switch for "stop entirely".
`baseline.newestUserMessageId` stays: it is part of the persisted
record, it is re-captured on every (re)schedule, and it is what makes a
stranded arm diagnosable from the state file.
2. `decide.ts` — narrow `already-pending`. A rejection naming a CONCRETE
reset time later than the pending one now supersedes it instead of
being dropped, so a `seven_day` limit landing on an armed `five_hour`
no longer leaves the arm firing into a window that is still shut.
Restricted to `windowOpensInFuture` on purpose: a ladder-derived time
is `nowMs + delay`, so it is later on every telemetry re-emit, and
letting those supersede would push the arm out forever and flood the
timeline. An earlier reset never supersedes either — the existing arm
is already the conservative choice.
A supersede re-writes the record with a fresh `captureBaseline`, so it
is also the re-arm path, and posts a distinct
`t3x.auto-resume.rescheduled` note rather than a second "scheduled".
The issue's third item — detecting the org monthly spend limit — is a
genuinely separate feature (no `account.rate-limits.updated` path reaches
it at all) and is deliberately not in scope here, per the triage.
Tests: 4 new behavioural cases, each verified to fail against the old
code — guards (a newer user message does not cancel; it still cancels
when that message is actually being worked on), decide (supersede,
earlier-window no-op, ladder-churn guard), and two reactor integration
tests covering the reported shape end to end. Reactor suite run 12x for
flake. `docs/t3x/loop/DESIGN.md` §6 gets a dated note: guard #9 still
stands, but the hazard it was defending against is gone.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Auto-resume never fires in practice: 'thread-advanced' false cancellation when the limited turn settles during the wait

1 participant

@radroid
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(t3x): auto-resume never fired — 'thread-advanced' false cancellation on settled turns - #9

Merged
radroid merged 1 commit into
mainfrom
t3x/fix-thread-advanced-false-cancel
Jul 25, 2026
Merged

fix(t3x): auto-resume never fired — 'thread-advanced' false cancellation on settled turns#9
radroid merged 1 commit into
mainfrom
t3x/fix-thread-advanced-false-cancel

Conversation

@radroid

Copy link
Copy Markdown
Owner

Fixes#6.

What was broken

Every real-world auto-resume was cancelled at wake time with "Auto-resume cancelled: thread-advanced." — observed 3/3 on 2026-07-25 (threads 33c06a5a, 8f31df69, bb66acc2, all cancelled within 2s of their shared window reset), and the durable state showed firedAtMs: [] everywhere: the feature had never successfully fired.

Root cause

ProjectionSnapshotQuery.getSnapshot() reads threads from the SQLite projection, where latestTurn is joined on projection_threads.latest_turn_id — a column populated only while a turn is active. A usage limit normally lands mid-turn, so the guard baseline captures the running turn's id; by wake time the turn has settled and the snapshot reports latestTurn: null. The guard

if((thread.latestTurn?.turnId??null)!==baseline.latestTurnId)return"thread-advanced";

read null !== "<turnId>" as advancement and cancelled — then nothing ever re-armed, because an idle thread emits no further rejection events.

The fix

thread-advanced now requires positive evidence — a different, non-null turn id:

constcurrentTurnId=thread.latestTurn?.turnId??null;if(currentTurnId!==null&&currentTurnId!==baseline.latestTurnId)return"thread-advanced";

null at fire time means "no active turn", the expected state after a settled limit. Genuine user takeovers remain covered by user-took-over (checked first, keyed on the newest user message); active work remains covered by progressing. Known residual: an auto-started turn that both runs and settles during the wait with no user message is no longer detected via turn id — acceptable against a guard that previously cancelled 100% of legitimate resumes.

Testing

  • New unit case (guards.test.ts): fire-time latestTurn: null vs non-null baseline → no cancellation; plus the converse (baseline null, turn present → still cancels).
  • New integration case (Reactor.test.ts): replays the production incident — rejection arrives mid-running-turn, thread settles to latestTurn: null + session stopped during the wait, wake must dispatch exactly one resume turn and post no thread-advanced note.
  • Both verified to fail against the unfixed guard (stash-run-restore), then pass with the fix.
  • Full src/t3x/autoResume suite: 69/69 across 4 consecutive runs. pnpm typecheck output identical to main (two pre-existing non-error diagnostics in untouched files).
  • Diagnosis evidence (timeline, event-store replay, live DB columns) recorded in Auto-resume never fires in practice: 'thread-advanced' false cancellation when the limited turn settles during the wait #6.

🤖 Generated with Claude Code

…d turn settles
The projection populates projection_threads.latest_turn_id only while a turn
is active, so a usage limit that lands mid-turn captures the running turn's id
in the guard baseline — and by wake time (turn settled, session idle) the
snapshot reports latestTurn: null. cancelReason treated that null as evidence
the thread had advanced and cancelled the pending resume, which killed every
real-world resume (observed 3/3 on 2026-07-25; firedAtMs was empty across all
threads — the feature had never fired).
Advancement now requires positive evidence: a different, NON-NULL turn id.
Null means "no active turn" — the expected state after a settled limit. User
takeovers stay covered by the user-took-over guard (checked first) and active
work by the progressing guard.
Regression coverage: a guards.test.ts unit case (null-at-fire vs non-null
baseline must not cancel) plus a Reactor.test.ts integration case replaying
the production shape (schedule mid-running-turn, settle to latestTurn: null,
assert the resume fires with no thread-advanced note). Both verified to fail
against the previous guard.
Fixes#6
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 9a36a907-aa9f-4ca4-bf0e-1e8885209e18

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch t3x/fix-thread-advanced-false-cancel

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@radroid
radroid merged commit 237ca66 into mainJul 25, 2026
1 check passed
@radroid
radroid deleted the t3x/fix-thread-advanced-false-cancel branch July 25, 2026 12:56
radroid added a commit that referenced this pull request Aug 17, 2026
Closes#39.
Two defects compounded into a permanently stranded thread: a usage limit
armed a resume, the user typed "keep going" at the banner, that message
was itself rejected so it started nothing, and the wake tick then threw
the arm away as `user-took-over`. Net pending afterwards: zero. Measured
on the reporting install at 4 of 17 armed resumes (~24%) since #6.
1. `guards.ts` — drop the `user-took-over` branch. It is the same
negative-evidence mistake #6 fixed on the line below it: "a message
exists that wasn't there when we armed" is not evidence the human took
the wheel, and in practice it is the opposite signal. Everything the
branch reached for is still covered — `progressing` (actively driving),
`awaiting-input` (blocked on a prompt), `thread-advanced` (a different
live turn), and the per-thread switch for "stop entirely".
`baseline.newestUserMessageId` stays: it is part of the persisted
record, it is re-captured on every (re)schedule, and it is what makes a
stranded arm diagnosable from the state file.
2. `decide.ts` — narrow `already-pending`. A rejection naming a CONCRETE
reset time later than the pending one now supersedes it instead of
being dropped, so a `seven_day` limit landing on an armed `five_hour`
no longer leaves the arm firing into a window that is still shut.
Restricted to `windowOpensInFuture` on purpose: a ladder-derived time
is `nowMs + delay`, so it is later on every telemetry re-emit, and
letting those supersede would push the arm out forever and flood the
timeline. An earlier reset never supersedes either — the existing arm
is already the conservative choice.
A supersede re-writes the record with a fresh `captureBaseline`, so it
is also the re-arm path, and posts a distinct
`t3x.auto-resume.rescheduled` note rather than a second "scheduled".
The issue's third item — detecting the org monthly spend limit — is a
genuinely separate feature (no `account.rate-limits.updated` path reaches
it at all) and is deliberately not in scope here, per the triage.
Tests: 4 new behavioural cases, each verified to fail against the old
code — guards (a newer user message does not cancel; it still cancels
when that message is actually being worked on), decide (supersede,
earlier-window no-op, ladder-churn guard), and two reactor integration
tests covering the reported shape end to end. Reactor suite run 12x for
flake. `docs/t3x/loop/DESIGN.md` §6 gets a dated note: guard #9 still
stands, but the hazard it was defending against is gone.
radroid added a commit that referenced this pull request Aug 18, 2026
Closes#39.
Two defects compounded into a permanently stranded thread: a usage limit
armed a resume, the user typed "keep going" at the banner, that message
was itself rejected so it started nothing, and the wake tick then threw
the arm away as `user-took-over`. Net pending afterwards: zero. Measured
on the reporting install at 4 of 17 armed resumes (~24%) since #6.
1. `guards.ts` — drop the `user-took-over` branch. It is the same
negative-evidence mistake #6 fixed on the line below it: "a message
exists that wasn't there when we armed" is not evidence the human took
the wheel, and in practice it is the opposite signal. Everything the
branch reached for is still covered — `progressing` (actively driving),
`awaiting-input` (blocked on a prompt), `thread-advanced` (a different
live turn), and the per-thread switch for "stop entirely".
`baseline.newestUserMessageId` stays: it is part of the persisted
record, it is re-captured on every (re)schedule, and it is what makes a
stranded arm diagnosable from the state file.
2. `decide.ts` — narrow `already-pending`. A rejection naming a CONCRETE
reset time later than the pending one now supersedes it instead of
being dropped, so a `seven_day` limit landing on an armed `five_hour`
no longer leaves the arm firing into a window that is still shut.
Restricted to `windowOpensInFuture` on purpose: a ladder-derived time
is `nowMs + delay`, so it is later on every telemetry re-emit, and
letting those supersede would push the arm out forever and flood the
timeline. An earlier reset never supersedes either — the existing arm
is already the conservative choice.
A supersede re-writes the record with a fresh `captureBaseline`, so it
is also the re-arm path, and posts a distinct
`t3x.auto-resume.rescheduled` note rather than a second "scheduled".
The issue's third item — detecting the org monthly spend limit — is a
genuinely separate feature (no `account.rate-limits.updated` path reaches
it at all) and is deliberately not in scope here, per the triage.
Tests: 4 new behavioural cases, each verified to fail against the old
code — guards (a newer user message does not cancel; it still cancels
when that message is actually being worked on), decide (supersede,
earlier-window no-op, ladder-churn guard), and two reactor integration
tests covering the reported shape end to end. Reactor suite run 12x for
flake. `docs/t3x/loop/DESIGN.md` §6 gets a dated note: guard #9 still
stands, but the hazard it was defending against is gone.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Auto-resume never fires in practice: 'thread-advanced' false cancellation when the limited turn settles during the wait

1 participant

@radroid
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(t3x): auto-resume never fired — 'thread-advanced' false cancellation on settled turns - #9

Merged
radroid merged 1 commit into
mainfrom
t3x/fix-thread-advanced-false-cancel
Jul 25, 2026
Merged

fix(t3x): auto-resume never fired — 'thread-advanced' false cancellation on settled turns#9
radroid merged 1 commit into
mainfrom
t3x/fix-thread-advanced-false-cancel

Conversation

@radroid

Copy link
Copy Markdown
Owner

Fixes#6.

What was broken

Every real-world auto-resume was cancelled at wake time with "Auto-resume cancelled: thread-advanced." — observed 3/3 on 2026-07-25 (threads 33c06a5a, 8f31df69, bb66acc2, all cancelled within 2s of their shared window reset), and the durable state showed firedAtMs: [] everywhere: the feature had never successfully fired.

Root cause

ProjectionSnapshotQuery.getSnapshot() reads threads from the SQLite projection, where latestTurn is joined on projection_threads.latest_turn_id — a column populated only while a turn is active. A usage limit normally lands mid-turn, so the guard baseline captures the running turn's id; by wake time the turn has settled and the snapshot reports latestTurn: null. The guard

if((thread.latestTurn?.turnId??null)!==baseline.latestTurnId)return"thread-advanced";

read null !== "<turnId>" as advancement and cancelled — then nothing ever re-armed, because an idle thread emits no further rejection events.

The fix

thread-advanced now requires positive evidence — a different, non-null turn id:

constcurrentTurnId=thread.latestTurn?.turnId??null;if(currentTurnId!==null&&currentTurnId!==baseline.latestTurnId)return"thread-advanced";

null at fire time means "no active turn", the expected state after a settled limit. Genuine user takeovers remain covered by user-took-over (checked first, keyed on the newest user message); active work remains covered by progressing. Known residual: an auto-started turn that both runs and settles during the wait with no user message is no longer detected via turn id — acceptable against a guard that previously cancelled 100% of legitimate resumes.

Testing

  • New unit case (guards.test.ts): fire-time latestTurn: null vs non-null baseline → no cancellation; plus the converse (baseline null, turn present → still cancels).
  • New integration case (Reactor.test.ts): replays the production incident — rejection arrives mid-running-turn, thread settles to latestTurn: null + session stopped during the wait, wake must dispatch exactly one resume turn and post no thread-advanced note.
  • Both verified to fail against the unfixed guard (stash-run-restore), then pass with the fix.
  • Full src/t3x/autoResume suite: 69/69 across 4 consecutive runs. pnpm typecheck output identical to main (two pre-existing non-error diagnostics in untouched files).
  • Diagnosis evidence (timeline, event-store replay, live DB columns) recorded in Auto-resume never fires in practice: 'thread-advanced' false cancellation when the limited turn settles during the wait #6.

🤖 Generated with Claude Code

…d turn settles
The projection populates projection_threads.latest_turn_id only while a turn
is active, so a usage limit that lands mid-turn captures the running turn's id
in the guard baseline — and by wake time (turn settled, session idle) the
snapshot reports latestTurn: null. cancelReason treated that null as evidence
the thread had advanced and cancelled the pending resume, which killed every
real-world resume (observed 3/3 on 2026-07-25; firedAtMs was empty across all
threads — the feature had never fired).
Advancement now requires positive evidence: a different, NON-NULL turn id.
Null means "no active turn" — the expected state after a settled limit. User
takeovers stay covered by the user-took-over guard (checked first) and active
work by the progressing guard.
Regression coverage: a guards.test.ts unit case (null-at-fire vs non-null
baseline must not cancel) plus a Reactor.test.ts integration case replaying
the production shape (schedule mid-running-turn, settle to latestTurn: null,
assert the resume fires with no thread-advanced note). Both verified to fail
against the previous guard.
Fixes#6
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 9a36a907-aa9f-4ca4-bf0e-1e8885209e18

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch t3x/fix-thread-advanced-false-cancel

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@radroid
radroid merged commit 237ca66 into mainJul 25, 2026
1 check passed
@radroid
radroid deleted the t3x/fix-thread-advanced-false-cancel branch July 25, 2026 12:56
radroid added a commit that referenced this pull request Aug 17, 2026
Closes#39.
Two defects compounded into a permanently stranded thread: a usage limit
armed a resume, the user typed "keep going" at the banner, that message
was itself rejected so it started nothing, and the wake tick then threw
the arm away as `user-took-over`. Net pending afterwards: zero. Measured
on the reporting install at 4 of 17 armed resumes (~24%) since #6.
1. `guards.ts` — drop the `user-took-over` branch. It is the same
negative-evidence mistake #6 fixed on the line below it: "a message
exists that wasn't there when we armed" is not evidence the human took
the wheel, and in practice it is the opposite signal. Everything the
branch reached for is still covered — `progressing` (actively driving),
`awaiting-input` (blocked on a prompt), `thread-advanced` (a different
live turn), and the per-thread switch for "stop entirely".
`baseline.newestUserMessageId` stays: it is part of the persisted
record, it is re-captured on every (re)schedule, and it is what makes a
stranded arm diagnosable from the state file.
2. `decide.ts` — narrow `already-pending`. A rejection naming a CONCRETE
reset time later than the pending one now supersedes it instead of
being dropped, so a `seven_day` limit landing on an armed `five_hour`
no longer leaves the arm firing into a window that is still shut.
Restricted to `windowOpensInFuture` on purpose: a ladder-derived time
is `nowMs + delay`, so it is later on every telemetry re-emit, and
letting those supersede would push the arm out forever and flood the
timeline. An earlier reset never supersedes either — the existing arm
is already the conservative choice.
A supersede re-writes the record with a fresh `captureBaseline`, so it
is also the re-arm path, and posts a distinct
`t3x.auto-resume.rescheduled` note rather than a second "scheduled".
The issue's third item — detecting the org monthly spend limit — is a
genuinely separate feature (no `account.rate-limits.updated` path reaches
it at all) and is deliberately not in scope here, per the triage.
Tests: 4 new behavioural cases, each verified to fail against the old
code — guards (a newer user message does not cancel; it still cancels
when that message is actually being worked on), decide (supersede,
earlier-window no-op, ladder-churn guard), and two reactor integration
tests covering the reported shape end to end. Reactor suite run 12x for
flake. `docs/t3x/loop/DESIGN.md` §6 gets a dated note: guard #9 still
stands, but the hazard it was defending against is gone.
radroid added a commit that referenced this pull request Aug 18, 2026
Closes#39.
Two defects compounded into a permanently stranded thread: a usage limit
armed a resume, the user typed "keep going" at the banner, that message
was itself rejected so it started nothing, and the wake tick then threw
the arm away as `user-took-over`. Net pending afterwards: zero. Measured
on the reporting install at 4 of 17 armed resumes (~24%) since #6.
1. `guards.ts` — drop the `user-took-over` branch. It is the same
negative-evidence mistake #6 fixed on the line below it: "a message
exists that wasn't there when we armed" is not evidence the human took
the wheel, and in practice it is the opposite signal. Everything the
branch reached for is still covered — `progressing` (actively driving),
`awaiting-input` (blocked on a prompt), `thread-advanced` (a different
live turn), and the per-thread switch for "stop entirely".
`baseline.newestUserMessageId` stays: it is part of the persisted
record, it is re-captured on every (re)schedule, and it is what makes a
stranded arm diagnosable from the state file.
2. `decide.ts` — narrow `already-pending`. A rejection naming a CONCRETE
reset time later than the pending one now supersedes it instead of
being dropped, so a `seven_day` limit landing on an armed `five_hour`
no longer leaves the arm firing into a window that is still shut.
Restricted to `windowOpensInFuture` on purpose: a ladder-derived time
is `nowMs + delay`, so it is later on every telemetry re-emit, and
letting those supersede would push the arm out forever and flood the
timeline. An earlier reset never supersedes either — the existing arm
is already the conservative choice.
A supersede re-writes the record with a fresh `captureBaseline`, so it
is also the re-arm path, and posts a distinct
`t3x.auto-resume.rescheduled` note rather than a second "scheduled".
The issue's third item — detecting the org monthly spend limit — is a
genuinely separate feature (no `account.rate-limits.updated` path reaches
it at all) and is deliberately not in scope here, per the triage.
Tests: 4 new behavioural cases, each verified to fail against the old
code — guards (a newer user message does not cancel; it still cancels
when that message is actually being worked on), decide (supersede,
earlier-window no-op, ladder-churn guard), and two reactor integration
tests covering the reported shape end to end. Reactor suite run 12x for
flake. `docs/t3x/loop/DESIGN.md` §6 gets a dated note: guard #9 still
stands, but the hazard it was defending against is gone.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Auto-resume never fires in practice: 'thread-advanced' false cancellation when the limited turn settles during the wait

1 participant

@radroid
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

fix(t3x): auto-resume never fired — 'thread-advanced' false cancellation on settled turns - #9

Merged
radroid merged 1 commit into
mainfrom
t3x/fix-thread-advanced-false-cancel
Jul 25, 2026
Merged

fix(t3x): auto-resume never fired — 'thread-advanced' false cancellation on settled turns#9
radroid merged 1 commit into
mainfrom
t3x/fix-thread-advanced-false-cancel

Conversation

@radroid

Copy link
Copy Markdown
Owner

Fixes#6.

What was broken

Every real-world auto-resume was cancelled at wake time with "Auto-resume cancelled: thread-advanced." — observed 3/3 on 2026-07-25 (threads 33c06a5a, 8f31df69, bb66acc2, all cancelled within 2s of their shared window reset), and the durable state showed firedAtMs: [] everywhere: the feature had never successfully fired.

Root cause

ProjectionSnapshotQuery.getSnapshot() reads threads from the SQLite projection, where latestTurn is joined on projection_threads.latest_turn_id — a column populated only while a turn is active. A usage limit normally lands mid-turn, so the guard baseline captures the running turn's id; by wake time the turn has settled and the snapshot reports latestTurn: null. The guard

if((thread.latestTurn?.turnId??null)!==baseline.latestTurnId)return"thread-advanced";

read null !== "<turnId>" as advancement and cancelled — then nothing ever re-armed, because an idle thread emits no further rejection events.

The fix

thread-advanced now requires positive evidence — a different, non-null turn id:

constcurrentTurnId=thread.latestTurn?.turnId??null;if(currentTurnId!==null&&currentTurnId!==baseline.latestTurnId)return"thread-advanced";

null at fire time means "no active turn", the expected state after a settled limit. Genuine user takeovers remain covered by user-took-over (checked first, keyed on the newest user message); active work remains covered by progressing. Known residual: an auto-started turn that both runs and settles during the wait with no user message is no longer detected via turn id — acceptable against a guard that previously cancelled 100% of legitimate resumes.

Testing

  • New unit case (guards.test.ts): fire-time latestTurn: null vs non-null baseline → no cancellation; plus the converse (baseline null, turn present → still cancels).
  • New integration case (Reactor.test.ts): replays the production incident — rejection arrives mid-running-turn, thread settles to latestTurn: null + session stopped during the wait, wake must dispatch exactly one resume turn and post no thread-advanced note.
  • Both verified to fail against the unfixed guard (stash-run-restore), then pass with the fix.
  • Full src/t3x/autoResume suite: 69/69 across 4 consecutive runs. pnpm typecheck output identical to main (two pre-existing non-error diagnostics in untouched files).
  • Diagnosis evidence (timeline, event-store replay, live DB columns) recorded in Auto-resume never fires in practice: 'thread-advanced' false cancellation when the limited turn settles during the wait #6.

🤖 Generated with Claude Code

…d turn settles
The projection populates projection_threads.latest_turn_id only while a turn
is active, so a usage limit that lands mid-turn captures the running turn's id
in the guard baseline — and by wake time (turn settled, session idle) the
snapshot reports latestTurn: null. cancelReason treated that null as evidence
the thread had advanced and cancelled the pending resume, which killed every
real-world resume (observed 3/3 on 2026-07-25; firedAtMs was empty across all
threads — the feature had never fired).
Advancement now requires positive evidence: a different, NON-NULL turn id.
Null means "no active turn" — the expected state after a settled limit. User
takeovers stay covered by the user-took-over guard (checked first) and active
work by the progressing guard.
Regression coverage: a guards.test.ts unit case (null-at-fire vs non-null
baseline must not cancel) plus a Reactor.test.ts integration case replaying
the production shape (schedule mid-running-turn, settle to latestTurn: null,
assert the resume fires with no thread-advanced note). Both verified to fail
against the previous guard.
Fixes#6
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 9a36a907-aa9f-4ca4-bf0e-1e8885209e18

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch t3x/fix-thread-advanced-false-cancel

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@radroid
radroid merged commit 237ca66 into mainJul 25, 2026
1 check passed
@radroid
radroid deleted the t3x/fix-thread-advanced-false-cancel branch July 25, 2026 12:56
radroid added a commit that referenced this pull request Aug 17, 2026
Closes#39.
Two defects compounded into a permanently stranded thread: a usage limit
armed a resume, the user typed "keep going" at the banner, that message
was itself rejected so it started nothing, and the wake tick then threw
the arm away as `user-took-over`. Net pending afterwards: zero. Measured
on the reporting install at 4 of 17 armed resumes (~24%) since #6.
1. `guards.ts` — drop the `user-took-over` branch. It is the same
negative-evidence mistake #6 fixed on the line below it: "a message
exists that wasn't there when we armed" is not evidence the human took
the wheel, and in practice it is the opposite signal. Everything the
branch reached for is still covered — `progressing` (actively driving),
`awaiting-input` (blocked on a prompt), `thread-advanced` (a different
live turn), and the per-thread switch for "stop entirely".
`baseline.newestUserMessageId` stays: it is part of the persisted
record, it is re-captured on every (re)schedule, and it is what makes a
stranded arm diagnosable from the state file.
2. `decide.ts` — narrow `already-pending`. A rejection naming a CONCRETE
reset time later than the pending one now supersedes it instead of
being dropped, so a `seven_day` limit landing on an armed `five_hour`
no longer leaves the arm firing into a window that is still shut.
Restricted to `windowOpensInFuture` on purpose: a ladder-derived time
is `nowMs + delay`, so it is later on every telemetry re-emit, and
letting those supersede would push the arm out forever and flood the
timeline. An earlier reset never supersedes either — the existing arm
is already the conservative choice.
A supersede re-writes the record with a fresh `captureBaseline`, so it
is also the re-arm path, and posts a distinct
`t3x.auto-resume.rescheduled` note rather than a second "scheduled".
The issue's third item — detecting the org monthly spend limit — is a
genuinely separate feature (no `account.rate-limits.updated` path reaches
it at all) and is deliberately not in scope here, per the triage.
Tests: 4 new behavioural cases, each verified to fail against the old
code — guards (a newer user message does not cancel; it still cancels
when that message is actually being worked on), decide (supersede,
earlier-window no-op, ladder-churn guard), and two reactor integration
tests covering the reported shape end to end. Reactor suite run 12x for
flake. `docs/t3x/loop/DESIGN.md` §6 gets a dated note: guard #9 still
stands, but the hazard it was defending against is gone.
radroid added a commit that referenced this pull request Aug 18, 2026
Closes#39.
Two defects compounded into a permanently stranded thread: a usage limit
armed a resume, the user typed "keep going" at the banner, that message
was itself rejected so it started nothing, and the wake tick then threw
the arm away as `user-took-over`. Net pending afterwards: zero. Measured
on the reporting install at 4 of 17 armed resumes (~24%) since #6.
1. `guards.ts` — drop the `user-took-over` branch. It is the same
negative-evidence mistake #6 fixed on the line below it: "a message
exists that wasn't there when we armed" is not evidence the human took
the wheel, and in practice it is the opposite signal. Everything the
branch reached for is still covered — `progressing` (actively driving),
`awaiting-input` (blocked on a prompt), `thread-advanced` (a different
live turn), and the per-thread switch for "stop entirely".
`baseline.newestUserMessageId` stays: it is part of the persisted
record, it is re-captured on every (re)schedule, and it is what makes a
stranded arm diagnosable from the state file.
2. `decide.ts` — narrow `already-pending`. A rejection naming a CONCRETE
reset time later than the pending one now supersedes it instead of
being dropped, so a `seven_day` limit landing on an armed `five_hour`
no longer leaves the arm firing into a window that is still shut.
Restricted to `windowOpensInFuture` on purpose: a ladder-derived time
is `nowMs + delay`, so it is later on every telemetry re-emit, and
letting those supersede would push the arm out forever and flood the
timeline. An earlier reset never supersedes either — the existing arm
is already the conservative choice.
A supersede re-writes the record with a fresh `captureBaseline`, so it
is also the re-arm path, and posts a distinct
`t3x.auto-resume.rescheduled` note rather than a second "scheduled".
The issue's third item — detecting the org monthly spend limit — is a
genuinely separate feature (no `account.rate-limits.updated` path reaches
it at all) and is deliberately not in scope here, per the triage.
Tests: 4 new behavioural cases, each verified to fail against the old
code — guards (a newer user message does not cancel; it still cancels
when that message is actually being worked on), decide (supersede,
earlier-window no-op, ladder-churn guard), and two reactor integration
tests covering the reported shape end to end. Reactor suite run 12x for
flake. `docs/t3x/loop/DESIGN.md` §6 gets a dated note: guard #9 still
stands, but the hazard it was defending against is gone.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Auto-resume never fires in practice: 'thread-advanced' false cancellation when the limited turn settles during the wait

1 participant

@radroid