fix(t3x): stop a user message destroying a pending auto-resume - #80

Merged
radroid merged 1 commit into
mainfrom
t3x/auto-resume-user-message
Aug 12, 2026
Merged

fix(t3x): stop a user message destroying a pending auto-resume#80
radroid merged 1 commit into
mainfrom
t3x/auto-resume-user-message

Conversation

@radroid

Copy link
Copy Markdown
Owner

Closes#39.

The bug

A usage limit arms a resume. The user types "keep going through the night" at the banner. That message is itself rejected by a limit, so it starts nothing — and the wake tick then throws the arm away as user-took-over. Pending count afterwards: zero. The thread never wakes.

Measured on the reporting install: 4 of 17 armed resumes (~24%) lost this way since #6 shipped.

The fix

1. guards.ts — drop the user-took-over branch.

It is the same negative-evidence mistake #6 fixed on the line directly below it. "A message exists that wasn't there when we armed" is not evidence the human took the wheel; in practice it is the opposite signal, because the banner is exactly when someone leaves instructions and steps away.

Everything the branch was reaching for is still covered:

caseguard
the user is actively driving right nowprogressing
the thread is blocked on a promptawaiting-input
a different turn is live at fire timethread-advanced
the user wants no resume at allthe per-thread switch, honoured in fireOne

baseline.newestUserMessageId stays in the shape: it is part of the persisted record, it is re-captured on every (re)schedule, and it is what makes a stranded arm diagnosable from t3x-auto-resume.json.

2. decide.ts — narrow already-pending.

A rejection naming a concrete reset time later than the pending one now supersedes it instead of being dropped, so a seven_day limit landing on top of an armed five_hour no longer leaves the arm firing into a window that is still shut.

Deliberately restricted to windowOpensInFuture. A ladder-derived time is nowMs + delay, so it is later on every telemetry re-emit — letting those supersede would push the arm out forever and flood the timeline with reschedule notes. An earlier reset time never supersedes either: the existing arm is already the conservative choice. Both are pinned by tests.

A supersede rewrites the record with a fresh captureBaseline, so it is also the re-arm path, and posts a distinct t3x.auto-resume.rescheduled note rather than a second "scheduled".

Scope

The issue's third item — detecting the org monthly spend limit — is not in this PR, per the triage on the issue: spend-limit detection does not exist anywhere in the repo (no account.rate-limits.updated path reaches it) and a monthly cap probably should not ride the backoff ladder at all. Worth its own issue. This PR is the two changes that stop the loss.

Verification

  • vp test run apps/server/src/t3x/autoResume — 76 passed (8 files).
  • The 4 new behavioural tests were each verified to fail against the old code (re-added the branch and disabled the supersede: exactly those 4 failed, 33 passed). No vacuous tests.
  • Reactor integration suite run 12× for flake — 0 failures. That file has a documented flake history, and the two new integration tests use the existing condition-based settleUntil/advanceUntil helpers rather than fixed spins.
  • vp lint apps/server/src/t3x/autoResume — clean. tsgo --noEmit -p apps/server — 0 errors.

Seam cost

None. Every changed source file is fork-owned under apps/server/src/t3x/autoResume/; git log <merge-base>..upstream/main -- apps/server/src/t3x/ is empty. No docs/t3x/SEAMS.md row changes.

docs/t3x/loop/DESIGN.md §6 gets a dated note: guard #9 still stands (nudging a thread inside a usage-limit window is pointless), but the hazard it was the last line of defence against is gone.

Closes#39.
Two defects compounded into a permanently stranded thread: a usage limit
armed a resume, the user typed "keep going" at the banner, that message
was itself rejected so it started nothing, and the wake tick then threw
the arm away as `user-took-over`. Net pending afterwards: zero. Measured
on the reporting install at 4 of 17 armed resumes (~24%) since #6.
1. `guards.ts` — drop the `user-took-over` branch. It is the same
negative-evidence mistake #6 fixed on the line below it: "a message
exists that wasn't there when we armed" is not evidence the human took
the wheel, and in practice it is the opposite signal. Everything the
branch reached for is still covered — `progressing` (actively driving),
`awaiting-input` (blocked on a prompt), `thread-advanced` (a different
live turn), and the per-thread switch for "stop entirely".
`baseline.newestUserMessageId` stays: it is part of the persisted
record, it is re-captured on every (re)schedule, and it is what makes a
stranded arm diagnosable from the state file.
2. `decide.ts` — narrow `already-pending`. A rejection naming a CONCRETE
reset time later than the pending one now supersedes it instead of
being dropped, so a `seven_day` limit landing on an armed `five_hour`
no longer leaves the arm firing into a window that is still shut.
Restricted to `windowOpensInFuture` on purpose: a ladder-derived time
is `nowMs + delay`, so it is later on every telemetry re-emit, and
letting those supersede would push the arm out forever and flood the
timeline. An earlier reset never supersedes either — the existing arm
is already the conservative choice.
A supersede re-writes the record with a fresh `captureBaseline`, so it
is also the re-arm path, and posts a distinct
`t3x.auto-resume.rescheduled` note rather than a second "scheduled".
The issue's third item — detecting the org monthly spend limit — is a
genuinely separate feature (no `account.rate-limits.updated` path reaches
it at all) and is deliberately not in scope here, per the triage.
Tests: 4 new behavioural cases, each verified to fail against the old
code — guards (a newer user message does not cancel; it still cancels
when that message is actually being worked on), decide (supersede,
earlier-window no-op, ladder-churn guard), and two reactor integration
tests covering the reported shape end to end. Reactor suite run 12x for
flake. `docs/t3x/loop/DESIGN.md` §6 gets a dated note: guard #9 still
stands, but the hazard it was defending against is gone.
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 7860f62f-a45e-4a99-b746-f40c66b5a874

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@radroid

Copy link
Copy Markdown
OwnerAuthor

Runtime verification — A/B on a real thread, not just unit tests

Ran this end to end in an isolated dev environment (worktree-local .t3, real claudeAgent session, real Claude turns). The fixture is the incident from the issue: a resume armed while e833f15e was the newest user message, then a newer user message fada9685 arrives, and no new turn runs — because in the real incident that message was itself rejected by a limit.

Identical seeded t3x-auto-resume.json in both runs. Only guards.ts differs.

With this PR's code — fires:

t3x-auto-resume.json -> "pending": null, "firedAtMs": [1786453083849]
activity -> t3x.auto-resume.resumed — "Resuming now (attempt 1 of 10)."
messages -> t3x-auto-resume:76a37e53-… role=user text="continue"
assistant:1b6114cb-… "41 42 43 44 45 …"

Claude genuinely picked the work back up — the thread had been counting to 40, and the resumed turn continued from 41.

With the user-took-over branch restored — cancels:

t3x-auto-resume.json -> "pending": null, "firedAtMs": []
activity -> t3x.auto-resume.cancelled — "Auto-resume cancelled: user-took-over."

Same fixture, same thread, nothing dispatched, attempt not even counted. That is the ~24% loss the issue measured, reproduced on demand and then fixed.

guards.ts was restored afterwards; git status is clean and the only remaining occurrence of the string user-took-over in the file is the explanatory comment.

Two side observations from the same session

1. thread-advanced still works — and I confirmed it by accident. My first attempt seeded baseline.latestTurnId: null while a turn had since run, and it correctly cancelled with thread-advanced. Removing the user-message branch did not weaken the guard beside it.

2. projection_threads.latest_turn_id does NOT appear to clear when a turn settles, which is contrary to what the #6 comment block in this file asserts:

latestTurn is joined on projection_threads.latest_turn_id, which is populated only while a turn is active

Observed directly: after the first turn completed, projection_thread_sessions.status = 'ready' and active_turn_id = NULL, but projection_threads.latest_turn_id was still 1592ce47-….

This does not change the correctness of either guard — #6's "require positive evidence" rule is right whether the column clears or not, and the existing regression test still pins the null case defensively. But the stated mechanism in that comment does not match this build, and since that comment is the documentation for why the guard looks the way it does, it is worth a proper look. Not fixing it in this PR; flagging it so the next person to read that comment does not trust it blindly.

@radroid
radroid merged commit a5e2abc into mainAug 12, 2026
2 checks passed
@radroid
radroid deleted the t3x/auto-resume-user-message branch August 12, 2026 02:06
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: auto-resume is cancelled as "user-took-over" by any user message — including one that was itself rejected by a limit

1 participant

@radroid
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

fix(t3x): stop a user message destroying a pending auto-resume - #80

Merged
radroid merged 1 commit into
mainfrom
t3x/auto-resume-user-message
Aug 12, 2026
Merged

fix(t3x): stop a user message destroying a pending auto-resume#80
radroid merged 1 commit into
mainfrom
t3x/auto-resume-user-message

Conversation

@radroid

Copy link
Copy Markdown
Owner

Closes#39.

The bug

A usage limit arms a resume. The user types "keep going through the night" at the banner. That message is itself rejected by a limit, so it starts nothing — and the wake tick then throws the arm away as user-took-over. Pending count afterwards: zero. The thread never wakes.

Measured on the reporting install: 4 of 17 armed resumes (~24%) lost this way since #6 shipped.

The fix

1. guards.ts — drop the user-took-over branch.

It is the same negative-evidence mistake #6 fixed on the line directly below it. "A message exists that wasn't there when we armed" is not evidence the human took the wheel; in practice it is the opposite signal, because the banner is exactly when someone leaves instructions and steps away.

Everything the branch was reaching for is still covered:

caseguard
the user is actively driving right nowprogressing
the thread is blocked on a promptawaiting-input
a different turn is live at fire timethread-advanced
the user wants no resume at allthe per-thread switch, honoured in fireOne

baseline.newestUserMessageId stays in the shape: it is part of the persisted record, it is re-captured on every (re)schedule, and it is what makes a stranded arm diagnosable from t3x-auto-resume.json.

2. decide.ts — narrow already-pending.

A rejection naming a concrete reset time later than the pending one now supersedes it instead of being dropped, so a seven_day limit landing on top of an armed five_hour no longer leaves the arm firing into a window that is still shut.

Deliberately restricted to windowOpensInFuture. A ladder-derived time is nowMs + delay, so it is later on every telemetry re-emit — letting those supersede would push the arm out forever and flood the timeline with reschedule notes. An earlier reset time never supersedes either: the existing arm is already the conservative choice. Both are pinned by tests.

A supersede rewrites the record with a fresh captureBaseline, so it is also the re-arm path, and posts a distinct t3x.auto-resume.rescheduled note rather than a second "scheduled".

Scope

The issue's third item — detecting the org monthly spend limit — is not in this PR, per the triage on the issue: spend-limit detection does not exist anywhere in the repo (no account.rate-limits.updated path reaches it) and a monthly cap probably should not ride the backoff ladder at all. Worth its own issue. This PR is the two changes that stop the loss.

Verification

  • vp test run apps/server/src/t3x/autoResume — 76 passed (8 files).
  • The 4 new behavioural tests were each verified to fail against the old code (re-added the branch and disabled the supersede: exactly those 4 failed, 33 passed). No vacuous tests.
  • Reactor integration suite run 12× for flake — 0 failures. That file has a documented flake history, and the two new integration tests use the existing condition-based settleUntil/advanceUntil helpers rather than fixed spins.
  • vp lint apps/server/src/t3x/autoResume — clean. tsgo --noEmit -p apps/server — 0 errors.

Seam cost

None. Every changed source file is fork-owned under apps/server/src/t3x/autoResume/; git log <merge-base>..upstream/main -- apps/server/src/t3x/ is empty. No docs/t3x/SEAMS.md row changes.

docs/t3x/loop/DESIGN.md §6 gets a dated note: guard #9 still stands (nudging a thread inside a usage-limit window is pointless), but the hazard it was the last line of defence against is gone.

Closes#39.
Two defects compounded into a permanently stranded thread: a usage limit
armed a resume, the user typed "keep going" at the banner, that message
was itself rejected so it started nothing, and the wake tick then threw
the arm away as `user-took-over`. Net pending afterwards: zero. Measured
on the reporting install at 4 of 17 armed resumes (~24%) since #6.
1. `guards.ts` — drop the `user-took-over` branch. It is the same
negative-evidence mistake #6 fixed on the line below it: "a message
exists that wasn't there when we armed" is not evidence the human took
the wheel, and in practice it is the opposite signal. Everything the
branch reached for is still covered — `progressing` (actively driving),
`awaiting-input` (blocked on a prompt), `thread-advanced` (a different
live turn), and the per-thread switch for "stop entirely".
`baseline.newestUserMessageId` stays: it is part of the persisted
record, it is re-captured on every (re)schedule, and it is what makes a
stranded arm diagnosable from the state file.
2. `decide.ts` — narrow `already-pending`. A rejection naming a CONCRETE
reset time later than the pending one now supersedes it instead of
being dropped, so a `seven_day` limit landing on an armed `five_hour`
no longer leaves the arm firing into a window that is still shut.
Restricted to `windowOpensInFuture` on purpose: a ladder-derived time
is `nowMs + delay`, so it is later on every telemetry re-emit, and
letting those supersede would push the arm out forever and flood the
timeline. An earlier reset never supersedes either — the existing arm
is already the conservative choice.
A supersede re-writes the record with a fresh `captureBaseline`, so it
is also the re-arm path, and posts a distinct
`t3x.auto-resume.rescheduled` note rather than a second "scheduled".
The issue's third item — detecting the org monthly spend limit — is a
genuinely separate feature (no `account.rate-limits.updated` path reaches
it at all) and is deliberately not in scope here, per the triage.
Tests: 4 new behavioural cases, each verified to fail against the old
code — guards (a newer user message does not cancel; it still cancels
when that message is actually being worked on), decide (supersede,
earlier-window no-op, ladder-churn guard), and two reactor integration
tests covering the reported shape end to end. Reactor suite run 12x for
flake. `docs/t3x/loop/DESIGN.md` §6 gets a dated note: guard #9 still
stands, but the hazard it was defending against is gone.
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 7860f62f-a45e-4a99-b746-f40c66b5a874

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@radroid

Copy link
Copy Markdown
OwnerAuthor

Runtime verification — A/B on a real thread, not just unit tests

Ran this end to end in an isolated dev environment (worktree-local .t3, real claudeAgent session, real Claude turns). The fixture is the incident from the issue: a resume armed while e833f15e was the newest user message, then a newer user message fada9685 arrives, and no new turn runs — because in the real incident that message was itself rejected by a limit.

Identical seeded t3x-auto-resume.json in both runs. Only guards.ts differs.

With this PR's code — fires:

t3x-auto-resume.json -> "pending": null, "firedAtMs": [1786453083849]
activity -> t3x.auto-resume.resumed — "Resuming now (attempt 1 of 10)."
messages -> t3x-auto-resume:76a37e53-… role=user text="continue"
assistant:1b6114cb-… "41 42 43 44 45 …"

Claude genuinely picked the work back up — the thread had been counting to 40, and the resumed turn continued from 41.

With the user-took-over branch restored — cancels:

t3x-auto-resume.json -> "pending": null, "firedAtMs": []
activity -> t3x.auto-resume.cancelled — "Auto-resume cancelled: user-took-over."

Same fixture, same thread, nothing dispatched, attempt not even counted. That is the ~24% loss the issue measured, reproduced on demand and then fixed.

guards.ts was restored afterwards; git status is clean and the only remaining occurrence of the string user-took-over in the file is the explanatory comment.

Two side observations from the same session

1. thread-advanced still works — and I confirmed it by accident. My first attempt seeded baseline.latestTurnId: null while a turn had since run, and it correctly cancelled with thread-advanced. Removing the user-message branch did not weaken the guard beside it.

2. projection_threads.latest_turn_id does NOT appear to clear when a turn settles, which is contrary to what the #6 comment block in this file asserts:

latestTurn is joined on projection_threads.latest_turn_id, which is populated only while a turn is active

Observed directly: after the first turn completed, projection_thread_sessions.status = 'ready' and active_turn_id = NULL, but projection_threads.latest_turn_id was still 1592ce47-….

This does not change the correctness of either guard — #6's "require positive evidence" rule is right whether the column clears or not, and the existing regression test still pins the null case defensively. But the stated mechanism in that comment does not match this build, and since that comment is the documentation for why the guard looks the way it does, it is worth a proper look. Not fixing it in this PR; flagging it so the next person to read that comment does not trust it blindly.

@radroid
radroid merged commit a5e2abc into mainAug 12, 2026
2 checks passed
@radroid
radroid deleted the t3x/auto-resume-user-message branch August 12, 2026 02:06
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: auto-resume is cancelled as "user-took-over" by any user message — including one that was itself rejected by a limit

1 participant

@radroid
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(t3x): stop a user message destroying a pending auto-resume - #80

Merged
radroid merged 1 commit into
mainfrom
t3x/auto-resume-user-message
Aug 12, 2026
Merged

fix(t3x): stop a user message destroying a pending auto-resume#80
radroid merged 1 commit into
mainfrom
t3x/auto-resume-user-message

Conversation

@radroid

Copy link
Copy Markdown
Owner

Closes#39.

The bug

A usage limit arms a resume. The user types "keep going through the night" at the banner. That message is itself rejected by a limit, so it starts nothing — and the wake tick then throws the arm away as user-took-over. Pending count afterwards: zero. The thread never wakes.

Measured on the reporting install: 4 of 17 armed resumes (~24%) lost this way since #6 shipped.

The fix

1. guards.ts — drop the user-took-over branch.

It is the same negative-evidence mistake #6 fixed on the line directly below it. "A message exists that wasn't there when we armed" is not evidence the human took the wheel; in practice it is the opposite signal, because the banner is exactly when someone leaves instructions and steps away.

Everything the branch was reaching for is still covered:

caseguard
the user is actively driving right nowprogressing
the thread is blocked on a promptawaiting-input
a different turn is live at fire timethread-advanced
the user wants no resume at allthe per-thread switch, honoured in fireOne

baseline.newestUserMessageId stays in the shape: it is part of the persisted record, it is re-captured on every (re)schedule, and it is what makes a stranded arm diagnosable from t3x-auto-resume.json.

2. decide.ts — narrow already-pending.

A rejection naming a concrete reset time later than the pending one now supersedes it instead of being dropped, so a seven_day limit landing on top of an armed five_hour no longer leaves the arm firing into a window that is still shut.

Deliberately restricted to windowOpensInFuture. A ladder-derived time is nowMs + delay, so it is later on every telemetry re-emit — letting those supersede would push the arm out forever and flood the timeline with reschedule notes. An earlier reset time never supersedes either: the existing arm is already the conservative choice. Both are pinned by tests.

A supersede rewrites the record with a fresh captureBaseline, so it is also the re-arm path, and posts a distinct t3x.auto-resume.rescheduled note rather than a second "scheduled".

Scope

The issue's third item — detecting the org monthly spend limit — is not in this PR, per the triage on the issue: spend-limit detection does not exist anywhere in the repo (no account.rate-limits.updated path reaches it) and a monthly cap probably should not ride the backoff ladder at all. Worth its own issue. This PR is the two changes that stop the loss.

Verification

  • vp test run apps/server/src/t3x/autoResume — 76 passed (8 files).
  • The 4 new behavioural tests were each verified to fail against the old code (re-added the branch and disabled the supersede: exactly those 4 failed, 33 passed). No vacuous tests.
  • Reactor integration suite run 12× for flake — 0 failures. That file has a documented flake history, and the two new integration tests use the existing condition-based settleUntil/advanceUntil helpers rather than fixed spins.
  • vp lint apps/server/src/t3x/autoResume — clean. tsgo --noEmit -p apps/server — 0 errors.

Seam cost

None. Every changed source file is fork-owned under apps/server/src/t3x/autoResume/; git log <merge-base>..upstream/main -- apps/server/src/t3x/ is empty. No docs/t3x/SEAMS.md row changes.

docs/t3x/loop/DESIGN.md §6 gets a dated note: guard #9 still stands (nudging a thread inside a usage-limit window is pointless), but the hazard it was the last line of defence against is gone.

Closes#39.
Two defects compounded into a permanently stranded thread: a usage limit
armed a resume, the user typed "keep going" at the banner, that message
was itself rejected so it started nothing, and the wake tick then threw
the arm away as `user-took-over`. Net pending afterwards: zero. Measured
on the reporting install at 4 of 17 armed resumes (~24%) since #6.
1. `guards.ts` — drop the `user-took-over` branch. It is the same
negative-evidence mistake #6 fixed on the line below it: "a message
exists that wasn't there when we armed" is not evidence the human took
the wheel, and in practice it is the opposite signal. Everything the
branch reached for is still covered — `progressing` (actively driving),
`awaiting-input` (blocked on a prompt), `thread-advanced` (a different
live turn), and the per-thread switch for "stop entirely".
`baseline.newestUserMessageId` stays: it is part of the persisted
record, it is re-captured on every (re)schedule, and it is what makes a
stranded arm diagnosable from the state file.
2. `decide.ts` — narrow `already-pending`. A rejection naming a CONCRETE
reset time later than the pending one now supersedes it instead of
being dropped, so a `seven_day` limit landing on an armed `five_hour`
no longer leaves the arm firing into a window that is still shut.
Restricted to `windowOpensInFuture` on purpose: a ladder-derived time
is `nowMs + delay`, so it is later on every telemetry re-emit, and
letting those supersede would push the arm out forever and flood the
timeline. An earlier reset never supersedes either — the existing arm
is already the conservative choice.
A supersede re-writes the record with a fresh `captureBaseline`, so it
is also the re-arm path, and posts a distinct
`t3x.auto-resume.rescheduled` note rather than a second "scheduled".
The issue's third item — detecting the org monthly spend limit — is a
genuinely separate feature (no `account.rate-limits.updated` path reaches
it at all) and is deliberately not in scope here, per the triage.
Tests: 4 new behavioural cases, each verified to fail against the old
code — guards (a newer user message does not cancel; it still cancels
when that message is actually being worked on), decide (supersede,
earlier-window no-op, ladder-churn guard), and two reactor integration
tests covering the reported shape end to end. Reactor suite run 12x for
flake. `docs/t3x/loop/DESIGN.md` §6 gets a dated note: guard #9 still
stands, but the hazard it was defending against is gone.
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 7860f62f-a45e-4a99-b746-f40c66b5a874

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@radroid

Copy link
Copy Markdown
OwnerAuthor

Runtime verification — A/B on a real thread, not just unit tests

Ran this end to end in an isolated dev environment (worktree-local .t3, real claudeAgent session, real Claude turns). The fixture is the incident from the issue: a resume armed while e833f15e was the newest user message, then a newer user message fada9685 arrives, and no new turn runs — because in the real incident that message was itself rejected by a limit.

Identical seeded t3x-auto-resume.json in both runs. Only guards.ts differs.

With this PR's code — fires:

t3x-auto-resume.json -> "pending": null, "firedAtMs": [1786453083849]
activity -> t3x.auto-resume.resumed — "Resuming now (attempt 1 of 10)."
messages -> t3x-auto-resume:76a37e53-… role=user text="continue"
assistant:1b6114cb-… "41 42 43 44 45 …"

Claude genuinely picked the work back up — the thread had been counting to 40, and the resumed turn continued from 41.

With the user-took-over branch restored — cancels:

t3x-auto-resume.json -> "pending": null, "firedAtMs": []
activity -> t3x.auto-resume.cancelled — "Auto-resume cancelled: user-took-over."

Same fixture, same thread, nothing dispatched, attempt not even counted. That is the ~24% loss the issue measured, reproduced on demand and then fixed.

guards.ts was restored afterwards; git status is clean and the only remaining occurrence of the string user-took-over in the file is the explanatory comment.

Two side observations from the same session

1. thread-advanced still works — and I confirmed it by accident. My first attempt seeded baseline.latestTurnId: null while a turn had since run, and it correctly cancelled with thread-advanced. Removing the user-message branch did not weaken the guard beside it.

2. projection_threads.latest_turn_id does NOT appear to clear when a turn settles, which is contrary to what the #6 comment block in this file asserts:

latestTurn is joined on projection_threads.latest_turn_id, which is populated only while a turn is active

Observed directly: after the first turn completed, projection_thread_sessions.status = 'ready' and active_turn_id = NULL, but projection_threads.latest_turn_id was still 1592ce47-….

This does not change the correctness of either guard — #6's "require positive evidence" rule is right whether the column clears or not, and the existing regression test still pins the null case defensively. But the stated mechanism in that comment does not match this build, and since that comment is the documentation for why the guard looks the way it does, it is worth a proper look. Not fixing it in this PR; flagging it so the next person to read that comment does not trust it blindly.

@radroid
radroid merged commit a5e2abc into mainAug 12, 2026
2 checks passed
@radroid
radroid deleted the t3x/auto-resume-user-message branch August 12, 2026 02:06
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: auto-resume is cancelled as "user-took-over" by any user message — including one that was itself rejected by a limit

1 participant

@radroid
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(t3x): stop a user message destroying a pending auto-resume - #80

Merged
radroid merged 1 commit into
mainfrom
t3x/auto-resume-user-message
Aug 12, 2026
Merged

fix(t3x): stop a user message destroying a pending auto-resume#80
radroid merged 1 commit into
mainfrom
t3x/auto-resume-user-message

Conversation

@radroid

Copy link
Copy Markdown
Owner

Closes#39.

The bug

A usage limit arms a resume. The user types "keep going through the night" at the banner. That message is itself rejected by a limit, so it starts nothing — and the wake tick then throws the arm away as user-took-over. Pending count afterwards: zero. The thread never wakes.

Measured on the reporting install: 4 of 17 armed resumes (~24%) lost this way since #6 shipped.

The fix

1. guards.ts — drop the user-took-over branch.

It is the same negative-evidence mistake #6 fixed on the line directly below it. "A message exists that wasn't there when we armed" is not evidence the human took the wheel; in practice it is the opposite signal, because the banner is exactly when someone leaves instructions and steps away.

Everything the branch was reaching for is still covered:

caseguard
the user is actively driving right nowprogressing
the thread is blocked on a promptawaiting-input
a different turn is live at fire timethread-advanced
the user wants no resume at allthe per-thread switch, honoured in fireOne

baseline.newestUserMessageId stays in the shape: it is part of the persisted record, it is re-captured on every (re)schedule, and it is what makes a stranded arm diagnosable from t3x-auto-resume.json.

2. decide.ts — narrow already-pending.

A rejection naming a concrete reset time later than the pending one now supersedes it instead of being dropped, so a seven_day limit landing on top of an armed five_hour no longer leaves the arm firing into a window that is still shut.

Deliberately restricted to windowOpensInFuture. A ladder-derived time is nowMs + delay, so it is later on every telemetry re-emit — letting those supersede would push the arm out forever and flood the timeline with reschedule notes. An earlier reset time never supersedes either: the existing arm is already the conservative choice. Both are pinned by tests.

A supersede rewrites the record with a fresh captureBaseline, so it is also the re-arm path, and posts a distinct t3x.auto-resume.rescheduled note rather than a second "scheduled".

Scope

The issue's third item — detecting the org monthly spend limit — is not in this PR, per the triage on the issue: spend-limit detection does not exist anywhere in the repo (no account.rate-limits.updated path reaches it) and a monthly cap probably should not ride the backoff ladder at all. Worth its own issue. This PR is the two changes that stop the loss.

Verification

  • vp test run apps/server/src/t3x/autoResume — 76 passed (8 files).
  • The 4 new behavioural tests were each verified to fail against the old code (re-added the branch and disabled the supersede: exactly those 4 failed, 33 passed). No vacuous tests.
  • Reactor integration suite run 12× for flake — 0 failures. That file has a documented flake history, and the two new integration tests use the existing condition-based settleUntil/advanceUntil helpers rather than fixed spins.
  • vp lint apps/server/src/t3x/autoResume — clean. tsgo --noEmit -p apps/server — 0 errors.

Seam cost

None. Every changed source file is fork-owned under apps/server/src/t3x/autoResume/; git log <merge-base>..upstream/main -- apps/server/src/t3x/ is empty. No docs/t3x/SEAMS.md row changes.

docs/t3x/loop/DESIGN.md §6 gets a dated note: guard #9 still stands (nudging a thread inside a usage-limit window is pointless), but the hazard it was the last line of defence against is gone.

Closes#39.
Two defects compounded into a permanently stranded thread: a usage limit
armed a resume, the user typed "keep going" at the banner, that message
was itself rejected so it started nothing, and the wake tick then threw
the arm away as `user-took-over`. Net pending afterwards: zero. Measured
on the reporting install at 4 of 17 armed resumes (~24%) since #6.
1. `guards.ts` — drop the `user-took-over` branch. It is the same
negative-evidence mistake #6 fixed on the line below it: "a message
exists that wasn't there when we armed" is not evidence the human took
the wheel, and in practice it is the opposite signal. Everything the
branch reached for is still covered — `progressing` (actively driving),
`awaiting-input` (blocked on a prompt), `thread-advanced` (a different
live turn), and the per-thread switch for "stop entirely".
`baseline.newestUserMessageId` stays: it is part of the persisted
record, it is re-captured on every (re)schedule, and it is what makes a
stranded arm diagnosable from the state file.
2. `decide.ts` — narrow `already-pending`. A rejection naming a CONCRETE
reset time later than the pending one now supersedes it instead of
being dropped, so a `seven_day` limit landing on an armed `five_hour`
no longer leaves the arm firing into a window that is still shut.
Restricted to `windowOpensInFuture` on purpose: a ladder-derived time
is `nowMs + delay`, so it is later on every telemetry re-emit, and
letting those supersede would push the arm out forever and flood the
timeline. An earlier reset never supersedes either — the existing arm
is already the conservative choice.
A supersede re-writes the record with a fresh `captureBaseline`, so it
is also the re-arm path, and posts a distinct
`t3x.auto-resume.rescheduled` note rather than a second "scheduled".
The issue's third item — detecting the org monthly spend limit — is a
genuinely separate feature (no `account.rate-limits.updated` path reaches
it at all) and is deliberately not in scope here, per the triage.
Tests: 4 new behavioural cases, each verified to fail against the old
code — guards (a newer user message does not cancel; it still cancels
when that message is actually being worked on), decide (supersede,
earlier-window no-op, ladder-churn guard), and two reactor integration
tests covering the reported shape end to end. Reactor suite run 12x for
flake. `docs/t3x/loop/DESIGN.md` §6 gets a dated note: guard #9 still
stands, but the hazard it was defending against is gone.
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 7860f62f-a45e-4a99-b746-f40c66b5a874

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@radroid

Copy link
Copy Markdown
OwnerAuthor

Runtime verification — A/B on a real thread, not just unit tests

Ran this end to end in an isolated dev environment (worktree-local .t3, real claudeAgent session, real Claude turns). The fixture is the incident from the issue: a resume armed while e833f15e was the newest user message, then a newer user message fada9685 arrives, and no new turn runs — because in the real incident that message was itself rejected by a limit.

Identical seeded t3x-auto-resume.json in both runs. Only guards.ts differs.

With this PR's code — fires:

t3x-auto-resume.json -> "pending": null, "firedAtMs": [1786453083849]
activity -> t3x.auto-resume.resumed — "Resuming now (attempt 1 of 10)."
messages -> t3x-auto-resume:76a37e53-… role=user text="continue"
assistant:1b6114cb-… "41 42 43 44 45 …"

Claude genuinely picked the work back up — the thread had been counting to 40, and the resumed turn continued from 41.

With the user-took-over branch restored — cancels:

t3x-auto-resume.json -> "pending": null, "firedAtMs": []
activity -> t3x.auto-resume.cancelled — "Auto-resume cancelled: user-took-over."

Same fixture, same thread, nothing dispatched, attempt not even counted. That is the ~24% loss the issue measured, reproduced on demand and then fixed.

guards.ts was restored afterwards; git status is clean and the only remaining occurrence of the string user-took-over in the file is the explanatory comment.

Two side observations from the same session

1. thread-advanced still works — and I confirmed it by accident. My first attempt seeded baseline.latestTurnId: null while a turn had since run, and it correctly cancelled with thread-advanced. Removing the user-message branch did not weaken the guard beside it.

2. projection_threads.latest_turn_id does NOT appear to clear when a turn settles, which is contrary to what the #6 comment block in this file asserts:

latestTurn is joined on projection_threads.latest_turn_id, which is populated only while a turn is active

Observed directly: after the first turn completed, projection_thread_sessions.status = 'ready' and active_turn_id = NULL, but projection_threads.latest_turn_id was still 1592ce47-….

This does not change the correctness of either guard — #6's "require positive evidence" rule is right whether the column clears or not, and the existing regression test still pins the null case defensively. But the stated mechanism in that comment does not match this build, and since that comment is the documentation for why the guard looks the way it does, it is worth a proper look. Not fixing it in this PR; flagging it so the next person to read that comment does not trust it blindly.

@radroid
radroid merged commit a5e2abc into mainAug 12, 2026
2 checks passed
@radroid
radroid deleted the t3x/auto-resume-user-message branch August 12, 2026 02:06
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: auto-resume is cancelled as "user-took-over" by any user message — including one that was itself rejected by a limit

1 participant

@radroid
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

fix(t3x): stop a user message destroying a pending auto-resume - #80

Merged
radroid merged 1 commit into
mainfrom
t3x/auto-resume-user-message
Aug 12, 2026
Merged

fix(t3x): stop a user message destroying a pending auto-resume#80
radroid merged 1 commit into
mainfrom
t3x/auto-resume-user-message

Conversation

@radroid

Copy link
Copy Markdown
Owner

Closes#39.

The bug

A usage limit arms a resume. The user types "keep going through the night" at the banner. That message is itself rejected by a limit, so it starts nothing — and the wake tick then throws the arm away as user-took-over. Pending count afterwards: zero. The thread never wakes.

Measured on the reporting install: 4 of 17 armed resumes (~24%) lost this way since #6 shipped.

The fix

1. guards.ts — drop the user-took-over branch.

It is the same negative-evidence mistake #6 fixed on the line directly below it. "A message exists that wasn't there when we armed" is not evidence the human took the wheel; in practice it is the opposite signal, because the banner is exactly when someone leaves instructions and steps away.

Everything the branch was reaching for is still covered:

caseguard
the user is actively driving right nowprogressing
the thread is blocked on a promptawaiting-input
a different turn is live at fire timethread-advanced
the user wants no resume at allthe per-thread switch, honoured in fireOne

baseline.newestUserMessageId stays in the shape: it is part of the persisted record, it is re-captured on every (re)schedule, and it is what makes a stranded arm diagnosable from t3x-auto-resume.json.

2. decide.ts — narrow already-pending.

A rejection naming a concrete reset time later than the pending one now supersedes it instead of being dropped, so a seven_day limit landing on top of an armed five_hour no longer leaves the arm firing into a window that is still shut.

Deliberately restricted to windowOpensInFuture. A ladder-derived time is nowMs + delay, so it is later on every telemetry re-emit — letting those supersede would push the arm out forever and flood the timeline with reschedule notes. An earlier reset time never supersedes either: the existing arm is already the conservative choice. Both are pinned by tests.

A supersede rewrites the record with a fresh captureBaseline, so it is also the re-arm path, and posts a distinct t3x.auto-resume.rescheduled note rather than a second "scheduled".

Scope

The issue's third item — detecting the org monthly spend limit — is not in this PR, per the triage on the issue: spend-limit detection does not exist anywhere in the repo (no account.rate-limits.updated path reaches it) and a monthly cap probably should not ride the backoff ladder at all. Worth its own issue. This PR is the two changes that stop the loss.

Verification

  • vp test run apps/server/src/t3x/autoResume — 76 passed (8 files).
  • The 4 new behavioural tests were each verified to fail against the old code (re-added the branch and disabled the supersede: exactly those 4 failed, 33 passed). No vacuous tests.
  • Reactor integration suite run 12× for flake — 0 failures. That file has a documented flake history, and the two new integration tests use the existing condition-based settleUntil/advanceUntil helpers rather than fixed spins.
  • vp lint apps/server/src/t3x/autoResume — clean. tsgo --noEmit -p apps/server — 0 errors.

Seam cost

None. Every changed source file is fork-owned under apps/server/src/t3x/autoResume/; git log <merge-base>..upstream/main -- apps/server/src/t3x/ is empty. No docs/t3x/SEAMS.md row changes.

docs/t3x/loop/DESIGN.md §6 gets a dated note: guard #9 still stands (nudging a thread inside a usage-limit window is pointless), but the hazard it was the last line of defence against is gone.

Closes#39.
Two defects compounded into a permanently stranded thread: a usage limit
armed a resume, the user typed "keep going" at the banner, that message
was itself rejected so it started nothing, and the wake tick then threw
the arm away as `user-took-over`. Net pending afterwards: zero. Measured
on the reporting install at 4 of 17 armed resumes (~24%) since #6.
1. `guards.ts` — drop the `user-took-over` branch. It is the same
negative-evidence mistake #6 fixed on the line below it: "a message
exists that wasn't there when we armed" is not evidence the human took
the wheel, and in practice it is the opposite signal. Everything the
branch reached for is still covered — `progressing` (actively driving),
`awaiting-input` (blocked on a prompt), `thread-advanced` (a different
live turn), and the per-thread switch for "stop entirely".
`baseline.newestUserMessageId` stays: it is part of the persisted
record, it is re-captured on every (re)schedule, and it is what makes a
stranded arm diagnosable from the state file.
2. `decide.ts` — narrow `already-pending`. A rejection naming a CONCRETE
reset time later than the pending one now supersedes it instead of
being dropped, so a `seven_day` limit landing on an armed `five_hour`
no longer leaves the arm firing into a window that is still shut.
Restricted to `windowOpensInFuture` on purpose: a ladder-derived time
is `nowMs + delay`, so it is later on every telemetry re-emit, and
letting those supersede would push the arm out forever and flood the
timeline. An earlier reset never supersedes either — the existing arm
is already the conservative choice.
A supersede re-writes the record with a fresh `captureBaseline`, so it
is also the re-arm path, and posts a distinct
`t3x.auto-resume.rescheduled` note rather than a second "scheduled".
The issue's third item — detecting the org monthly spend limit — is a
genuinely separate feature (no `account.rate-limits.updated` path reaches
it at all) and is deliberately not in scope here, per the triage.
Tests: 4 new behavioural cases, each verified to fail against the old
code — guards (a newer user message does not cancel; it still cancels
when that message is actually being worked on), decide (supersede,
earlier-window no-op, ladder-churn guard), and two reactor integration
tests covering the reported shape end to end. Reactor suite run 12x for
flake. `docs/t3x/loop/DESIGN.md` §6 gets a dated note: guard #9 still
stands, but the hazard it was defending against is gone.
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 7860f62f-a45e-4a99-b746-f40c66b5a874

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@radroid

Copy link
Copy Markdown
OwnerAuthor

Runtime verification — A/B on a real thread, not just unit tests

Ran this end to end in an isolated dev environment (worktree-local .t3, real claudeAgent session, real Claude turns). The fixture is the incident from the issue: a resume armed while e833f15e was the newest user message, then a newer user message fada9685 arrives, and no new turn runs — because in the real incident that message was itself rejected by a limit.

Identical seeded t3x-auto-resume.json in both runs. Only guards.ts differs.

With this PR's code — fires:

t3x-auto-resume.json -> "pending": null, "firedAtMs": [1786453083849]
activity -> t3x.auto-resume.resumed — "Resuming now (attempt 1 of 10)."
messages -> t3x-auto-resume:76a37e53-… role=user text="continue"
assistant:1b6114cb-… "41 42 43 44 45 …"

Claude genuinely picked the work back up — the thread had been counting to 40, and the resumed turn continued from 41.

With the user-took-over branch restored — cancels:

t3x-auto-resume.json -> "pending": null, "firedAtMs": []
activity -> t3x.auto-resume.cancelled — "Auto-resume cancelled: user-took-over."

Same fixture, same thread, nothing dispatched, attempt not even counted. That is the ~24% loss the issue measured, reproduced on demand and then fixed.

guards.ts was restored afterwards; git status is clean and the only remaining occurrence of the string user-took-over in the file is the explanatory comment.

Two side observations from the same session

1. thread-advanced still works — and I confirmed it by accident. My first attempt seeded baseline.latestTurnId: null while a turn had since run, and it correctly cancelled with thread-advanced. Removing the user-message branch did not weaken the guard beside it.

2. projection_threads.latest_turn_id does NOT appear to clear when a turn settles, which is contrary to what the #6 comment block in this file asserts:

latestTurn is joined on projection_threads.latest_turn_id, which is populated only while a turn is active

Observed directly: after the first turn completed, projection_thread_sessions.status = 'ready' and active_turn_id = NULL, but projection_threads.latest_turn_id was still 1592ce47-….

This does not change the correctness of either guard — #6's "require positive evidence" rule is right whether the column clears or not, and the existing regression test still pins the null case defensively. But the stated mechanism in that comment does not match this build, and since that comment is the documentation for why the guard looks the way it does, it is worth a proper look. Not fixing it in this PR; flagging it so the next person to read that comment does not trust it blindly.

@radroid
radroid merged commit a5e2abc into mainAug 12, 2026
2 checks passed
@radroid
radroid deleted the t3x/auto-resume-user-message branch August 12, 2026 02:06
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: auto-resume is cancelled as "user-took-over" by any user message — including one that was itself rejected by a limit

1 participant

@radroid
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(t3x): stop a user message destroying a pending auto-resume - #80

Merged
radroid merged 1 commit into
mainfrom
t3x/auto-resume-user-message
Aug 12, 2026
Merged

fix(t3x): stop a user message destroying a pending auto-resume#80
radroid merged 1 commit into
mainfrom
t3x/auto-resume-user-message

Conversation

@radroid

Copy link
Copy Markdown
Owner

Closes#39.

The bug

A usage limit arms a resume. The user types "keep going through the night" at the banner. That message is itself rejected by a limit, so it starts nothing — and the wake tick then throws the arm away as user-took-over. Pending count afterwards: zero. The thread never wakes.

Measured on the reporting install: 4 of 17 armed resumes (~24%) lost this way since #6 shipped.

The fix

1. guards.ts — drop the user-took-over branch.

It is the same negative-evidence mistake #6 fixed on the line directly below it. "A message exists that wasn't there when we armed" is not evidence the human took the wheel; in practice it is the opposite signal, because the banner is exactly when someone leaves instructions and steps away.

Everything the branch was reaching for is still covered:

caseguard
the user is actively driving right nowprogressing
the thread is blocked on a promptawaiting-input
a different turn is live at fire timethread-advanced
the user wants no resume at allthe per-thread switch, honoured in fireOne

baseline.newestUserMessageId stays in the shape: it is part of the persisted record, it is re-captured on every (re)schedule, and it is what makes a stranded arm diagnosable from t3x-auto-resume.json.

2. decide.ts — narrow already-pending.

A rejection naming a concrete reset time later than the pending one now supersedes it instead of being dropped, so a seven_day limit landing on top of an armed five_hour no longer leaves the arm firing into a window that is still shut.

Deliberately restricted to windowOpensInFuture. A ladder-derived time is nowMs + delay, so it is later on every telemetry re-emit — letting those supersede would push the arm out forever and flood the timeline with reschedule notes. An earlier reset time never supersedes either: the existing arm is already the conservative choice. Both are pinned by tests.

A supersede rewrites the record with a fresh captureBaseline, so it is also the re-arm path, and posts a distinct t3x.auto-resume.rescheduled note rather than a second "scheduled".

Scope

The issue's third item — detecting the org monthly spend limit — is not in this PR, per the triage on the issue: spend-limit detection does not exist anywhere in the repo (no account.rate-limits.updated path reaches it) and a monthly cap probably should not ride the backoff ladder at all. Worth its own issue. This PR is the two changes that stop the loss.

Verification

  • vp test run apps/server/src/t3x/autoResume — 76 passed (8 files).
  • The 4 new behavioural tests were each verified to fail against the old code (re-added the branch and disabled the supersede: exactly those 4 failed, 33 passed). No vacuous tests.
  • Reactor integration suite run 12× for flake — 0 failures. That file has a documented flake history, and the two new integration tests use the existing condition-based settleUntil/advanceUntil helpers rather than fixed spins.
  • vp lint apps/server/src/t3x/autoResume — clean. tsgo --noEmit -p apps/server — 0 errors.

Seam cost

None. Every changed source file is fork-owned under apps/server/src/t3x/autoResume/; git log <merge-base>..upstream/main -- apps/server/src/t3x/ is empty. No docs/t3x/SEAMS.md row changes.

docs/t3x/loop/DESIGN.md §6 gets a dated note: guard #9 still stands (nudging a thread inside a usage-limit window is pointless), but the hazard it was the last line of defence against is gone.

Closes#39.
Two defects compounded into a permanently stranded thread: a usage limit
armed a resume, the user typed "keep going" at the banner, that message
was itself rejected so it started nothing, and the wake tick then threw
the arm away as `user-took-over`. Net pending afterwards: zero. Measured
on the reporting install at 4 of 17 armed resumes (~24%) since #6.
1. `guards.ts` — drop the `user-took-over` branch. It is the same
negative-evidence mistake #6 fixed on the line below it: "a message
exists that wasn't there when we armed" is not evidence the human took
the wheel, and in practice it is the opposite signal. Everything the
branch reached for is still covered — `progressing` (actively driving),
`awaiting-input` (blocked on a prompt), `thread-advanced` (a different
live turn), and the per-thread switch for "stop entirely".
`baseline.newestUserMessageId` stays: it is part of the persisted
record, it is re-captured on every (re)schedule, and it is what makes a
stranded arm diagnosable from the state file.
2. `decide.ts` — narrow `already-pending`. A rejection naming a CONCRETE
reset time later than the pending one now supersedes it instead of
being dropped, so a `seven_day` limit landing on an armed `five_hour`
no longer leaves the arm firing into a window that is still shut.
Restricted to `windowOpensInFuture` on purpose: a ladder-derived time
is `nowMs + delay`, so it is later on every telemetry re-emit, and
letting those supersede would push the arm out forever and flood the
timeline. An earlier reset never supersedes either — the existing arm
is already the conservative choice.
A supersede re-writes the record with a fresh `captureBaseline`, so it
is also the re-arm path, and posts a distinct
`t3x.auto-resume.rescheduled` note rather than a second "scheduled".
The issue's third item — detecting the org monthly spend limit — is a
genuinely separate feature (no `account.rate-limits.updated` path reaches
it at all) and is deliberately not in scope here, per the triage.
Tests: 4 new behavioural cases, each verified to fail against the old
code — guards (a newer user message does not cancel; it still cancels
when that message is actually being worked on), decide (supersede,
earlier-window no-op, ladder-churn guard), and two reactor integration
tests covering the reported shape end to end. Reactor suite run 12x for
flake. `docs/t3x/loop/DESIGN.md` §6 gets a dated note: guard #9 still
stands, but the hazard it was defending against is gone.
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 7860f62f-a45e-4a99-b746-f40c66b5a874

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@radroid

Copy link
Copy Markdown
OwnerAuthor

Runtime verification — A/B on a real thread, not just unit tests

Ran this end to end in an isolated dev environment (worktree-local .t3, real claudeAgent session, real Claude turns). The fixture is the incident from the issue: a resume armed while e833f15e was the newest user message, then a newer user message fada9685 arrives, and no new turn runs — because in the real incident that message was itself rejected by a limit.

Identical seeded t3x-auto-resume.json in both runs. Only guards.ts differs.

With this PR's code — fires:

t3x-auto-resume.json -> "pending": null, "firedAtMs": [1786453083849]
activity -> t3x.auto-resume.resumed — "Resuming now (attempt 1 of 10)."
messages -> t3x-auto-resume:76a37e53-… role=user text="continue"
assistant:1b6114cb-… "41 42 43 44 45 …"

Claude genuinely picked the work back up — the thread had been counting to 40, and the resumed turn continued from 41.

With the user-took-over branch restored — cancels:

t3x-auto-resume.json -> "pending": null, "firedAtMs": []
activity -> t3x.auto-resume.cancelled — "Auto-resume cancelled: user-took-over."

Same fixture, same thread, nothing dispatched, attempt not even counted. That is the ~24% loss the issue measured, reproduced on demand and then fixed.

guards.ts was restored afterwards; git status is clean and the only remaining occurrence of the string user-took-over in the file is the explanatory comment.

Two side observations from the same session

1. thread-advanced still works — and I confirmed it by accident. My first attempt seeded baseline.latestTurnId: null while a turn had since run, and it correctly cancelled with thread-advanced. Removing the user-message branch did not weaken the guard beside it.

2. projection_threads.latest_turn_id does NOT appear to clear when a turn settles, which is contrary to what the #6 comment block in this file asserts:

latestTurn is joined on projection_threads.latest_turn_id, which is populated only while a turn is active

Observed directly: after the first turn completed, projection_thread_sessions.status = 'ready' and active_turn_id = NULL, but projection_threads.latest_turn_id was still 1592ce47-….

This does not change the correctness of either guard — #6's "require positive evidence" rule is right whether the column clears or not, and the existing regression test still pins the null case defensively. But the stated mechanism in that comment does not match this build, and since that comment is the documentation for why the guard looks the way it does, it is worth a proper look. Not fixing it in this PR; flagging it so the next person to read that comment does not trust it blindly.

@radroid
radroid merged commit a5e2abc into mainAug 12, 2026
2 checks passed
@radroid
radroid deleted the t3x/auto-resume-user-message branch August 12, 2026 02:06
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: auto-resume is cancelled as "user-took-over" by any user message — including one that was itself rejected by a limit

1 participant

@radroid
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(t3x): stop a user message destroying a pending auto-resume - #80

Merged
radroid merged 1 commit into
mainfrom
t3x/auto-resume-user-message
Aug 12, 2026
Merged

fix(t3x): stop a user message destroying a pending auto-resume#80
radroid merged 1 commit into
mainfrom
t3x/auto-resume-user-message

Conversation

@radroid

Copy link
Copy Markdown
Owner

Closes#39.

The bug

A usage limit arms a resume. The user types "keep going through the night" at the banner. That message is itself rejected by a limit, so it starts nothing — and the wake tick then throws the arm away as user-took-over. Pending count afterwards: zero. The thread never wakes.

Measured on the reporting install: 4 of 17 armed resumes (~24%) lost this way since #6 shipped.

The fix

1. guards.ts — drop the user-took-over branch.

It is the same negative-evidence mistake #6 fixed on the line directly below it. "A message exists that wasn't there when we armed" is not evidence the human took the wheel; in practice it is the opposite signal, because the banner is exactly when someone leaves instructions and steps away.

Everything the branch was reaching for is still covered:

caseguard
the user is actively driving right nowprogressing
the thread is blocked on a promptawaiting-input
a different turn is live at fire timethread-advanced
the user wants no resume at allthe per-thread switch, honoured in fireOne

baseline.newestUserMessageId stays in the shape: it is part of the persisted record, it is re-captured on every (re)schedule, and it is what makes a stranded arm diagnosable from t3x-auto-resume.json.

2. decide.ts — narrow already-pending.

A rejection naming a concrete reset time later than the pending one now supersedes it instead of being dropped, so a seven_day limit landing on top of an armed five_hour no longer leaves the arm firing into a window that is still shut.

Deliberately restricted to windowOpensInFuture. A ladder-derived time is nowMs + delay, so it is later on every telemetry re-emit — letting those supersede would push the arm out forever and flood the timeline with reschedule notes. An earlier reset time never supersedes either: the existing arm is already the conservative choice. Both are pinned by tests.

A supersede rewrites the record with a fresh captureBaseline, so it is also the re-arm path, and posts a distinct t3x.auto-resume.rescheduled note rather than a second "scheduled".

Scope

The issue's third item — detecting the org monthly spend limit — is not in this PR, per the triage on the issue: spend-limit detection does not exist anywhere in the repo (no account.rate-limits.updated path reaches it) and a monthly cap probably should not ride the backoff ladder at all. Worth its own issue. This PR is the two changes that stop the loss.

Verification

  • vp test run apps/server/src/t3x/autoResume — 76 passed (8 files).
  • The 4 new behavioural tests were each verified to fail against the old code (re-added the branch and disabled the supersede: exactly those 4 failed, 33 passed). No vacuous tests.
  • Reactor integration suite run 12× for flake — 0 failures. That file has a documented flake history, and the two new integration tests use the existing condition-based settleUntil/advanceUntil helpers rather than fixed spins.
  • vp lint apps/server/src/t3x/autoResume — clean. tsgo --noEmit -p apps/server — 0 errors.

Seam cost

None. Every changed source file is fork-owned under apps/server/src/t3x/autoResume/; git log <merge-base>..upstream/main -- apps/server/src/t3x/ is empty. No docs/t3x/SEAMS.md row changes.

docs/t3x/loop/DESIGN.md §6 gets a dated note: guard #9 still stands (nudging a thread inside a usage-limit window is pointless), but the hazard it was the last line of defence against is gone.

Closes#39.
Two defects compounded into a permanently stranded thread: a usage limit
armed a resume, the user typed "keep going" at the banner, that message
was itself rejected so it started nothing, and the wake tick then threw
the arm away as `user-took-over`. Net pending afterwards: zero. Measured
on the reporting install at 4 of 17 armed resumes (~24%) since #6.
1. `guards.ts` — drop the `user-took-over` branch. It is the same
negative-evidence mistake #6 fixed on the line below it: "a message
exists that wasn't there when we armed" is not evidence the human took
the wheel, and in practice it is the opposite signal. Everything the
branch reached for is still covered — `progressing` (actively driving),
`awaiting-input` (blocked on a prompt), `thread-advanced` (a different
live turn), and the per-thread switch for "stop entirely".
`baseline.newestUserMessageId` stays: it is part of the persisted
record, it is re-captured on every (re)schedule, and it is what makes a
stranded arm diagnosable from the state file.
2. `decide.ts` — narrow `already-pending`. A rejection naming a CONCRETE
reset time later than the pending one now supersedes it instead of
being dropped, so a `seven_day` limit landing on an armed `five_hour`
no longer leaves the arm firing into a window that is still shut.
Restricted to `windowOpensInFuture` on purpose: a ladder-derived time
is `nowMs + delay`, so it is later on every telemetry re-emit, and
letting those supersede would push the arm out forever and flood the
timeline. An earlier reset never supersedes either — the existing arm
is already the conservative choice.
A supersede re-writes the record with a fresh `captureBaseline`, so it
is also the re-arm path, and posts a distinct
`t3x.auto-resume.rescheduled` note rather than a second "scheduled".
The issue's third item — detecting the org monthly spend limit — is a
genuinely separate feature (no `account.rate-limits.updated` path reaches
it at all) and is deliberately not in scope here, per the triage.
Tests: 4 new behavioural cases, each verified to fail against the old
code — guards (a newer user message does not cancel; it still cancels
when that message is actually being worked on), decide (supersede,
earlier-window no-op, ladder-churn guard), and two reactor integration
tests covering the reported shape end to end. Reactor suite run 12x for
flake. `docs/t3x/loop/DESIGN.md` §6 gets a dated note: guard #9 still
stands, but the hazard it was defending against is gone.
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 7860f62f-a45e-4a99-b746-f40c66b5a874

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@radroid

Copy link
Copy Markdown
OwnerAuthor

Runtime verification — A/B on a real thread, not just unit tests

Ran this end to end in an isolated dev environment (worktree-local .t3, real claudeAgent session, real Claude turns). The fixture is the incident from the issue: a resume armed while e833f15e was the newest user message, then a newer user message fada9685 arrives, and no new turn runs — because in the real incident that message was itself rejected by a limit.

Identical seeded t3x-auto-resume.json in both runs. Only guards.ts differs.

With this PR's code — fires:

t3x-auto-resume.json -> "pending": null, "firedAtMs": [1786453083849]
activity -> t3x.auto-resume.resumed — "Resuming now (attempt 1 of 10)."
messages -> t3x-auto-resume:76a37e53-… role=user text="continue"
assistant:1b6114cb-… "41 42 43 44 45 …"

Claude genuinely picked the work back up — the thread had been counting to 40, and the resumed turn continued from 41.

With the user-took-over branch restored — cancels:

t3x-auto-resume.json -> "pending": null, "firedAtMs": []
activity -> t3x.auto-resume.cancelled — "Auto-resume cancelled: user-took-over."

Same fixture, same thread, nothing dispatched, attempt not even counted. That is the ~24% loss the issue measured, reproduced on demand and then fixed.

guards.ts was restored afterwards; git status is clean and the only remaining occurrence of the string user-took-over in the file is the explanatory comment.

Two side observations from the same session

1. thread-advanced still works — and I confirmed it by accident. My first attempt seeded baseline.latestTurnId: null while a turn had since run, and it correctly cancelled with thread-advanced. Removing the user-message branch did not weaken the guard beside it.

2. projection_threads.latest_turn_id does NOT appear to clear when a turn settles, which is contrary to what the #6 comment block in this file asserts:

latestTurn is joined on projection_threads.latest_turn_id, which is populated only while a turn is active

Observed directly: after the first turn completed, projection_thread_sessions.status = 'ready' and active_turn_id = NULL, but projection_threads.latest_turn_id was still 1592ce47-….

This does not change the correctness of either guard — #6's "require positive evidence" rule is right whether the column clears or not, and the existing regression test still pins the null case defensively. But the stated mechanism in that comment does not match this build, and since that comment is the documentation for why the guard looks the way it does, it is worth a proper look. Not fixing it in this PR; flagging it so the next person to read that comment does not trust it blindly.

@radroid
radroid merged commit a5e2abc into mainAug 12, 2026
2 checks passed
@radroid
radroid deleted the t3x/auto-resume-user-message branch August 12, 2026 02:06
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: auto-resume is cancelled as "user-took-over" by any user message — including one that was itself rejected by a limit

1 participant

@radroid
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

fix(t3x): stop a user message destroying a pending auto-resume - #80

Merged
radroid merged 1 commit into
mainfrom
t3x/auto-resume-user-message
Aug 12, 2026
Merged

fix(t3x): stop a user message destroying a pending auto-resume#80
radroid merged 1 commit into
mainfrom
t3x/auto-resume-user-message

Conversation

@radroid

Copy link
Copy Markdown
Owner

Closes#39.

The bug

A usage limit arms a resume. The user types "keep going through the night" at the banner. That message is itself rejected by a limit, so it starts nothing — and the wake tick then throws the arm away as user-took-over. Pending count afterwards: zero. The thread never wakes.

Measured on the reporting install: 4 of 17 armed resumes (~24%) lost this way since #6 shipped.

The fix

1. guards.ts — drop the user-took-over branch.

It is the same negative-evidence mistake #6 fixed on the line directly below it. "A message exists that wasn't there when we armed" is not evidence the human took the wheel; in practice it is the opposite signal, because the banner is exactly when someone leaves instructions and steps away.

Everything the branch was reaching for is still covered:

caseguard
the user is actively driving right nowprogressing
the thread is blocked on a promptawaiting-input
a different turn is live at fire timethread-advanced
the user wants no resume at allthe per-thread switch, honoured in fireOne

baseline.newestUserMessageId stays in the shape: it is part of the persisted record, it is re-captured on every (re)schedule, and it is what makes a stranded arm diagnosable from t3x-auto-resume.json.

2. decide.ts — narrow already-pending.

A rejection naming a concrete reset time later than the pending one now supersedes it instead of being dropped, so a seven_day limit landing on top of an armed five_hour no longer leaves the arm firing into a window that is still shut.

Deliberately restricted to windowOpensInFuture. A ladder-derived time is nowMs + delay, so it is later on every telemetry re-emit — letting those supersede would push the arm out forever and flood the timeline with reschedule notes. An earlier reset time never supersedes either: the existing arm is already the conservative choice. Both are pinned by tests.

A supersede rewrites the record with a fresh captureBaseline, so it is also the re-arm path, and posts a distinct t3x.auto-resume.rescheduled note rather than a second "scheduled".

Scope

The issue's third item — detecting the org monthly spend limit — is not in this PR, per the triage on the issue: spend-limit detection does not exist anywhere in the repo (no account.rate-limits.updated path reaches it) and a monthly cap probably should not ride the backoff ladder at all. Worth its own issue. This PR is the two changes that stop the loss.

Verification

  • vp test run apps/server/src/t3x/autoResume — 76 passed (8 files).
  • The 4 new behavioural tests were each verified to fail against the old code (re-added the branch and disabled the supersede: exactly those 4 failed, 33 passed). No vacuous tests.
  • Reactor integration suite run 12× for flake — 0 failures. That file has a documented flake history, and the two new integration tests use the existing condition-based settleUntil/advanceUntil helpers rather than fixed spins.
  • vp lint apps/server/src/t3x/autoResume — clean. tsgo --noEmit -p apps/server — 0 errors.

Seam cost

None. Every changed source file is fork-owned under apps/server/src/t3x/autoResume/; git log <merge-base>..upstream/main -- apps/server/src/t3x/ is empty. No docs/t3x/SEAMS.md row changes.

docs/t3x/loop/DESIGN.md §6 gets a dated note: guard #9 still stands (nudging a thread inside a usage-limit window is pointless), but the hazard it was the last line of defence against is gone.

Closes#39.
Two defects compounded into a permanently stranded thread: a usage limit
armed a resume, the user typed "keep going" at the banner, that message
was itself rejected so it started nothing, and the wake tick then threw
the arm away as `user-took-over`. Net pending afterwards: zero. Measured
on the reporting install at 4 of 17 armed resumes (~24%) since #6.
1. `guards.ts` — drop the `user-took-over` branch. It is the same
negative-evidence mistake #6 fixed on the line below it: "a message
exists that wasn't there when we armed" is not evidence the human took
the wheel, and in practice it is the opposite signal. Everything the
branch reached for is still covered — `progressing` (actively driving),
`awaiting-input` (blocked on a prompt), `thread-advanced` (a different
live turn), and the per-thread switch for "stop entirely".
`baseline.newestUserMessageId` stays: it is part of the persisted
record, it is re-captured on every (re)schedule, and it is what makes a
stranded arm diagnosable from the state file.
2. `decide.ts` — narrow `already-pending`. A rejection naming a CONCRETE
reset time later than the pending one now supersedes it instead of
being dropped, so a `seven_day` limit landing on an armed `five_hour`
no longer leaves the arm firing into a window that is still shut.
Restricted to `windowOpensInFuture` on purpose: a ladder-derived time
is `nowMs + delay`, so it is later on every telemetry re-emit, and
letting those supersede would push the arm out forever and flood the
timeline. An earlier reset never supersedes either — the existing arm
is already the conservative choice.
A supersede re-writes the record with a fresh `captureBaseline`, so it
is also the re-arm path, and posts a distinct
`t3x.auto-resume.rescheduled` note rather than a second "scheduled".
The issue's third item — detecting the org monthly spend limit — is a
genuinely separate feature (no `account.rate-limits.updated` path reaches
it at all) and is deliberately not in scope here, per the triage.
Tests: 4 new behavioural cases, each verified to fail against the old
code — guards (a newer user message does not cancel; it still cancels
when that message is actually being worked on), decide (supersede,
earlier-window no-op, ladder-churn guard), and two reactor integration
tests covering the reported shape end to end. Reactor suite run 12x for
flake. `docs/t3x/loop/DESIGN.md` §6 gets a dated note: guard #9 still
stands, but the hazard it was defending against is gone.
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 7860f62f-a45e-4a99-b746-f40c66b5a874

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@radroid

Copy link
Copy Markdown
OwnerAuthor

Runtime verification — A/B on a real thread, not just unit tests

Ran this end to end in an isolated dev environment (worktree-local .t3, real claudeAgent session, real Claude turns). The fixture is the incident from the issue: a resume armed while e833f15e was the newest user message, then a newer user message fada9685 arrives, and no new turn runs — because in the real incident that message was itself rejected by a limit.

Identical seeded t3x-auto-resume.json in both runs. Only guards.ts differs.

With this PR's code — fires:

t3x-auto-resume.json -> "pending": null, "firedAtMs": [1786453083849]
activity -> t3x.auto-resume.resumed — "Resuming now (attempt 1 of 10)."
messages -> t3x-auto-resume:76a37e53-… role=user text="continue"
assistant:1b6114cb-… "41 42 43 44 45 …"

Claude genuinely picked the work back up — the thread had been counting to 40, and the resumed turn continued from 41.

With the user-took-over branch restored — cancels:

t3x-auto-resume.json -> "pending": null, "firedAtMs": []
activity -> t3x.auto-resume.cancelled — "Auto-resume cancelled: user-took-over."

Same fixture, same thread, nothing dispatched, attempt not even counted. That is the ~24% loss the issue measured, reproduced on demand and then fixed.

guards.ts was restored afterwards; git status is clean and the only remaining occurrence of the string user-took-over in the file is the explanatory comment.

Two side observations from the same session

1. thread-advanced still works — and I confirmed it by accident. My first attempt seeded baseline.latestTurnId: null while a turn had since run, and it correctly cancelled with thread-advanced. Removing the user-message branch did not weaken the guard beside it.

2. projection_threads.latest_turn_id does NOT appear to clear when a turn settles, which is contrary to what the #6 comment block in this file asserts:

latestTurn is joined on projection_threads.latest_turn_id, which is populated only while a turn is active

Observed directly: after the first turn completed, projection_thread_sessions.status = 'ready' and active_turn_id = NULL, but projection_threads.latest_turn_id was still 1592ce47-….

This does not change the correctness of either guard — #6's "require positive evidence" rule is right whether the column clears or not, and the existing regression test still pins the null case defensively. But the stated mechanism in that comment does not match this build, and since that comment is the documentation for why the guard looks the way it does, it is worth a proper look. Not fixing it in this PR; flagging it so the next person to read that comment does not trust it blindly.

@radroid
radroid merged commit a5e2abc into mainAug 12, 2026
2 checks passed
@radroid
radroid deleted the t3x/auto-resume-user-message branch August 12, 2026 02:06
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: auto-resume is cancelled as "user-took-over" by any user message — including one that was itself rejected by a limit

1 participant

@radroid