fix(codex): recover turns after usage limits - #7308

Open
estensen wants to merge 1 commit into
pingdotgg:mainfrom
estensen:fix/codex-usage-limit-failover
Open

fix(codex): recover turns after usage limits#7308
estensen wants to merge 1 commit into
pingdotgg:mainfrom
estensen:fix/codex-usage-limit-failover

Conversation

@estensen

@estensenestensen commented Aug 17, 2026

Copy link
Copy Markdown

Decision

Codex App Server keeps one account per process, so a T3 thread stops when that account runs out. This change restarts the configured executable, strictly resumes the same provider thread, and retries a safe failed turn once. Account-selecting launchers can then continue the conversation with available quota. Other providers, contracts, clients, and UI behavior do not change.

What Changed

  • Retain each in-flight Codex turn's mapped input and launch settings in the adapter.
  • Detect the structured usageLimitExceeded completion emitted by Codex App Server.
  • Restart the configured executable and require the original provider thread. Never fall back to a fresh thread during recovery.
  • Read the resumed thread. Roll back only a failed user-only turn, then retry the exact input once.
  • Preserve normal event forwarding, including the original usage-limit failure.
  • Document automatic Codex recovery for account-selecting executables.

Production Behavior And Risk

The recovery path runs only for the first usage-limit failure. It refuses to retry after assistant output, tool requests, commands, hooks, diffs, plans, collaboration activity, or queued turns. It verifies the resumed provider thread and requested model before retrying.

Manual session starts, stops, and interrupts cancel an in-flight recovery. A recovery process cannot overwrite a newer user-started session. A second usage-limit failure remains visible and does not launch a third process.

The worst case is a configured executable that restarts but cannot resume the original provider thread. T3 Code closes that recovery session and exposes the existing failure. It never creates a disconnected conversation. T3 Code adds no migrations or new persistent state. Its only provider mutation is rolling back the failed user-only turn before retrying it.

Reviewer Focus

  • Verify event ordering when turn/completed arrives before the turn/start response.
  • Verify that the side-effect guard is fail-closed for provider activity while allowing Codex's user-message lifecycle events.
  • Verify cancellation when manual session replacement races automatic recovery.

Test Plan

  • vp test run apps/server/src/provider/Layers/CodexAdapter.test.ts apps/server/src/provider/Layers/CodexSessionRuntime.test.ts — 57 tests passed.
  • Targeted vp lint for the four changed TypeScript files.
  • vp run --filter t3 typecheck — passed with existing suggestions in unrelated files.
  • vp fmt --check for all five changed files.

Checklist

  • This PR is small and focused
  • I explained what changed and why

Built with GPT-5.6 Sol in Codex.


Note

Medium Risk
Changes Codex session lifecycle (process restart, thread resume, rollback, and async recovery) on a critical provider path, but scope is narrow with fail-closed guards and extensive tests.

Overview
Codex now handles usageLimitExceeded turn failures by giving account-selecting launchers one automatic retry without leaving the conversation.

When a turn fails with that error and produced no assistant/tool activity (only user-message lifecycle events are allowed), the adapter restarts the configured Codex process, opens the session with requireResume so thread/resume cannot fall back to a fresh thread, verifies the same provider thread and model, optionally rolls back a user-only failed turn, then re-sends the exact prompt once. The original failure still flows through the event stream; a second limit hit is not retried again.

Guards and races: pending-turn tracking handles turn/completed arriving before turn/start; unsafeTurnIds blocks retry after any real provider output (including items only on completion); queued turns, manual startSession/stopSession/interruptTurn, and in-flight recovery use recovery tokens so recovery cannot replace a newer user session. sendTurn is rejected while recovery is active.

Tests cover happy path, race with manual session start, no-retry with activity, and strict resume behavior. User docs describe when auto-recovery runs and when it stops.

Reviewed by Cursor Bugbot for commit 92a148c. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Add automatic recovery from usage-limit failures in CodexAdapter

  • When a Codex turn fails with usageLimitExceeded and no provider activity has occurred in that turn, the adapter automatically restarts the runtime, resumes the same thread, and retries the turn exactly once.
  • A new requireResume flag in CodexSessionRuntimeOptions prevents fallback to a fresh thread when strict resume is required; recoverable resume errors are surfaced instead of silently starting over.
  • sendTurn is blocked while recovery is in progress or a recovery token is active; tokens are cleared on explicit starts, interrupts, and stops.
  • New tests in CodexAdapter.test.ts verify recovery semantics using observable TestQueue utilities for deterministic event ordering.
  • Risk: Any provider activity (non-userMessage items, tool calls, hooks, diffs) disables retry for that turn, aborting recovery silently.
📊 Macroscope summarized 92a148c. 3 files reviewed, 0 issues evaluated, 0 issues filtered, 0 comments posted

🗂️ Filtered Issues

No issues evaluated.

@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 5024b059-11d5-4ed6-974c-a7656c1888cb

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:L 100-499 changed lines (additions + deletions). labels Aug 17, 2026
return;
}

const resumedSession = yield* startSessionInternal(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟠 HighLayers/CodexAdapter.ts:1965

A stale usage-limit recovery can stop a newer manual session and leave the thread with no active session. recoverUsageLimitTurn calls startSessionInternal before validating that input.token is still current; if startSession has already deleted the token and installed a replacement, the recovery replaces and then stops that replacement, and finally closes its own runtime. Validate the token before replacing the registered session and atomically guard the replacement.

🤖 Copy this AI Prompt to have your agent fix this:
In file @apps/server/src/provider/Layers/CodexAdapter.ts around line 1965:
A stale usage-limit recovery can stop a newer manual session and leave the thread with no active session. `recoverUsageLimitTurn` calls `startSessionInternal` before validating that `input.token` is still current; if `startSession` has already deleted the token and installed a replacement, the recovery replaces and then stops that replacement, and finally closes its own runtime. Validate the token before replacing the registered session and atomically guard the replacement.

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Want fixes drafted automatically? Bugbot Autofix can create code changes for findings. A team admin can enable Autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 92a148c. Configure here.

failedTurnId: input.failedTurnId,
},
);
return;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Aborted recovery leaves session open

Medium Severity

When post-resume readThread finds provider activity on the failed turn, recoverUsageLimitTurn returns without calling stopSessionInternal. By then startSessionInternal has already stopped the original session and registered the recovery process. Sibling abort paths for wrong thread or model do close that session. An account-selecting executable can therefore leave the thread on a replaced process after refusing to retry, instead of closing the recovery session and surfacing the original usage-limit failure.

Additional Locations (1)
Fix in CursorFix in Web

Reviewed by Cursor Bugbot for commit 92a148c. Configure here.

@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Needs human review

1 blocking correctness issue found. This PR introduces a new automatic recovery feature for usage-limit failures with significant state management complexity (recovery tokens, pending turn tracking, deferred completions). A High-severity finding identifies a potential race condition where stale recovery can interfere with newer sessions.

You can customize Macroscope's approvability policy. Learn more.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L100-499 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@estensen
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

fix(codex): recover turns after usage limits - #7308

Open
estensen wants to merge 1 commit into
pingdotgg:mainfrom
estensen:fix/codex-usage-limit-failover
Open

fix(codex): recover turns after usage limits#7308
estensen wants to merge 1 commit into
pingdotgg:mainfrom
estensen:fix/codex-usage-limit-failover

Conversation

@estensen

@estensenestensen commented Aug 17, 2026

Copy link
Copy Markdown

Decision

Codex App Server keeps one account per process, so a T3 thread stops when that account runs out. This change restarts the configured executable, strictly resumes the same provider thread, and retries a safe failed turn once. Account-selecting launchers can then continue the conversation with available quota. Other providers, contracts, clients, and UI behavior do not change.

What Changed

  • Retain each in-flight Codex turn's mapped input and launch settings in the adapter.
  • Detect the structured usageLimitExceeded completion emitted by Codex App Server.
  • Restart the configured executable and require the original provider thread. Never fall back to a fresh thread during recovery.
  • Read the resumed thread. Roll back only a failed user-only turn, then retry the exact input once.
  • Preserve normal event forwarding, including the original usage-limit failure.
  • Document automatic Codex recovery for account-selecting executables.

Production Behavior And Risk

The recovery path runs only for the first usage-limit failure. It refuses to retry after assistant output, tool requests, commands, hooks, diffs, plans, collaboration activity, or queued turns. It verifies the resumed provider thread and requested model before retrying.

Manual session starts, stops, and interrupts cancel an in-flight recovery. A recovery process cannot overwrite a newer user-started session. A second usage-limit failure remains visible and does not launch a third process.

The worst case is a configured executable that restarts but cannot resume the original provider thread. T3 Code closes that recovery session and exposes the existing failure. It never creates a disconnected conversation. T3 Code adds no migrations or new persistent state. Its only provider mutation is rolling back the failed user-only turn before retrying it.

Reviewer Focus

  • Verify event ordering when turn/completed arrives before the turn/start response.
  • Verify that the side-effect guard is fail-closed for provider activity while allowing Codex's user-message lifecycle events.
  • Verify cancellation when manual session replacement races automatic recovery.

Test Plan

  • vp test run apps/server/src/provider/Layers/CodexAdapter.test.ts apps/server/src/provider/Layers/CodexSessionRuntime.test.ts — 57 tests passed.
  • Targeted vp lint for the four changed TypeScript files.
  • vp run --filter t3 typecheck — passed with existing suggestions in unrelated files.
  • vp fmt --check for all five changed files.

Checklist

  • This PR is small and focused
  • I explained what changed and why

Built with GPT-5.6 Sol in Codex.


Note

Medium Risk
Changes Codex session lifecycle (process restart, thread resume, rollback, and async recovery) on a critical provider path, but scope is narrow with fail-closed guards and extensive tests.

Overview
Codex now handles usageLimitExceeded turn failures by giving account-selecting launchers one automatic retry without leaving the conversation.

When a turn fails with that error and produced no assistant/tool activity (only user-message lifecycle events are allowed), the adapter restarts the configured Codex process, opens the session with requireResume so thread/resume cannot fall back to a fresh thread, verifies the same provider thread and model, optionally rolls back a user-only failed turn, then re-sends the exact prompt once. The original failure still flows through the event stream; a second limit hit is not retried again.

Guards and races: pending-turn tracking handles turn/completed arriving before turn/start; unsafeTurnIds blocks retry after any real provider output (including items only on completion); queued turns, manual startSession/stopSession/interruptTurn, and in-flight recovery use recovery tokens so recovery cannot replace a newer user session. sendTurn is rejected while recovery is active.

Tests cover happy path, race with manual session start, no-retry with activity, and strict resume behavior. User docs describe when auto-recovery runs and when it stops.

Reviewed by Cursor Bugbot for commit 92a148c. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Add automatic recovery from usage-limit failures in CodexAdapter

  • When a Codex turn fails with usageLimitExceeded and no provider activity has occurred in that turn, the adapter automatically restarts the runtime, resumes the same thread, and retries the turn exactly once.
  • A new requireResume flag in CodexSessionRuntimeOptions prevents fallback to a fresh thread when strict resume is required; recoverable resume errors are surfaced instead of silently starting over.
  • sendTurn is blocked while recovery is in progress or a recovery token is active; tokens are cleared on explicit starts, interrupts, and stops.
  • New tests in CodexAdapter.test.ts verify recovery semantics using observable TestQueue utilities for deterministic event ordering.
  • Risk: Any provider activity (non-userMessage items, tool calls, hooks, diffs) disables retry for that turn, aborting recovery silently.
📊 Macroscope summarized 92a148c. 3 files reviewed, 0 issues evaluated, 0 issues filtered, 0 comments posted

🗂️ Filtered Issues

No issues evaluated.

@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 5024b059-11d5-4ed6-974c-a7656c1888cb

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:L 100-499 changed lines (additions + deletions). labels Aug 17, 2026
return;
}

const resumedSession = yield* startSessionInternal(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟠 HighLayers/CodexAdapter.ts:1965

A stale usage-limit recovery can stop a newer manual session and leave the thread with no active session. recoverUsageLimitTurn calls startSessionInternal before validating that input.token is still current; if startSession has already deleted the token and installed a replacement, the recovery replaces and then stops that replacement, and finally closes its own runtime. Validate the token before replacing the registered session and atomically guard the replacement.

🤖 Copy this AI Prompt to have your agent fix this:
In file @apps/server/src/provider/Layers/CodexAdapter.ts around line 1965:
A stale usage-limit recovery can stop a newer manual session and leave the thread with no active session. `recoverUsageLimitTurn` calls `startSessionInternal` before validating that `input.token` is still current; if `startSession` has already deleted the token and installed a replacement, the recovery replaces and then stops that replacement, and finally closes its own runtime. Validate the token before replacing the registered session and atomically guard the replacement.

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Want fixes drafted automatically? Bugbot Autofix can create code changes for findings. A team admin can enable Autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 92a148c. Configure here.

failedTurnId: input.failedTurnId,
},
);
return;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Aborted recovery leaves session open

Medium Severity

When post-resume readThread finds provider activity on the failed turn, recoverUsageLimitTurn returns without calling stopSessionInternal. By then startSessionInternal has already stopped the original session and registered the recovery process. Sibling abort paths for wrong thread or model do close that session. An account-selecting executable can therefore leave the thread on a replaced process after refusing to retry, instead of closing the recovery session and surfacing the original usage-limit failure.

Additional Locations (1)
Fix in CursorFix in Web

Reviewed by Cursor Bugbot for commit 92a148c. Configure here.

@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Needs human review

1 blocking correctness issue found. This PR introduces a new automatic recovery feature for usage-limit failures with significant state management complexity (recovery tokens, pending turn tracking, deferred completions). A High-severity finding identifies a potential race condition where stale recovery can interfere with newer sessions.

You can customize Macroscope's approvability policy. Learn more.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L100-499 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@estensen
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(codex): recover turns after usage limits - #7308

Open
estensen wants to merge 1 commit into
pingdotgg:mainfrom
estensen:fix/codex-usage-limit-failover
Open

fix(codex): recover turns after usage limits#7308
estensen wants to merge 1 commit into
pingdotgg:mainfrom
estensen:fix/codex-usage-limit-failover

Conversation

@estensen

@estensenestensen commented Aug 17, 2026

Copy link
Copy Markdown

Decision

Codex App Server keeps one account per process, so a T3 thread stops when that account runs out. This change restarts the configured executable, strictly resumes the same provider thread, and retries a safe failed turn once. Account-selecting launchers can then continue the conversation with available quota. Other providers, contracts, clients, and UI behavior do not change.

What Changed

  • Retain each in-flight Codex turn's mapped input and launch settings in the adapter.
  • Detect the structured usageLimitExceeded completion emitted by Codex App Server.
  • Restart the configured executable and require the original provider thread. Never fall back to a fresh thread during recovery.
  • Read the resumed thread. Roll back only a failed user-only turn, then retry the exact input once.
  • Preserve normal event forwarding, including the original usage-limit failure.
  • Document automatic Codex recovery for account-selecting executables.

Production Behavior And Risk

The recovery path runs only for the first usage-limit failure. It refuses to retry after assistant output, tool requests, commands, hooks, diffs, plans, collaboration activity, or queued turns. It verifies the resumed provider thread and requested model before retrying.

Manual session starts, stops, and interrupts cancel an in-flight recovery. A recovery process cannot overwrite a newer user-started session. A second usage-limit failure remains visible and does not launch a third process.

The worst case is a configured executable that restarts but cannot resume the original provider thread. T3 Code closes that recovery session and exposes the existing failure. It never creates a disconnected conversation. T3 Code adds no migrations or new persistent state. Its only provider mutation is rolling back the failed user-only turn before retrying it.

Reviewer Focus

  • Verify event ordering when turn/completed arrives before the turn/start response.
  • Verify that the side-effect guard is fail-closed for provider activity while allowing Codex's user-message lifecycle events.
  • Verify cancellation when manual session replacement races automatic recovery.

Test Plan

  • vp test run apps/server/src/provider/Layers/CodexAdapter.test.ts apps/server/src/provider/Layers/CodexSessionRuntime.test.ts — 57 tests passed.
  • Targeted vp lint for the four changed TypeScript files.
  • vp run --filter t3 typecheck — passed with existing suggestions in unrelated files.
  • vp fmt --check for all five changed files.

Checklist

  • This PR is small and focused
  • I explained what changed and why

Built with GPT-5.6 Sol in Codex.


Note

Medium Risk
Changes Codex session lifecycle (process restart, thread resume, rollback, and async recovery) on a critical provider path, but scope is narrow with fail-closed guards and extensive tests.

Overview
Codex now handles usageLimitExceeded turn failures by giving account-selecting launchers one automatic retry without leaving the conversation.

When a turn fails with that error and produced no assistant/tool activity (only user-message lifecycle events are allowed), the adapter restarts the configured Codex process, opens the session with requireResume so thread/resume cannot fall back to a fresh thread, verifies the same provider thread and model, optionally rolls back a user-only failed turn, then re-sends the exact prompt once. The original failure still flows through the event stream; a second limit hit is not retried again.

Guards and races: pending-turn tracking handles turn/completed arriving before turn/start; unsafeTurnIds blocks retry after any real provider output (including items only on completion); queued turns, manual startSession/stopSession/interruptTurn, and in-flight recovery use recovery tokens so recovery cannot replace a newer user session. sendTurn is rejected while recovery is active.

Tests cover happy path, race with manual session start, no-retry with activity, and strict resume behavior. User docs describe when auto-recovery runs and when it stops.

Reviewed by Cursor Bugbot for commit 92a148c. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Add automatic recovery from usage-limit failures in CodexAdapter

  • When a Codex turn fails with usageLimitExceeded and no provider activity has occurred in that turn, the adapter automatically restarts the runtime, resumes the same thread, and retries the turn exactly once.
  • A new requireResume flag in CodexSessionRuntimeOptions prevents fallback to a fresh thread when strict resume is required; recoverable resume errors are surfaced instead of silently starting over.
  • sendTurn is blocked while recovery is in progress or a recovery token is active; tokens are cleared on explicit starts, interrupts, and stops.
  • New tests in CodexAdapter.test.ts verify recovery semantics using observable TestQueue utilities for deterministic event ordering.
  • Risk: Any provider activity (non-userMessage items, tool calls, hooks, diffs) disables retry for that turn, aborting recovery silently.
📊 Macroscope summarized 92a148c. 3 files reviewed, 0 issues evaluated, 0 issues filtered, 0 comments posted

🗂️ Filtered Issues

No issues evaluated.

@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 5024b059-11d5-4ed6-974c-a7656c1888cb

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:L 100-499 changed lines (additions + deletions). labels Aug 17, 2026
return;
}

const resumedSession = yield* startSessionInternal(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟠 HighLayers/CodexAdapter.ts:1965

A stale usage-limit recovery can stop a newer manual session and leave the thread with no active session. recoverUsageLimitTurn calls startSessionInternal before validating that input.token is still current; if startSession has already deleted the token and installed a replacement, the recovery replaces and then stops that replacement, and finally closes its own runtime. Validate the token before replacing the registered session and atomically guard the replacement.

🤖 Copy this AI Prompt to have your agent fix this:
In file @apps/server/src/provider/Layers/CodexAdapter.ts around line 1965:
A stale usage-limit recovery can stop a newer manual session and leave the thread with no active session. `recoverUsageLimitTurn` calls `startSessionInternal` before validating that `input.token` is still current; if `startSession` has already deleted the token and installed a replacement, the recovery replaces and then stops that replacement, and finally closes its own runtime. Validate the token before replacing the registered session and atomically guard the replacement.

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Want fixes drafted automatically? Bugbot Autofix can create code changes for findings. A team admin can enable Autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 92a148c. Configure here.

failedTurnId: input.failedTurnId,
},
);
return;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Aborted recovery leaves session open

Medium Severity

When post-resume readThread finds provider activity on the failed turn, recoverUsageLimitTurn returns without calling stopSessionInternal. By then startSessionInternal has already stopped the original session and registered the recovery process. Sibling abort paths for wrong thread or model do close that session. An account-selecting executable can therefore leave the thread on a replaced process after refusing to retry, instead of closing the recovery session and surfacing the original usage-limit failure.

Additional Locations (1)
Fix in CursorFix in Web

Reviewed by Cursor Bugbot for commit 92a148c. Configure here.

@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Needs human review

1 blocking correctness issue found. This PR introduces a new automatic recovery feature for usage-limit failures with significant state management complexity (recovery tokens, pending turn tracking, deferred completions). A High-severity finding identifies a potential race condition where stale recovery can interfere with newer sessions.

You can customize Macroscope's approvability policy. Learn more.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L100-499 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@estensen
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(codex): recover turns after usage limits - #7308

Open
estensen wants to merge 1 commit into
pingdotgg:mainfrom
estensen:fix/codex-usage-limit-failover
Open

fix(codex): recover turns after usage limits#7308
estensen wants to merge 1 commit into
pingdotgg:mainfrom
estensen:fix/codex-usage-limit-failover

Conversation

@estensen

@estensenestensen commented Aug 17, 2026

Copy link
Copy Markdown

Decision

Codex App Server keeps one account per process, so a T3 thread stops when that account runs out. This change restarts the configured executable, strictly resumes the same provider thread, and retries a safe failed turn once. Account-selecting launchers can then continue the conversation with available quota. Other providers, contracts, clients, and UI behavior do not change.

What Changed

  • Retain each in-flight Codex turn's mapped input and launch settings in the adapter.
  • Detect the structured usageLimitExceeded completion emitted by Codex App Server.
  • Restart the configured executable and require the original provider thread. Never fall back to a fresh thread during recovery.
  • Read the resumed thread. Roll back only a failed user-only turn, then retry the exact input once.
  • Preserve normal event forwarding, including the original usage-limit failure.
  • Document automatic Codex recovery for account-selecting executables.

Production Behavior And Risk

The recovery path runs only for the first usage-limit failure. It refuses to retry after assistant output, tool requests, commands, hooks, diffs, plans, collaboration activity, or queued turns. It verifies the resumed provider thread and requested model before retrying.

Manual session starts, stops, and interrupts cancel an in-flight recovery. A recovery process cannot overwrite a newer user-started session. A second usage-limit failure remains visible and does not launch a third process.

The worst case is a configured executable that restarts but cannot resume the original provider thread. T3 Code closes that recovery session and exposes the existing failure. It never creates a disconnected conversation. T3 Code adds no migrations or new persistent state. Its only provider mutation is rolling back the failed user-only turn before retrying it.

Reviewer Focus

  • Verify event ordering when turn/completed arrives before the turn/start response.
  • Verify that the side-effect guard is fail-closed for provider activity while allowing Codex's user-message lifecycle events.
  • Verify cancellation when manual session replacement races automatic recovery.

Test Plan

  • vp test run apps/server/src/provider/Layers/CodexAdapter.test.ts apps/server/src/provider/Layers/CodexSessionRuntime.test.ts — 57 tests passed.
  • Targeted vp lint for the four changed TypeScript files.
  • vp run --filter t3 typecheck — passed with existing suggestions in unrelated files.
  • vp fmt --check for all five changed files.

Checklist

  • This PR is small and focused
  • I explained what changed and why

Built with GPT-5.6 Sol in Codex.


Note

Medium Risk
Changes Codex session lifecycle (process restart, thread resume, rollback, and async recovery) on a critical provider path, but scope is narrow with fail-closed guards and extensive tests.

Overview
Codex now handles usageLimitExceeded turn failures by giving account-selecting launchers one automatic retry without leaving the conversation.

When a turn fails with that error and produced no assistant/tool activity (only user-message lifecycle events are allowed), the adapter restarts the configured Codex process, opens the session with requireResume so thread/resume cannot fall back to a fresh thread, verifies the same provider thread and model, optionally rolls back a user-only failed turn, then re-sends the exact prompt once. The original failure still flows through the event stream; a second limit hit is not retried again.

Guards and races: pending-turn tracking handles turn/completed arriving before turn/start; unsafeTurnIds blocks retry after any real provider output (including items only on completion); queued turns, manual startSession/stopSession/interruptTurn, and in-flight recovery use recovery tokens so recovery cannot replace a newer user session. sendTurn is rejected while recovery is active.

Tests cover happy path, race with manual session start, no-retry with activity, and strict resume behavior. User docs describe when auto-recovery runs and when it stops.

Reviewed by Cursor Bugbot for commit 92a148c. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Add automatic recovery from usage-limit failures in CodexAdapter

  • When a Codex turn fails with usageLimitExceeded and no provider activity has occurred in that turn, the adapter automatically restarts the runtime, resumes the same thread, and retries the turn exactly once.
  • A new requireResume flag in CodexSessionRuntimeOptions prevents fallback to a fresh thread when strict resume is required; recoverable resume errors are surfaced instead of silently starting over.
  • sendTurn is blocked while recovery is in progress or a recovery token is active; tokens are cleared on explicit starts, interrupts, and stops.
  • New tests in CodexAdapter.test.ts verify recovery semantics using observable TestQueue utilities for deterministic event ordering.
  • Risk: Any provider activity (non-userMessage items, tool calls, hooks, diffs) disables retry for that turn, aborting recovery silently.
📊 Macroscope summarized 92a148c. 3 files reviewed, 0 issues evaluated, 0 issues filtered, 0 comments posted

🗂️ Filtered Issues

No issues evaluated.

@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 5024b059-11d5-4ed6-974c-a7656c1888cb

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:L 100-499 changed lines (additions + deletions). labels Aug 17, 2026
return;
}

const resumedSession = yield* startSessionInternal(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟠 HighLayers/CodexAdapter.ts:1965

A stale usage-limit recovery can stop a newer manual session and leave the thread with no active session. recoverUsageLimitTurn calls startSessionInternal before validating that input.token is still current; if startSession has already deleted the token and installed a replacement, the recovery replaces and then stops that replacement, and finally closes its own runtime. Validate the token before replacing the registered session and atomically guard the replacement.

🤖 Copy this AI Prompt to have your agent fix this:
In file @apps/server/src/provider/Layers/CodexAdapter.ts around line 1965:
A stale usage-limit recovery can stop a newer manual session and leave the thread with no active session. `recoverUsageLimitTurn` calls `startSessionInternal` before validating that `input.token` is still current; if `startSession` has already deleted the token and installed a replacement, the recovery replaces and then stops that replacement, and finally closes its own runtime. Validate the token before replacing the registered session and atomically guard the replacement.

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Want fixes drafted automatically? Bugbot Autofix can create code changes for findings. A team admin can enable Autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 92a148c. Configure here.

failedTurnId: input.failedTurnId,
},
);
return;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Aborted recovery leaves session open

Medium Severity

When post-resume readThread finds provider activity on the failed turn, recoverUsageLimitTurn returns without calling stopSessionInternal. By then startSessionInternal has already stopped the original session and registered the recovery process. Sibling abort paths for wrong thread or model do close that session. An account-selecting executable can therefore leave the thread on a replaced process after refusing to retry, instead of closing the recovery session and surfacing the original usage-limit failure.

Additional Locations (1)
Fix in CursorFix in Web

Reviewed by Cursor Bugbot for commit 92a148c. Configure here.

@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Needs human review

1 blocking correctness issue found. This PR introduces a new automatic recovery feature for usage-limit failures with significant state management complexity (recovery tokens, pending turn tracking, deferred completions). A High-severity finding identifies a potential race condition where stale recovery can interfere with newer sessions.

You can customize Macroscope's approvability policy. Learn more.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L100-499 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@estensen
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

fix(codex): recover turns after usage limits - #7308

Open
estensen wants to merge 1 commit into
pingdotgg:mainfrom
estensen:fix/codex-usage-limit-failover
Open

fix(codex): recover turns after usage limits#7308
estensen wants to merge 1 commit into
pingdotgg:mainfrom
estensen:fix/codex-usage-limit-failover

Conversation

@estensen

@estensenestensen commented Aug 17, 2026

Copy link
Copy Markdown

Decision

Codex App Server keeps one account per process, so a T3 thread stops when that account runs out. This change restarts the configured executable, strictly resumes the same provider thread, and retries a safe failed turn once. Account-selecting launchers can then continue the conversation with available quota. Other providers, contracts, clients, and UI behavior do not change.

What Changed

  • Retain each in-flight Codex turn's mapped input and launch settings in the adapter.
  • Detect the structured usageLimitExceeded completion emitted by Codex App Server.
  • Restart the configured executable and require the original provider thread. Never fall back to a fresh thread during recovery.
  • Read the resumed thread. Roll back only a failed user-only turn, then retry the exact input once.
  • Preserve normal event forwarding, including the original usage-limit failure.
  • Document automatic Codex recovery for account-selecting executables.

Production Behavior And Risk

The recovery path runs only for the first usage-limit failure. It refuses to retry after assistant output, tool requests, commands, hooks, diffs, plans, collaboration activity, or queued turns. It verifies the resumed provider thread and requested model before retrying.

Manual session starts, stops, and interrupts cancel an in-flight recovery. A recovery process cannot overwrite a newer user-started session. A second usage-limit failure remains visible and does not launch a third process.

The worst case is a configured executable that restarts but cannot resume the original provider thread. T3 Code closes that recovery session and exposes the existing failure. It never creates a disconnected conversation. T3 Code adds no migrations or new persistent state. Its only provider mutation is rolling back the failed user-only turn before retrying it.

Reviewer Focus

  • Verify event ordering when turn/completed arrives before the turn/start response.
  • Verify that the side-effect guard is fail-closed for provider activity while allowing Codex's user-message lifecycle events.
  • Verify cancellation when manual session replacement races automatic recovery.

Test Plan

  • vp test run apps/server/src/provider/Layers/CodexAdapter.test.ts apps/server/src/provider/Layers/CodexSessionRuntime.test.ts — 57 tests passed.
  • Targeted vp lint for the four changed TypeScript files.
  • vp run --filter t3 typecheck — passed with existing suggestions in unrelated files.
  • vp fmt --check for all five changed files.

Checklist

  • This PR is small and focused
  • I explained what changed and why

Built with GPT-5.6 Sol in Codex.


Note

Medium Risk
Changes Codex session lifecycle (process restart, thread resume, rollback, and async recovery) on a critical provider path, but scope is narrow with fail-closed guards and extensive tests.

Overview
Codex now handles usageLimitExceeded turn failures by giving account-selecting launchers one automatic retry without leaving the conversation.

When a turn fails with that error and produced no assistant/tool activity (only user-message lifecycle events are allowed), the adapter restarts the configured Codex process, opens the session with requireResume so thread/resume cannot fall back to a fresh thread, verifies the same provider thread and model, optionally rolls back a user-only failed turn, then re-sends the exact prompt once. The original failure still flows through the event stream; a second limit hit is not retried again.

Guards and races: pending-turn tracking handles turn/completed arriving before turn/start; unsafeTurnIds blocks retry after any real provider output (including items only on completion); queued turns, manual startSession/stopSession/interruptTurn, and in-flight recovery use recovery tokens so recovery cannot replace a newer user session. sendTurn is rejected while recovery is active.

Tests cover happy path, race with manual session start, no-retry with activity, and strict resume behavior. User docs describe when auto-recovery runs and when it stops.

Reviewed by Cursor Bugbot for commit 92a148c. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Add automatic recovery from usage-limit failures in CodexAdapter

  • When a Codex turn fails with usageLimitExceeded and no provider activity has occurred in that turn, the adapter automatically restarts the runtime, resumes the same thread, and retries the turn exactly once.
  • A new requireResume flag in CodexSessionRuntimeOptions prevents fallback to a fresh thread when strict resume is required; recoverable resume errors are surfaced instead of silently starting over.
  • sendTurn is blocked while recovery is in progress or a recovery token is active; tokens are cleared on explicit starts, interrupts, and stops.
  • New tests in CodexAdapter.test.ts verify recovery semantics using observable TestQueue utilities for deterministic event ordering.
  • Risk: Any provider activity (non-userMessage items, tool calls, hooks, diffs) disables retry for that turn, aborting recovery silently.
📊 Macroscope summarized 92a148c. 3 files reviewed, 0 issues evaluated, 0 issues filtered, 0 comments posted

🗂️ Filtered Issues

No issues evaluated.

@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 5024b059-11d5-4ed6-974c-a7656c1888cb

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:L 100-499 changed lines (additions + deletions). labels Aug 17, 2026
return;
}

const resumedSession = yield* startSessionInternal(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟠 HighLayers/CodexAdapter.ts:1965

A stale usage-limit recovery can stop a newer manual session and leave the thread with no active session. recoverUsageLimitTurn calls startSessionInternal before validating that input.token is still current; if startSession has already deleted the token and installed a replacement, the recovery replaces and then stops that replacement, and finally closes its own runtime. Validate the token before replacing the registered session and atomically guard the replacement.

🤖 Copy this AI Prompt to have your agent fix this:
In file @apps/server/src/provider/Layers/CodexAdapter.ts around line 1965:
A stale usage-limit recovery can stop a newer manual session and leave the thread with no active session. `recoverUsageLimitTurn` calls `startSessionInternal` before validating that `input.token` is still current; if `startSession` has already deleted the token and installed a replacement, the recovery replaces and then stops that replacement, and finally closes its own runtime. Validate the token before replacing the registered session and atomically guard the replacement.

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Want fixes drafted automatically? Bugbot Autofix can create code changes for findings. A team admin can enable Autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 92a148c. Configure here.

failedTurnId: input.failedTurnId,
},
);
return;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Aborted recovery leaves session open

Medium Severity

When post-resume readThread finds provider activity on the failed turn, recoverUsageLimitTurn returns without calling stopSessionInternal. By then startSessionInternal has already stopped the original session and registered the recovery process. Sibling abort paths for wrong thread or model do close that session. An account-selecting executable can therefore leave the thread on a replaced process after refusing to retry, instead of closing the recovery session and surfacing the original usage-limit failure.

Additional Locations (1)
Fix in CursorFix in Web

Reviewed by Cursor Bugbot for commit 92a148c. Configure here.

@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Needs human review

1 blocking correctness issue found. This PR introduces a new automatic recovery feature for usage-limit failures with significant state management complexity (recovery tokens, pending turn tracking, deferred completions). A High-severity finding identifies a potential race condition where stale recovery can interfere with newer sessions.

You can customize Macroscope's approvability policy. Learn more.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L100-499 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@estensen
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(codex): recover turns after usage limits - #7308

Open
estensen wants to merge 1 commit into
pingdotgg:mainfrom
estensen:fix/codex-usage-limit-failover
Open

fix(codex): recover turns after usage limits#7308
estensen wants to merge 1 commit into
pingdotgg:mainfrom
estensen:fix/codex-usage-limit-failover

Conversation

@estensen

@estensenestensen commented Aug 17, 2026

Copy link
Copy Markdown

Decision

Codex App Server keeps one account per process, so a T3 thread stops when that account runs out. This change restarts the configured executable, strictly resumes the same provider thread, and retries a safe failed turn once. Account-selecting launchers can then continue the conversation with available quota. Other providers, contracts, clients, and UI behavior do not change.

What Changed

  • Retain each in-flight Codex turn's mapped input and launch settings in the adapter.
  • Detect the structured usageLimitExceeded completion emitted by Codex App Server.
  • Restart the configured executable and require the original provider thread. Never fall back to a fresh thread during recovery.
  • Read the resumed thread. Roll back only a failed user-only turn, then retry the exact input once.
  • Preserve normal event forwarding, including the original usage-limit failure.
  • Document automatic Codex recovery for account-selecting executables.

Production Behavior And Risk

The recovery path runs only for the first usage-limit failure. It refuses to retry after assistant output, tool requests, commands, hooks, diffs, plans, collaboration activity, or queued turns. It verifies the resumed provider thread and requested model before retrying.

Manual session starts, stops, and interrupts cancel an in-flight recovery. A recovery process cannot overwrite a newer user-started session. A second usage-limit failure remains visible and does not launch a third process.

The worst case is a configured executable that restarts but cannot resume the original provider thread. T3 Code closes that recovery session and exposes the existing failure. It never creates a disconnected conversation. T3 Code adds no migrations or new persistent state. Its only provider mutation is rolling back the failed user-only turn before retrying it.

Reviewer Focus

  • Verify event ordering when turn/completed arrives before the turn/start response.
  • Verify that the side-effect guard is fail-closed for provider activity while allowing Codex's user-message lifecycle events.
  • Verify cancellation when manual session replacement races automatic recovery.

Test Plan

  • vp test run apps/server/src/provider/Layers/CodexAdapter.test.ts apps/server/src/provider/Layers/CodexSessionRuntime.test.ts — 57 tests passed.
  • Targeted vp lint for the four changed TypeScript files.
  • vp run --filter t3 typecheck — passed with existing suggestions in unrelated files.
  • vp fmt --check for all five changed files.

Checklist

  • This PR is small and focused
  • I explained what changed and why

Built with GPT-5.6 Sol in Codex.


Note

Medium Risk
Changes Codex session lifecycle (process restart, thread resume, rollback, and async recovery) on a critical provider path, but scope is narrow with fail-closed guards and extensive tests.

Overview
Codex now handles usageLimitExceeded turn failures by giving account-selecting launchers one automatic retry without leaving the conversation.

When a turn fails with that error and produced no assistant/tool activity (only user-message lifecycle events are allowed), the adapter restarts the configured Codex process, opens the session with requireResume so thread/resume cannot fall back to a fresh thread, verifies the same provider thread and model, optionally rolls back a user-only failed turn, then re-sends the exact prompt once. The original failure still flows through the event stream; a second limit hit is not retried again.

Guards and races: pending-turn tracking handles turn/completed arriving before turn/start; unsafeTurnIds blocks retry after any real provider output (including items only on completion); queued turns, manual startSession/stopSession/interruptTurn, and in-flight recovery use recovery tokens so recovery cannot replace a newer user session. sendTurn is rejected while recovery is active.

Tests cover happy path, race with manual session start, no-retry with activity, and strict resume behavior. User docs describe when auto-recovery runs and when it stops.

Reviewed by Cursor Bugbot for commit 92a148c. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Add automatic recovery from usage-limit failures in CodexAdapter

  • When a Codex turn fails with usageLimitExceeded and no provider activity has occurred in that turn, the adapter automatically restarts the runtime, resumes the same thread, and retries the turn exactly once.
  • A new requireResume flag in CodexSessionRuntimeOptions prevents fallback to a fresh thread when strict resume is required; recoverable resume errors are surfaced instead of silently starting over.
  • sendTurn is blocked while recovery is in progress or a recovery token is active; tokens are cleared on explicit starts, interrupts, and stops.
  • New tests in CodexAdapter.test.ts verify recovery semantics using observable TestQueue utilities for deterministic event ordering.
  • Risk: Any provider activity (non-userMessage items, tool calls, hooks, diffs) disables retry for that turn, aborting recovery silently.
📊 Macroscope summarized 92a148c. 3 files reviewed, 0 issues evaluated, 0 issues filtered, 0 comments posted

🗂️ Filtered Issues

No issues evaluated.

@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 5024b059-11d5-4ed6-974c-a7656c1888cb

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:L 100-499 changed lines (additions + deletions). labels Aug 17, 2026
return;
}

const resumedSession = yield* startSessionInternal(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟠 HighLayers/CodexAdapter.ts:1965

A stale usage-limit recovery can stop a newer manual session and leave the thread with no active session. recoverUsageLimitTurn calls startSessionInternal before validating that input.token is still current; if startSession has already deleted the token and installed a replacement, the recovery replaces and then stops that replacement, and finally closes its own runtime. Validate the token before replacing the registered session and atomically guard the replacement.

🤖 Copy this AI Prompt to have your agent fix this:
In file @apps/server/src/provider/Layers/CodexAdapter.ts around line 1965:
A stale usage-limit recovery can stop a newer manual session and leave the thread with no active session. `recoverUsageLimitTurn` calls `startSessionInternal` before validating that `input.token` is still current; if `startSession` has already deleted the token and installed a replacement, the recovery replaces and then stops that replacement, and finally closes its own runtime. Validate the token before replacing the registered session and atomically guard the replacement.

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Want fixes drafted automatically? Bugbot Autofix can create code changes for findings. A team admin can enable Autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 92a148c. Configure here.

failedTurnId: input.failedTurnId,
},
);
return;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Aborted recovery leaves session open

Medium Severity

When post-resume readThread finds provider activity on the failed turn, recoverUsageLimitTurn returns without calling stopSessionInternal. By then startSessionInternal has already stopped the original session and registered the recovery process. Sibling abort paths for wrong thread or model do close that session. An account-selecting executable can therefore leave the thread on a replaced process after refusing to retry, instead of closing the recovery session and surfacing the original usage-limit failure.

Additional Locations (1)
Fix in CursorFix in Web

Reviewed by Cursor Bugbot for commit 92a148c. Configure here.

@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Needs human review

1 blocking correctness issue found. This PR introduces a new automatic recovery feature for usage-limit failures with significant state management complexity (recovery tokens, pending turn tracking, deferred completions). A High-severity finding identifies a potential race condition where stale recovery can interfere with newer sessions.

You can customize Macroscope's approvability policy. Learn more.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L100-499 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@estensen
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(codex): recover turns after usage limits - #7308

Open
estensen wants to merge 1 commit into
pingdotgg:mainfrom
estensen:fix/codex-usage-limit-failover
Open

fix(codex): recover turns after usage limits#7308
estensen wants to merge 1 commit into
pingdotgg:mainfrom
estensen:fix/codex-usage-limit-failover

Conversation

@estensen

@estensenestensen commented Aug 17, 2026

Copy link
Copy Markdown

Decision

Codex App Server keeps one account per process, so a T3 thread stops when that account runs out. This change restarts the configured executable, strictly resumes the same provider thread, and retries a safe failed turn once. Account-selecting launchers can then continue the conversation with available quota. Other providers, contracts, clients, and UI behavior do not change.

What Changed

  • Retain each in-flight Codex turn's mapped input and launch settings in the adapter.
  • Detect the structured usageLimitExceeded completion emitted by Codex App Server.
  • Restart the configured executable and require the original provider thread. Never fall back to a fresh thread during recovery.
  • Read the resumed thread. Roll back only a failed user-only turn, then retry the exact input once.
  • Preserve normal event forwarding, including the original usage-limit failure.
  • Document automatic Codex recovery for account-selecting executables.

Production Behavior And Risk

The recovery path runs only for the first usage-limit failure. It refuses to retry after assistant output, tool requests, commands, hooks, diffs, plans, collaboration activity, or queued turns. It verifies the resumed provider thread and requested model before retrying.

Manual session starts, stops, and interrupts cancel an in-flight recovery. A recovery process cannot overwrite a newer user-started session. A second usage-limit failure remains visible and does not launch a third process.

The worst case is a configured executable that restarts but cannot resume the original provider thread. T3 Code closes that recovery session and exposes the existing failure. It never creates a disconnected conversation. T3 Code adds no migrations or new persistent state. Its only provider mutation is rolling back the failed user-only turn before retrying it.

Reviewer Focus

  • Verify event ordering when turn/completed arrives before the turn/start response.
  • Verify that the side-effect guard is fail-closed for provider activity while allowing Codex's user-message lifecycle events.
  • Verify cancellation when manual session replacement races automatic recovery.

Test Plan

  • vp test run apps/server/src/provider/Layers/CodexAdapter.test.ts apps/server/src/provider/Layers/CodexSessionRuntime.test.ts — 57 tests passed.
  • Targeted vp lint for the four changed TypeScript files.
  • vp run --filter t3 typecheck — passed with existing suggestions in unrelated files.
  • vp fmt --check for all five changed files.

Checklist

  • This PR is small and focused
  • I explained what changed and why

Built with GPT-5.6 Sol in Codex.


Note

Medium Risk
Changes Codex session lifecycle (process restart, thread resume, rollback, and async recovery) on a critical provider path, but scope is narrow with fail-closed guards and extensive tests.

Overview
Codex now handles usageLimitExceeded turn failures by giving account-selecting launchers one automatic retry without leaving the conversation.

When a turn fails with that error and produced no assistant/tool activity (only user-message lifecycle events are allowed), the adapter restarts the configured Codex process, opens the session with requireResume so thread/resume cannot fall back to a fresh thread, verifies the same provider thread and model, optionally rolls back a user-only failed turn, then re-sends the exact prompt once. The original failure still flows through the event stream; a second limit hit is not retried again.

Guards and races: pending-turn tracking handles turn/completed arriving before turn/start; unsafeTurnIds blocks retry after any real provider output (including items only on completion); queued turns, manual startSession/stopSession/interruptTurn, and in-flight recovery use recovery tokens so recovery cannot replace a newer user session. sendTurn is rejected while recovery is active.

Tests cover happy path, race with manual session start, no-retry with activity, and strict resume behavior. User docs describe when auto-recovery runs and when it stops.

Reviewed by Cursor Bugbot for commit 92a148c. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Add automatic recovery from usage-limit failures in CodexAdapter

  • When a Codex turn fails with usageLimitExceeded and no provider activity has occurred in that turn, the adapter automatically restarts the runtime, resumes the same thread, and retries the turn exactly once.
  • A new requireResume flag in CodexSessionRuntimeOptions prevents fallback to a fresh thread when strict resume is required; recoverable resume errors are surfaced instead of silently starting over.
  • sendTurn is blocked while recovery is in progress or a recovery token is active; tokens are cleared on explicit starts, interrupts, and stops.
  • New tests in CodexAdapter.test.ts verify recovery semantics using observable TestQueue utilities for deterministic event ordering.
  • Risk: Any provider activity (non-userMessage items, tool calls, hooks, diffs) disables retry for that turn, aborting recovery silently.
📊 Macroscope summarized 92a148c. 3 files reviewed, 0 issues evaluated, 0 issues filtered, 0 comments posted

🗂️ Filtered Issues

No issues evaluated.

@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 5024b059-11d5-4ed6-974c-a7656c1888cb

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:L 100-499 changed lines (additions + deletions). labels Aug 17, 2026
return;
}

const resumedSession = yield* startSessionInternal(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟠 HighLayers/CodexAdapter.ts:1965

A stale usage-limit recovery can stop a newer manual session and leave the thread with no active session. recoverUsageLimitTurn calls startSessionInternal before validating that input.token is still current; if startSession has already deleted the token and installed a replacement, the recovery replaces and then stops that replacement, and finally closes its own runtime. Validate the token before replacing the registered session and atomically guard the replacement.

🤖 Copy this AI Prompt to have your agent fix this:
In file @apps/server/src/provider/Layers/CodexAdapter.ts around line 1965:
A stale usage-limit recovery can stop a newer manual session and leave the thread with no active session. `recoverUsageLimitTurn` calls `startSessionInternal` before validating that `input.token` is still current; if `startSession` has already deleted the token and installed a replacement, the recovery replaces and then stops that replacement, and finally closes its own runtime. Validate the token before replacing the registered session and atomically guard the replacement.

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Want fixes drafted automatically? Bugbot Autofix can create code changes for findings. A team admin can enable Autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 92a148c. Configure here.

failedTurnId: input.failedTurnId,
},
);
return;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Aborted recovery leaves session open

Medium Severity

When post-resume readThread finds provider activity on the failed turn, recoverUsageLimitTurn returns without calling stopSessionInternal. By then startSessionInternal has already stopped the original session and registered the recovery process. Sibling abort paths for wrong thread or model do close that session. An account-selecting executable can therefore leave the thread on a replaced process after refusing to retry, instead of closing the recovery session and surfacing the original usage-limit failure.

Additional Locations (1)
Fix in CursorFix in Web

Reviewed by Cursor Bugbot for commit 92a148c. Configure here.

@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Needs human review

1 blocking correctness issue found. This PR introduces a new automatic recovery feature for usage-limit failures with significant state management complexity (recovery tokens, pending turn tracking, deferred completions). A High-severity finding identifies a potential race condition where stale recovery can interfere with newer sessions.

You can customize Macroscope's approvability policy. Learn more.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L100-499 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@estensen
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

fix(codex): recover turns after usage limits - #7308

Open
estensen wants to merge 1 commit into
pingdotgg:mainfrom
estensen:fix/codex-usage-limit-failover
Open

fix(codex): recover turns after usage limits#7308
estensen wants to merge 1 commit into
pingdotgg:mainfrom
estensen:fix/codex-usage-limit-failover

Conversation

@estensen

@estensenestensen commented Aug 17, 2026

Copy link
Copy Markdown

Decision

Codex App Server keeps one account per process, so a T3 thread stops when that account runs out. This change restarts the configured executable, strictly resumes the same provider thread, and retries a safe failed turn once. Account-selecting launchers can then continue the conversation with available quota. Other providers, contracts, clients, and UI behavior do not change.

What Changed

  • Retain each in-flight Codex turn's mapped input and launch settings in the adapter.
  • Detect the structured usageLimitExceeded completion emitted by Codex App Server.
  • Restart the configured executable and require the original provider thread. Never fall back to a fresh thread during recovery.
  • Read the resumed thread. Roll back only a failed user-only turn, then retry the exact input once.
  • Preserve normal event forwarding, including the original usage-limit failure.
  • Document automatic Codex recovery for account-selecting executables.

Production Behavior And Risk

The recovery path runs only for the first usage-limit failure. It refuses to retry after assistant output, tool requests, commands, hooks, diffs, plans, collaboration activity, or queued turns. It verifies the resumed provider thread and requested model before retrying.

Manual session starts, stops, and interrupts cancel an in-flight recovery. A recovery process cannot overwrite a newer user-started session. A second usage-limit failure remains visible and does not launch a third process.

The worst case is a configured executable that restarts but cannot resume the original provider thread. T3 Code closes that recovery session and exposes the existing failure. It never creates a disconnected conversation. T3 Code adds no migrations or new persistent state. Its only provider mutation is rolling back the failed user-only turn before retrying it.

Reviewer Focus

  • Verify event ordering when turn/completed arrives before the turn/start response.
  • Verify that the side-effect guard is fail-closed for provider activity while allowing Codex's user-message lifecycle events.
  • Verify cancellation when manual session replacement races automatic recovery.

Test Plan

  • vp test run apps/server/src/provider/Layers/CodexAdapter.test.ts apps/server/src/provider/Layers/CodexSessionRuntime.test.ts — 57 tests passed.
  • Targeted vp lint for the four changed TypeScript files.
  • vp run --filter t3 typecheck — passed with existing suggestions in unrelated files.
  • vp fmt --check for all five changed files.

Checklist

  • This PR is small and focused
  • I explained what changed and why

Built with GPT-5.6 Sol in Codex.


Note

Medium Risk
Changes Codex session lifecycle (process restart, thread resume, rollback, and async recovery) on a critical provider path, but scope is narrow with fail-closed guards and extensive tests.

Overview
Codex now handles usageLimitExceeded turn failures by giving account-selecting launchers one automatic retry without leaving the conversation.

When a turn fails with that error and produced no assistant/tool activity (only user-message lifecycle events are allowed), the adapter restarts the configured Codex process, opens the session with requireResume so thread/resume cannot fall back to a fresh thread, verifies the same provider thread and model, optionally rolls back a user-only failed turn, then re-sends the exact prompt once. The original failure still flows through the event stream; a second limit hit is not retried again.

Guards and races: pending-turn tracking handles turn/completed arriving before turn/start; unsafeTurnIds blocks retry after any real provider output (including items only on completion); queued turns, manual startSession/stopSession/interruptTurn, and in-flight recovery use recovery tokens so recovery cannot replace a newer user session. sendTurn is rejected while recovery is active.

Tests cover happy path, race with manual session start, no-retry with activity, and strict resume behavior. User docs describe when auto-recovery runs and when it stops.

Reviewed by Cursor Bugbot for commit 92a148c. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Add automatic recovery from usage-limit failures in CodexAdapter

  • When a Codex turn fails with usageLimitExceeded and no provider activity has occurred in that turn, the adapter automatically restarts the runtime, resumes the same thread, and retries the turn exactly once.
  • A new requireResume flag in CodexSessionRuntimeOptions prevents fallback to a fresh thread when strict resume is required; recoverable resume errors are surfaced instead of silently starting over.
  • sendTurn is blocked while recovery is in progress or a recovery token is active; tokens are cleared on explicit starts, interrupts, and stops.
  • New tests in CodexAdapter.test.ts verify recovery semantics using observable TestQueue utilities for deterministic event ordering.
  • Risk: Any provider activity (non-userMessage items, tool calls, hooks, diffs) disables retry for that turn, aborting recovery silently.
📊 Macroscope summarized 92a148c. 3 files reviewed, 0 issues evaluated, 0 issues filtered, 0 comments posted

🗂️ Filtered Issues

No issues evaluated.

@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 5024b059-11d5-4ed6-974c-a7656c1888cb

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:L 100-499 changed lines (additions + deletions). labels Aug 17, 2026
return;
}

const resumedSession = yield* startSessionInternal(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟠 HighLayers/CodexAdapter.ts:1965

A stale usage-limit recovery can stop a newer manual session and leave the thread with no active session. recoverUsageLimitTurn calls startSessionInternal before validating that input.token is still current; if startSession has already deleted the token and installed a replacement, the recovery replaces and then stops that replacement, and finally closes its own runtime. Validate the token before replacing the registered session and atomically guard the replacement.

🤖 Copy this AI Prompt to have your agent fix this:
In file @apps/server/src/provider/Layers/CodexAdapter.ts around line 1965:
A stale usage-limit recovery can stop a newer manual session and leave the thread with no active session. `recoverUsageLimitTurn` calls `startSessionInternal` before validating that `input.token` is still current; if `startSession` has already deleted the token and installed a replacement, the recovery replaces and then stops that replacement, and finally closes its own runtime. Validate the token before replacing the registered session and atomically guard the replacement.

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Want fixes drafted automatically? Bugbot Autofix can create code changes for findings. A team admin can enable Autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 92a148c. Configure here.

failedTurnId: input.failedTurnId,
},
);
return;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Aborted recovery leaves session open

Medium Severity

When post-resume readThread finds provider activity on the failed turn, recoverUsageLimitTurn returns without calling stopSessionInternal. By then startSessionInternal has already stopped the original session and registered the recovery process. Sibling abort paths for wrong thread or model do close that session. An account-selecting executable can therefore leave the thread on a replaced process after refusing to retry, instead of closing the recovery session and surfacing the original usage-limit failure.

Additional Locations (1)
Fix in CursorFix in Web

Reviewed by Cursor Bugbot for commit 92a148c. Configure here.

@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Needs human review

1 blocking correctness issue found. This PR introduces a new automatic recovery feature for usage-limit failures with significant state management complexity (recovery tokens, pending turn tracking, deferred completions). A High-severity finding identifies a potential race condition where stale recovery can interfere with newer sessions.

You can customize Macroscope's approvability policy. Learn more.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L100-499 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@estensen