feat: add cross-platform voice supervisor - #6206

Closed
duncan-vc wants to merge 24 commits into
pingdotgg:mainfrom
duncan-vc:feat/voice-supervisor
Closed

feat: add cross-platform voice supervisor#6206
duncan-vc wants to merge 24 commits into
pingdotgg:mainfrom
duncan-vc:feat/voice-supervisor

Conversation

@duncan-vc

@duncan-vcduncan-vc commented Aug 11, 2026

Copy link
Copy Markdown
Contributor

What changed

  • Add typed OpenAI Realtime session and credential handling on the T3 server.
  • Add a generation-scoped Voice Supervisor core with bounded transcripts, replay protection, local mutation confirmations, and exact environment/thread revalidation.
  • Add web and desktop Voice Supervisor UI, Voice settings, entry points, keybinding support, and macOS microphone permission handling.
  • Add foreground-only iOS and Android support using WebRTC plus a narrowly scoped native audio-session module.
  • Share platform-neutral tools, repository, transport, host, event, and state logic through @t3tools/client-runtime.
  • Document user behavior, architecture, remote/multi-environment semantics, and privacy disclosures.

Why

T3 can coordinate multiple coding agents and environments, but doing so currently requires staying at the keyboard. Voice Supervisor adds an explicit, reviewable speech interface for listing work, checking status, opening threads, and proposing confirmed commands while preserving T3's existing environment authority and receipt-backed execution model.

This is a full Realtime conversation rather than dictation. It uses the host environment's OpenAI API credential or OPENAI_API_KEY; it does not reuse ChatGPT subscription entitlements.

Behavior and safety

  • Opening the panel or mobile route never requests microphone access or mints a client secret; only explicit Start does.
  • API credentials remain on the selected T3 environment. Clients receive short-lived Realtime secrets and connect directly to OpenAI over WebRTC.
  • Mutating tools require local button confirmation tied to the exact voice generation, call, environment, target, and captured version.
  • Stop, host loss, terminal transport failure, mobile backgrounding, and native audio interruption tear down owned media and pending work without automatic reconnect.
  • Tool context is bounded, IDs and raw workspace paths stay opaque, and disconnected/stale targets are limited to list context until live revalidation succeeds.

Verification

  • vp i --frozen-lockfile --force — dependency and supply-chain verification passed; the installed WebRTC package contains the exact Android and iOS native pins.
  • vp test run <54 changed test files> — 54 files, 682 tests passed.
  • Targeted typechecks passed for contracts, shared, client-runtime, server, web, desktop, mobile, and marketing.
  • Targeted lint passed with zero diagnostics.
  • Formatting passed across 136 supported changed files.
  • git diff --check origin/main...HEAD passed.

Draft status

The source and focused verification gates are green. A rebuilt forked desktop/mobile client, real OpenAI Realtime smoke, and final UI screenshots are intentionally still pending and will be added before requesting maintainer review.

Related prior Realtime panel exploration: #3997.

Implemented with GPT-5.6 Sol through the Codex harness in T3 Code.

Note

Add cross-platform Voice Supervisor for real-time AI voice conversations

  • Introduces a full Voice Supervisor feature across web, desktop, and mobile that lets users start voice conversations with coding agents via the OpenAI Realtime API using WebRTC.
  • Adds server-side voice HTTP endpoints (/api/voice/openai/credential, /api/voice/realtime/client-secret) for credential management and ephemeral client-secret minting, with rate limiting, scope enforcement, and no-store headers.
  • Adds shared client-runtime modules for realtime session transport, event decoding, supervisor state reduction, tool schemas, and thread/voice supervisor repository operations.
  • Implements platform-specific realtime session adapters: a browser WebRTC adapter for web and a react-native-webrtc-backed adapter for mobile, each with audio session lifecycle management.
  • Adds native Expo modules (T3VoiceAudioSession) for iOS and Android to observe audio interruptions, route loss, and media services reset events without taking audio session ownership.
  • Adds a /settings/voice page on web and a Voice Supervisor settings row on mobile for configuring the host environment and OpenAI API key.
  • Adds voice.toggle as a built-in keybinding command, wired into the command palette and sidebar on web.
  • Risk: react-native-webrtc@124.0.7 is added as a new mobile dependency and patched via pnpm-workspace.yaml; the Android manifest explicitly removes MediaProjectionService from the WebRTC module.
📊 Macroscope summarized 90f076e. 86 files reviewed, 0 issues evaluated, 0 issues filtered, 0 comments posted

🗂️ Filtered Issues

No issues evaluated.

@coderabbitai

coderabbitaiBot commented Aug 11, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: ae2bef0b-6f3e-432f-91b8-d00cf2f1697a

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:XXL 1,000+ changed lines (additions + deletions). labels Aug 11, 2026
Comment threadpackages/client-runtime/src/voice/voiceSupervisorState.ts Outdated
Comment threadpackages/client-runtime/src/operations/threadSupervisor.ts
Comment threadapps/mobile/src/voice/voiceStartDefaults.ts
Comment on lines +107 to +111
const upstreamRetryAfterSeconds = (value: string | undefined): number => {
if (value === undefined) return 1;
const parsed = Number.parseFloat(value);
return Number.isFinite(parsed) && parsed > 0 ? Math.min(120, Math.ceil(parsed)) : 1;
};

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Mediumvoice/OpenAiRealtime.ts:107

When a 429 response includes an HTTP-date in the Retry-After header (e.g. Tue, 11 Aug 2026 12:00:00 GMT), upstreamRetryAfterSeconds parses it with Number.parseFloat, which returns NaN. The function then falls back to 1, so callers receive retryAfterSeconds: 1 and retry before the upstream-specified time. HTTP permits Retry-After to be either a delta-seconds value or an HTTP-date; consider parsing both forms (e.g. falling back to Date.parse when the float parse fails).

 const upstreamRetryAfterSeconds = (value: string | undefined): number => {
if (value === undefined) return 1;
const parsed = Number.parseFloat(value);
- return Number.isFinite(parsed) && parsed > 0 ? Math.min(120, Math.ceil(parsed)) : 1;+ if (Number.isFinite(parsed) && parsed > 0) return Math.min(120, Math.ceil(parsed));+ const dateMs = Date.parse(value);+ if (Number.isFinite(dateMs)) {+ const delta = Math.ceil((dateMs - Date.now()) / 1_000);+ return delta > 0 ? Math.min(120, delta) : 1;+ }+ return 1;
};
🤖 Copy this AI Prompt to have your agent fix this:
In file @apps/server/src/voice/OpenAiRealtime.ts around lines 107-111:
When a 429 response includes an HTTP-date in the `Retry-After` header (e.g. `Tue, 11 Aug 2026 12:00:00 GMT`), `upstreamRetryAfterSeconds` parses it with `Number.parseFloat`, which returns `NaN`. The function then falls back to `1`, so callers receive `retryAfterSeconds: 1` and retry before the upstream-specified time. HTTP permits `Retry-After` to be either a delta-seconds value or an HTTP-date; consider parsing both forms (e.g. falling back to `Date.parse` when the float parse fails).

projectModelSelection: project.defaultModelSelection,
draft,
});
const primaryDefaults = dependencies.readPrimaryThreadDefaults();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟠 Highvoice/voiceStartDefaults.ts:253

primaryDefaults is read from readPrimaryThreadDefaults(), which returns the primary environment's settings. When project.environmentId targets a different environment, defaultThreadEnvMode (line 264) and newWorktreesStartFromOrigin (line 307) are resolved against the wrong environment. This can cause a voice command on a remote project to create a local checkout instead of a worktree (or vice versa), or to apply the wrong start-from-origin policy. The target environment's settings are already available in environment.settings — use that instead.

🤖 Copy this AI Prompt to have your agent fix this:
In file @apps/web/src/voice/voiceStartDefaults.ts around line 253:
`primaryDefaults` is read from `readPrimaryThreadDefaults()`, which returns the *primary* environment's settings. When `project.environmentId` targets a different environment, `defaultThreadEnvMode` (line 264) and `newWorktreesStartFromOrigin` (line 307) are resolved against the wrong environment. This can cause a voice command on a remote project to create a local checkout instead of a worktree (or vice versa), or to apply the wrong start-from-origin policy. The target environment's settings are already available in `environment.settings` — use that instead.

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the new/changed Effect service code in this PR (apps/server/src/voice/*, apps/desktop/src/electron/ElectronSystemPreferences.ts, packages/client-runtime/src/state/voiceHttp.ts, packages/contracts/src/voice.ts, and the web/mobile voice adapters).

Most of it follows the conventions: ElectronSystemPreferences.ts, OpenAiRealtime.ts, and OpenAiRealtimeCredential.ts all use the canonical imports/Context.Service/make/layer shape with inline interfaces, Schema.TaggedErrorClass failures, structured attributes, exported Schema.is predicates, and dependencies acquired from the environment; ManagedRuntime/runPromise stay at React, Expo, and app-runtime boundaries.

Four convention violations in the changed scope are noted inline: Effect.catchTag in the two new server voice modules, two pure compatibility re-export shims left behind by the web-to-client-runtime move, and a non-canonical layer export name for a new Context.Service.

Posted via Macroscope — Effect Service Conventions

Comment threadpackages/client-runtime/src/state/voiceHttp.ts Outdated
Comment threadapps/server/src/voice/http.ts
Comment threadapps/web/src/voice/voiceTools.ts Outdated
Comment threadapps/web/src/voice/realtimeEvents.ts Outdated
Comment threadapps/server/src/voice/OpenAiRealtime.ts Outdated
@t3dotgg

Copy link
Copy Markdown
Member

Note

🤖 GPT-5.6 Sol responding on behalf of Theo

Closing this PR after an automated pass over open pull requests. Duplicates the maintainer-owned voice implementation in #5213.

@t3dotggt3dotgg closed this Aug 23, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXL1,000+ changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@duncan-vc@t3dotgg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

feat: add cross-platform voice supervisor - #6206

Closed
duncan-vc wants to merge 24 commits into
pingdotgg:mainfrom
duncan-vc:feat/voice-supervisor
Closed

feat: add cross-platform voice supervisor#6206
duncan-vc wants to merge 24 commits into
pingdotgg:mainfrom
duncan-vc:feat/voice-supervisor

Conversation

@duncan-vc

@duncan-vcduncan-vc commented Aug 11, 2026

Copy link
Copy Markdown
Contributor

What changed

  • Add typed OpenAI Realtime session and credential handling on the T3 server.
  • Add a generation-scoped Voice Supervisor core with bounded transcripts, replay protection, local mutation confirmations, and exact environment/thread revalidation.
  • Add web and desktop Voice Supervisor UI, Voice settings, entry points, keybinding support, and macOS microphone permission handling.
  • Add foreground-only iOS and Android support using WebRTC plus a narrowly scoped native audio-session module.
  • Share platform-neutral tools, repository, transport, host, event, and state logic through @t3tools/client-runtime.
  • Document user behavior, architecture, remote/multi-environment semantics, and privacy disclosures.

Why

T3 can coordinate multiple coding agents and environments, but doing so currently requires staying at the keyboard. Voice Supervisor adds an explicit, reviewable speech interface for listing work, checking status, opening threads, and proposing confirmed commands while preserving T3's existing environment authority and receipt-backed execution model.

This is a full Realtime conversation rather than dictation. It uses the host environment's OpenAI API credential or OPENAI_API_KEY; it does not reuse ChatGPT subscription entitlements.

Behavior and safety

  • Opening the panel or mobile route never requests microphone access or mints a client secret; only explicit Start does.
  • API credentials remain on the selected T3 environment. Clients receive short-lived Realtime secrets and connect directly to OpenAI over WebRTC.
  • Mutating tools require local button confirmation tied to the exact voice generation, call, environment, target, and captured version.
  • Stop, host loss, terminal transport failure, mobile backgrounding, and native audio interruption tear down owned media and pending work without automatic reconnect.
  • Tool context is bounded, IDs and raw workspace paths stay opaque, and disconnected/stale targets are limited to list context until live revalidation succeeds.

Verification

  • vp i --frozen-lockfile --force — dependency and supply-chain verification passed; the installed WebRTC package contains the exact Android and iOS native pins.
  • vp test run <54 changed test files> — 54 files, 682 tests passed.
  • Targeted typechecks passed for contracts, shared, client-runtime, server, web, desktop, mobile, and marketing.
  • Targeted lint passed with zero diagnostics.
  • Formatting passed across 136 supported changed files.
  • git diff --check origin/main...HEAD passed.

Draft status

The source and focused verification gates are green. A rebuilt forked desktop/mobile client, real OpenAI Realtime smoke, and final UI screenshots are intentionally still pending and will be added before requesting maintainer review.

Related prior Realtime panel exploration: #3997.

Implemented with GPT-5.6 Sol through the Codex harness in T3 Code.

Note

Add cross-platform Voice Supervisor for real-time AI voice conversations

  • Introduces a full Voice Supervisor feature across web, desktop, and mobile that lets users start voice conversations with coding agents via the OpenAI Realtime API using WebRTC.
  • Adds server-side voice HTTP endpoints (/api/voice/openai/credential, /api/voice/realtime/client-secret) for credential management and ephemeral client-secret minting, with rate limiting, scope enforcement, and no-store headers.
  • Adds shared client-runtime modules for realtime session transport, event decoding, supervisor state reduction, tool schemas, and thread/voice supervisor repository operations.
  • Implements platform-specific realtime session adapters: a browser WebRTC adapter for web and a react-native-webrtc-backed adapter for mobile, each with audio session lifecycle management.
  • Adds native Expo modules (T3VoiceAudioSession) for iOS and Android to observe audio interruptions, route loss, and media services reset events without taking audio session ownership.
  • Adds a /settings/voice page on web and a Voice Supervisor settings row on mobile for configuring the host environment and OpenAI API key.
  • Adds voice.toggle as a built-in keybinding command, wired into the command palette and sidebar on web.
  • Risk: react-native-webrtc@124.0.7 is added as a new mobile dependency and patched via pnpm-workspace.yaml; the Android manifest explicitly removes MediaProjectionService from the WebRTC module.
📊 Macroscope summarized 90f076e. 86 files reviewed, 0 issues evaluated, 0 issues filtered, 0 comments posted

🗂️ Filtered Issues

No issues evaluated.

@coderabbitai

coderabbitaiBot commented Aug 11, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: ae2bef0b-6f3e-432f-91b8-d00cf2f1697a

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:XXL 1,000+ changed lines (additions + deletions). labels Aug 11, 2026
Comment threadpackages/client-runtime/src/voice/voiceSupervisorState.ts Outdated
Comment threadpackages/client-runtime/src/operations/threadSupervisor.ts
Comment threadapps/mobile/src/voice/voiceStartDefaults.ts
Comment on lines +107 to +111
const upstreamRetryAfterSeconds = (value: string | undefined): number => {
if (value === undefined) return 1;
const parsed = Number.parseFloat(value);
return Number.isFinite(parsed) && parsed > 0 ? Math.min(120, Math.ceil(parsed)) : 1;
};

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Mediumvoice/OpenAiRealtime.ts:107

When a 429 response includes an HTTP-date in the Retry-After header (e.g. Tue, 11 Aug 2026 12:00:00 GMT), upstreamRetryAfterSeconds parses it with Number.parseFloat, which returns NaN. The function then falls back to 1, so callers receive retryAfterSeconds: 1 and retry before the upstream-specified time. HTTP permits Retry-After to be either a delta-seconds value or an HTTP-date; consider parsing both forms (e.g. falling back to Date.parse when the float parse fails).

 const upstreamRetryAfterSeconds = (value: string | undefined): number => {
if (value === undefined) return 1;
const parsed = Number.parseFloat(value);
- return Number.isFinite(parsed) && parsed > 0 ? Math.min(120, Math.ceil(parsed)) : 1;+ if (Number.isFinite(parsed) && parsed > 0) return Math.min(120, Math.ceil(parsed));+ const dateMs = Date.parse(value);+ if (Number.isFinite(dateMs)) {+ const delta = Math.ceil((dateMs - Date.now()) / 1_000);+ return delta > 0 ? Math.min(120, delta) : 1;+ }+ return 1;
};
🤖 Copy this AI Prompt to have your agent fix this:
In file @apps/server/src/voice/OpenAiRealtime.ts around lines 107-111:
When a 429 response includes an HTTP-date in the `Retry-After` header (e.g. `Tue, 11 Aug 2026 12:00:00 GMT`), `upstreamRetryAfterSeconds` parses it with `Number.parseFloat`, which returns `NaN`. The function then falls back to `1`, so callers receive `retryAfterSeconds: 1` and retry before the upstream-specified time. HTTP permits `Retry-After` to be either a delta-seconds value or an HTTP-date; consider parsing both forms (e.g. falling back to `Date.parse` when the float parse fails).

projectModelSelection: project.defaultModelSelection,
draft,
});
const primaryDefaults = dependencies.readPrimaryThreadDefaults();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟠 Highvoice/voiceStartDefaults.ts:253

primaryDefaults is read from readPrimaryThreadDefaults(), which returns the primary environment's settings. When project.environmentId targets a different environment, defaultThreadEnvMode (line 264) and newWorktreesStartFromOrigin (line 307) are resolved against the wrong environment. This can cause a voice command on a remote project to create a local checkout instead of a worktree (or vice versa), or to apply the wrong start-from-origin policy. The target environment's settings are already available in environment.settings — use that instead.

🤖 Copy this AI Prompt to have your agent fix this:
In file @apps/web/src/voice/voiceStartDefaults.ts around line 253:
`primaryDefaults` is read from `readPrimaryThreadDefaults()`, which returns the *primary* environment's settings. When `project.environmentId` targets a different environment, `defaultThreadEnvMode` (line 264) and `newWorktreesStartFromOrigin` (line 307) are resolved against the wrong environment. This can cause a voice command on a remote project to create a local checkout instead of a worktree (or vice versa), or to apply the wrong start-from-origin policy. The target environment's settings are already available in `environment.settings` — use that instead.

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the new/changed Effect service code in this PR (apps/server/src/voice/*, apps/desktop/src/electron/ElectronSystemPreferences.ts, packages/client-runtime/src/state/voiceHttp.ts, packages/contracts/src/voice.ts, and the web/mobile voice adapters).

Most of it follows the conventions: ElectronSystemPreferences.ts, OpenAiRealtime.ts, and OpenAiRealtimeCredential.ts all use the canonical imports/Context.Service/make/layer shape with inline interfaces, Schema.TaggedErrorClass failures, structured attributes, exported Schema.is predicates, and dependencies acquired from the environment; ManagedRuntime/runPromise stay at React, Expo, and app-runtime boundaries.

Four convention violations in the changed scope are noted inline: Effect.catchTag in the two new server voice modules, two pure compatibility re-export shims left behind by the web-to-client-runtime move, and a non-canonical layer export name for a new Context.Service.

Posted via Macroscope — Effect Service Conventions

Comment threadpackages/client-runtime/src/state/voiceHttp.ts Outdated
Comment threadapps/server/src/voice/http.ts
Comment threadapps/web/src/voice/voiceTools.ts Outdated
Comment threadapps/web/src/voice/realtimeEvents.ts Outdated
Comment threadapps/server/src/voice/OpenAiRealtime.ts Outdated
@t3dotgg

Copy link
Copy Markdown
Member

Note

🤖 GPT-5.6 Sol responding on behalf of Theo

Closing this PR after an automated pass over open pull requests. Duplicates the maintainer-owned voice implementation in #5213.

@t3dotggt3dotgg closed this Aug 23, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXL1,000+ changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@duncan-vc@t3dotgg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat: add cross-platform voice supervisor - #6206

Closed
duncan-vc wants to merge 24 commits into
pingdotgg:mainfrom
duncan-vc:feat/voice-supervisor
Closed

feat: add cross-platform voice supervisor#6206
duncan-vc wants to merge 24 commits into
pingdotgg:mainfrom
duncan-vc:feat/voice-supervisor

Conversation

@duncan-vc

@duncan-vcduncan-vc commented Aug 11, 2026

Copy link
Copy Markdown
Contributor

What changed

  • Add typed OpenAI Realtime session and credential handling on the T3 server.
  • Add a generation-scoped Voice Supervisor core with bounded transcripts, replay protection, local mutation confirmations, and exact environment/thread revalidation.
  • Add web and desktop Voice Supervisor UI, Voice settings, entry points, keybinding support, and macOS microphone permission handling.
  • Add foreground-only iOS and Android support using WebRTC plus a narrowly scoped native audio-session module.
  • Share platform-neutral tools, repository, transport, host, event, and state logic through @t3tools/client-runtime.
  • Document user behavior, architecture, remote/multi-environment semantics, and privacy disclosures.

Why

T3 can coordinate multiple coding agents and environments, but doing so currently requires staying at the keyboard. Voice Supervisor adds an explicit, reviewable speech interface for listing work, checking status, opening threads, and proposing confirmed commands while preserving T3's existing environment authority and receipt-backed execution model.

This is a full Realtime conversation rather than dictation. It uses the host environment's OpenAI API credential or OPENAI_API_KEY; it does not reuse ChatGPT subscription entitlements.

Behavior and safety

  • Opening the panel or mobile route never requests microphone access or mints a client secret; only explicit Start does.
  • API credentials remain on the selected T3 environment. Clients receive short-lived Realtime secrets and connect directly to OpenAI over WebRTC.
  • Mutating tools require local button confirmation tied to the exact voice generation, call, environment, target, and captured version.
  • Stop, host loss, terminal transport failure, mobile backgrounding, and native audio interruption tear down owned media and pending work without automatic reconnect.
  • Tool context is bounded, IDs and raw workspace paths stay opaque, and disconnected/stale targets are limited to list context until live revalidation succeeds.

Verification

  • vp i --frozen-lockfile --force — dependency and supply-chain verification passed; the installed WebRTC package contains the exact Android and iOS native pins.
  • vp test run <54 changed test files> — 54 files, 682 tests passed.
  • Targeted typechecks passed for contracts, shared, client-runtime, server, web, desktop, mobile, and marketing.
  • Targeted lint passed with zero diagnostics.
  • Formatting passed across 136 supported changed files.
  • git diff --check origin/main...HEAD passed.

Draft status

The source and focused verification gates are green. A rebuilt forked desktop/mobile client, real OpenAI Realtime smoke, and final UI screenshots are intentionally still pending and will be added before requesting maintainer review.

Related prior Realtime panel exploration: #3997.

Implemented with GPT-5.6 Sol through the Codex harness in T3 Code.

Note

Add cross-platform Voice Supervisor for real-time AI voice conversations

  • Introduces a full Voice Supervisor feature across web, desktop, and mobile that lets users start voice conversations with coding agents via the OpenAI Realtime API using WebRTC.
  • Adds server-side voice HTTP endpoints (/api/voice/openai/credential, /api/voice/realtime/client-secret) for credential management and ephemeral client-secret minting, with rate limiting, scope enforcement, and no-store headers.
  • Adds shared client-runtime modules for realtime session transport, event decoding, supervisor state reduction, tool schemas, and thread/voice supervisor repository operations.
  • Implements platform-specific realtime session adapters: a browser WebRTC adapter for web and a react-native-webrtc-backed adapter for mobile, each with audio session lifecycle management.
  • Adds native Expo modules (T3VoiceAudioSession) for iOS and Android to observe audio interruptions, route loss, and media services reset events without taking audio session ownership.
  • Adds a /settings/voice page on web and a Voice Supervisor settings row on mobile for configuring the host environment and OpenAI API key.
  • Adds voice.toggle as a built-in keybinding command, wired into the command palette and sidebar on web.
  • Risk: react-native-webrtc@124.0.7 is added as a new mobile dependency and patched via pnpm-workspace.yaml; the Android manifest explicitly removes MediaProjectionService from the WebRTC module.
📊 Macroscope summarized 90f076e. 86 files reviewed, 0 issues evaluated, 0 issues filtered, 0 comments posted

🗂️ Filtered Issues

No issues evaluated.

@coderabbitai

coderabbitaiBot commented Aug 11, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: ae2bef0b-6f3e-432f-91b8-d00cf2f1697a

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:XXL 1,000+ changed lines (additions + deletions). labels Aug 11, 2026
Comment threadpackages/client-runtime/src/voice/voiceSupervisorState.ts Outdated
Comment threadpackages/client-runtime/src/operations/threadSupervisor.ts
Comment threadapps/mobile/src/voice/voiceStartDefaults.ts
Comment on lines +107 to +111
const upstreamRetryAfterSeconds = (value: string | undefined): number => {
if (value === undefined) return 1;
const parsed = Number.parseFloat(value);
return Number.isFinite(parsed) && parsed > 0 ? Math.min(120, Math.ceil(parsed)) : 1;
};

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Mediumvoice/OpenAiRealtime.ts:107

When a 429 response includes an HTTP-date in the Retry-After header (e.g. Tue, 11 Aug 2026 12:00:00 GMT), upstreamRetryAfterSeconds parses it with Number.parseFloat, which returns NaN. The function then falls back to 1, so callers receive retryAfterSeconds: 1 and retry before the upstream-specified time. HTTP permits Retry-After to be either a delta-seconds value or an HTTP-date; consider parsing both forms (e.g. falling back to Date.parse when the float parse fails).

 const upstreamRetryAfterSeconds = (value: string | undefined): number => {
if (value === undefined) return 1;
const parsed = Number.parseFloat(value);
- return Number.isFinite(parsed) && parsed > 0 ? Math.min(120, Math.ceil(parsed)) : 1;+ if (Number.isFinite(parsed) && parsed > 0) return Math.min(120, Math.ceil(parsed));+ const dateMs = Date.parse(value);+ if (Number.isFinite(dateMs)) {+ const delta = Math.ceil((dateMs - Date.now()) / 1_000);+ return delta > 0 ? Math.min(120, delta) : 1;+ }+ return 1;
};
🤖 Copy this AI Prompt to have your agent fix this:
In file @apps/server/src/voice/OpenAiRealtime.ts around lines 107-111:
When a 429 response includes an HTTP-date in the `Retry-After` header (e.g. `Tue, 11 Aug 2026 12:00:00 GMT`), `upstreamRetryAfterSeconds` parses it with `Number.parseFloat`, which returns `NaN`. The function then falls back to `1`, so callers receive `retryAfterSeconds: 1` and retry before the upstream-specified time. HTTP permits `Retry-After` to be either a delta-seconds value or an HTTP-date; consider parsing both forms (e.g. falling back to `Date.parse` when the float parse fails).

projectModelSelection: project.defaultModelSelection,
draft,
});
const primaryDefaults = dependencies.readPrimaryThreadDefaults();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟠 Highvoice/voiceStartDefaults.ts:253

primaryDefaults is read from readPrimaryThreadDefaults(), which returns the primary environment's settings. When project.environmentId targets a different environment, defaultThreadEnvMode (line 264) and newWorktreesStartFromOrigin (line 307) are resolved against the wrong environment. This can cause a voice command on a remote project to create a local checkout instead of a worktree (or vice versa), or to apply the wrong start-from-origin policy. The target environment's settings are already available in environment.settings — use that instead.

🤖 Copy this AI Prompt to have your agent fix this:
In file @apps/web/src/voice/voiceStartDefaults.ts around line 253:
`primaryDefaults` is read from `readPrimaryThreadDefaults()`, which returns the *primary* environment's settings. When `project.environmentId` targets a different environment, `defaultThreadEnvMode` (line 264) and `newWorktreesStartFromOrigin` (line 307) are resolved against the wrong environment. This can cause a voice command on a remote project to create a local checkout instead of a worktree (or vice versa), or to apply the wrong start-from-origin policy. The target environment's settings are already available in `environment.settings` — use that instead.

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the new/changed Effect service code in this PR (apps/server/src/voice/*, apps/desktop/src/electron/ElectronSystemPreferences.ts, packages/client-runtime/src/state/voiceHttp.ts, packages/contracts/src/voice.ts, and the web/mobile voice adapters).

Most of it follows the conventions: ElectronSystemPreferences.ts, OpenAiRealtime.ts, and OpenAiRealtimeCredential.ts all use the canonical imports/Context.Service/make/layer shape with inline interfaces, Schema.TaggedErrorClass failures, structured attributes, exported Schema.is predicates, and dependencies acquired from the environment; ManagedRuntime/runPromise stay at React, Expo, and app-runtime boundaries.

Four convention violations in the changed scope are noted inline: Effect.catchTag in the two new server voice modules, two pure compatibility re-export shims left behind by the web-to-client-runtime move, and a non-canonical layer export name for a new Context.Service.

Posted via Macroscope — Effect Service Conventions

Comment threadpackages/client-runtime/src/state/voiceHttp.ts Outdated
Comment threadapps/server/src/voice/http.ts
Comment threadapps/web/src/voice/voiceTools.ts Outdated
Comment threadapps/web/src/voice/realtimeEvents.ts Outdated
Comment threadapps/server/src/voice/OpenAiRealtime.ts Outdated
@t3dotgg

Copy link
Copy Markdown
Member

Note

🤖 GPT-5.6 Sol responding on behalf of Theo

Closing this PR after an automated pass over open pull requests. Duplicates the maintainer-owned voice implementation in #5213.

@t3dotggt3dotgg closed this Aug 23, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXL1,000+ changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@duncan-vc@t3dotgg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat: add cross-platform voice supervisor - #6206

Closed
duncan-vc wants to merge 24 commits into
pingdotgg:mainfrom
duncan-vc:feat/voice-supervisor
Closed

feat: add cross-platform voice supervisor#6206
duncan-vc wants to merge 24 commits into
pingdotgg:mainfrom
duncan-vc:feat/voice-supervisor

Conversation

@duncan-vc

@duncan-vcduncan-vc commented Aug 11, 2026

Copy link
Copy Markdown
Contributor

What changed

  • Add typed OpenAI Realtime session and credential handling on the T3 server.
  • Add a generation-scoped Voice Supervisor core with bounded transcripts, replay protection, local mutation confirmations, and exact environment/thread revalidation.
  • Add web and desktop Voice Supervisor UI, Voice settings, entry points, keybinding support, and macOS microphone permission handling.
  • Add foreground-only iOS and Android support using WebRTC plus a narrowly scoped native audio-session module.
  • Share platform-neutral tools, repository, transport, host, event, and state logic through @t3tools/client-runtime.
  • Document user behavior, architecture, remote/multi-environment semantics, and privacy disclosures.

Why

T3 can coordinate multiple coding agents and environments, but doing so currently requires staying at the keyboard. Voice Supervisor adds an explicit, reviewable speech interface for listing work, checking status, opening threads, and proposing confirmed commands while preserving T3's existing environment authority and receipt-backed execution model.

This is a full Realtime conversation rather than dictation. It uses the host environment's OpenAI API credential or OPENAI_API_KEY; it does not reuse ChatGPT subscription entitlements.

Behavior and safety

  • Opening the panel or mobile route never requests microphone access or mints a client secret; only explicit Start does.
  • API credentials remain on the selected T3 environment. Clients receive short-lived Realtime secrets and connect directly to OpenAI over WebRTC.
  • Mutating tools require local button confirmation tied to the exact voice generation, call, environment, target, and captured version.
  • Stop, host loss, terminal transport failure, mobile backgrounding, and native audio interruption tear down owned media and pending work without automatic reconnect.
  • Tool context is bounded, IDs and raw workspace paths stay opaque, and disconnected/stale targets are limited to list context until live revalidation succeeds.

Verification

  • vp i --frozen-lockfile --force — dependency and supply-chain verification passed; the installed WebRTC package contains the exact Android and iOS native pins.
  • vp test run <54 changed test files> — 54 files, 682 tests passed.
  • Targeted typechecks passed for contracts, shared, client-runtime, server, web, desktop, mobile, and marketing.
  • Targeted lint passed with zero diagnostics.
  • Formatting passed across 136 supported changed files.
  • git diff --check origin/main...HEAD passed.

Draft status

The source and focused verification gates are green. A rebuilt forked desktop/mobile client, real OpenAI Realtime smoke, and final UI screenshots are intentionally still pending and will be added before requesting maintainer review.

Related prior Realtime panel exploration: #3997.

Implemented with GPT-5.6 Sol through the Codex harness in T3 Code.

Note

Add cross-platform Voice Supervisor for real-time AI voice conversations

  • Introduces a full Voice Supervisor feature across web, desktop, and mobile that lets users start voice conversations with coding agents via the OpenAI Realtime API using WebRTC.
  • Adds server-side voice HTTP endpoints (/api/voice/openai/credential, /api/voice/realtime/client-secret) for credential management and ephemeral client-secret minting, with rate limiting, scope enforcement, and no-store headers.
  • Adds shared client-runtime modules for realtime session transport, event decoding, supervisor state reduction, tool schemas, and thread/voice supervisor repository operations.
  • Implements platform-specific realtime session adapters: a browser WebRTC adapter for web and a react-native-webrtc-backed adapter for mobile, each with audio session lifecycle management.
  • Adds native Expo modules (T3VoiceAudioSession) for iOS and Android to observe audio interruptions, route loss, and media services reset events without taking audio session ownership.
  • Adds a /settings/voice page on web and a Voice Supervisor settings row on mobile for configuring the host environment and OpenAI API key.
  • Adds voice.toggle as a built-in keybinding command, wired into the command palette and sidebar on web.
  • Risk: react-native-webrtc@124.0.7 is added as a new mobile dependency and patched via pnpm-workspace.yaml; the Android manifest explicitly removes MediaProjectionService from the WebRTC module.
📊 Macroscope summarized 90f076e. 86 files reviewed, 0 issues evaluated, 0 issues filtered, 0 comments posted

🗂️ Filtered Issues

No issues evaluated.

@coderabbitai

coderabbitaiBot commented Aug 11, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: ae2bef0b-6f3e-432f-91b8-d00cf2f1697a

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:XXL 1,000+ changed lines (additions + deletions). labels Aug 11, 2026
Comment threadpackages/client-runtime/src/voice/voiceSupervisorState.ts Outdated
Comment threadpackages/client-runtime/src/operations/threadSupervisor.ts
Comment threadapps/mobile/src/voice/voiceStartDefaults.ts
Comment on lines +107 to +111
const upstreamRetryAfterSeconds = (value: string | undefined): number => {
if (value === undefined) return 1;
const parsed = Number.parseFloat(value);
return Number.isFinite(parsed) && parsed > 0 ? Math.min(120, Math.ceil(parsed)) : 1;
};

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Mediumvoice/OpenAiRealtime.ts:107

When a 429 response includes an HTTP-date in the Retry-After header (e.g. Tue, 11 Aug 2026 12:00:00 GMT), upstreamRetryAfterSeconds parses it with Number.parseFloat, which returns NaN. The function then falls back to 1, so callers receive retryAfterSeconds: 1 and retry before the upstream-specified time. HTTP permits Retry-After to be either a delta-seconds value or an HTTP-date; consider parsing both forms (e.g. falling back to Date.parse when the float parse fails).

 const upstreamRetryAfterSeconds = (value: string | undefined): number => {
if (value === undefined) return 1;
const parsed = Number.parseFloat(value);
- return Number.isFinite(parsed) && parsed > 0 ? Math.min(120, Math.ceil(parsed)) : 1;+ if (Number.isFinite(parsed) && parsed > 0) return Math.min(120, Math.ceil(parsed));+ const dateMs = Date.parse(value);+ if (Number.isFinite(dateMs)) {+ const delta = Math.ceil((dateMs - Date.now()) / 1_000);+ return delta > 0 ? Math.min(120, delta) : 1;+ }+ return 1;
};
🤖 Copy this AI Prompt to have your agent fix this:
In file @apps/server/src/voice/OpenAiRealtime.ts around lines 107-111:
When a 429 response includes an HTTP-date in the `Retry-After` header (e.g. `Tue, 11 Aug 2026 12:00:00 GMT`), `upstreamRetryAfterSeconds` parses it with `Number.parseFloat`, which returns `NaN`. The function then falls back to `1`, so callers receive `retryAfterSeconds: 1` and retry before the upstream-specified time. HTTP permits `Retry-After` to be either a delta-seconds value or an HTTP-date; consider parsing both forms (e.g. falling back to `Date.parse` when the float parse fails).

projectModelSelection: project.defaultModelSelection,
draft,
});
const primaryDefaults = dependencies.readPrimaryThreadDefaults();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟠 Highvoice/voiceStartDefaults.ts:253

primaryDefaults is read from readPrimaryThreadDefaults(), which returns the primary environment's settings. When project.environmentId targets a different environment, defaultThreadEnvMode (line 264) and newWorktreesStartFromOrigin (line 307) are resolved against the wrong environment. This can cause a voice command on a remote project to create a local checkout instead of a worktree (or vice versa), or to apply the wrong start-from-origin policy. The target environment's settings are already available in environment.settings — use that instead.

🤖 Copy this AI Prompt to have your agent fix this:
In file @apps/web/src/voice/voiceStartDefaults.ts around line 253:
`primaryDefaults` is read from `readPrimaryThreadDefaults()`, which returns the *primary* environment's settings. When `project.environmentId` targets a different environment, `defaultThreadEnvMode` (line 264) and `newWorktreesStartFromOrigin` (line 307) are resolved against the wrong environment. This can cause a voice command on a remote project to create a local checkout instead of a worktree (or vice versa), or to apply the wrong start-from-origin policy. The target environment's settings are already available in `environment.settings` — use that instead.

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the new/changed Effect service code in this PR (apps/server/src/voice/*, apps/desktop/src/electron/ElectronSystemPreferences.ts, packages/client-runtime/src/state/voiceHttp.ts, packages/contracts/src/voice.ts, and the web/mobile voice adapters).

Most of it follows the conventions: ElectronSystemPreferences.ts, OpenAiRealtime.ts, and OpenAiRealtimeCredential.ts all use the canonical imports/Context.Service/make/layer shape with inline interfaces, Schema.TaggedErrorClass failures, structured attributes, exported Schema.is predicates, and dependencies acquired from the environment; ManagedRuntime/runPromise stay at React, Expo, and app-runtime boundaries.

Four convention violations in the changed scope are noted inline: Effect.catchTag in the two new server voice modules, two pure compatibility re-export shims left behind by the web-to-client-runtime move, and a non-canonical layer export name for a new Context.Service.

Posted via Macroscope — Effect Service Conventions

Comment threadpackages/client-runtime/src/state/voiceHttp.ts Outdated
Comment threadapps/server/src/voice/http.ts
Comment threadapps/web/src/voice/voiceTools.ts Outdated
Comment threadapps/web/src/voice/realtimeEvents.ts Outdated
Comment threadapps/server/src/voice/OpenAiRealtime.ts Outdated
@t3dotgg

Copy link
Copy Markdown
Member

Note

🤖 GPT-5.6 Sol responding on behalf of Theo

Closing this PR after an automated pass over open pull requests. Duplicates the maintainer-owned voice implementation in #5213.

@t3dotggt3dotgg closed this Aug 23, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXL1,000+ changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@duncan-vc@t3dotgg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

feat: add cross-platform voice supervisor - #6206

Closed
duncan-vc wants to merge 24 commits into
pingdotgg:mainfrom
duncan-vc:feat/voice-supervisor
Closed

feat: add cross-platform voice supervisor#6206
duncan-vc wants to merge 24 commits into
pingdotgg:mainfrom
duncan-vc:feat/voice-supervisor

Conversation

@duncan-vc

@duncan-vcduncan-vc commented Aug 11, 2026

Copy link
Copy Markdown
Contributor

What changed

  • Add typed OpenAI Realtime session and credential handling on the T3 server.
  • Add a generation-scoped Voice Supervisor core with bounded transcripts, replay protection, local mutation confirmations, and exact environment/thread revalidation.
  • Add web and desktop Voice Supervisor UI, Voice settings, entry points, keybinding support, and macOS microphone permission handling.
  • Add foreground-only iOS and Android support using WebRTC plus a narrowly scoped native audio-session module.
  • Share platform-neutral tools, repository, transport, host, event, and state logic through @t3tools/client-runtime.
  • Document user behavior, architecture, remote/multi-environment semantics, and privacy disclosures.

Why

T3 can coordinate multiple coding agents and environments, but doing so currently requires staying at the keyboard. Voice Supervisor adds an explicit, reviewable speech interface for listing work, checking status, opening threads, and proposing confirmed commands while preserving T3's existing environment authority and receipt-backed execution model.

This is a full Realtime conversation rather than dictation. It uses the host environment's OpenAI API credential or OPENAI_API_KEY; it does not reuse ChatGPT subscription entitlements.

Behavior and safety

  • Opening the panel or mobile route never requests microphone access or mints a client secret; only explicit Start does.
  • API credentials remain on the selected T3 environment. Clients receive short-lived Realtime secrets and connect directly to OpenAI over WebRTC.
  • Mutating tools require local button confirmation tied to the exact voice generation, call, environment, target, and captured version.
  • Stop, host loss, terminal transport failure, mobile backgrounding, and native audio interruption tear down owned media and pending work without automatic reconnect.
  • Tool context is bounded, IDs and raw workspace paths stay opaque, and disconnected/stale targets are limited to list context until live revalidation succeeds.

Verification

  • vp i --frozen-lockfile --force — dependency and supply-chain verification passed; the installed WebRTC package contains the exact Android and iOS native pins.
  • vp test run <54 changed test files> — 54 files, 682 tests passed.
  • Targeted typechecks passed for contracts, shared, client-runtime, server, web, desktop, mobile, and marketing.
  • Targeted lint passed with zero diagnostics.
  • Formatting passed across 136 supported changed files.
  • git diff --check origin/main...HEAD passed.

Draft status

The source and focused verification gates are green. A rebuilt forked desktop/mobile client, real OpenAI Realtime smoke, and final UI screenshots are intentionally still pending and will be added before requesting maintainer review.

Related prior Realtime panel exploration: #3997.

Implemented with GPT-5.6 Sol through the Codex harness in T3 Code.

Note

Add cross-platform Voice Supervisor for real-time AI voice conversations

  • Introduces a full Voice Supervisor feature across web, desktop, and mobile that lets users start voice conversations with coding agents via the OpenAI Realtime API using WebRTC.
  • Adds server-side voice HTTP endpoints (/api/voice/openai/credential, /api/voice/realtime/client-secret) for credential management and ephemeral client-secret minting, with rate limiting, scope enforcement, and no-store headers.
  • Adds shared client-runtime modules for realtime session transport, event decoding, supervisor state reduction, tool schemas, and thread/voice supervisor repository operations.
  • Implements platform-specific realtime session adapters: a browser WebRTC adapter for web and a react-native-webrtc-backed adapter for mobile, each with audio session lifecycle management.
  • Adds native Expo modules (T3VoiceAudioSession) for iOS and Android to observe audio interruptions, route loss, and media services reset events without taking audio session ownership.
  • Adds a /settings/voice page on web and a Voice Supervisor settings row on mobile for configuring the host environment and OpenAI API key.
  • Adds voice.toggle as a built-in keybinding command, wired into the command palette and sidebar on web.
  • Risk: react-native-webrtc@124.0.7 is added as a new mobile dependency and patched via pnpm-workspace.yaml; the Android manifest explicitly removes MediaProjectionService from the WebRTC module.
📊 Macroscope summarized 90f076e. 86 files reviewed, 0 issues evaluated, 0 issues filtered, 0 comments posted

🗂️ Filtered Issues

No issues evaluated.

@coderabbitai

coderabbitaiBot commented Aug 11, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: ae2bef0b-6f3e-432f-91b8-d00cf2f1697a

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:XXL 1,000+ changed lines (additions + deletions). labels Aug 11, 2026
Comment threadpackages/client-runtime/src/voice/voiceSupervisorState.ts Outdated
Comment threadpackages/client-runtime/src/operations/threadSupervisor.ts
Comment threadapps/mobile/src/voice/voiceStartDefaults.ts
Comment on lines +107 to +111
const upstreamRetryAfterSeconds = (value: string | undefined): number => {
if (value === undefined) return 1;
const parsed = Number.parseFloat(value);
return Number.isFinite(parsed) && parsed > 0 ? Math.min(120, Math.ceil(parsed)) : 1;
};

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Mediumvoice/OpenAiRealtime.ts:107

When a 429 response includes an HTTP-date in the Retry-After header (e.g. Tue, 11 Aug 2026 12:00:00 GMT), upstreamRetryAfterSeconds parses it with Number.parseFloat, which returns NaN. The function then falls back to 1, so callers receive retryAfterSeconds: 1 and retry before the upstream-specified time. HTTP permits Retry-After to be either a delta-seconds value or an HTTP-date; consider parsing both forms (e.g. falling back to Date.parse when the float parse fails).

 const upstreamRetryAfterSeconds = (value: string | undefined): number => {
if (value === undefined) return 1;
const parsed = Number.parseFloat(value);
- return Number.isFinite(parsed) && parsed > 0 ? Math.min(120, Math.ceil(parsed)) : 1;+ if (Number.isFinite(parsed) && parsed > 0) return Math.min(120, Math.ceil(parsed));+ const dateMs = Date.parse(value);+ if (Number.isFinite(dateMs)) {+ const delta = Math.ceil((dateMs - Date.now()) / 1_000);+ return delta > 0 ? Math.min(120, delta) : 1;+ }+ return 1;
};
🤖 Copy this AI Prompt to have your agent fix this:
In file @apps/server/src/voice/OpenAiRealtime.ts around lines 107-111:
When a 429 response includes an HTTP-date in the `Retry-After` header (e.g. `Tue, 11 Aug 2026 12:00:00 GMT`), `upstreamRetryAfterSeconds` parses it with `Number.parseFloat`, which returns `NaN`. The function then falls back to `1`, so callers receive `retryAfterSeconds: 1` and retry before the upstream-specified time. HTTP permits `Retry-After` to be either a delta-seconds value or an HTTP-date; consider parsing both forms (e.g. falling back to `Date.parse` when the float parse fails).

projectModelSelection: project.defaultModelSelection,
draft,
});
const primaryDefaults = dependencies.readPrimaryThreadDefaults();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟠 Highvoice/voiceStartDefaults.ts:253

primaryDefaults is read from readPrimaryThreadDefaults(), which returns the primary environment's settings. When project.environmentId targets a different environment, defaultThreadEnvMode (line 264) and newWorktreesStartFromOrigin (line 307) are resolved against the wrong environment. This can cause a voice command on a remote project to create a local checkout instead of a worktree (or vice versa), or to apply the wrong start-from-origin policy. The target environment's settings are already available in environment.settings — use that instead.

🤖 Copy this AI Prompt to have your agent fix this:
In file @apps/web/src/voice/voiceStartDefaults.ts around line 253:
`primaryDefaults` is read from `readPrimaryThreadDefaults()`, which returns the *primary* environment's settings. When `project.environmentId` targets a different environment, `defaultThreadEnvMode` (line 264) and `newWorktreesStartFromOrigin` (line 307) are resolved against the wrong environment. This can cause a voice command on a remote project to create a local checkout instead of a worktree (or vice versa), or to apply the wrong start-from-origin policy. The target environment's settings are already available in `environment.settings` — use that instead.

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the new/changed Effect service code in this PR (apps/server/src/voice/*, apps/desktop/src/electron/ElectronSystemPreferences.ts, packages/client-runtime/src/state/voiceHttp.ts, packages/contracts/src/voice.ts, and the web/mobile voice adapters).

Most of it follows the conventions: ElectronSystemPreferences.ts, OpenAiRealtime.ts, and OpenAiRealtimeCredential.ts all use the canonical imports/Context.Service/make/layer shape with inline interfaces, Schema.TaggedErrorClass failures, structured attributes, exported Schema.is predicates, and dependencies acquired from the environment; ManagedRuntime/runPromise stay at React, Expo, and app-runtime boundaries.

Four convention violations in the changed scope are noted inline: Effect.catchTag in the two new server voice modules, two pure compatibility re-export shims left behind by the web-to-client-runtime move, and a non-canonical layer export name for a new Context.Service.

Posted via Macroscope — Effect Service Conventions

Comment threadpackages/client-runtime/src/state/voiceHttp.ts Outdated
Comment threadapps/server/src/voice/http.ts
Comment threadapps/web/src/voice/voiceTools.ts Outdated
Comment threadapps/web/src/voice/realtimeEvents.ts Outdated
Comment threadapps/server/src/voice/OpenAiRealtime.ts Outdated
@t3dotgg

Copy link
Copy Markdown
Member

Note

🤖 GPT-5.6 Sol responding on behalf of Theo

Closing this PR after an automated pass over open pull requests. Duplicates the maintainer-owned voice implementation in #5213.

@t3dotggt3dotgg closed this Aug 23, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXL1,000+ changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@duncan-vc@t3dotgg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat: add cross-platform voice supervisor - #6206

Closed
duncan-vc wants to merge 24 commits into
pingdotgg:mainfrom
duncan-vc:feat/voice-supervisor
Closed

feat: add cross-platform voice supervisor#6206
duncan-vc wants to merge 24 commits into
pingdotgg:mainfrom
duncan-vc:feat/voice-supervisor

Conversation

@duncan-vc

@duncan-vcduncan-vc commented Aug 11, 2026

Copy link
Copy Markdown
Contributor

What changed

  • Add typed OpenAI Realtime session and credential handling on the T3 server.
  • Add a generation-scoped Voice Supervisor core with bounded transcripts, replay protection, local mutation confirmations, and exact environment/thread revalidation.
  • Add web and desktop Voice Supervisor UI, Voice settings, entry points, keybinding support, and macOS microphone permission handling.
  • Add foreground-only iOS and Android support using WebRTC plus a narrowly scoped native audio-session module.
  • Share platform-neutral tools, repository, transport, host, event, and state logic through @t3tools/client-runtime.
  • Document user behavior, architecture, remote/multi-environment semantics, and privacy disclosures.

Why

T3 can coordinate multiple coding agents and environments, but doing so currently requires staying at the keyboard. Voice Supervisor adds an explicit, reviewable speech interface for listing work, checking status, opening threads, and proposing confirmed commands while preserving T3's existing environment authority and receipt-backed execution model.

This is a full Realtime conversation rather than dictation. It uses the host environment's OpenAI API credential or OPENAI_API_KEY; it does not reuse ChatGPT subscription entitlements.

Behavior and safety

  • Opening the panel or mobile route never requests microphone access or mints a client secret; only explicit Start does.
  • API credentials remain on the selected T3 environment. Clients receive short-lived Realtime secrets and connect directly to OpenAI over WebRTC.
  • Mutating tools require local button confirmation tied to the exact voice generation, call, environment, target, and captured version.
  • Stop, host loss, terminal transport failure, mobile backgrounding, and native audio interruption tear down owned media and pending work without automatic reconnect.
  • Tool context is bounded, IDs and raw workspace paths stay opaque, and disconnected/stale targets are limited to list context until live revalidation succeeds.

Verification

  • vp i --frozen-lockfile --force — dependency and supply-chain verification passed; the installed WebRTC package contains the exact Android and iOS native pins.
  • vp test run <54 changed test files> — 54 files, 682 tests passed.
  • Targeted typechecks passed for contracts, shared, client-runtime, server, web, desktop, mobile, and marketing.
  • Targeted lint passed with zero diagnostics.
  • Formatting passed across 136 supported changed files.
  • git diff --check origin/main...HEAD passed.

Draft status

The source and focused verification gates are green. A rebuilt forked desktop/mobile client, real OpenAI Realtime smoke, and final UI screenshots are intentionally still pending and will be added before requesting maintainer review.

Related prior Realtime panel exploration: #3997.

Implemented with GPT-5.6 Sol through the Codex harness in T3 Code.

Note

Add cross-platform Voice Supervisor for real-time AI voice conversations

  • Introduces a full Voice Supervisor feature across web, desktop, and mobile that lets users start voice conversations with coding agents via the OpenAI Realtime API using WebRTC.
  • Adds server-side voice HTTP endpoints (/api/voice/openai/credential, /api/voice/realtime/client-secret) for credential management and ephemeral client-secret minting, with rate limiting, scope enforcement, and no-store headers.
  • Adds shared client-runtime modules for realtime session transport, event decoding, supervisor state reduction, tool schemas, and thread/voice supervisor repository operations.
  • Implements platform-specific realtime session adapters: a browser WebRTC adapter for web and a react-native-webrtc-backed adapter for mobile, each with audio session lifecycle management.
  • Adds native Expo modules (T3VoiceAudioSession) for iOS and Android to observe audio interruptions, route loss, and media services reset events without taking audio session ownership.
  • Adds a /settings/voice page on web and a Voice Supervisor settings row on mobile for configuring the host environment and OpenAI API key.
  • Adds voice.toggle as a built-in keybinding command, wired into the command palette and sidebar on web.
  • Risk: react-native-webrtc@124.0.7 is added as a new mobile dependency and patched via pnpm-workspace.yaml; the Android manifest explicitly removes MediaProjectionService from the WebRTC module.
📊 Macroscope summarized 90f076e. 86 files reviewed, 0 issues evaluated, 0 issues filtered, 0 comments posted

🗂️ Filtered Issues

No issues evaluated.

@coderabbitai

coderabbitaiBot commented Aug 11, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: ae2bef0b-6f3e-432f-91b8-d00cf2f1697a

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:XXL 1,000+ changed lines (additions + deletions). labels Aug 11, 2026
Comment threadpackages/client-runtime/src/voice/voiceSupervisorState.ts Outdated
Comment threadpackages/client-runtime/src/operations/threadSupervisor.ts
Comment threadapps/mobile/src/voice/voiceStartDefaults.ts
Comment on lines +107 to +111
const upstreamRetryAfterSeconds = (value: string | undefined): number => {
if (value === undefined) return 1;
const parsed = Number.parseFloat(value);
return Number.isFinite(parsed) && parsed > 0 ? Math.min(120, Math.ceil(parsed)) : 1;
};

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Mediumvoice/OpenAiRealtime.ts:107

When a 429 response includes an HTTP-date in the Retry-After header (e.g. Tue, 11 Aug 2026 12:00:00 GMT), upstreamRetryAfterSeconds parses it with Number.parseFloat, which returns NaN. The function then falls back to 1, so callers receive retryAfterSeconds: 1 and retry before the upstream-specified time. HTTP permits Retry-After to be either a delta-seconds value or an HTTP-date; consider parsing both forms (e.g. falling back to Date.parse when the float parse fails).

 const upstreamRetryAfterSeconds = (value: string | undefined): number => {
if (value === undefined) return 1;
const parsed = Number.parseFloat(value);
- return Number.isFinite(parsed) && parsed > 0 ? Math.min(120, Math.ceil(parsed)) : 1;+ if (Number.isFinite(parsed) && parsed > 0) return Math.min(120, Math.ceil(parsed));+ const dateMs = Date.parse(value);+ if (Number.isFinite(dateMs)) {+ const delta = Math.ceil((dateMs - Date.now()) / 1_000);+ return delta > 0 ? Math.min(120, delta) : 1;+ }+ return 1;
};
🤖 Copy this AI Prompt to have your agent fix this:
In file @apps/server/src/voice/OpenAiRealtime.ts around lines 107-111:
When a 429 response includes an HTTP-date in the `Retry-After` header (e.g. `Tue, 11 Aug 2026 12:00:00 GMT`), `upstreamRetryAfterSeconds` parses it with `Number.parseFloat`, which returns `NaN`. The function then falls back to `1`, so callers receive `retryAfterSeconds: 1` and retry before the upstream-specified time. HTTP permits `Retry-After` to be either a delta-seconds value or an HTTP-date; consider parsing both forms (e.g. falling back to `Date.parse` when the float parse fails).

projectModelSelection: project.defaultModelSelection,
draft,
});
const primaryDefaults = dependencies.readPrimaryThreadDefaults();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟠 Highvoice/voiceStartDefaults.ts:253

primaryDefaults is read from readPrimaryThreadDefaults(), which returns the primary environment's settings. When project.environmentId targets a different environment, defaultThreadEnvMode (line 264) and newWorktreesStartFromOrigin (line 307) are resolved against the wrong environment. This can cause a voice command on a remote project to create a local checkout instead of a worktree (or vice versa), or to apply the wrong start-from-origin policy. The target environment's settings are already available in environment.settings — use that instead.

🤖 Copy this AI Prompt to have your agent fix this:
In file @apps/web/src/voice/voiceStartDefaults.ts around line 253:
`primaryDefaults` is read from `readPrimaryThreadDefaults()`, which returns the *primary* environment's settings. When `project.environmentId` targets a different environment, `defaultThreadEnvMode` (line 264) and `newWorktreesStartFromOrigin` (line 307) are resolved against the wrong environment. This can cause a voice command on a remote project to create a local checkout instead of a worktree (or vice versa), or to apply the wrong start-from-origin policy. The target environment's settings are already available in `environment.settings` — use that instead.

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the new/changed Effect service code in this PR (apps/server/src/voice/*, apps/desktop/src/electron/ElectronSystemPreferences.ts, packages/client-runtime/src/state/voiceHttp.ts, packages/contracts/src/voice.ts, and the web/mobile voice adapters).

Most of it follows the conventions: ElectronSystemPreferences.ts, OpenAiRealtime.ts, and OpenAiRealtimeCredential.ts all use the canonical imports/Context.Service/make/layer shape with inline interfaces, Schema.TaggedErrorClass failures, structured attributes, exported Schema.is predicates, and dependencies acquired from the environment; ManagedRuntime/runPromise stay at React, Expo, and app-runtime boundaries.

Four convention violations in the changed scope are noted inline: Effect.catchTag in the two new server voice modules, two pure compatibility re-export shims left behind by the web-to-client-runtime move, and a non-canonical layer export name for a new Context.Service.

Posted via Macroscope — Effect Service Conventions

Comment threadpackages/client-runtime/src/state/voiceHttp.ts Outdated
Comment threadapps/server/src/voice/http.ts
Comment threadapps/web/src/voice/voiceTools.ts Outdated
Comment threadapps/web/src/voice/realtimeEvents.ts Outdated
Comment threadapps/server/src/voice/OpenAiRealtime.ts Outdated
@t3dotgg

Copy link
Copy Markdown
Member

Note

🤖 GPT-5.6 Sol responding on behalf of Theo

Closing this PR after an automated pass over open pull requests. Duplicates the maintainer-owned voice implementation in #5213.

@t3dotggt3dotgg closed this Aug 23, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXL1,000+ changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@duncan-vc@t3dotgg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat: add cross-platform voice supervisor - #6206

Closed
duncan-vc wants to merge 24 commits into
pingdotgg:mainfrom
duncan-vc:feat/voice-supervisor
Closed

feat: add cross-platform voice supervisor#6206
duncan-vc wants to merge 24 commits into
pingdotgg:mainfrom
duncan-vc:feat/voice-supervisor

Conversation

@duncan-vc

@duncan-vcduncan-vc commented Aug 11, 2026

Copy link
Copy Markdown
Contributor

What changed

  • Add typed OpenAI Realtime session and credential handling on the T3 server.
  • Add a generation-scoped Voice Supervisor core with bounded transcripts, replay protection, local mutation confirmations, and exact environment/thread revalidation.
  • Add web and desktop Voice Supervisor UI, Voice settings, entry points, keybinding support, and macOS microphone permission handling.
  • Add foreground-only iOS and Android support using WebRTC plus a narrowly scoped native audio-session module.
  • Share platform-neutral tools, repository, transport, host, event, and state logic through @t3tools/client-runtime.
  • Document user behavior, architecture, remote/multi-environment semantics, and privacy disclosures.

Why

T3 can coordinate multiple coding agents and environments, but doing so currently requires staying at the keyboard. Voice Supervisor adds an explicit, reviewable speech interface for listing work, checking status, opening threads, and proposing confirmed commands while preserving T3's existing environment authority and receipt-backed execution model.

This is a full Realtime conversation rather than dictation. It uses the host environment's OpenAI API credential or OPENAI_API_KEY; it does not reuse ChatGPT subscription entitlements.

Behavior and safety

  • Opening the panel or mobile route never requests microphone access or mints a client secret; only explicit Start does.
  • API credentials remain on the selected T3 environment. Clients receive short-lived Realtime secrets and connect directly to OpenAI over WebRTC.
  • Mutating tools require local button confirmation tied to the exact voice generation, call, environment, target, and captured version.
  • Stop, host loss, terminal transport failure, mobile backgrounding, and native audio interruption tear down owned media and pending work without automatic reconnect.
  • Tool context is bounded, IDs and raw workspace paths stay opaque, and disconnected/stale targets are limited to list context until live revalidation succeeds.

Verification

  • vp i --frozen-lockfile --force — dependency and supply-chain verification passed; the installed WebRTC package contains the exact Android and iOS native pins.
  • vp test run <54 changed test files> — 54 files, 682 tests passed.
  • Targeted typechecks passed for contracts, shared, client-runtime, server, web, desktop, mobile, and marketing.
  • Targeted lint passed with zero diagnostics.
  • Formatting passed across 136 supported changed files.
  • git diff --check origin/main...HEAD passed.

Draft status

The source and focused verification gates are green. A rebuilt forked desktop/mobile client, real OpenAI Realtime smoke, and final UI screenshots are intentionally still pending and will be added before requesting maintainer review.

Related prior Realtime panel exploration: #3997.

Implemented with GPT-5.6 Sol through the Codex harness in T3 Code.

Note

Add cross-platform Voice Supervisor for real-time AI voice conversations

  • Introduces a full Voice Supervisor feature across web, desktop, and mobile that lets users start voice conversations with coding agents via the OpenAI Realtime API using WebRTC.
  • Adds server-side voice HTTP endpoints (/api/voice/openai/credential, /api/voice/realtime/client-secret) for credential management and ephemeral client-secret minting, with rate limiting, scope enforcement, and no-store headers.
  • Adds shared client-runtime modules for realtime session transport, event decoding, supervisor state reduction, tool schemas, and thread/voice supervisor repository operations.
  • Implements platform-specific realtime session adapters: a browser WebRTC adapter for web and a react-native-webrtc-backed adapter for mobile, each with audio session lifecycle management.
  • Adds native Expo modules (T3VoiceAudioSession) for iOS and Android to observe audio interruptions, route loss, and media services reset events without taking audio session ownership.
  • Adds a /settings/voice page on web and a Voice Supervisor settings row on mobile for configuring the host environment and OpenAI API key.
  • Adds voice.toggle as a built-in keybinding command, wired into the command palette and sidebar on web.
  • Risk: react-native-webrtc@124.0.7 is added as a new mobile dependency and patched via pnpm-workspace.yaml; the Android manifest explicitly removes MediaProjectionService from the WebRTC module.
📊 Macroscope summarized 90f076e. 86 files reviewed, 0 issues evaluated, 0 issues filtered, 0 comments posted

🗂️ Filtered Issues

No issues evaluated.

@coderabbitai

coderabbitaiBot commented Aug 11, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: ae2bef0b-6f3e-432f-91b8-d00cf2f1697a

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:XXL 1,000+ changed lines (additions + deletions). labels Aug 11, 2026
Comment threadpackages/client-runtime/src/voice/voiceSupervisorState.ts Outdated
Comment threadpackages/client-runtime/src/operations/threadSupervisor.ts
Comment threadapps/mobile/src/voice/voiceStartDefaults.ts
Comment on lines +107 to +111
const upstreamRetryAfterSeconds = (value: string | undefined): number => {
if (value === undefined) return 1;
const parsed = Number.parseFloat(value);
return Number.isFinite(parsed) && parsed > 0 ? Math.min(120, Math.ceil(parsed)) : 1;
};

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Mediumvoice/OpenAiRealtime.ts:107

When a 429 response includes an HTTP-date in the Retry-After header (e.g. Tue, 11 Aug 2026 12:00:00 GMT), upstreamRetryAfterSeconds parses it with Number.parseFloat, which returns NaN. The function then falls back to 1, so callers receive retryAfterSeconds: 1 and retry before the upstream-specified time. HTTP permits Retry-After to be either a delta-seconds value or an HTTP-date; consider parsing both forms (e.g. falling back to Date.parse when the float parse fails).

 const upstreamRetryAfterSeconds = (value: string | undefined): number => {
if (value === undefined) return 1;
const parsed = Number.parseFloat(value);
- return Number.isFinite(parsed) && parsed > 0 ? Math.min(120, Math.ceil(parsed)) : 1;+ if (Number.isFinite(parsed) && parsed > 0) return Math.min(120, Math.ceil(parsed));+ const dateMs = Date.parse(value);+ if (Number.isFinite(dateMs)) {+ const delta = Math.ceil((dateMs - Date.now()) / 1_000);+ return delta > 0 ? Math.min(120, delta) : 1;+ }+ return 1;
};
🤖 Copy this AI Prompt to have your agent fix this:
In file @apps/server/src/voice/OpenAiRealtime.ts around lines 107-111:
When a 429 response includes an HTTP-date in the `Retry-After` header (e.g. `Tue, 11 Aug 2026 12:00:00 GMT`), `upstreamRetryAfterSeconds` parses it with `Number.parseFloat`, which returns `NaN`. The function then falls back to `1`, so callers receive `retryAfterSeconds: 1` and retry before the upstream-specified time. HTTP permits `Retry-After` to be either a delta-seconds value or an HTTP-date; consider parsing both forms (e.g. falling back to `Date.parse` when the float parse fails).

projectModelSelection: project.defaultModelSelection,
draft,
});
const primaryDefaults = dependencies.readPrimaryThreadDefaults();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟠 Highvoice/voiceStartDefaults.ts:253

primaryDefaults is read from readPrimaryThreadDefaults(), which returns the primary environment's settings. When project.environmentId targets a different environment, defaultThreadEnvMode (line 264) and newWorktreesStartFromOrigin (line 307) are resolved against the wrong environment. This can cause a voice command on a remote project to create a local checkout instead of a worktree (or vice versa), or to apply the wrong start-from-origin policy. The target environment's settings are already available in environment.settings — use that instead.

🤖 Copy this AI Prompt to have your agent fix this:
In file @apps/web/src/voice/voiceStartDefaults.ts around line 253:
`primaryDefaults` is read from `readPrimaryThreadDefaults()`, which returns the *primary* environment's settings. When `project.environmentId` targets a different environment, `defaultThreadEnvMode` (line 264) and `newWorktreesStartFromOrigin` (line 307) are resolved against the wrong environment. This can cause a voice command on a remote project to create a local checkout instead of a worktree (or vice versa), or to apply the wrong start-from-origin policy. The target environment's settings are already available in `environment.settings` — use that instead.

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the new/changed Effect service code in this PR (apps/server/src/voice/*, apps/desktop/src/electron/ElectronSystemPreferences.ts, packages/client-runtime/src/state/voiceHttp.ts, packages/contracts/src/voice.ts, and the web/mobile voice adapters).

Most of it follows the conventions: ElectronSystemPreferences.ts, OpenAiRealtime.ts, and OpenAiRealtimeCredential.ts all use the canonical imports/Context.Service/make/layer shape with inline interfaces, Schema.TaggedErrorClass failures, structured attributes, exported Schema.is predicates, and dependencies acquired from the environment; ManagedRuntime/runPromise stay at React, Expo, and app-runtime boundaries.

Four convention violations in the changed scope are noted inline: Effect.catchTag in the two new server voice modules, two pure compatibility re-export shims left behind by the web-to-client-runtime move, and a non-canonical layer export name for a new Context.Service.

Posted via Macroscope — Effect Service Conventions

Comment threadpackages/client-runtime/src/state/voiceHttp.ts Outdated
Comment threadapps/server/src/voice/http.ts
Comment threadapps/web/src/voice/voiceTools.ts Outdated
Comment threadapps/web/src/voice/realtimeEvents.ts Outdated
Comment threadapps/server/src/voice/OpenAiRealtime.ts Outdated
@t3dotgg

Copy link
Copy Markdown
Member

Note

🤖 GPT-5.6 Sol responding on behalf of Theo

Closing this PR after an automated pass over open pull requests. Duplicates the maintainer-owned voice implementation in #5213.

@t3dotggt3dotgg closed this Aug 23, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXL1,000+ changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@duncan-vc@t3dotgg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

feat: add cross-platform voice supervisor - #6206

Closed
duncan-vc wants to merge 24 commits into
pingdotgg:mainfrom
duncan-vc:feat/voice-supervisor
Closed

feat: add cross-platform voice supervisor#6206
duncan-vc wants to merge 24 commits into
pingdotgg:mainfrom
duncan-vc:feat/voice-supervisor

Conversation

@duncan-vc

@duncan-vcduncan-vc commented Aug 11, 2026

Copy link
Copy Markdown
Contributor

What changed

  • Add typed OpenAI Realtime session and credential handling on the T3 server.
  • Add a generation-scoped Voice Supervisor core with bounded transcripts, replay protection, local mutation confirmations, and exact environment/thread revalidation.
  • Add web and desktop Voice Supervisor UI, Voice settings, entry points, keybinding support, and macOS microphone permission handling.
  • Add foreground-only iOS and Android support using WebRTC plus a narrowly scoped native audio-session module.
  • Share platform-neutral tools, repository, transport, host, event, and state logic through @t3tools/client-runtime.
  • Document user behavior, architecture, remote/multi-environment semantics, and privacy disclosures.

Why

T3 can coordinate multiple coding agents and environments, but doing so currently requires staying at the keyboard. Voice Supervisor adds an explicit, reviewable speech interface for listing work, checking status, opening threads, and proposing confirmed commands while preserving T3's existing environment authority and receipt-backed execution model.

This is a full Realtime conversation rather than dictation. It uses the host environment's OpenAI API credential or OPENAI_API_KEY; it does not reuse ChatGPT subscription entitlements.

Behavior and safety

  • Opening the panel or mobile route never requests microphone access or mints a client secret; only explicit Start does.
  • API credentials remain on the selected T3 environment. Clients receive short-lived Realtime secrets and connect directly to OpenAI over WebRTC.
  • Mutating tools require local button confirmation tied to the exact voice generation, call, environment, target, and captured version.
  • Stop, host loss, terminal transport failure, mobile backgrounding, and native audio interruption tear down owned media and pending work without automatic reconnect.
  • Tool context is bounded, IDs and raw workspace paths stay opaque, and disconnected/stale targets are limited to list context until live revalidation succeeds.

Verification

  • vp i --frozen-lockfile --force — dependency and supply-chain verification passed; the installed WebRTC package contains the exact Android and iOS native pins.
  • vp test run <54 changed test files> — 54 files, 682 tests passed.
  • Targeted typechecks passed for contracts, shared, client-runtime, server, web, desktop, mobile, and marketing.
  • Targeted lint passed with zero diagnostics.
  • Formatting passed across 136 supported changed files.
  • git diff --check origin/main...HEAD passed.

Draft status

The source and focused verification gates are green. A rebuilt forked desktop/mobile client, real OpenAI Realtime smoke, and final UI screenshots are intentionally still pending and will be added before requesting maintainer review.

Related prior Realtime panel exploration: #3997.

Implemented with GPT-5.6 Sol through the Codex harness in T3 Code.

Note

Add cross-platform Voice Supervisor for real-time AI voice conversations

  • Introduces a full Voice Supervisor feature across web, desktop, and mobile that lets users start voice conversations with coding agents via the OpenAI Realtime API using WebRTC.
  • Adds server-side voice HTTP endpoints (/api/voice/openai/credential, /api/voice/realtime/client-secret) for credential management and ephemeral client-secret minting, with rate limiting, scope enforcement, and no-store headers.
  • Adds shared client-runtime modules for realtime session transport, event decoding, supervisor state reduction, tool schemas, and thread/voice supervisor repository operations.
  • Implements platform-specific realtime session adapters: a browser WebRTC adapter for web and a react-native-webrtc-backed adapter for mobile, each with audio session lifecycle management.
  • Adds native Expo modules (T3VoiceAudioSession) for iOS and Android to observe audio interruptions, route loss, and media services reset events without taking audio session ownership.
  • Adds a /settings/voice page on web and a Voice Supervisor settings row on mobile for configuring the host environment and OpenAI API key.
  • Adds voice.toggle as a built-in keybinding command, wired into the command palette and sidebar on web.
  • Risk: react-native-webrtc@124.0.7 is added as a new mobile dependency and patched via pnpm-workspace.yaml; the Android manifest explicitly removes MediaProjectionService from the WebRTC module.
📊 Macroscope summarized 90f076e. 86 files reviewed, 0 issues evaluated, 0 issues filtered, 0 comments posted

🗂️ Filtered Issues

No issues evaluated.

@coderabbitai

coderabbitaiBot commented Aug 11, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: ae2bef0b-6f3e-432f-91b8-d00cf2f1697a

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:XXL 1,000+ changed lines (additions + deletions). labels Aug 11, 2026
Comment threadpackages/client-runtime/src/voice/voiceSupervisorState.ts Outdated
Comment threadpackages/client-runtime/src/operations/threadSupervisor.ts
Comment threadapps/mobile/src/voice/voiceStartDefaults.ts
Comment on lines +107 to +111
const upstreamRetryAfterSeconds = (value: string | undefined): number => {
if (value === undefined) return 1;
const parsed = Number.parseFloat(value);
return Number.isFinite(parsed) && parsed > 0 ? Math.min(120, Math.ceil(parsed)) : 1;
};

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Mediumvoice/OpenAiRealtime.ts:107

When a 429 response includes an HTTP-date in the Retry-After header (e.g. Tue, 11 Aug 2026 12:00:00 GMT), upstreamRetryAfterSeconds parses it with Number.parseFloat, which returns NaN. The function then falls back to 1, so callers receive retryAfterSeconds: 1 and retry before the upstream-specified time. HTTP permits Retry-After to be either a delta-seconds value or an HTTP-date; consider parsing both forms (e.g. falling back to Date.parse when the float parse fails).

 const upstreamRetryAfterSeconds = (value: string | undefined): number => {
if (value === undefined) return 1;
const parsed = Number.parseFloat(value);
- return Number.isFinite(parsed) && parsed > 0 ? Math.min(120, Math.ceil(parsed)) : 1;+ if (Number.isFinite(parsed) && parsed > 0) return Math.min(120, Math.ceil(parsed));+ const dateMs = Date.parse(value);+ if (Number.isFinite(dateMs)) {+ const delta = Math.ceil((dateMs - Date.now()) / 1_000);+ return delta > 0 ? Math.min(120, delta) : 1;+ }+ return 1;
};
🤖 Copy this AI Prompt to have your agent fix this:
In file @apps/server/src/voice/OpenAiRealtime.ts around lines 107-111:
When a 429 response includes an HTTP-date in the `Retry-After` header (e.g. `Tue, 11 Aug 2026 12:00:00 GMT`), `upstreamRetryAfterSeconds` parses it with `Number.parseFloat`, which returns `NaN`. The function then falls back to `1`, so callers receive `retryAfterSeconds: 1` and retry before the upstream-specified time. HTTP permits `Retry-After` to be either a delta-seconds value or an HTTP-date; consider parsing both forms (e.g. falling back to `Date.parse` when the float parse fails).

projectModelSelection: project.defaultModelSelection,
draft,
});
const primaryDefaults = dependencies.readPrimaryThreadDefaults();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟠 Highvoice/voiceStartDefaults.ts:253

primaryDefaults is read from readPrimaryThreadDefaults(), which returns the primary environment's settings. When project.environmentId targets a different environment, defaultThreadEnvMode (line 264) and newWorktreesStartFromOrigin (line 307) are resolved against the wrong environment. This can cause a voice command on a remote project to create a local checkout instead of a worktree (or vice versa), or to apply the wrong start-from-origin policy. The target environment's settings are already available in environment.settings — use that instead.

🤖 Copy this AI Prompt to have your agent fix this:
In file @apps/web/src/voice/voiceStartDefaults.ts around line 253:
`primaryDefaults` is read from `readPrimaryThreadDefaults()`, which returns the *primary* environment's settings. When `project.environmentId` targets a different environment, `defaultThreadEnvMode` (line 264) and `newWorktreesStartFromOrigin` (line 307) are resolved against the wrong environment. This can cause a voice command on a remote project to create a local checkout instead of a worktree (or vice versa), or to apply the wrong start-from-origin policy. The target environment's settings are already available in `environment.settings` — use that instead.

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the new/changed Effect service code in this PR (apps/server/src/voice/*, apps/desktop/src/electron/ElectronSystemPreferences.ts, packages/client-runtime/src/state/voiceHttp.ts, packages/contracts/src/voice.ts, and the web/mobile voice adapters).

Most of it follows the conventions: ElectronSystemPreferences.ts, OpenAiRealtime.ts, and OpenAiRealtimeCredential.ts all use the canonical imports/Context.Service/make/layer shape with inline interfaces, Schema.TaggedErrorClass failures, structured attributes, exported Schema.is predicates, and dependencies acquired from the environment; ManagedRuntime/runPromise stay at React, Expo, and app-runtime boundaries.

Four convention violations in the changed scope are noted inline: Effect.catchTag in the two new server voice modules, two pure compatibility re-export shims left behind by the web-to-client-runtime move, and a non-canonical layer export name for a new Context.Service.

Posted via Macroscope — Effect Service Conventions

Comment threadpackages/client-runtime/src/state/voiceHttp.ts Outdated
Comment threadapps/server/src/voice/http.ts
Comment threadapps/web/src/voice/voiceTools.ts Outdated
Comment threadapps/web/src/voice/realtimeEvents.ts Outdated
Comment threadapps/server/src/voice/OpenAiRealtime.ts Outdated
@t3dotgg

Copy link
Copy Markdown
Member

Note

🤖 GPT-5.6 Sol responding on behalf of Theo

Closing this PR after an automated pass over open pull requests. Duplicates the maintainer-owned voice implementation in #5213.

@t3dotggt3dotgg closed this Aug 23, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXL1,000+ changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@duncan-vc@t3dotgg