feat: add Codex-style voice dictation to web and mobile - #6625

Closed
KachurPro wants to merge 20 commits into
pingdotgg:mainfrom
KachurPro:codex/voice-dictation-codex-ux
Closed

feat: add Codex-style voice dictation to web and mobile#6625
KachurPro wants to merge 20 commits into
pingdotgg:mainfrom
KachurPro:codex/voice-dictation-codex-ux

Conversation

@KachurPro

@KachurProKachurPro commented Aug 14, 2026

Copy link
Copy Markdown

Problem

The existing voice dictation PR (#5213) only covered web/desktop and had become stale. Native mobile dictation was still missing, and the first mobile version was limited to a hard-coded OpenAI endpoint and model.

Solution

  • rebase the voice transcription work onto the current main
  • match the Codex-style actions: cancel/discard, stop/transcribe/insert, and transcribe/insert/send
  • always append the normalized transcript to the end of the current draft
  • add native iOS/Android recording with expo-audio
  • let mobile users choose OpenAI or Groq
  • store a separate API key and model for each provider in expo-secure-store, with migration of the legacy OpenAI key
  • default OpenAI to gpt-4o-transcribe and Groq to whisper-large-v3-turbo
  • offer curated model choices while also accepting a custom model ID
  • send mobile recordings directly to the official provider endpoint
  • preserve the existing OpenAI/Groq server proxy and model picker on web/desktop
  • share transcript append and terminal-action rules between web and mobile
  • fully reset recorder, stream, timer, and audio-mode state after every terminal path

This supersedes #5213 and carries its web/desktop behavior forward.

Interaction

  • X: discard the recording
  • stop: transcribe and append to the draft
  • send: transcribe, append, then use the normal composer send path
  • existing draft text is preserved; speech is appended at the end

Verification

  • all 716 mobile tests passed
  • 4 focused web transcription tests passed
  • mobile and web typechecks passed
  • full repository lint and formatting checks passed
  • mobile native static analysis passed
  • Expo production config resolves expo-audio, the microphone permission, and Secure Store

The current development machine does not have Xcode or an iOS Simulator installed, so a local simulator screenshot is not included. The native flow still needs the EAS preview/device pass produced by CI.

Built with GPT-5.6 Sol in Codex Desktop.


Note

Medium Risk
New authenticated transcription proxy accepts client-supplied API keys in custom headers and forwards audio to third parties; composer send/stop paths gain async recording state, and mobile sends audio and keys directly to providers.

Overview
Adds Codex-style voice dictation across web/desktop chat, native mobile thread composer, and server-backed transcription.

Web/desktop: New Voice dictation settings (OpenAI or Groq, client API key or server OPENAI_API_KEY / GROQ_API_KEY, model picker). The composer gets a mic, a recording UI (waveform, cancel / stop / transcribe-and-send), and send behavior that waits on transcription when recording. Audio goes through authenticated /api/transcription and /api/transcription/models on the connected T3 server (25 MB cap). Shared appendVoiceTranscript and terminal-action rules live in @t3tools/shared. Desktop gains macOS microphone usage text.

Mobile:expo-audio recording, per-provider keys/models in Secure Store (with legacy OpenAI key migration), a Voice Dictation settings screen, and direct uploads to OpenAI/Groq (not the server proxy). Composer integration mirrors web (thread-scoped append, cancel on navigation).

Contracts: Client settings add voiceTranscription* fields; reset-to-defaults includes voice dictation when dirty.

Reviewed by Cursor Bugbot for commit ff3706c. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Add Codex-style voice dictation to web and mobile chat composers

  • Adds microphone-based voice dictation to the web (ChatComposer) and mobile (ThreadComposer) chat composers, replacing the toolbar with a live waveform panel during recording/transcribing and appending the transcript to the draft on stop.
  • Supports OpenAI and Groq as transcription providers; users configure their provider, API key, and model in General settings (web) or a new Voice Dictation settings screen (mobile).
  • Server-side transcription proxy routes (GET/POST /api/transcription, GET /api/transcription/models) forward audio to the selected provider with a 25 MiB size limit and 2-minute timeout, supporting both environment-level and user-supplied API keys.
  • Shared utilities in @t3tools/shared handle transcript appending (whitespace-aware) and terminal action resolution (abort > send > insert).
  • macOS desktop build now includes NSMicrophoneUsageDescription in Info.plist; the mobile Expo app adds the expo-audio plugin with microphone permission.
  • Risk: voice transcription defaults to disabled (voiceTranscriptionEnabled=false); users must configure a model to activate the mic button. Existing client settings are migrated from a legacy OpenAI key format on first load.

Macroscope summarized ff3706c.

@coderabbitai

coderabbitaiBot commented Aug 14, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 35fa00aa-c405-43c0-b5aa-e7bb99d0e03d

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:XXL 1,000+ changed lines (additions + deletions). labels Aug 14, 2026
Comment threadapps/web/src/components/chat/ChatComposer.tsx
Comment threadapps/mobile/src/features/voice-dictation/voiceTranscriptionSettings.ts Outdated
Comment threadapps/web/src/components/chat/ChatComposer.tsx
Comment threadapps/web/src/components/settings/SettingsPanels.tsx
@macroscopeapp

macroscopeappBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Needs human review

This PR introduces a complete new voice dictation feature across web, mobile, and server, including new external API integrations (OpenAI/Groq), new server endpoints, audio recording infrastructure, and API key handling. New features of this scope and complexity require human review.

You can customize Macroscope's approvability policy. Learn more.

apiKey: voiceTranscriptionSettings.apiKey,
onTranscriptInsert: appendVoiceTranscriptToDraft,
onTranscriptSend: (transcript) => {
if (appendVoiceTranscriptToDraft(transcript)) void sendCurrentDraft();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟠 Highthreads/ThreadComposer.tsx:585

The voice “send” action sends the previous draftMessage snapshot instead of the transcript, so the transcript remains in the composer (or nothing is sent). appendVoiceTranscriptToDraft only schedules onChangeDraftMessage, while sendCurrentDraft() immediately calls onSendMessage() without the new text; pass the updated draft through the send path or wait for the state update before sending.

🤖 Copy this AI Prompt to have your agent fix this:
In file @apps/mobile/src/features/threads/ThreadComposer.tsx around line 585:
The voice “send” action sends the previous `draftMessage` snapshot instead of the transcript, so the transcript remains in the composer (or nothing is sent). `appendVoiceTranscriptToDraft` only schedules `onChangeDraftMessage`, while `sendCurrentDraft()` immediately calls `onSendMessage()` without the new text; pass the updated draft through the send path or wait for the state update before sending.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Verified as a false positive. onChangeDraftMessage does not schedule React state: useThreadComposerState calls setComposerDraftText, which synchronously calls appAtomRegistry.set. onSendMessage then calls getComposerDraftSnapshot from that same registry, so the just-appended transcript is present before enqueue. The existing mobile composer-drafts test also covers synchronous latest-draft reads for send.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, I'm unable to act on this request because you do not have permissions within this repository.

@KachurPro
KachurProforce-pushed the codex/voice-dictation-codex-ux branch from 196ba3a to 5d1e3e3CompareAugust 14, 2026 20:13
Comment threadapps/web/src/hooks/useVoiceTranscription.ts Outdated
Comment threadapps/web/src/components/settings/SettingsPanels.tsx
Comment threadapps/web/src/hooks/useVoiceTranscription.ts
Comment threadapps/web/src/hooks/useVoiceTranscription.ts Outdated
Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx
Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated
@brvale97

Copy link
Copy Markdown

Really looking forward to this — mobile dictation is exactly what I've been missing on iPhone.

One request: the web/desktop implementation from #5213 supports both OpenAI and Groq with model selection, but the mobile path hardcodes the OpenAI endpoint and gpt-4o-mini-transcribe in mobileVoiceTranscription.ts. Groq's transcription API is OpenAI-compatible (https://api.groq.com/openai/v1/audio/transcriptions, e.g. whisper-large-v3-turbo), so carrying the provider choice over to mobile would mostly mean making the base URL + model configurable alongside the stored key. Would you consider that?

@KachurPro
KachurProforce-pushed the codex/voice-dictation-codex-ux branch from 59ba2fc to 495d058CompareAugust 14, 2026 21:37
@KachurPro

Copy link
Copy Markdown
Author

@brvale97 Implemented in 495d058bc. Mobile now supports both OpenAI and Groq with fixed official transcription endpoints, separate securely stored keys and model choices per provider, a model picker, and a custom model ID field so future compatible models do not require an app update. OpenAI now defaults to gpt-4o-transcribe; Groq defaults to whisper-large-v3-turbo. Existing saved OpenAI keys are migrated automatically. All 716 mobile tests and the mobile typecheck pass. Thanks for catching this.

Comment threadapps/mobile/src/features/voice-dictation/useMobileVoiceTranscription.ts Outdated
Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

UI consistency review of the web changes (apps/web/src/**). Three findings, all in the new dictation UI: a ticking live region that conflicts with the documented pattern in Sidebar.tsx, and two call sites that rebuild the existing ghost-muted button variant with a class override.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated
Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx Outdated
Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx Outdated
Comment threadapps/web/src/components/chat/ChatComposer.tsx

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the web-scope changes (ChatComposer.tsx, VoiceTranscriptionPanel.tsx, SettingsPanels.tsx, settingsSearch.ts, hooks/useVoiceTranscription.ts, lib/voiceTranscription.ts). The earlier findings on the panel (muted-ghost variant, live-region timer) are resolved. Two remaining consistency issues in the composer.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated
Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

UI consistency review: one finding on the web composer dictation panel. Previous findings (muted-ghost variant usage, live-region timer, responsive error inset, mic pointer-focus) are resolved in this head.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/ChatComposer.tsx
Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One finding on the dictation panel's pointer-focus contract in the composer footer. Prior comments (mic variant/pointer focus, error-row inset, ghost-muted, aria-hidden timer) are addressed in this revision.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the new Effect code (apps/server/src/transcription.ts, the new route layers in apps/server/src/http.ts, and the contracts/shared additions) against the service conventions. Error classes use Schema.TaggedErrorClass with structural attributes and preserved cause, failures are recovered with Effect.catchTags, HttpClient is acquired from the environment, and Effect modules are imported from their subpaths — all consistent with the conventions. One finding on the web client below.

Posted via Macroscope — Effect Service Conventions

Comment threadapps/web/src/lib/voiceTranscription.ts Outdated
@KachurPro
KachurProforce-pushed the codex/voice-dictation-codex-ux branch from 9f92517 to 4732d3aCompareAugust 14, 2026 22:03
Comment threadapps/web/src/components/settings/SettingsPanels.tsx Outdated

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 2 potential issues.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 4732d3a. Configure here.

Comment threadapps/mobile/src/features/threads/ThreadComposer.tsx Outdated
Comment threadapps/web/src/components/settings/SettingsPanels.tsx Outdated

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One finding: the dictation send action bypasses the composer's themeable send tokens. Everything else (pointer-focus contract, ghost-muted variants, aria-hidden timer, responsive error inset) looks consistent with the composer's existing patterns.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx Outdated
@t3dotgg

Copy link
Copy Markdown
Member

Note

🤖 GPT-5.6 Sol responding on behalf of Theo

Closing this PR after an automated pass over open pull requests. Duplicates the maintainer-owned voice implementation in #5213.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXL1,000+ changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@KachurPro@brvale97@t3dotgg@maria-rcks
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

feat: add Codex-style voice dictation to web and mobile - #6625

Closed
KachurPro wants to merge 20 commits into
pingdotgg:mainfrom
KachurPro:codex/voice-dictation-codex-ux
Closed

feat: add Codex-style voice dictation to web and mobile#6625
KachurPro wants to merge 20 commits into
pingdotgg:mainfrom
KachurPro:codex/voice-dictation-codex-ux

Conversation

@KachurPro

@KachurProKachurPro commented Aug 14, 2026

Copy link
Copy Markdown

Problem

The existing voice dictation PR (#5213) only covered web/desktop and had become stale. Native mobile dictation was still missing, and the first mobile version was limited to a hard-coded OpenAI endpoint and model.

Solution

  • rebase the voice transcription work onto the current main
  • match the Codex-style actions: cancel/discard, stop/transcribe/insert, and transcribe/insert/send
  • always append the normalized transcript to the end of the current draft
  • add native iOS/Android recording with expo-audio
  • let mobile users choose OpenAI or Groq
  • store a separate API key and model for each provider in expo-secure-store, with migration of the legacy OpenAI key
  • default OpenAI to gpt-4o-transcribe and Groq to whisper-large-v3-turbo
  • offer curated model choices while also accepting a custom model ID
  • send mobile recordings directly to the official provider endpoint
  • preserve the existing OpenAI/Groq server proxy and model picker on web/desktop
  • share transcript append and terminal-action rules between web and mobile
  • fully reset recorder, stream, timer, and audio-mode state after every terminal path

This supersedes #5213 and carries its web/desktop behavior forward.

Interaction

  • X: discard the recording
  • stop: transcribe and append to the draft
  • send: transcribe, append, then use the normal composer send path
  • existing draft text is preserved; speech is appended at the end

Verification

  • all 716 mobile tests passed
  • 4 focused web transcription tests passed
  • mobile and web typechecks passed
  • full repository lint and formatting checks passed
  • mobile native static analysis passed
  • Expo production config resolves expo-audio, the microphone permission, and Secure Store

The current development machine does not have Xcode or an iOS Simulator installed, so a local simulator screenshot is not included. The native flow still needs the EAS preview/device pass produced by CI.

Built with GPT-5.6 Sol in Codex Desktop.


Note

Medium Risk
New authenticated transcription proxy accepts client-supplied API keys in custom headers and forwards audio to third parties; composer send/stop paths gain async recording state, and mobile sends audio and keys directly to providers.

Overview
Adds Codex-style voice dictation across web/desktop chat, native mobile thread composer, and server-backed transcription.

Web/desktop: New Voice dictation settings (OpenAI or Groq, client API key or server OPENAI_API_KEY / GROQ_API_KEY, model picker). The composer gets a mic, a recording UI (waveform, cancel / stop / transcribe-and-send), and send behavior that waits on transcription when recording. Audio goes through authenticated /api/transcription and /api/transcription/models on the connected T3 server (25 MB cap). Shared appendVoiceTranscript and terminal-action rules live in @t3tools/shared. Desktop gains macOS microphone usage text.

Mobile:expo-audio recording, per-provider keys/models in Secure Store (with legacy OpenAI key migration), a Voice Dictation settings screen, and direct uploads to OpenAI/Groq (not the server proxy). Composer integration mirrors web (thread-scoped append, cancel on navigation).

Contracts: Client settings add voiceTranscription* fields; reset-to-defaults includes voice dictation when dirty.

Reviewed by Cursor Bugbot for commit ff3706c. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Add Codex-style voice dictation to web and mobile chat composers

  • Adds microphone-based voice dictation to the web (ChatComposer) and mobile (ThreadComposer) chat composers, replacing the toolbar with a live waveform panel during recording/transcribing and appending the transcript to the draft on stop.
  • Supports OpenAI and Groq as transcription providers; users configure their provider, API key, and model in General settings (web) or a new Voice Dictation settings screen (mobile).
  • Server-side transcription proxy routes (GET/POST /api/transcription, GET /api/transcription/models) forward audio to the selected provider with a 25 MiB size limit and 2-minute timeout, supporting both environment-level and user-supplied API keys.
  • Shared utilities in @t3tools/shared handle transcript appending (whitespace-aware) and terminal action resolution (abort > send > insert).
  • macOS desktop build now includes NSMicrophoneUsageDescription in Info.plist; the mobile Expo app adds the expo-audio plugin with microphone permission.
  • Risk: voice transcription defaults to disabled (voiceTranscriptionEnabled=false); users must configure a model to activate the mic button. Existing client settings are migrated from a legacy OpenAI key format on first load.

Macroscope summarized ff3706c.

@coderabbitai

coderabbitaiBot commented Aug 14, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 35fa00aa-c405-43c0-b5aa-e7bb99d0e03d

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:XXL 1,000+ changed lines (additions + deletions). labels Aug 14, 2026
Comment threadapps/web/src/components/chat/ChatComposer.tsx
Comment threadapps/mobile/src/features/voice-dictation/voiceTranscriptionSettings.ts Outdated
Comment threadapps/web/src/components/chat/ChatComposer.tsx
Comment threadapps/web/src/components/settings/SettingsPanels.tsx
@macroscopeapp

macroscopeappBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Needs human review

This PR introduces a complete new voice dictation feature across web, mobile, and server, including new external API integrations (OpenAI/Groq), new server endpoints, audio recording infrastructure, and API key handling. New features of this scope and complexity require human review.

You can customize Macroscope's approvability policy. Learn more.

apiKey: voiceTranscriptionSettings.apiKey,
onTranscriptInsert: appendVoiceTranscriptToDraft,
onTranscriptSend: (transcript) => {
if (appendVoiceTranscriptToDraft(transcript)) void sendCurrentDraft();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟠 Highthreads/ThreadComposer.tsx:585

The voice “send” action sends the previous draftMessage snapshot instead of the transcript, so the transcript remains in the composer (or nothing is sent). appendVoiceTranscriptToDraft only schedules onChangeDraftMessage, while sendCurrentDraft() immediately calls onSendMessage() without the new text; pass the updated draft through the send path or wait for the state update before sending.

🤖 Copy this AI Prompt to have your agent fix this:
In file @apps/mobile/src/features/threads/ThreadComposer.tsx around line 585:
The voice “send” action sends the previous `draftMessage` snapshot instead of the transcript, so the transcript remains in the composer (or nothing is sent). `appendVoiceTranscriptToDraft` only schedules `onChangeDraftMessage`, while `sendCurrentDraft()` immediately calls `onSendMessage()` without the new text; pass the updated draft through the send path or wait for the state update before sending.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Verified as a false positive. onChangeDraftMessage does not schedule React state: useThreadComposerState calls setComposerDraftText, which synchronously calls appAtomRegistry.set. onSendMessage then calls getComposerDraftSnapshot from that same registry, so the just-appended transcript is present before enqueue. The existing mobile composer-drafts test also covers synchronous latest-draft reads for send.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, I'm unable to act on this request because you do not have permissions within this repository.

@KachurPro
KachurProforce-pushed the codex/voice-dictation-codex-ux branch from 196ba3a to 5d1e3e3CompareAugust 14, 2026 20:13
Comment threadapps/web/src/hooks/useVoiceTranscription.ts Outdated
Comment threadapps/web/src/components/settings/SettingsPanels.tsx
Comment threadapps/web/src/hooks/useVoiceTranscription.ts
Comment threadapps/web/src/hooks/useVoiceTranscription.ts Outdated
Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx
Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated
@brvale97

Copy link
Copy Markdown

Really looking forward to this — mobile dictation is exactly what I've been missing on iPhone.

One request: the web/desktop implementation from #5213 supports both OpenAI and Groq with model selection, but the mobile path hardcodes the OpenAI endpoint and gpt-4o-mini-transcribe in mobileVoiceTranscription.ts. Groq's transcription API is OpenAI-compatible (https://api.groq.com/openai/v1/audio/transcriptions, e.g. whisper-large-v3-turbo), so carrying the provider choice over to mobile would mostly mean making the base URL + model configurable alongside the stored key. Would you consider that?

@KachurPro
KachurProforce-pushed the codex/voice-dictation-codex-ux branch from 59ba2fc to 495d058CompareAugust 14, 2026 21:37
@KachurPro

Copy link
Copy Markdown
Author

@brvale97 Implemented in 495d058bc. Mobile now supports both OpenAI and Groq with fixed official transcription endpoints, separate securely stored keys and model choices per provider, a model picker, and a custom model ID field so future compatible models do not require an app update. OpenAI now defaults to gpt-4o-transcribe; Groq defaults to whisper-large-v3-turbo. Existing saved OpenAI keys are migrated automatically. All 716 mobile tests and the mobile typecheck pass. Thanks for catching this.

Comment threadapps/mobile/src/features/voice-dictation/useMobileVoiceTranscription.ts Outdated
Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

UI consistency review of the web changes (apps/web/src/**). Three findings, all in the new dictation UI: a ticking live region that conflicts with the documented pattern in Sidebar.tsx, and two call sites that rebuild the existing ghost-muted button variant with a class override.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated
Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx Outdated
Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx Outdated
Comment threadapps/web/src/components/chat/ChatComposer.tsx

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the web-scope changes (ChatComposer.tsx, VoiceTranscriptionPanel.tsx, SettingsPanels.tsx, settingsSearch.ts, hooks/useVoiceTranscription.ts, lib/voiceTranscription.ts). The earlier findings on the panel (muted-ghost variant, live-region timer) are resolved. Two remaining consistency issues in the composer.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated
Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

UI consistency review: one finding on the web composer dictation panel. Previous findings (muted-ghost variant usage, live-region timer, responsive error inset, mic pointer-focus) are resolved in this head.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/ChatComposer.tsx
Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One finding on the dictation panel's pointer-focus contract in the composer footer. Prior comments (mic variant/pointer focus, error-row inset, ghost-muted, aria-hidden timer) are addressed in this revision.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the new Effect code (apps/server/src/transcription.ts, the new route layers in apps/server/src/http.ts, and the contracts/shared additions) against the service conventions. Error classes use Schema.TaggedErrorClass with structural attributes and preserved cause, failures are recovered with Effect.catchTags, HttpClient is acquired from the environment, and Effect modules are imported from their subpaths — all consistent with the conventions. One finding on the web client below.

Posted via Macroscope — Effect Service Conventions

Comment threadapps/web/src/lib/voiceTranscription.ts Outdated
@KachurPro
KachurProforce-pushed the codex/voice-dictation-codex-ux branch from 9f92517 to 4732d3aCompareAugust 14, 2026 22:03
Comment threadapps/web/src/components/settings/SettingsPanels.tsx Outdated

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 2 potential issues.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 4732d3a. Configure here.

Comment threadapps/mobile/src/features/threads/ThreadComposer.tsx Outdated
Comment threadapps/web/src/components/settings/SettingsPanels.tsx Outdated

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One finding: the dictation send action bypasses the composer's themeable send tokens. Everything else (pointer-focus contract, ghost-muted variants, aria-hidden timer, responsive error inset) looks consistent with the composer's existing patterns.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx Outdated
@t3dotgg

Copy link
Copy Markdown
Member

Note

🤖 GPT-5.6 Sol responding on behalf of Theo

Closing this PR after an automated pass over open pull requests. Duplicates the maintainer-owned voice implementation in #5213.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXL1,000+ changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@KachurPro@brvale97@t3dotgg@maria-rcks
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat: add Codex-style voice dictation to web and mobile - #6625

Closed
KachurPro wants to merge 20 commits into
pingdotgg:mainfrom
KachurPro:codex/voice-dictation-codex-ux
Closed

feat: add Codex-style voice dictation to web and mobile#6625
KachurPro wants to merge 20 commits into
pingdotgg:mainfrom
KachurPro:codex/voice-dictation-codex-ux

Conversation

@KachurPro

@KachurProKachurPro commented Aug 14, 2026

Copy link
Copy Markdown

Problem

The existing voice dictation PR (#5213) only covered web/desktop and had become stale. Native mobile dictation was still missing, and the first mobile version was limited to a hard-coded OpenAI endpoint and model.

Solution

  • rebase the voice transcription work onto the current main
  • match the Codex-style actions: cancel/discard, stop/transcribe/insert, and transcribe/insert/send
  • always append the normalized transcript to the end of the current draft
  • add native iOS/Android recording with expo-audio
  • let mobile users choose OpenAI or Groq
  • store a separate API key and model for each provider in expo-secure-store, with migration of the legacy OpenAI key
  • default OpenAI to gpt-4o-transcribe and Groq to whisper-large-v3-turbo
  • offer curated model choices while also accepting a custom model ID
  • send mobile recordings directly to the official provider endpoint
  • preserve the existing OpenAI/Groq server proxy and model picker on web/desktop
  • share transcript append and terminal-action rules between web and mobile
  • fully reset recorder, stream, timer, and audio-mode state after every terminal path

This supersedes #5213 and carries its web/desktop behavior forward.

Interaction

  • X: discard the recording
  • stop: transcribe and append to the draft
  • send: transcribe, append, then use the normal composer send path
  • existing draft text is preserved; speech is appended at the end

Verification

  • all 716 mobile tests passed
  • 4 focused web transcription tests passed
  • mobile and web typechecks passed
  • full repository lint and formatting checks passed
  • mobile native static analysis passed
  • Expo production config resolves expo-audio, the microphone permission, and Secure Store

The current development machine does not have Xcode or an iOS Simulator installed, so a local simulator screenshot is not included. The native flow still needs the EAS preview/device pass produced by CI.

Built with GPT-5.6 Sol in Codex Desktop.


Note

Medium Risk
New authenticated transcription proxy accepts client-supplied API keys in custom headers and forwards audio to third parties; composer send/stop paths gain async recording state, and mobile sends audio and keys directly to providers.

Overview
Adds Codex-style voice dictation across web/desktop chat, native mobile thread composer, and server-backed transcription.

Web/desktop: New Voice dictation settings (OpenAI or Groq, client API key or server OPENAI_API_KEY / GROQ_API_KEY, model picker). The composer gets a mic, a recording UI (waveform, cancel / stop / transcribe-and-send), and send behavior that waits on transcription when recording. Audio goes through authenticated /api/transcription and /api/transcription/models on the connected T3 server (25 MB cap). Shared appendVoiceTranscript and terminal-action rules live in @t3tools/shared. Desktop gains macOS microphone usage text.

Mobile:expo-audio recording, per-provider keys/models in Secure Store (with legacy OpenAI key migration), a Voice Dictation settings screen, and direct uploads to OpenAI/Groq (not the server proxy). Composer integration mirrors web (thread-scoped append, cancel on navigation).

Contracts: Client settings add voiceTranscription* fields; reset-to-defaults includes voice dictation when dirty.

Reviewed by Cursor Bugbot for commit ff3706c. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Add Codex-style voice dictation to web and mobile chat composers

  • Adds microphone-based voice dictation to the web (ChatComposer) and mobile (ThreadComposer) chat composers, replacing the toolbar with a live waveform panel during recording/transcribing and appending the transcript to the draft on stop.
  • Supports OpenAI and Groq as transcription providers; users configure their provider, API key, and model in General settings (web) or a new Voice Dictation settings screen (mobile).
  • Server-side transcription proxy routes (GET/POST /api/transcription, GET /api/transcription/models) forward audio to the selected provider with a 25 MiB size limit and 2-minute timeout, supporting both environment-level and user-supplied API keys.
  • Shared utilities in @t3tools/shared handle transcript appending (whitespace-aware) and terminal action resolution (abort > send > insert).
  • macOS desktop build now includes NSMicrophoneUsageDescription in Info.plist; the mobile Expo app adds the expo-audio plugin with microphone permission.
  • Risk: voice transcription defaults to disabled (voiceTranscriptionEnabled=false); users must configure a model to activate the mic button. Existing client settings are migrated from a legacy OpenAI key format on first load.

Macroscope summarized ff3706c.

@coderabbitai

coderabbitaiBot commented Aug 14, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 35fa00aa-c405-43c0-b5aa-e7bb99d0e03d

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:XXL 1,000+ changed lines (additions + deletions). labels Aug 14, 2026
Comment threadapps/web/src/components/chat/ChatComposer.tsx
Comment threadapps/mobile/src/features/voice-dictation/voiceTranscriptionSettings.ts Outdated
Comment threadapps/web/src/components/chat/ChatComposer.tsx
Comment threadapps/web/src/components/settings/SettingsPanels.tsx
@macroscopeapp

macroscopeappBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Needs human review

This PR introduces a complete new voice dictation feature across web, mobile, and server, including new external API integrations (OpenAI/Groq), new server endpoints, audio recording infrastructure, and API key handling. New features of this scope and complexity require human review.

You can customize Macroscope's approvability policy. Learn more.

apiKey: voiceTranscriptionSettings.apiKey,
onTranscriptInsert: appendVoiceTranscriptToDraft,
onTranscriptSend: (transcript) => {
if (appendVoiceTranscriptToDraft(transcript)) void sendCurrentDraft();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟠 Highthreads/ThreadComposer.tsx:585

The voice “send” action sends the previous draftMessage snapshot instead of the transcript, so the transcript remains in the composer (or nothing is sent). appendVoiceTranscriptToDraft only schedules onChangeDraftMessage, while sendCurrentDraft() immediately calls onSendMessage() without the new text; pass the updated draft through the send path or wait for the state update before sending.

🤖 Copy this AI Prompt to have your agent fix this:
In file @apps/mobile/src/features/threads/ThreadComposer.tsx around line 585:
The voice “send” action sends the previous `draftMessage` snapshot instead of the transcript, so the transcript remains in the composer (or nothing is sent). `appendVoiceTranscriptToDraft` only schedules `onChangeDraftMessage`, while `sendCurrentDraft()` immediately calls `onSendMessage()` without the new text; pass the updated draft through the send path or wait for the state update before sending.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Verified as a false positive. onChangeDraftMessage does not schedule React state: useThreadComposerState calls setComposerDraftText, which synchronously calls appAtomRegistry.set. onSendMessage then calls getComposerDraftSnapshot from that same registry, so the just-appended transcript is present before enqueue. The existing mobile composer-drafts test also covers synchronous latest-draft reads for send.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, I'm unable to act on this request because you do not have permissions within this repository.

@KachurPro
KachurProforce-pushed the codex/voice-dictation-codex-ux branch from 196ba3a to 5d1e3e3CompareAugust 14, 2026 20:13
Comment threadapps/web/src/hooks/useVoiceTranscription.ts Outdated
Comment threadapps/web/src/components/settings/SettingsPanels.tsx
Comment threadapps/web/src/hooks/useVoiceTranscription.ts
Comment threadapps/web/src/hooks/useVoiceTranscription.ts Outdated
Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx
Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated
@brvale97

Copy link
Copy Markdown

Really looking forward to this — mobile dictation is exactly what I've been missing on iPhone.

One request: the web/desktop implementation from #5213 supports both OpenAI and Groq with model selection, but the mobile path hardcodes the OpenAI endpoint and gpt-4o-mini-transcribe in mobileVoiceTranscription.ts. Groq's transcription API is OpenAI-compatible (https://api.groq.com/openai/v1/audio/transcriptions, e.g. whisper-large-v3-turbo), so carrying the provider choice over to mobile would mostly mean making the base URL + model configurable alongside the stored key. Would you consider that?

@KachurPro
KachurProforce-pushed the codex/voice-dictation-codex-ux branch from 59ba2fc to 495d058CompareAugust 14, 2026 21:37
@KachurPro

Copy link
Copy Markdown
Author

@brvale97 Implemented in 495d058bc. Mobile now supports both OpenAI and Groq with fixed official transcription endpoints, separate securely stored keys and model choices per provider, a model picker, and a custom model ID field so future compatible models do not require an app update. OpenAI now defaults to gpt-4o-transcribe; Groq defaults to whisper-large-v3-turbo. Existing saved OpenAI keys are migrated automatically. All 716 mobile tests and the mobile typecheck pass. Thanks for catching this.

Comment threadapps/mobile/src/features/voice-dictation/useMobileVoiceTranscription.ts Outdated
Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

UI consistency review of the web changes (apps/web/src/**). Three findings, all in the new dictation UI: a ticking live region that conflicts with the documented pattern in Sidebar.tsx, and two call sites that rebuild the existing ghost-muted button variant with a class override.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated
Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx Outdated
Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx Outdated
Comment threadapps/web/src/components/chat/ChatComposer.tsx

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the web-scope changes (ChatComposer.tsx, VoiceTranscriptionPanel.tsx, SettingsPanels.tsx, settingsSearch.ts, hooks/useVoiceTranscription.ts, lib/voiceTranscription.ts). The earlier findings on the panel (muted-ghost variant, live-region timer) are resolved. Two remaining consistency issues in the composer.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated
Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

UI consistency review: one finding on the web composer dictation panel. Previous findings (muted-ghost variant usage, live-region timer, responsive error inset, mic pointer-focus) are resolved in this head.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/ChatComposer.tsx
Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One finding on the dictation panel's pointer-focus contract in the composer footer. Prior comments (mic variant/pointer focus, error-row inset, ghost-muted, aria-hidden timer) are addressed in this revision.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the new Effect code (apps/server/src/transcription.ts, the new route layers in apps/server/src/http.ts, and the contracts/shared additions) against the service conventions. Error classes use Schema.TaggedErrorClass with structural attributes and preserved cause, failures are recovered with Effect.catchTags, HttpClient is acquired from the environment, and Effect modules are imported from their subpaths — all consistent with the conventions. One finding on the web client below.

Posted via Macroscope — Effect Service Conventions

Comment threadapps/web/src/lib/voiceTranscription.ts Outdated
@KachurPro
KachurProforce-pushed the codex/voice-dictation-codex-ux branch from 9f92517 to 4732d3aCompareAugust 14, 2026 22:03
Comment threadapps/web/src/components/settings/SettingsPanels.tsx Outdated

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 2 potential issues.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 4732d3a. Configure here.

Comment threadapps/mobile/src/features/threads/ThreadComposer.tsx Outdated
Comment threadapps/web/src/components/settings/SettingsPanels.tsx Outdated

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One finding: the dictation send action bypasses the composer's themeable send tokens. Everything else (pointer-focus contract, ghost-muted variants, aria-hidden timer, responsive error inset) looks consistent with the composer's existing patterns.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx Outdated
@t3dotgg

Copy link
Copy Markdown
Member

Note

🤖 GPT-5.6 Sol responding on behalf of Theo

Closing this PR after an automated pass over open pull requests. Duplicates the maintainer-owned voice implementation in #5213.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXL1,000+ changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@KachurPro@brvale97@t3dotgg@maria-rcks
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat: add Codex-style voice dictation to web and mobile - #6625

Closed
KachurPro wants to merge 20 commits into
pingdotgg:mainfrom
KachurPro:codex/voice-dictation-codex-ux
Closed

feat: add Codex-style voice dictation to web and mobile#6625
KachurPro wants to merge 20 commits into
pingdotgg:mainfrom
KachurPro:codex/voice-dictation-codex-ux

Conversation

@KachurPro

@KachurProKachurPro commented Aug 14, 2026

Copy link
Copy Markdown

Problem

The existing voice dictation PR (#5213) only covered web/desktop and had become stale. Native mobile dictation was still missing, and the first mobile version was limited to a hard-coded OpenAI endpoint and model.

Solution

  • rebase the voice transcription work onto the current main
  • match the Codex-style actions: cancel/discard, stop/transcribe/insert, and transcribe/insert/send
  • always append the normalized transcript to the end of the current draft
  • add native iOS/Android recording with expo-audio
  • let mobile users choose OpenAI or Groq
  • store a separate API key and model for each provider in expo-secure-store, with migration of the legacy OpenAI key
  • default OpenAI to gpt-4o-transcribe and Groq to whisper-large-v3-turbo
  • offer curated model choices while also accepting a custom model ID
  • send mobile recordings directly to the official provider endpoint
  • preserve the existing OpenAI/Groq server proxy and model picker on web/desktop
  • share transcript append and terminal-action rules between web and mobile
  • fully reset recorder, stream, timer, and audio-mode state after every terminal path

This supersedes #5213 and carries its web/desktop behavior forward.

Interaction

  • X: discard the recording
  • stop: transcribe and append to the draft
  • send: transcribe, append, then use the normal composer send path
  • existing draft text is preserved; speech is appended at the end

Verification

  • all 716 mobile tests passed
  • 4 focused web transcription tests passed
  • mobile and web typechecks passed
  • full repository lint and formatting checks passed
  • mobile native static analysis passed
  • Expo production config resolves expo-audio, the microphone permission, and Secure Store

The current development machine does not have Xcode or an iOS Simulator installed, so a local simulator screenshot is not included. The native flow still needs the EAS preview/device pass produced by CI.

Built with GPT-5.6 Sol in Codex Desktop.


Note

Medium Risk
New authenticated transcription proxy accepts client-supplied API keys in custom headers and forwards audio to third parties; composer send/stop paths gain async recording state, and mobile sends audio and keys directly to providers.

Overview
Adds Codex-style voice dictation across web/desktop chat, native mobile thread composer, and server-backed transcription.

Web/desktop: New Voice dictation settings (OpenAI or Groq, client API key or server OPENAI_API_KEY / GROQ_API_KEY, model picker). The composer gets a mic, a recording UI (waveform, cancel / stop / transcribe-and-send), and send behavior that waits on transcription when recording. Audio goes through authenticated /api/transcription and /api/transcription/models on the connected T3 server (25 MB cap). Shared appendVoiceTranscript and terminal-action rules live in @t3tools/shared. Desktop gains macOS microphone usage text.

Mobile:expo-audio recording, per-provider keys/models in Secure Store (with legacy OpenAI key migration), a Voice Dictation settings screen, and direct uploads to OpenAI/Groq (not the server proxy). Composer integration mirrors web (thread-scoped append, cancel on navigation).

Contracts: Client settings add voiceTranscription* fields; reset-to-defaults includes voice dictation when dirty.

Reviewed by Cursor Bugbot for commit ff3706c. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Add Codex-style voice dictation to web and mobile chat composers

  • Adds microphone-based voice dictation to the web (ChatComposer) and mobile (ThreadComposer) chat composers, replacing the toolbar with a live waveform panel during recording/transcribing and appending the transcript to the draft on stop.
  • Supports OpenAI and Groq as transcription providers; users configure their provider, API key, and model in General settings (web) or a new Voice Dictation settings screen (mobile).
  • Server-side transcription proxy routes (GET/POST /api/transcription, GET /api/transcription/models) forward audio to the selected provider with a 25 MiB size limit and 2-minute timeout, supporting both environment-level and user-supplied API keys.
  • Shared utilities in @t3tools/shared handle transcript appending (whitespace-aware) and terminal action resolution (abort > send > insert).
  • macOS desktop build now includes NSMicrophoneUsageDescription in Info.plist; the mobile Expo app adds the expo-audio plugin with microphone permission.
  • Risk: voice transcription defaults to disabled (voiceTranscriptionEnabled=false); users must configure a model to activate the mic button. Existing client settings are migrated from a legacy OpenAI key format on first load.

Macroscope summarized ff3706c.

@coderabbitai

coderabbitaiBot commented Aug 14, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 35fa00aa-c405-43c0-b5aa-e7bb99d0e03d

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:XXL 1,000+ changed lines (additions + deletions). labels Aug 14, 2026
Comment threadapps/web/src/components/chat/ChatComposer.tsx
Comment threadapps/mobile/src/features/voice-dictation/voiceTranscriptionSettings.ts Outdated
Comment threadapps/web/src/components/chat/ChatComposer.tsx
Comment threadapps/web/src/components/settings/SettingsPanels.tsx
@macroscopeapp

macroscopeappBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Needs human review

This PR introduces a complete new voice dictation feature across web, mobile, and server, including new external API integrations (OpenAI/Groq), new server endpoints, audio recording infrastructure, and API key handling. New features of this scope and complexity require human review.

You can customize Macroscope's approvability policy. Learn more.

apiKey: voiceTranscriptionSettings.apiKey,
onTranscriptInsert: appendVoiceTranscriptToDraft,
onTranscriptSend: (transcript) => {
if (appendVoiceTranscriptToDraft(transcript)) void sendCurrentDraft();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟠 Highthreads/ThreadComposer.tsx:585

The voice “send” action sends the previous draftMessage snapshot instead of the transcript, so the transcript remains in the composer (or nothing is sent). appendVoiceTranscriptToDraft only schedules onChangeDraftMessage, while sendCurrentDraft() immediately calls onSendMessage() without the new text; pass the updated draft through the send path or wait for the state update before sending.

🤖 Copy this AI Prompt to have your agent fix this:
In file @apps/mobile/src/features/threads/ThreadComposer.tsx around line 585:
The voice “send” action sends the previous `draftMessage` snapshot instead of the transcript, so the transcript remains in the composer (or nothing is sent). `appendVoiceTranscriptToDraft` only schedules `onChangeDraftMessage`, while `sendCurrentDraft()` immediately calls `onSendMessage()` without the new text; pass the updated draft through the send path or wait for the state update before sending.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Verified as a false positive. onChangeDraftMessage does not schedule React state: useThreadComposerState calls setComposerDraftText, which synchronously calls appAtomRegistry.set. onSendMessage then calls getComposerDraftSnapshot from that same registry, so the just-appended transcript is present before enqueue. The existing mobile composer-drafts test also covers synchronous latest-draft reads for send.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, I'm unable to act on this request because you do not have permissions within this repository.

@KachurPro
KachurProforce-pushed the codex/voice-dictation-codex-ux branch from 196ba3a to 5d1e3e3CompareAugust 14, 2026 20:13
Comment threadapps/web/src/hooks/useVoiceTranscription.ts Outdated
Comment threadapps/web/src/components/settings/SettingsPanels.tsx
Comment threadapps/web/src/hooks/useVoiceTranscription.ts
Comment threadapps/web/src/hooks/useVoiceTranscription.ts Outdated
Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx
Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated
@brvale97

Copy link
Copy Markdown

Really looking forward to this — mobile dictation is exactly what I've been missing on iPhone.

One request: the web/desktop implementation from #5213 supports both OpenAI and Groq with model selection, but the mobile path hardcodes the OpenAI endpoint and gpt-4o-mini-transcribe in mobileVoiceTranscription.ts. Groq's transcription API is OpenAI-compatible (https://api.groq.com/openai/v1/audio/transcriptions, e.g. whisper-large-v3-turbo), so carrying the provider choice over to mobile would mostly mean making the base URL + model configurable alongside the stored key. Would you consider that?

@KachurPro
KachurProforce-pushed the codex/voice-dictation-codex-ux branch from 59ba2fc to 495d058CompareAugust 14, 2026 21:37
@KachurPro

Copy link
Copy Markdown
Author

@brvale97 Implemented in 495d058bc. Mobile now supports both OpenAI and Groq with fixed official transcription endpoints, separate securely stored keys and model choices per provider, a model picker, and a custom model ID field so future compatible models do not require an app update. OpenAI now defaults to gpt-4o-transcribe; Groq defaults to whisper-large-v3-turbo. Existing saved OpenAI keys are migrated automatically. All 716 mobile tests and the mobile typecheck pass. Thanks for catching this.

Comment threadapps/mobile/src/features/voice-dictation/useMobileVoiceTranscription.ts Outdated
Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

UI consistency review of the web changes (apps/web/src/**). Three findings, all in the new dictation UI: a ticking live region that conflicts with the documented pattern in Sidebar.tsx, and two call sites that rebuild the existing ghost-muted button variant with a class override.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated
Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx Outdated
Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx Outdated
Comment threadapps/web/src/components/chat/ChatComposer.tsx

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the web-scope changes (ChatComposer.tsx, VoiceTranscriptionPanel.tsx, SettingsPanels.tsx, settingsSearch.ts, hooks/useVoiceTranscription.ts, lib/voiceTranscription.ts). The earlier findings on the panel (muted-ghost variant, live-region timer) are resolved. Two remaining consistency issues in the composer.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated
Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

UI consistency review: one finding on the web composer dictation panel. Previous findings (muted-ghost variant usage, live-region timer, responsive error inset, mic pointer-focus) are resolved in this head.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/ChatComposer.tsx
Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One finding on the dictation panel's pointer-focus contract in the composer footer. Prior comments (mic variant/pointer focus, error-row inset, ghost-muted, aria-hidden timer) are addressed in this revision.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the new Effect code (apps/server/src/transcription.ts, the new route layers in apps/server/src/http.ts, and the contracts/shared additions) against the service conventions. Error classes use Schema.TaggedErrorClass with structural attributes and preserved cause, failures are recovered with Effect.catchTags, HttpClient is acquired from the environment, and Effect modules are imported from their subpaths — all consistent with the conventions. One finding on the web client below.

Posted via Macroscope — Effect Service Conventions

Comment threadapps/web/src/lib/voiceTranscription.ts Outdated
@KachurPro
KachurProforce-pushed the codex/voice-dictation-codex-ux branch from 9f92517 to 4732d3aCompareAugust 14, 2026 22:03
Comment threadapps/web/src/components/settings/SettingsPanels.tsx Outdated

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 2 potential issues.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 4732d3a. Configure here.

Comment threadapps/mobile/src/features/threads/ThreadComposer.tsx Outdated
Comment threadapps/web/src/components/settings/SettingsPanels.tsx Outdated

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One finding: the dictation send action bypasses the composer's themeable send tokens. Everything else (pointer-focus contract, ghost-muted variants, aria-hidden timer, responsive error inset) looks consistent with the composer's existing patterns.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx Outdated
@t3dotgg

Copy link
Copy Markdown
Member

Note

🤖 GPT-5.6 Sol responding on behalf of Theo

Closing this PR after an automated pass over open pull requests. Duplicates the maintainer-owned voice implementation in #5213.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXL1,000+ changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@KachurPro@brvale97@t3dotgg@maria-rcks
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

feat: add Codex-style voice dictation to web and mobile - #6625

Closed
KachurPro wants to merge 20 commits into
pingdotgg:mainfrom
KachurPro:codex/voice-dictation-codex-ux
Closed

feat: add Codex-style voice dictation to web and mobile#6625
KachurPro wants to merge 20 commits into
pingdotgg:mainfrom
KachurPro:codex/voice-dictation-codex-ux

Conversation

@KachurPro

@KachurProKachurPro commented Aug 14, 2026

Copy link
Copy Markdown

Problem

The existing voice dictation PR (#5213) only covered web/desktop and had become stale. Native mobile dictation was still missing, and the first mobile version was limited to a hard-coded OpenAI endpoint and model.

Solution

  • rebase the voice transcription work onto the current main
  • match the Codex-style actions: cancel/discard, stop/transcribe/insert, and transcribe/insert/send
  • always append the normalized transcript to the end of the current draft
  • add native iOS/Android recording with expo-audio
  • let mobile users choose OpenAI or Groq
  • store a separate API key and model for each provider in expo-secure-store, with migration of the legacy OpenAI key
  • default OpenAI to gpt-4o-transcribe and Groq to whisper-large-v3-turbo
  • offer curated model choices while also accepting a custom model ID
  • send mobile recordings directly to the official provider endpoint
  • preserve the existing OpenAI/Groq server proxy and model picker on web/desktop
  • share transcript append and terminal-action rules between web and mobile
  • fully reset recorder, stream, timer, and audio-mode state after every terminal path

This supersedes #5213 and carries its web/desktop behavior forward.

Interaction

  • X: discard the recording
  • stop: transcribe and append to the draft
  • send: transcribe, append, then use the normal composer send path
  • existing draft text is preserved; speech is appended at the end

Verification

  • all 716 mobile tests passed
  • 4 focused web transcription tests passed
  • mobile and web typechecks passed
  • full repository lint and formatting checks passed
  • mobile native static analysis passed
  • Expo production config resolves expo-audio, the microphone permission, and Secure Store

The current development machine does not have Xcode or an iOS Simulator installed, so a local simulator screenshot is not included. The native flow still needs the EAS preview/device pass produced by CI.

Built with GPT-5.6 Sol in Codex Desktop.


Note

Medium Risk
New authenticated transcription proxy accepts client-supplied API keys in custom headers and forwards audio to third parties; composer send/stop paths gain async recording state, and mobile sends audio and keys directly to providers.

Overview
Adds Codex-style voice dictation across web/desktop chat, native mobile thread composer, and server-backed transcription.

Web/desktop: New Voice dictation settings (OpenAI or Groq, client API key or server OPENAI_API_KEY / GROQ_API_KEY, model picker). The composer gets a mic, a recording UI (waveform, cancel / stop / transcribe-and-send), and send behavior that waits on transcription when recording. Audio goes through authenticated /api/transcription and /api/transcription/models on the connected T3 server (25 MB cap). Shared appendVoiceTranscript and terminal-action rules live in @t3tools/shared. Desktop gains macOS microphone usage text.

Mobile:expo-audio recording, per-provider keys/models in Secure Store (with legacy OpenAI key migration), a Voice Dictation settings screen, and direct uploads to OpenAI/Groq (not the server proxy). Composer integration mirrors web (thread-scoped append, cancel on navigation).

Contracts: Client settings add voiceTranscription* fields; reset-to-defaults includes voice dictation when dirty.

Reviewed by Cursor Bugbot for commit ff3706c. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Add Codex-style voice dictation to web and mobile chat composers

  • Adds microphone-based voice dictation to the web (ChatComposer) and mobile (ThreadComposer) chat composers, replacing the toolbar with a live waveform panel during recording/transcribing and appending the transcript to the draft on stop.
  • Supports OpenAI and Groq as transcription providers; users configure their provider, API key, and model in General settings (web) or a new Voice Dictation settings screen (mobile).
  • Server-side transcription proxy routes (GET/POST /api/transcription, GET /api/transcription/models) forward audio to the selected provider with a 25 MiB size limit and 2-minute timeout, supporting both environment-level and user-supplied API keys.
  • Shared utilities in @t3tools/shared handle transcript appending (whitespace-aware) and terminal action resolution (abort > send > insert).
  • macOS desktop build now includes NSMicrophoneUsageDescription in Info.plist; the mobile Expo app adds the expo-audio plugin with microphone permission.
  • Risk: voice transcription defaults to disabled (voiceTranscriptionEnabled=false); users must configure a model to activate the mic button. Existing client settings are migrated from a legacy OpenAI key format on first load.

Macroscope summarized ff3706c.

@coderabbitai

coderabbitaiBot commented Aug 14, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 35fa00aa-c405-43c0-b5aa-e7bb99d0e03d

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:XXL 1,000+ changed lines (additions + deletions). labels Aug 14, 2026
Comment threadapps/web/src/components/chat/ChatComposer.tsx
Comment threadapps/mobile/src/features/voice-dictation/voiceTranscriptionSettings.ts Outdated
Comment threadapps/web/src/components/chat/ChatComposer.tsx
Comment threadapps/web/src/components/settings/SettingsPanels.tsx
@macroscopeapp

macroscopeappBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Needs human review

This PR introduces a complete new voice dictation feature across web, mobile, and server, including new external API integrations (OpenAI/Groq), new server endpoints, audio recording infrastructure, and API key handling. New features of this scope and complexity require human review.

You can customize Macroscope's approvability policy. Learn more.

apiKey: voiceTranscriptionSettings.apiKey,
onTranscriptInsert: appendVoiceTranscriptToDraft,
onTranscriptSend: (transcript) => {
if (appendVoiceTranscriptToDraft(transcript)) void sendCurrentDraft();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟠 Highthreads/ThreadComposer.tsx:585

The voice “send” action sends the previous draftMessage snapshot instead of the transcript, so the transcript remains in the composer (or nothing is sent). appendVoiceTranscriptToDraft only schedules onChangeDraftMessage, while sendCurrentDraft() immediately calls onSendMessage() without the new text; pass the updated draft through the send path or wait for the state update before sending.

🤖 Copy this AI Prompt to have your agent fix this:
In file @apps/mobile/src/features/threads/ThreadComposer.tsx around line 585:
The voice “send” action sends the previous `draftMessage` snapshot instead of the transcript, so the transcript remains in the composer (or nothing is sent). `appendVoiceTranscriptToDraft` only schedules `onChangeDraftMessage`, while `sendCurrentDraft()` immediately calls `onSendMessage()` without the new text; pass the updated draft through the send path or wait for the state update before sending.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Verified as a false positive. onChangeDraftMessage does not schedule React state: useThreadComposerState calls setComposerDraftText, which synchronously calls appAtomRegistry.set. onSendMessage then calls getComposerDraftSnapshot from that same registry, so the just-appended transcript is present before enqueue. The existing mobile composer-drafts test also covers synchronous latest-draft reads for send.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, I'm unable to act on this request because you do not have permissions within this repository.

@KachurPro
KachurProforce-pushed the codex/voice-dictation-codex-ux branch from 196ba3a to 5d1e3e3CompareAugust 14, 2026 20:13
Comment threadapps/web/src/hooks/useVoiceTranscription.ts Outdated
Comment threadapps/web/src/components/settings/SettingsPanels.tsx
Comment threadapps/web/src/hooks/useVoiceTranscription.ts
Comment threadapps/web/src/hooks/useVoiceTranscription.ts Outdated
Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx
Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated
@brvale97

Copy link
Copy Markdown

Really looking forward to this — mobile dictation is exactly what I've been missing on iPhone.

One request: the web/desktop implementation from #5213 supports both OpenAI and Groq with model selection, but the mobile path hardcodes the OpenAI endpoint and gpt-4o-mini-transcribe in mobileVoiceTranscription.ts. Groq's transcription API is OpenAI-compatible (https://api.groq.com/openai/v1/audio/transcriptions, e.g. whisper-large-v3-turbo), so carrying the provider choice over to mobile would mostly mean making the base URL + model configurable alongside the stored key. Would you consider that?

@KachurPro
KachurProforce-pushed the codex/voice-dictation-codex-ux branch from 59ba2fc to 495d058CompareAugust 14, 2026 21:37
@KachurPro

Copy link
Copy Markdown
Author

@brvale97 Implemented in 495d058bc. Mobile now supports both OpenAI and Groq with fixed official transcription endpoints, separate securely stored keys and model choices per provider, a model picker, and a custom model ID field so future compatible models do not require an app update. OpenAI now defaults to gpt-4o-transcribe; Groq defaults to whisper-large-v3-turbo. Existing saved OpenAI keys are migrated automatically. All 716 mobile tests and the mobile typecheck pass. Thanks for catching this.

Comment threadapps/mobile/src/features/voice-dictation/useMobileVoiceTranscription.ts Outdated
Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

UI consistency review of the web changes (apps/web/src/**). Three findings, all in the new dictation UI: a ticking live region that conflicts with the documented pattern in Sidebar.tsx, and two call sites that rebuild the existing ghost-muted button variant with a class override.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated
Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx Outdated
Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx Outdated
Comment threadapps/web/src/components/chat/ChatComposer.tsx

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the web-scope changes (ChatComposer.tsx, VoiceTranscriptionPanel.tsx, SettingsPanels.tsx, settingsSearch.ts, hooks/useVoiceTranscription.ts, lib/voiceTranscription.ts). The earlier findings on the panel (muted-ghost variant, live-region timer) are resolved. Two remaining consistency issues in the composer.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated
Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

UI consistency review: one finding on the web composer dictation panel. Previous findings (muted-ghost variant usage, live-region timer, responsive error inset, mic pointer-focus) are resolved in this head.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/ChatComposer.tsx
Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One finding on the dictation panel's pointer-focus contract in the composer footer. Prior comments (mic variant/pointer focus, error-row inset, ghost-muted, aria-hidden timer) are addressed in this revision.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the new Effect code (apps/server/src/transcription.ts, the new route layers in apps/server/src/http.ts, and the contracts/shared additions) against the service conventions. Error classes use Schema.TaggedErrorClass with structural attributes and preserved cause, failures are recovered with Effect.catchTags, HttpClient is acquired from the environment, and Effect modules are imported from their subpaths — all consistent with the conventions. One finding on the web client below.

Posted via Macroscope — Effect Service Conventions

Comment threadapps/web/src/lib/voiceTranscription.ts Outdated
@KachurPro
KachurProforce-pushed the codex/voice-dictation-codex-ux branch from 9f92517 to 4732d3aCompareAugust 14, 2026 22:03
Comment threadapps/web/src/components/settings/SettingsPanels.tsx Outdated

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 2 potential issues.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 4732d3a. Configure here.

Comment threadapps/mobile/src/features/threads/ThreadComposer.tsx Outdated
Comment threadapps/web/src/components/settings/SettingsPanels.tsx Outdated

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One finding: the dictation send action bypasses the composer's themeable send tokens. Everything else (pointer-focus contract, ghost-muted variants, aria-hidden timer, responsive error inset) looks consistent with the composer's existing patterns.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx Outdated
@t3dotgg

Copy link
Copy Markdown
Member

Note

🤖 GPT-5.6 Sol responding on behalf of Theo

Closing this PR after an automated pass over open pull requests. Duplicates the maintainer-owned voice implementation in #5213.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXL1,000+ changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@KachurPro@brvale97@t3dotgg@maria-rcks
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat: add Codex-style voice dictation to web and mobile - #6625

Closed
KachurPro wants to merge 20 commits into
pingdotgg:mainfrom
KachurPro:codex/voice-dictation-codex-ux
Closed

feat: add Codex-style voice dictation to web and mobile#6625
KachurPro wants to merge 20 commits into
pingdotgg:mainfrom
KachurPro:codex/voice-dictation-codex-ux

Conversation

@KachurPro

@KachurProKachurPro commented Aug 14, 2026

Copy link
Copy Markdown

Problem

The existing voice dictation PR (#5213) only covered web/desktop and had become stale. Native mobile dictation was still missing, and the first mobile version was limited to a hard-coded OpenAI endpoint and model.

Solution

  • rebase the voice transcription work onto the current main
  • match the Codex-style actions: cancel/discard, stop/transcribe/insert, and transcribe/insert/send
  • always append the normalized transcript to the end of the current draft
  • add native iOS/Android recording with expo-audio
  • let mobile users choose OpenAI or Groq
  • store a separate API key and model for each provider in expo-secure-store, with migration of the legacy OpenAI key
  • default OpenAI to gpt-4o-transcribe and Groq to whisper-large-v3-turbo
  • offer curated model choices while also accepting a custom model ID
  • send mobile recordings directly to the official provider endpoint
  • preserve the existing OpenAI/Groq server proxy and model picker on web/desktop
  • share transcript append and terminal-action rules between web and mobile
  • fully reset recorder, stream, timer, and audio-mode state after every terminal path

This supersedes #5213 and carries its web/desktop behavior forward.

Interaction

  • X: discard the recording
  • stop: transcribe and append to the draft
  • send: transcribe, append, then use the normal composer send path
  • existing draft text is preserved; speech is appended at the end

Verification

  • all 716 mobile tests passed
  • 4 focused web transcription tests passed
  • mobile and web typechecks passed
  • full repository lint and formatting checks passed
  • mobile native static analysis passed
  • Expo production config resolves expo-audio, the microphone permission, and Secure Store

The current development machine does not have Xcode or an iOS Simulator installed, so a local simulator screenshot is not included. The native flow still needs the EAS preview/device pass produced by CI.

Built with GPT-5.6 Sol in Codex Desktop.


Note

Medium Risk
New authenticated transcription proxy accepts client-supplied API keys in custom headers and forwards audio to third parties; composer send/stop paths gain async recording state, and mobile sends audio and keys directly to providers.

Overview
Adds Codex-style voice dictation across web/desktop chat, native mobile thread composer, and server-backed transcription.

Web/desktop: New Voice dictation settings (OpenAI or Groq, client API key or server OPENAI_API_KEY / GROQ_API_KEY, model picker). The composer gets a mic, a recording UI (waveform, cancel / stop / transcribe-and-send), and send behavior that waits on transcription when recording. Audio goes through authenticated /api/transcription and /api/transcription/models on the connected T3 server (25 MB cap). Shared appendVoiceTranscript and terminal-action rules live in @t3tools/shared. Desktop gains macOS microphone usage text.

Mobile:expo-audio recording, per-provider keys/models in Secure Store (with legacy OpenAI key migration), a Voice Dictation settings screen, and direct uploads to OpenAI/Groq (not the server proxy). Composer integration mirrors web (thread-scoped append, cancel on navigation).

Contracts: Client settings add voiceTranscription* fields; reset-to-defaults includes voice dictation when dirty.

Reviewed by Cursor Bugbot for commit ff3706c. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Add Codex-style voice dictation to web and mobile chat composers

  • Adds microphone-based voice dictation to the web (ChatComposer) and mobile (ThreadComposer) chat composers, replacing the toolbar with a live waveform panel during recording/transcribing and appending the transcript to the draft on stop.
  • Supports OpenAI and Groq as transcription providers; users configure their provider, API key, and model in General settings (web) or a new Voice Dictation settings screen (mobile).
  • Server-side transcription proxy routes (GET/POST /api/transcription, GET /api/transcription/models) forward audio to the selected provider with a 25 MiB size limit and 2-minute timeout, supporting both environment-level and user-supplied API keys.
  • Shared utilities in @t3tools/shared handle transcript appending (whitespace-aware) and terminal action resolution (abort > send > insert).
  • macOS desktop build now includes NSMicrophoneUsageDescription in Info.plist; the mobile Expo app adds the expo-audio plugin with microphone permission.
  • Risk: voice transcription defaults to disabled (voiceTranscriptionEnabled=false); users must configure a model to activate the mic button. Existing client settings are migrated from a legacy OpenAI key format on first load.

Macroscope summarized ff3706c.

@coderabbitai

coderabbitaiBot commented Aug 14, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 35fa00aa-c405-43c0-b5aa-e7bb99d0e03d

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:XXL 1,000+ changed lines (additions + deletions). labels Aug 14, 2026
Comment threadapps/web/src/components/chat/ChatComposer.tsx
Comment threadapps/mobile/src/features/voice-dictation/voiceTranscriptionSettings.ts Outdated
Comment threadapps/web/src/components/chat/ChatComposer.tsx
Comment threadapps/web/src/components/settings/SettingsPanels.tsx
@macroscopeapp

macroscopeappBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Needs human review

This PR introduces a complete new voice dictation feature across web, mobile, and server, including new external API integrations (OpenAI/Groq), new server endpoints, audio recording infrastructure, and API key handling. New features of this scope and complexity require human review.

You can customize Macroscope's approvability policy. Learn more.

apiKey: voiceTranscriptionSettings.apiKey,
onTranscriptInsert: appendVoiceTranscriptToDraft,
onTranscriptSend: (transcript) => {
if (appendVoiceTranscriptToDraft(transcript)) void sendCurrentDraft();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟠 Highthreads/ThreadComposer.tsx:585

The voice “send” action sends the previous draftMessage snapshot instead of the transcript, so the transcript remains in the composer (or nothing is sent). appendVoiceTranscriptToDraft only schedules onChangeDraftMessage, while sendCurrentDraft() immediately calls onSendMessage() without the new text; pass the updated draft through the send path or wait for the state update before sending.

🤖 Copy this AI Prompt to have your agent fix this:
In file @apps/mobile/src/features/threads/ThreadComposer.tsx around line 585:
The voice “send” action sends the previous `draftMessage` snapshot instead of the transcript, so the transcript remains in the composer (or nothing is sent). `appendVoiceTranscriptToDraft` only schedules `onChangeDraftMessage`, while `sendCurrentDraft()` immediately calls `onSendMessage()` without the new text; pass the updated draft through the send path or wait for the state update before sending.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Verified as a false positive. onChangeDraftMessage does not schedule React state: useThreadComposerState calls setComposerDraftText, which synchronously calls appAtomRegistry.set. onSendMessage then calls getComposerDraftSnapshot from that same registry, so the just-appended transcript is present before enqueue. The existing mobile composer-drafts test also covers synchronous latest-draft reads for send.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, I'm unable to act on this request because you do not have permissions within this repository.

@KachurPro
KachurProforce-pushed the codex/voice-dictation-codex-ux branch from 196ba3a to 5d1e3e3CompareAugust 14, 2026 20:13
Comment threadapps/web/src/hooks/useVoiceTranscription.ts Outdated
Comment threadapps/web/src/components/settings/SettingsPanels.tsx
Comment threadapps/web/src/hooks/useVoiceTranscription.ts
Comment threadapps/web/src/hooks/useVoiceTranscription.ts Outdated
Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx
Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated
@brvale97

Copy link
Copy Markdown

Really looking forward to this — mobile dictation is exactly what I've been missing on iPhone.

One request: the web/desktop implementation from #5213 supports both OpenAI and Groq with model selection, but the mobile path hardcodes the OpenAI endpoint and gpt-4o-mini-transcribe in mobileVoiceTranscription.ts. Groq's transcription API is OpenAI-compatible (https://api.groq.com/openai/v1/audio/transcriptions, e.g. whisper-large-v3-turbo), so carrying the provider choice over to mobile would mostly mean making the base URL + model configurable alongside the stored key. Would you consider that?

@KachurPro
KachurProforce-pushed the codex/voice-dictation-codex-ux branch from 59ba2fc to 495d058CompareAugust 14, 2026 21:37
@KachurPro

Copy link
Copy Markdown
Author

@brvale97 Implemented in 495d058bc. Mobile now supports both OpenAI and Groq with fixed official transcription endpoints, separate securely stored keys and model choices per provider, a model picker, and a custom model ID field so future compatible models do not require an app update. OpenAI now defaults to gpt-4o-transcribe; Groq defaults to whisper-large-v3-turbo. Existing saved OpenAI keys are migrated automatically. All 716 mobile tests and the mobile typecheck pass. Thanks for catching this.

Comment threadapps/mobile/src/features/voice-dictation/useMobileVoiceTranscription.ts Outdated
Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

UI consistency review of the web changes (apps/web/src/**). Three findings, all in the new dictation UI: a ticking live region that conflicts with the documented pattern in Sidebar.tsx, and two call sites that rebuild the existing ghost-muted button variant with a class override.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated
Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx Outdated
Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx Outdated
Comment threadapps/web/src/components/chat/ChatComposer.tsx

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the web-scope changes (ChatComposer.tsx, VoiceTranscriptionPanel.tsx, SettingsPanels.tsx, settingsSearch.ts, hooks/useVoiceTranscription.ts, lib/voiceTranscription.ts). The earlier findings on the panel (muted-ghost variant, live-region timer) are resolved. Two remaining consistency issues in the composer.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated
Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

UI consistency review: one finding on the web composer dictation panel. Previous findings (muted-ghost variant usage, live-region timer, responsive error inset, mic pointer-focus) are resolved in this head.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/ChatComposer.tsx
Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One finding on the dictation panel's pointer-focus contract in the composer footer. Prior comments (mic variant/pointer focus, error-row inset, ghost-muted, aria-hidden timer) are addressed in this revision.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the new Effect code (apps/server/src/transcription.ts, the new route layers in apps/server/src/http.ts, and the contracts/shared additions) against the service conventions. Error classes use Schema.TaggedErrorClass with structural attributes and preserved cause, failures are recovered with Effect.catchTags, HttpClient is acquired from the environment, and Effect modules are imported from their subpaths — all consistent with the conventions. One finding on the web client below.

Posted via Macroscope — Effect Service Conventions

Comment threadapps/web/src/lib/voiceTranscription.ts Outdated
@KachurPro
KachurProforce-pushed the codex/voice-dictation-codex-ux branch from 9f92517 to 4732d3aCompareAugust 14, 2026 22:03
Comment threadapps/web/src/components/settings/SettingsPanels.tsx Outdated

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 2 potential issues.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 4732d3a. Configure here.

Comment threadapps/mobile/src/features/threads/ThreadComposer.tsx Outdated
Comment threadapps/web/src/components/settings/SettingsPanels.tsx Outdated

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One finding: the dictation send action bypasses the composer's themeable send tokens. Everything else (pointer-focus contract, ghost-muted variants, aria-hidden timer, responsive error inset) looks consistent with the composer's existing patterns.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx Outdated
@t3dotgg

Copy link
Copy Markdown
Member

Note

🤖 GPT-5.6 Sol responding on behalf of Theo

Closing this PR after an automated pass over open pull requests. Duplicates the maintainer-owned voice implementation in #5213.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXL1,000+ changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@KachurPro@brvale97@t3dotgg@maria-rcks
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat: add Codex-style voice dictation to web and mobile - #6625

Closed
KachurPro wants to merge 20 commits into
pingdotgg:mainfrom
KachurPro:codex/voice-dictation-codex-ux
Closed

feat: add Codex-style voice dictation to web and mobile#6625
KachurPro wants to merge 20 commits into
pingdotgg:mainfrom
KachurPro:codex/voice-dictation-codex-ux

Conversation

@KachurPro

@KachurProKachurPro commented Aug 14, 2026

Copy link
Copy Markdown

Problem

The existing voice dictation PR (#5213) only covered web/desktop and had become stale. Native mobile dictation was still missing, and the first mobile version was limited to a hard-coded OpenAI endpoint and model.

Solution

  • rebase the voice transcription work onto the current main
  • match the Codex-style actions: cancel/discard, stop/transcribe/insert, and transcribe/insert/send
  • always append the normalized transcript to the end of the current draft
  • add native iOS/Android recording with expo-audio
  • let mobile users choose OpenAI or Groq
  • store a separate API key and model for each provider in expo-secure-store, with migration of the legacy OpenAI key
  • default OpenAI to gpt-4o-transcribe and Groq to whisper-large-v3-turbo
  • offer curated model choices while also accepting a custom model ID
  • send mobile recordings directly to the official provider endpoint
  • preserve the existing OpenAI/Groq server proxy and model picker on web/desktop
  • share transcript append and terminal-action rules between web and mobile
  • fully reset recorder, stream, timer, and audio-mode state after every terminal path

This supersedes #5213 and carries its web/desktop behavior forward.

Interaction

  • X: discard the recording
  • stop: transcribe and append to the draft
  • send: transcribe, append, then use the normal composer send path
  • existing draft text is preserved; speech is appended at the end

Verification

  • all 716 mobile tests passed
  • 4 focused web transcription tests passed
  • mobile and web typechecks passed
  • full repository lint and formatting checks passed
  • mobile native static analysis passed
  • Expo production config resolves expo-audio, the microphone permission, and Secure Store

The current development machine does not have Xcode or an iOS Simulator installed, so a local simulator screenshot is not included. The native flow still needs the EAS preview/device pass produced by CI.

Built with GPT-5.6 Sol in Codex Desktop.


Note

Medium Risk
New authenticated transcription proxy accepts client-supplied API keys in custom headers and forwards audio to third parties; composer send/stop paths gain async recording state, and mobile sends audio and keys directly to providers.

Overview
Adds Codex-style voice dictation across web/desktop chat, native mobile thread composer, and server-backed transcription.

Web/desktop: New Voice dictation settings (OpenAI or Groq, client API key or server OPENAI_API_KEY / GROQ_API_KEY, model picker). The composer gets a mic, a recording UI (waveform, cancel / stop / transcribe-and-send), and send behavior that waits on transcription when recording. Audio goes through authenticated /api/transcription and /api/transcription/models on the connected T3 server (25 MB cap). Shared appendVoiceTranscript and terminal-action rules live in @t3tools/shared. Desktop gains macOS microphone usage text.

Mobile:expo-audio recording, per-provider keys/models in Secure Store (with legacy OpenAI key migration), a Voice Dictation settings screen, and direct uploads to OpenAI/Groq (not the server proxy). Composer integration mirrors web (thread-scoped append, cancel on navigation).

Contracts: Client settings add voiceTranscription* fields; reset-to-defaults includes voice dictation when dirty.

Reviewed by Cursor Bugbot for commit ff3706c. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Add Codex-style voice dictation to web and mobile chat composers

  • Adds microphone-based voice dictation to the web (ChatComposer) and mobile (ThreadComposer) chat composers, replacing the toolbar with a live waveform panel during recording/transcribing and appending the transcript to the draft on stop.
  • Supports OpenAI and Groq as transcription providers; users configure their provider, API key, and model in General settings (web) or a new Voice Dictation settings screen (mobile).
  • Server-side transcription proxy routes (GET/POST /api/transcription, GET /api/transcription/models) forward audio to the selected provider with a 25 MiB size limit and 2-minute timeout, supporting both environment-level and user-supplied API keys.
  • Shared utilities in @t3tools/shared handle transcript appending (whitespace-aware) and terminal action resolution (abort > send > insert).
  • macOS desktop build now includes NSMicrophoneUsageDescription in Info.plist; the mobile Expo app adds the expo-audio plugin with microphone permission.
  • Risk: voice transcription defaults to disabled (voiceTranscriptionEnabled=false); users must configure a model to activate the mic button. Existing client settings are migrated from a legacy OpenAI key format on first load.

Macroscope summarized ff3706c.

@coderabbitai

coderabbitaiBot commented Aug 14, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 35fa00aa-c405-43c0-b5aa-e7bb99d0e03d

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:XXL 1,000+ changed lines (additions + deletions). labels Aug 14, 2026
Comment threadapps/web/src/components/chat/ChatComposer.tsx
Comment threadapps/mobile/src/features/voice-dictation/voiceTranscriptionSettings.ts Outdated
Comment threadapps/web/src/components/chat/ChatComposer.tsx
Comment threadapps/web/src/components/settings/SettingsPanels.tsx
@macroscopeapp

macroscopeappBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Needs human review

This PR introduces a complete new voice dictation feature across web, mobile, and server, including new external API integrations (OpenAI/Groq), new server endpoints, audio recording infrastructure, and API key handling. New features of this scope and complexity require human review.

You can customize Macroscope's approvability policy. Learn more.

apiKey: voiceTranscriptionSettings.apiKey,
onTranscriptInsert: appendVoiceTranscriptToDraft,
onTranscriptSend: (transcript) => {
if (appendVoiceTranscriptToDraft(transcript)) void sendCurrentDraft();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟠 Highthreads/ThreadComposer.tsx:585

The voice “send” action sends the previous draftMessage snapshot instead of the transcript, so the transcript remains in the composer (or nothing is sent). appendVoiceTranscriptToDraft only schedules onChangeDraftMessage, while sendCurrentDraft() immediately calls onSendMessage() without the new text; pass the updated draft through the send path or wait for the state update before sending.

🤖 Copy this AI Prompt to have your agent fix this:
In file @apps/mobile/src/features/threads/ThreadComposer.tsx around line 585:
The voice “send” action sends the previous `draftMessage` snapshot instead of the transcript, so the transcript remains in the composer (or nothing is sent). `appendVoiceTranscriptToDraft` only schedules `onChangeDraftMessage`, while `sendCurrentDraft()` immediately calls `onSendMessage()` without the new text; pass the updated draft through the send path or wait for the state update before sending.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Verified as a false positive. onChangeDraftMessage does not schedule React state: useThreadComposerState calls setComposerDraftText, which synchronously calls appAtomRegistry.set. onSendMessage then calls getComposerDraftSnapshot from that same registry, so the just-appended transcript is present before enqueue. The existing mobile composer-drafts test also covers synchronous latest-draft reads for send.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, I'm unable to act on this request because you do not have permissions within this repository.

@KachurPro
KachurProforce-pushed the codex/voice-dictation-codex-ux branch from 196ba3a to 5d1e3e3CompareAugust 14, 2026 20:13
Comment threadapps/web/src/hooks/useVoiceTranscription.ts Outdated
Comment threadapps/web/src/components/settings/SettingsPanels.tsx
Comment threadapps/web/src/hooks/useVoiceTranscription.ts
Comment threadapps/web/src/hooks/useVoiceTranscription.ts Outdated
Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx
Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated
@brvale97

Copy link
Copy Markdown

Really looking forward to this — mobile dictation is exactly what I've been missing on iPhone.

One request: the web/desktop implementation from #5213 supports both OpenAI and Groq with model selection, but the mobile path hardcodes the OpenAI endpoint and gpt-4o-mini-transcribe in mobileVoiceTranscription.ts. Groq's transcription API is OpenAI-compatible (https://api.groq.com/openai/v1/audio/transcriptions, e.g. whisper-large-v3-turbo), so carrying the provider choice over to mobile would mostly mean making the base URL + model configurable alongside the stored key. Would you consider that?

@KachurPro
KachurProforce-pushed the codex/voice-dictation-codex-ux branch from 59ba2fc to 495d058CompareAugust 14, 2026 21:37
@KachurPro

Copy link
Copy Markdown
Author

@brvale97 Implemented in 495d058bc. Mobile now supports both OpenAI and Groq with fixed official transcription endpoints, separate securely stored keys and model choices per provider, a model picker, and a custom model ID field so future compatible models do not require an app update. OpenAI now defaults to gpt-4o-transcribe; Groq defaults to whisper-large-v3-turbo. Existing saved OpenAI keys are migrated automatically. All 716 mobile tests and the mobile typecheck pass. Thanks for catching this.

Comment threadapps/mobile/src/features/voice-dictation/useMobileVoiceTranscription.ts Outdated
Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

UI consistency review of the web changes (apps/web/src/**). Three findings, all in the new dictation UI: a ticking live region that conflicts with the documented pattern in Sidebar.tsx, and two call sites that rebuild the existing ghost-muted button variant with a class override.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated
Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx Outdated
Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx Outdated
Comment threadapps/web/src/components/chat/ChatComposer.tsx

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the web-scope changes (ChatComposer.tsx, VoiceTranscriptionPanel.tsx, SettingsPanels.tsx, settingsSearch.ts, hooks/useVoiceTranscription.ts, lib/voiceTranscription.ts). The earlier findings on the panel (muted-ghost variant, live-region timer) are resolved. Two remaining consistency issues in the composer.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated
Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

UI consistency review: one finding on the web composer dictation panel. Previous findings (muted-ghost variant usage, live-region timer, responsive error inset, mic pointer-focus) are resolved in this head.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/ChatComposer.tsx
Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One finding on the dictation panel's pointer-focus contract in the composer footer. Prior comments (mic variant/pointer focus, error-row inset, ghost-muted, aria-hidden timer) are addressed in this revision.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the new Effect code (apps/server/src/transcription.ts, the new route layers in apps/server/src/http.ts, and the contracts/shared additions) against the service conventions. Error classes use Schema.TaggedErrorClass with structural attributes and preserved cause, failures are recovered with Effect.catchTags, HttpClient is acquired from the environment, and Effect modules are imported from their subpaths — all consistent with the conventions. One finding on the web client below.

Posted via Macroscope — Effect Service Conventions

Comment threadapps/web/src/lib/voiceTranscription.ts Outdated
@KachurPro
KachurProforce-pushed the codex/voice-dictation-codex-ux branch from 9f92517 to 4732d3aCompareAugust 14, 2026 22:03
Comment threadapps/web/src/components/settings/SettingsPanels.tsx Outdated

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 2 potential issues.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 4732d3a. Configure here.

Comment threadapps/mobile/src/features/threads/ThreadComposer.tsx Outdated
Comment threadapps/web/src/components/settings/SettingsPanels.tsx Outdated

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One finding: the dictation send action bypasses the composer's themeable send tokens. Everything else (pointer-focus contract, ghost-muted variants, aria-hidden timer, responsive error inset) looks consistent with the composer's existing patterns.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx Outdated
@t3dotgg

Copy link
Copy Markdown
Member

Note

🤖 GPT-5.6 Sol responding on behalf of Theo

Closing this PR after an automated pass over open pull requests. Duplicates the maintainer-owned voice implementation in #5213.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXL1,000+ changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@KachurPro@brvale97@t3dotgg@maria-rcks
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

feat: add Codex-style voice dictation to web and mobile - #6625

Closed
KachurPro wants to merge 20 commits into
pingdotgg:mainfrom
KachurPro:codex/voice-dictation-codex-ux
Closed

feat: add Codex-style voice dictation to web and mobile#6625
KachurPro wants to merge 20 commits into
pingdotgg:mainfrom
KachurPro:codex/voice-dictation-codex-ux

Conversation

@KachurPro

@KachurProKachurPro commented Aug 14, 2026

Copy link
Copy Markdown

Problem

The existing voice dictation PR (#5213) only covered web/desktop and had become stale. Native mobile dictation was still missing, and the first mobile version was limited to a hard-coded OpenAI endpoint and model.

Solution

  • rebase the voice transcription work onto the current main
  • match the Codex-style actions: cancel/discard, stop/transcribe/insert, and transcribe/insert/send
  • always append the normalized transcript to the end of the current draft
  • add native iOS/Android recording with expo-audio
  • let mobile users choose OpenAI or Groq
  • store a separate API key and model for each provider in expo-secure-store, with migration of the legacy OpenAI key
  • default OpenAI to gpt-4o-transcribe and Groq to whisper-large-v3-turbo
  • offer curated model choices while also accepting a custom model ID
  • send mobile recordings directly to the official provider endpoint
  • preserve the existing OpenAI/Groq server proxy and model picker on web/desktop
  • share transcript append and terminal-action rules between web and mobile
  • fully reset recorder, stream, timer, and audio-mode state after every terminal path

This supersedes #5213 and carries its web/desktop behavior forward.

Interaction

  • X: discard the recording
  • stop: transcribe and append to the draft
  • send: transcribe, append, then use the normal composer send path
  • existing draft text is preserved; speech is appended at the end

Verification

  • all 716 mobile tests passed
  • 4 focused web transcription tests passed
  • mobile and web typechecks passed
  • full repository lint and formatting checks passed
  • mobile native static analysis passed
  • Expo production config resolves expo-audio, the microphone permission, and Secure Store

The current development machine does not have Xcode or an iOS Simulator installed, so a local simulator screenshot is not included. The native flow still needs the EAS preview/device pass produced by CI.

Built with GPT-5.6 Sol in Codex Desktop.


Note

Medium Risk
New authenticated transcription proxy accepts client-supplied API keys in custom headers and forwards audio to third parties; composer send/stop paths gain async recording state, and mobile sends audio and keys directly to providers.

Overview
Adds Codex-style voice dictation across web/desktop chat, native mobile thread composer, and server-backed transcription.

Web/desktop: New Voice dictation settings (OpenAI or Groq, client API key or server OPENAI_API_KEY / GROQ_API_KEY, model picker). The composer gets a mic, a recording UI (waveform, cancel / stop / transcribe-and-send), and send behavior that waits on transcription when recording. Audio goes through authenticated /api/transcription and /api/transcription/models on the connected T3 server (25 MB cap). Shared appendVoiceTranscript and terminal-action rules live in @t3tools/shared. Desktop gains macOS microphone usage text.

Mobile:expo-audio recording, per-provider keys/models in Secure Store (with legacy OpenAI key migration), a Voice Dictation settings screen, and direct uploads to OpenAI/Groq (not the server proxy). Composer integration mirrors web (thread-scoped append, cancel on navigation).

Contracts: Client settings add voiceTranscription* fields; reset-to-defaults includes voice dictation when dirty.

Reviewed by Cursor Bugbot for commit ff3706c. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Add Codex-style voice dictation to web and mobile chat composers

  • Adds microphone-based voice dictation to the web (ChatComposer) and mobile (ThreadComposer) chat composers, replacing the toolbar with a live waveform panel during recording/transcribing and appending the transcript to the draft on stop.
  • Supports OpenAI and Groq as transcription providers; users configure their provider, API key, and model in General settings (web) or a new Voice Dictation settings screen (mobile).
  • Server-side transcription proxy routes (GET/POST /api/transcription, GET /api/transcription/models) forward audio to the selected provider with a 25 MiB size limit and 2-minute timeout, supporting both environment-level and user-supplied API keys.
  • Shared utilities in @t3tools/shared handle transcript appending (whitespace-aware) and terminal action resolution (abort > send > insert).
  • macOS desktop build now includes NSMicrophoneUsageDescription in Info.plist; the mobile Expo app adds the expo-audio plugin with microphone permission.
  • Risk: voice transcription defaults to disabled (voiceTranscriptionEnabled=false); users must configure a model to activate the mic button. Existing client settings are migrated from a legacy OpenAI key format on first load.

Macroscope summarized ff3706c.

@coderabbitai

coderabbitaiBot commented Aug 14, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 35fa00aa-c405-43c0-b5aa-e7bb99d0e03d

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:XXL 1,000+ changed lines (additions + deletions). labels Aug 14, 2026
Comment threadapps/web/src/components/chat/ChatComposer.tsx
Comment threadapps/mobile/src/features/voice-dictation/voiceTranscriptionSettings.ts Outdated
Comment threadapps/web/src/components/chat/ChatComposer.tsx
Comment threadapps/web/src/components/settings/SettingsPanels.tsx
@macroscopeapp

macroscopeappBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Needs human review

This PR introduces a complete new voice dictation feature across web, mobile, and server, including new external API integrations (OpenAI/Groq), new server endpoints, audio recording infrastructure, and API key handling. New features of this scope and complexity require human review.

You can customize Macroscope's approvability policy. Learn more.

apiKey: voiceTranscriptionSettings.apiKey,
onTranscriptInsert: appendVoiceTranscriptToDraft,
onTranscriptSend: (transcript) => {
if (appendVoiceTranscriptToDraft(transcript)) void sendCurrentDraft();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟠 Highthreads/ThreadComposer.tsx:585

The voice “send” action sends the previous draftMessage snapshot instead of the transcript, so the transcript remains in the composer (or nothing is sent). appendVoiceTranscriptToDraft only schedules onChangeDraftMessage, while sendCurrentDraft() immediately calls onSendMessage() without the new text; pass the updated draft through the send path or wait for the state update before sending.

🤖 Copy this AI Prompt to have your agent fix this:
In file @apps/mobile/src/features/threads/ThreadComposer.tsx around line 585:
The voice “send” action sends the previous `draftMessage` snapshot instead of the transcript, so the transcript remains in the composer (or nothing is sent). `appendVoiceTranscriptToDraft` only schedules `onChangeDraftMessage`, while `sendCurrentDraft()` immediately calls `onSendMessage()` without the new text; pass the updated draft through the send path or wait for the state update before sending.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Verified as a false positive. onChangeDraftMessage does not schedule React state: useThreadComposerState calls setComposerDraftText, which synchronously calls appAtomRegistry.set. onSendMessage then calls getComposerDraftSnapshot from that same registry, so the just-appended transcript is present before enqueue. The existing mobile composer-drafts test also covers synchronous latest-draft reads for send.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, I'm unable to act on this request because you do not have permissions within this repository.

@KachurPro
KachurProforce-pushed the codex/voice-dictation-codex-ux branch from 196ba3a to 5d1e3e3CompareAugust 14, 2026 20:13
Comment threadapps/web/src/hooks/useVoiceTranscription.ts Outdated
Comment threadapps/web/src/components/settings/SettingsPanels.tsx
Comment threadapps/web/src/hooks/useVoiceTranscription.ts
Comment threadapps/web/src/hooks/useVoiceTranscription.ts Outdated
Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx
Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated
@brvale97

Copy link
Copy Markdown

Really looking forward to this — mobile dictation is exactly what I've been missing on iPhone.

One request: the web/desktop implementation from #5213 supports both OpenAI and Groq with model selection, but the mobile path hardcodes the OpenAI endpoint and gpt-4o-mini-transcribe in mobileVoiceTranscription.ts. Groq's transcription API is OpenAI-compatible (https://api.groq.com/openai/v1/audio/transcriptions, e.g. whisper-large-v3-turbo), so carrying the provider choice over to mobile would mostly mean making the base URL + model configurable alongside the stored key. Would you consider that?

@KachurPro
KachurProforce-pushed the codex/voice-dictation-codex-ux branch from 59ba2fc to 495d058CompareAugust 14, 2026 21:37
@KachurPro

Copy link
Copy Markdown
Author

@brvale97 Implemented in 495d058bc. Mobile now supports both OpenAI and Groq with fixed official transcription endpoints, separate securely stored keys and model choices per provider, a model picker, and a custom model ID field so future compatible models do not require an app update. OpenAI now defaults to gpt-4o-transcribe; Groq defaults to whisper-large-v3-turbo. Existing saved OpenAI keys are migrated automatically. All 716 mobile tests and the mobile typecheck pass. Thanks for catching this.

Comment threadapps/mobile/src/features/voice-dictation/useMobileVoiceTranscription.ts Outdated
Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

UI consistency review of the web changes (apps/web/src/**). Three findings, all in the new dictation UI: a ticking live region that conflicts with the documented pattern in Sidebar.tsx, and two call sites that rebuild the existing ghost-muted button variant with a class override.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated
Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx Outdated
Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx Outdated
Comment threadapps/web/src/components/chat/ChatComposer.tsx

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the web-scope changes (ChatComposer.tsx, VoiceTranscriptionPanel.tsx, SettingsPanels.tsx, settingsSearch.ts, hooks/useVoiceTranscription.ts, lib/voiceTranscription.ts). The earlier findings on the panel (muted-ghost variant, live-region timer) are resolved. Two remaining consistency issues in the composer.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated
Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

UI consistency review: one finding on the web composer dictation panel. Previous findings (muted-ghost variant usage, live-region timer, responsive error inset, mic pointer-focus) are resolved in this head.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/ChatComposer.tsx
Comment threadapps/web/src/components/chat/ChatComposer.tsx Outdated

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One finding on the dictation panel's pointer-focus contract in the composer footer. Prior comments (mic variant/pointer focus, error-row inset, ghost-muted, aria-hidden timer) are addressed in this revision.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the new Effect code (apps/server/src/transcription.ts, the new route layers in apps/server/src/http.ts, and the contracts/shared additions) against the service conventions. Error classes use Schema.TaggedErrorClass with structural attributes and preserved cause, failures are recovered with Effect.catchTags, HttpClient is acquired from the environment, and Effect modules are imported from their subpaths — all consistent with the conventions. One finding on the web client below.

Posted via Macroscope — Effect Service Conventions

Comment threadapps/web/src/lib/voiceTranscription.ts Outdated
@KachurPro
KachurProforce-pushed the codex/voice-dictation-codex-ux branch from 9f92517 to 4732d3aCompareAugust 14, 2026 22:03
Comment threadapps/web/src/components/settings/SettingsPanels.tsx Outdated

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 2 potential issues.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 4732d3a. Configure here.

Comment threadapps/mobile/src/features/threads/ThreadComposer.tsx Outdated
Comment threadapps/web/src/components/settings/SettingsPanels.tsx Outdated

@macroscopeappmacroscopeappBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One finding: the dictation send action bypasses the composer's themeable send tokens. Everything else (pointer-focus contract, ghost-muted variants, aria-hidden timer, responsive error inset) looks consistent with the composer's existing patterns.

Posted via Macroscope — UI Consistency

Comment threadapps/web/src/components/chat/VoiceTranscriptionPanel.tsx Outdated
@t3dotgg

Copy link
Copy Markdown
Member

Note

🤖 GPT-5.6 Sol responding on behalf of Theo

Closing this PR after an automated pass over open pull requests. Duplicates the maintainer-owned voice implementation in #5213.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXL1,000+ changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@KachurPro@brvale97@t3dotgg@maria-rcks