feat(mobile): add offline iPhone voice input - #233

Merged
rynfar merged 5 commits into
pylonfrom
upstream/2026-09-01-voice-input
Sep 1, 2026
Merged

feat(mobile): add offline iPhone voice input#233
rynfar merged 5 commits into
pylonfrom
upstream/2026-09-01-voice-input

Conversation

@rynfar

@rynfarrynfar commented Sep 1, 2026

Copy link
Copy Markdown
Collaborator

Fourth of the mobile batch. Adopted from upstream pingdotgg/t3code#8614 (352710d49). Follows #223, #224, #231.

Mobile gains on-device voice input: hold the mic in the composer, speak, and the transcript commits into the draft. Brings a voice-input controller to client-runtime, native transcription, the dictation UI, and two patched native deps.

The composer restructure, settled

This batch has cost a conflict in every cherry-pick because upstream's #8793 restructure arrives as context in each one. It turns out to be two independent changes, and treating them as one decision was the mistake:

  • ComposerSurface — animate borderRadius on a shared value, absolute glass layer, bounded collapsed radius. Touches nothing in the toolbar. Adopted (ceaae36aa); our file is now structurally identical to upstream there, so these hunks should stop conflicting.
  • The toolbar row — drop ComposerToolbarScroller for a fixed flex row. Declined, permanently. ComposerToolbarScroller is upstream's own component and they still ship it; they stopped using it because their toolbar holds four controls. Ours holds sixteen, because ControlPillMenu — which does not exist upstream at all — carries Refine, session goal, context window, agent count, the input queue, depth, resources, and reload.

Counts unchanged throughout: scroller 3, ControlPillMenu 13, QuickQuestionTrigger 2, ContextWindowIndicator 3.

Three defects I introduced, and fixed

The dictation controls had to be placed by hand rather than ported, since upstream places them inside the restructure we declined. Review found three ways that stranded the user. All 36 ported files were byte-identical to upstream — every defect was in the hand-written JSX, and none was caught by typecheck, lint, or 1042 tests.

The editor was never frozen.NewTaskDraftScreen passes readOnly={voiceInput.freezesEditor}; the thread composer did not. The keyboard stayed live during recording, and one keystroke makes resolveTranscriptCommit see a changed draft and discard the whole transcript as stale — up to five minutes of speech, silently. It also made every native read-only guard this branch adds dead code on that surface.

Send vanished on any dictation error. Gated on isVoiceInputPresented rather than voicePresentation.showsSend. Those look interchangeable but diverge in exactly one phase: error has showsSend: trueand a non-null statusLabel. So a denied permission, "no speech detected", or the stale-draft error above left no send control. Worse on entry, since dispose() no-ops in the error phase and a stale error survives navigating away and back.

Stop was unreachable for the whole dictation window. Placed inside ComposerToolbarScroller, which is the else branch of the dictation ternary — so across preparing → recording → transcribing → error there was no way to stop a running agent.

Also fixed: blocksSubmission guards on both submission entry points (canSend is derived above voiceInput and structurally cannot include it); a 4px spacer overflowing ComposerDictationToolbar's fixed 44px box and clipping the collapsed strip; the dropped pointerEvents="none" on the glass layer.

The showsSend divergence ships with a regression test — it is the one defect that is a pure predicate rather than JSX placement, and the one most likely to be reintroduced by someone reasoning that the two predicates look the same.

Pylon branding

Two leaks caught, both user-visible. app.config.ts carried upstream's "Allow T3 Code to use your microphone for voice input" — iOS renders that verbatim in the permission dialog, and the camera permission two lines below already said Pylon. Plus two T3 Code strings in docs/user/composer.md and docs/internals/voice-input.md.

No speech-recognition permission is needed, verified rather than assumed: @react-native-ai/apple uses SpeechAnalyzer/SpeechTranscriber, Apple's on-device framework, not SFSpeechRecognizer. That is what makes this offline.

Verification

Typecheck clean, lint clean, 1044 tests passing. Native build green with AppleLLM 0.12.0 and ExpoAudio.

This needs a native rebuild, not an OTA — two new native modules plus native Swift changes in t3-composer-editor/ios. runtimeVersion.policy is fingerprint, which covers new native deps and patches/, so a JS update cannot land on a stale binary.

Simulator pass verified the unavailable path only: with voice unavailable the composer renders correctly, full toolbar, no gap where the mic would be. The dictation UI itself has zero simulator coverageAppleTranscription.isAvailable() is a native check whose own error string reads "requires a supported device with iOS 26 or later", so the on-device model is absent on a simulator. Mic, status row, cancel, and the toolbar swap need a device pass before this merges.

Reviewed and integrated with Claude Opus 5 in Claude Code.


View with [code]smithAutofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

NOT READY TO MERGE. Parked as a branch commit so the conflict resolution
is not lost.
Done: all five conflicts resolved, the two new patched native deps
(@react-native-ai/apple, expo-audio) installed, the voice-input feature
directories and client-runtime module landed intact, and ThreadComposer
compiles with the controller wired (composerOwnerKey, useVoiceInputController,
resolveVoiceComposerPresentation, showsCompactDictation).
Also resolved use-composer-command-menu.ts Pylon-first: upstream's file is far
smaller than Pylon's, so every Pylon export (buildComposerCommandItems,
resolveComposerProviderSlashCommands, and the ranking helpers) is kept and only
upstream's composerSelectionAtEnd helper and owner-key ref are added.
Not done, and the reason this is parked: the dictation UI is imported but not
rendered. ComposerDictationToolbar, ComposerDictationPrimaryAction,
ComposerDictationStatus and ComposerDictationCancelAction are all unused, so the
mic never appears. Wiring them means restructuring Pylon's composer rather than
patching it — upstream wraps its toolbar row directly, while Pylon's is
ComposerToolbarRow > ComposerToolbarScroller with thirteen controls. Pylon also
declares canSend roughly 750 lines above where voiceInput can exist, so even
`canSend && !voiceInput.blocksSubmission` needs the declaration order changed.
Shipping it compiling-but-inert would look done and do nothing, so it waits.
…lbar
#8793 changed two independent things in the mobile composer, and treating
them as one decision has cost a conflict in every cherry-pick from
upstream's 2026-08-30 batch since.
The first is ComposerSurface: animate borderRadius on a shared value, put
the glass on an absolute layer, render children in their own animated view,
and bound the collapsed pill radius so the morph interpolates instead of
travelling from 999. Nothing in it touches the toolbar row.
The second is the toolbar row: drop ComposerToolbarScroller for a fixed flex
row. That one genuinely conflicts. ComposerToolbarScroller is upstream's own
component and they still ship it; they stopped using it because their
toolbar holds four controls. Pylon's holds sixteen, because ControlPillMenu
- which does not exist upstream at all - carries Refine, session goal,
context window, agent count, the input queue, depth, resources, and reload.
Those need the scroller.
So take the first, decline the second. ComposerSurface is now structurally
identical to upstream (animatedBorderRadius, AnimatedGlassSurface,
layoutTransition, animatedShapeStyle, and the bounded radius all match), and
the toolbar is untouched: scroller, 13 ControlPillMenus, QuickQuestionTrigger
and ContextWindowIndicator all at their previous counts.
Pylon's shadow wrapper survives with its comment; upstream has no equivalent,
and the radius it carries is now animated alongside the surface.
Completes the #8614 port's UI half. Upstream places the dictation controls
inside its restructured collapsed row and fixed toolbar; Pylon declined that
restructure, so they are placed into Pylon's own structure instead.
The toolbar now shows whenever isToolbarVisible rather than only when
expanded, so dictation stays reachable from the collapsed pill, and it is
wrapped in ComposerDictationToolbar. The cancel action leads the row; while
dictating, ComposerDictationStatus replaces the toolbar scroller rather than
upstream's fixed left group, so Pylon's thirteen ControlPillMenu controls
keep their scroller when not dictating. The mic sits beside send in both the
collapsed row and the toolbar, and send is hidden while dictation owns the
row.
Scroller, ControlPillMenu, QuickQuestionTrigger and ContextWindowIndicator
are all at unchanged counts.
The #8614 port carried upstream's string verbatim: "Allow T3 Code to use
your microphone for voice input." iOS shows that text in the permission
dialog, so it is product copy, not a compatibility identifier. The camera
permission two lines below already reads "Allow Pylon to access your
camera", so this was purely adoption drift.
Verified while checking permissions that no speech-recognition key is
needed: @react-native-ai/apple uses SpeechAnalyzer and SpeechTranscriber,
Apple's on-device Speech framework, rather than SFSpeechRecognizer. Only
NSMicrophoneUsageDescription applies, and it is present.
Adversarial review of the hand-placed dictation UI found three ways the
composer strands the user. All three are in code written by hand rather
than ported, and none was caught by typecheck, lint, or 1042 tests.
The editor was never frozen. NewTaskDraftScreen passes
readOnly={voiceInput.freezesEditor}; the thread composer did not, so the
keyboard stayed live during recording. One keystroke makes
resolveTranscriptCommit see a changed draft and discard the entire
transcript as stale - up to five minutes of speech, silently. It also made
every native read-only guard this branch adds dead code on this surface.
Send was gated on isVoiceInputPresented rather than
voicePresentation.showsSend. Those look equivalent but diverge in exactly
one phase: error shows a status label AND keeps send. Any dictation failure
- denied permission, no speech detected, the stale-draft error above - left
the composer with no send control until the user found the dismiss button.
Worse on entry, since dispose() no-ops in the error phase, so a stale error
survives navigating away and back.
Stop sat inside ComposerToolbarScroller, which is the else branch of the
dictation ternary, so an agent was unstoppable for the whole recording and
transcription window. Moved to the always-rendered right cluster.
Also: guard both submission entry points on blocksSubmission, since canSend
is derived above voiceInput and cannot include it; move the 4px spacer
outside ComposerDictationToolbar's fixed 44px box, where it was overflowing
and clipping the collapsed dictation strip; restore pointerEvents="none" on
the glass layer with a comment matching the new sibling structure; and
rebrand two T3 Code strings in docs.
The showsSend divergence now has a regression test. It is the one defect
here that is a pure predicate rather than JSX placement, and it is the one
most likely to be reintroduced.
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:XXL labels Sep 1, 2026
@github-actions

Copy link
Copy Markdown

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ProviderMetricMain baselineThis PRImpactPR ceiling
CodexTotal thread wire13.7 KiB13.7 KiB+8 B (+0.1%)15.1 KiB
CodexThread snapshot wire6.9 KiB6.9 KiB−3 B (−0.0%)7.3 KiB
CodexLive turn WebSocket wire6.8 KiB6.8 KiB+11 B (+0.2%)7.8 KiB
CodexLive turn WebSocket decoded58.7 KiB58.7 KiB0 B (0.0%)66.4 KiB
CodexLive turn messages11110 (0.0%)21
ClaudeTotal thread wire13.5 KiB13.8 KiB+230 B (+1.7%)15.1 KiB
ClaudeThread snapshot wire6.9 KiB6.9 KiB+2 B (+0.0%)7.3 KiB
ClaudeLive turn WebSocket wire6.6 KiB6.8 KiB+228 B (+3.4%)7.8 KiB
ClaudeLive turn WebSocket decoded58.0 KiB59.5 KiB+1.5 KiB (+2.6%)66.4 KiB
ClaudeLive turn messages911+2 (+22.2%)21

Baseline: c908bb2 · PR result: b21cb8f · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 109.5 KiB
  • Claude decoded thread snapshot: 110.2 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

@rynfar
rynfar merged commit 74377fd into pylonSep 1, 2026
19 checks passed
@rynfar
rynfar deleted the upstream/2026-09-01-voice-input branch September 1, 2026 23:28
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXLvouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@rynfar
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

feat(mobile): add offline iPhone voice input - #233

Merged
rynfar merged 5 commits into
pylonfrom
upstream/2026-09-01-voice-input
Sep 1, 2026
Merged

feat(mobile): add offline iPhone voice input#233
rynfar merged 5 commits into
pylonfrom
upstream/2026-09-01-voice-input

Conversation

@rynfar

@rynfarrynfar commented Sep 1, 2026

Copy link
Copy Markdown
Collaborator

Fourth of the mobile batch. Adopted from upstream pingdotgg/t3code#8614 (352710d49). Follows #223, #224, #231.

Mobile gains on-device voice input: hold the mic in the composer, speak, and the transcript commits into the draft. Brings a voice-input controller to client-runtime, native transcription, the dictation UI, and two patched native deps.

The composer restructure, settled

This batch has cost a conflict in every cherry-pick because upstream's #8793 restructure arrives as context in each one. It turns out to be two independent changes, and treating them as one decision was the mistake:

  • ComposerSurface — animate borderRadius on a shared value, absolute glass layer, bounded collapsed radius. Touches nothing in the toolbar. Adopted (ceaae36aa); our file is now structurally identical to upstream there, so these hunks should stop conflicting.
  • The toolbar row — drop ComposerToolbarScroller for a fixed flex row. Declined, permanently. ComposerToolbarScroller is upstream's own component and they still ship it; they stopped using it because their toolbar holds four controls. Ours holds sixteen, because ControlPillMenu — which does not exist upstream at all — carries Refine, session goal, context window, agent count, the input queue, depth, resources, and reload.

Counts unchanged throughout: scroller 3, ControlPillMenu 13, QuickQuestionTrigger 2, ContextWindowIndicator 3.

Three defects I introduced, and fixed

The dictation controls had to be placed by hand rather than ported, since upstream places them inside the restructure we declined. Review found three ways that stranded the user. All 36 ported files were byte-identical to upstream — every defect was in the hand-written JSX, and none was caught by typecheck, lint, or 1042 tests.

The editor was never frozen.NewTaskDraftScreen passes readOnly={voiceInput.freezesEditor}; the thread composer did not. The keyboard stayed live during recording, and one keystroke makes resolveTranscriptCommit see a changed draft and discard the whole transcript as stale — up to five minutes of speech, silently. It also made every native read-only guard this branch adds dead code on that surface.

Send vanished on any dictation error. Gated on isVoiceInputPresented rather than voicePresentation.showsSend. Those look interchangeable but diverge in exactly one phase: error has showsSend: trueand a non-null statusLabel. So a denied permission, "no speech detected", or the stale-draft error above left no send control. Worse on entry, since dispose() no-ops in the error phase and a stale error survives navigating away and back.

Stop was unreachable for the whole dictation window. Placed inside ComposerToolbarScroller, which is the else branch of the dictation ternary — so across preparing → recording → transcribing → error there was no way to stop a running agent.

Also fixed: blocksSubmission guards on both submission entry points (canSend is derived above voiceInput and structurally cannot include it); a 4px spacer overflowing ComposerDictationToolbar's fixed 44px box and clipping the collapsed strip; the dropped pointerEvents="none" on the glass layer.

The showsSend divergence ships with a regression test — it is the one defect that is a pure predicate rather than JSX placement, and the one most likely to be reintroduced by someone reasoning that the two predicates look the same.

Pylon branding

Two leaks caught, both user-visible. app.config.ts carried upstream's "Allow T3 Code to use your microphone for voice input" — iOS renders that verbatim in the permission dialog, and the camera permission two lines below already said Pylon. Plus two T3 Code strings in docs/user/composer.md and docs/internals/voice-input.md.

No speech-recognition permission is needed, verified rather than assumed: @react-native-ai/apple uses SpeechAnalyzer/SpeechTranscriber, Apple's on-device framework, not SFSpeechRecognizer. That is what makes this offline.

Verification

Typecheck clean, lint clean, 1044 tests passing. Native build green with AppleLLM 0.12.0 and ExpoAudio.

This needs a native rebuild, not an OTA — two new native modules plus native Swift changes in t3-composer-editor/ios. runtimeVersion.policy is fingerprint, which covers new native deps and patches/, so a JS update cannot land on a stale binary.

Simulator pass verified the unavailable path only: with voice unavailable the composer renders correctly, full toolbar, no gap where the mic would be. The dictation UI itself has zero simulator coverageAppleTranscription.isAvailable() is a native check whose own error string reads "requires a supported device with iOS 26 or later", so the on-device model is absent on a simulator. Mic, status row, cancel, and the toolbar swap need a device pass before this merges.

Reviewed and integrated with Claude Opus 5 in Claude Code.


View with [code]smithAutofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

NOT READY TO MERGE. Parked as a branch commit so the conflict resolution
is not lost.
Done: all five conflicts resolved, the two new patched native deps
(@react-native-ai/apple, expo-audio) installed, the voice-input feature
directories and client-runtime module landed intact, and ThreadComposer
compiles with the controller wired (composerOwnerKey, useVoiceInputController,
resolveVoiceComposerPresentation, showsCompactDictation).
Also resolved use-composer-command-menu.ts Pylon-first: upstream's file is far
smaller than Pylon's, so every Pylon export (buildComposerCommandItems,
resolveComposerProviderSlashCommands, and the ranking helpers) is kept and only
upstream's composerSelectionAtEnd helper and owner-key ref are added.
Not done, and the reason this is parked: the dictation UI is imported but not
rendered. ComposerDictationToolbar, ComposerDictationPrimaryAction,
ComposerDictationStatus and ComposerDictationCancelAction are all unused, so the
mic never appears. Wiring them means restructuring Pylon's composer rather than
patching it — upstream wraps its toolbar row directly, while Pylon's is
ComposerToolbarRow > ComposerToolbarScroller with thirteen controls. Pylon also
declares canSend roughly 750 lines above where voiceInput can exist, so even
`canSend && !voiceInput.blocksSubmission` needs the declaration order changed.
Shipping it compiling-but-inert would look done and do nothing, so it waits.
…lbar
#8793 changed two independent things in the mobile composer, and treating
them as one decision has cost a conflict in every cherry-pick from
upstream's 2026-08-30 batch since.
The first is ComposerSurface: animate borderRadius on a shared value, put
the glass on an absolute layer, render children in their own animated view,
and bound the collapsed pill radius so the morph interpolates instead of
travelling from 999. Nothing in it touches the toolbar row.
The second is the toolbar row: drop ComposerToolbarScroller for a fixed flex
row. That one genuinely conflicts. ComposerToolbarScroller is upstream's own
component and they still ship it; they stopped using it because their
toolbar holds four controls. Pylon's holds sixteen, because ControlPillMenu
- which does not exist upstream at all - carries Refine, session goal,
context window, agent count, the input queue, depth, resources, and reload.
Those need the scroller.
So take the first, decline the second. ComposerSurface is now structurally
identical to upstream (animatedBorderRadius, AnimatedGlassSurface,
layoutTransition, animatedShapeStyle, and the bounded radius all match), and
the toolbar is untouched: scroller, 13 ControlPillMenus, QuickQuestionTrigger
and ContextWindowIndicator all at their previous counts.
Pylon's shadow wrapper survives with its comment; upstream has no equivalent,
and the radius it carries is now animated alongside the surface.
Completes the #8614 port's UI half. Upstream places the dictation controls
inside its restructured collapsed row and fixed toolbar; Pylon declined that
restructure, so they are placed into Pylon's own structure instead.
The toolbar now shows whenever isToolbarVisible rather than only when
expanded, so dictation stays reachable from the collapsed pill, and it is
wrapped in ComposerDictationToolbar. The cancel action leads the row; while
dictating, ComposerDictationStatus replaces the toolbar scroller rather than
upstream's fixed left group, so Pylon's thirteen ControlPillMenu controls
keep their scroller when not dictating. The mic sits beside send in both the
collapsed row and the toolbar, and send is hidden while dictation owns the
row.
Scroller, ControlPillMenu, QuickQuestionTrigger and ContextWindowIndicator
are all at unchanged counts.
The #8614 port carried upstream's string verbatim: "Allow T3 Code to use
your microphone for voice input." iOS shows that text in the permission
dialog, so it is product copy, not a compatibility identifier. The camera
permission two lines below already reads "Allow Pylon to access your
camera", so this was purely adoption drift.
Verified while checking permissions that no speech-recognition key is
needed: @react-native-ai/apple uses SpeechAnalyzer and SpeechTranscriber,
Apple's on-device Speech framework, rather than SFSpeechRecognizer. Only
NSMicrophoneUsageDescription applies, and it is present.
Adversarial review of the hand-placed dictation UI found three ways the
composer strands the user. All three are in code written by hand rather
than ported, and none was caught by typecheck, lint, or 1042 tests.
The editor was never frozen. NewTaskDraftScreen passes
readOnly={voiceInput.freezesEditor}; the thread composer did not, so the
keyboard stayed live during recording. One keystroke makes
resolveTranscriptCommit see a changed draft and discard the entire
transcript as stale - up to five minutes of speech, silently. It also made
every native read-only guard this branch adds dead code on this surface.
Send was gated on isVoiceInputPresented rather than
voicePresentation.showsSend. Those look equivalent but diverge in exactly
one phase: error shows a status label AND keeps send. Any dictation failure
- denied permission, no speech detected, the stale-draft error above - left
the composer with no send control until the user found the dismiss button.
Worse on entry, since dispose() no-ops in the error phase, so a stale error
survives navigating away and back.
Stop sat inside ComposerToolbarScroller, which is the else branch of the
dictation ternary, so an agent was unstoppable for the whole recording and
transcription window. Moved to the always-rendered right cluster.
Also: guard both submission entry points on blocksSubmission, since canSend
is derived above voiceInput and cannot include it; move the 4px spacer
outside ComposerDictationToolbar's fixed 44px box, where it was overflowing
and clipping the collapsed dictation strip; restore pointerEvents="none" on
the glass layer with a comment matching the new sibling structure; and
rebrand two T3 Code strings in docs.
The showsSend divergence now has a regression test. It is the one defect
here that is a pure predicate rather than JSX placement, and it is the one
most likely to be reintroduced.
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:XXL labels Sep 1, 2026
@github-actions

Copy link
Copy Markdown

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ProviderMetricMain baselineThis PRImpactPR ceiling
CodexTotal thread wire13.7 KiB13.7 KiB+8 B (+0.1%)15.1 KiB
CodexThread snapshot wire6.9 KiB6.9 KiB−3 B (−0.0%)7.3 KiB
CodexLive turn WebSocket wire6.8 KiB6.8 KiB+11 B (+0.2%)7.8 KiB
CodexLive turn WebSocket decoded58.7 KiB58.7 KiB0 B (0.0%)66.4 KiB
CodexLive turn messages11110 (0.0%)21
ClaudeTotal thread wire13.5 KiB13.8 KiB+230 B (+1.7%)15.1 KiB
ClaudeThread snapshot wire6.9 KiB6.9 KiB+2 B (+0.0%)7.3 KiB
ClaudeLive turn WebSocket wire6.6 KiB6.8 KiB+228 B (+3.4%)7.8 KiB
ClaudeLive turn WebSocket decoded58.0 KiB59.5 KiB+1.5 KiB (+2.6%)66.4 KiB
ClaudeLive turn messages911+2 (+22.2%)21

Baseline: c908bb2 · PR result: b21cb8f · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 109.5 KiB
  • Claude decoded thread snapshot: 110.2 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

@rynfar
rynfar merged commit 74377fd into pylonSep 1, 2026
19 checks passed
@rynfar
rynfar deleted the upstream/2026-09-01-voice-input branch September 1, 2026 23:28
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXLvouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@rynfar
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(mobile): add offline iPhone voice input - #233

Merged
rynfar merged 5 commits into
pylonfrom
upstream/2026-09-01-voice-input
Sep 1, 2026
Merged

feat(mobile): add offline iPhone voice input#233
rynfar merged 5 commits into
pylonfrom
upstream/2026-09-01-voice-input

Conversation

@rynfar

@rynfarrynfar commented Sep 1, 2026

Copy link
Copy Markdown
Collaborator

Fourth of the mobile batch. Adopted from upstream pingdotgg/t3code#8614 (352710d49). Follows #223, #224, #231.

Mobile gains on-device voice input: hold the mic in the composer, speak, and the transcript commits into the draft. Brings a voice-input controller to client-runtime, native transcription, the dictation UI, and two patched native deps.

The composer restructure, settled

This batch has cost a conflict in every cherry-pick because upstream's #8793 restructure arrives as context in each one. It turns out to be two independent changes, and treating them as one decision was the mistake:

  • ComposerSurface — animate borderRadius on a shared value, absolute glass layer, bounded collapsed radius. Touches nothing in the toolbar. Adopted (ceaae36aa); our file is now structurally identical to upstream there, so these hunks should stop conflicting.
  • The toolbar row — drop ComposerToolbarScroller for a fixed flex row. Declined, permanently. ComposerToolbarScroller is upstream's own component and they still ship it; they stopped using it because their toolbar holds four controls. Ours holds sixteen, because ControlPillMenu — which does not exist upstream at all — carries Refine, session goal, context window, agent count, the input queue, depth, resources, and reload.

Counts unchanged throughout: scroller 3, ControlPillMenu 13, QuickQuestionTrigger 2, ContextWindowIndicator 3.

Three defects I introduced, and fixed

The dictation controls had to be placed by hand rather than ported, since upstream places them inside the restructure we declined. Review found three ways that stranded the user. All 36 ported files were byte-identical to upstream — every defect was in the hand-written JSX, and none was caught by typecheck, lint, or 1042 tests.

The editor was never frozen.NewTaskDraftScreen passes readOnly={voiceInput.freezesEditor}; the thread composer did not. The keyboard stayed live during recording, and one keystroke makes resolveTranscriptCommit see a changed draft and discard the whole transcript as stale — up to five minutes of speech, silently. It also made every native read-only guard this branch adds dead code on that surface.

Send vanished on any dictation error. Gated on isVoiceInputPresented rather than voicePresentation.showsSend. Those look interchangeable but diverge in exactly one phase: error has showsSend: trueand a non-null statusLabel. So a denied permission, "no speech detected", or the stale-draft error above left no send control. Worse on entry, since dispose() no-ops in the error phase and a stale error survives navigating away and back.

Stop was unreachable for the whole dictation window. Placed inside ComposerToolbarScroller, which is the else branch of the dictation ternary — so across preparing → recording → transcribing → error there was no way to stop a running agent.

Also fixed: blocksSubmission guards on both submission entry points (canSend is derived above voiceInput and structurally cannot include it); a 4px spacer overflowing ComposerDictationToolbar's fixed 44px box and clipping the collapsed strip; the dropped pointerEvents="none" on the glass layer.

The showsSend divergence ships with a regression test — it is the one defect that is a pure predicate rather than JSX placement, and the one most likely to be reintroduced by someone reasoning that the two predicates look the same.

Pylon branding

Two leaks caught, both user-visible. app.config.ts carried upstream's "Allow T3 Code to use your microphone for voice input" — iOS renders that verbatim in the permission dialog, and the camera permission two lines below already said Pylon. Plus two T3 Code strings in docs/user/composer.md and docs/internals/voice-input.md.

No speech-recognition permission is needed, verified rather than assumed: @react-native-ai/apple uses SpeechAnalyzer/SpeechTranscriber, Apple's on-device framework, not SFSpeechRecognizer. That is what makes this offline.

Verification

Typecheck clean, lint clean, 1044 tests passing. Native build green with AppleLLM 0.12.0 and ExpoAudio.

This needs a native rebuild, not an OTA — two new native modules plus native Swift changes in t3-composer-editor/ios. runtimeVersion.policy is fingerprint, which covers new native deps and patches/, so a JS update cannot land on a stale binary.

Simulator pass verified the unavailable path only: with voice unavailable the composer renders correctly, full toolbar, no gap where the mic would be. The dictation UI itself has zero simulator coverageAppleTranscription.isAvailable() is a native check whose own error string reads "requires a supported device with iOS 26 or later", so the on-device model is absent on a simulator. Mic, status row, cancel, and the toolbar swap need a device pass before this merges.

Reviewed and integrated with Claude Opus 5 in Claude Code.


View with [code]smithAutofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

NOT READY TO MERGE. Parked as a branch commit so the conflict resolution
is not lost.
Done: all five conflicts resolved, the two new patched native deps
(@react-native-ai/apple, expo-audio) installed, the voice-input feature
directories and client-runtime module landed intact, and ThreadComposer
compiles with the controller wired (composerOwnerKey, useVoiceInputController,
resolveVoiceComposerPresentation, showsCompactDictation).
Also resolved use-composer-command-menu.ts Pylon-first: upstream's file is far
smaller than Pylon's, so every Pylon export (buildComposerCommandItems,
resolveComposerProviderSlashCommands, and the ranking helpers) is kept and only
upstream's composerSelectionAtEnd helper and owner-key ref are added.
Not done, and the reason this is parked: the dictation UI is imported but not
rendered. ComposerDictationToolbar, ComposerDictationPrimaryAction,
ComposerDictationStatus and ComposerDictationCancelAction are all unused, so the
mic never appears. Wiring them means restructuring Pylon's composer rather than
patching it — upstream wraps its toolbar row directly, while Pylon's is
ComposerToolbarRow > ComposerToolbarScroller with thirteen controls. Pylon also
declares canSend roughly 750 lines above where voiceInput can exist, so even
`canSend && !voiceInput.blocksSubmission` needs the declaration order changed.
Shipping it compiling-but-inert would look done and do nothing, so it waits.
…lbar
#8793 changed two independent things in the mobile composer, and treating
them as one decision has cost a conflict in every cherry-pick from
upstream's 2026-08-30 batch since.
The first is ComposerSurface: animate borderRadius on a shared value, put
the glass on an absolute layer, render children in their own animated view,
and bound the collapsed pill radius so the morph interpolates instead of
travelling from 999. Nothing in it touches the toolbar row.
The second is the toolbar row: drop ComposerToolbarScroller for a fixed flex
row. That one genuinely conflicts. ComposerToolbarScroller is upstream's own
component and they still ship it; they stopped using it because their
toolbar holds four controls. Pylon's holds sixteen, because ControlPillMenu
- which does not exist upstream at all - carries Refine, session goal,
context window, agent count, the input queue, depth, resources, and reload.
Those need the scroller.
So take the first, decline the second. ComposerSurface is now structurally
identical to upstream (animatedBorderRadius, AnimatedGlassSurface,
layoutTransition, animatedShapeStyle, and the bounded radius all match), and
the toolbar is untouched: scroller, 13 ControlPillMenus, QuickQuestionTrigger
and ContextWindowIndicator all at their previous counts.
Pylon's shadow wrapper survives with its comment; upstream has no equivalent,
and the radius it carries is now animated alongside the surface.
Completes the #8614 port's UI half. Upstream places the dictation controls
inside its restructured collapsed row and fixed toolbar; Pylon declined that
restructure, so they are placed into Pylon's own structure instead.
The toolbar now shows whenever isToolbarVisible rather than only when
expanded, so dictation stays reachable from the collapsed pill, and it is
wrapped in ComposerDictationToolbar. The cancel action leads the row; while
dictating, ComposerDictationStatus replaces the toolbar scroller rather than
upstream's fixed left group, so Pylon's thirteen ControlPillMenu controls
keep their scroller when not dictating. The mic sits beside send in both the
collapsed row and the toolbar, and send is hidden while dictation owns the
row.
Scroller, ControlPillMenu, QuickQuestionTrigger and ContextWindowIndicator
are all at unchanged counts.
The #8614 port carried upstream's string verbatim: "Allow T3 Code to use
your microphone for voice input." iOS shows that text in the permission
dialog, so it is product copy, not a compatibility identifier. The camera
permission two lines below already reads "Allow Pylon to access your
camera", so this was purely adoption drift.
Verified while checking permissions that no speech-recognition key is
needed: @react-native-ai/apple uses SpeechAnalyzer and SpeechTranscriber,
Apple's on-device Speech framework, rather than SFSpeechRecognizer. Only
NSMicrophoneUsageDescription applies, and it is present.
Adversarial review of the hand-placed dictation UI found three ways the
composer strands the user. All three are in code written by hand rather
than ported, and none was caught by typecheck, lint, or 1042 tests.
The editor was never frozen. NewTaskDraftScreen passes
readOnly={voiceInput.freezesEditor}; the thread composer did not, so the
keyboard stayed live during recording. One keystroke makes
resolveTranscriptCommit see a changed draft and discard the entire
transcript as stale - up to five minutes of speech, silently. It also made
every native read-only guard this branch adds dead code on this surface.
Send was gated on isVoiceInputPresented rather than
voicePresentation.showsSend. Those look equivalent but diverge in exactly
one phase: error shows a status label AND keeps send. Any dictation failure
- denied permission, no speech detected, the stale-draft error above - left
the composer with no send control until the user found the dismiss button.
Worse on entry, since dispose() no-ops in the error phase, so a stale error
survives navigating away and back.
Stop sat inside ComposerToolbarScroller, which is the else branch of the
dictation ternary, so an agent was unstoppable for the whole recording and
transcription window. Moved to the always-rendered right cluster.
Also: guard both submission entry points on blocksSubmission, since canSend
is derived above voiceInput and cannot include it; move the 4px spacer
outside ComposerDictationToolbar's fixed 44px box, where it was overflowing
and clipping the collapsed dictation strip; restore pointerEvents="none" on
the glass layer with a comment matching the new sibling structure; and
rebrand two T3 Code strings in docs.
The showsSend divergence now has a regression test. It is the one defect
here that is a pure predicate rather than JSX placement, and it is the one
most likely to be reintroduced.
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:XXL labels Sep 1, 2026
@github-actions

Copy link
Copy Markdown

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ProviderMetricMain baselineThis PRImpactPR ceiling
CodexTotal thread wire13.7 KiB13.7 KiB+8 B (+0.1%)15.1 KiB
CodexThread snapshot wire6.9 KiB6.9 KiB−3 B (−0.0%)7.3 KiB
CodexLive turn WebSocket wire6.8 KiB6.8 KiB+11 B (+0.2%)7.8 KiB
CodexLive turn WebSocket decoded58.7 KiB58.7 KiB0 B (0.0%)66.4 KiB
CodexLive turn messages11110 (0.0%)21
ClaudeTotal thread wire13.5 KiB13.8 KiB+230 B (+1.7%)15.1 KiB
ClaudeThread snapshot wire6.9 KiB6.9 KiB+2 B (+0.0%)7.3 KiB
ClaudeLive turn WebSocket wire6.6 KiB6.8 KiB+228 B (+3.4%)7.8 KiB
ClaudeLive turn WebSocket decoded58.0 KiB59.5 KiB+1.5 KiB (+2.6%)66.4 KiB
ClaudeLive turn messages911+2 (+22.2%)21

Baseline: c908bb2 · PR result: b21cb8f · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 109.5 KiB
  • Claude decoded thread snapshot: 110.2 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

@rynfar
rynfar merged commit 74377fd into pylonSep 1, 2026
19 checks passed
@rynfar
rynfar deleted the upstream/2026-09-01-voice-input branch September 1, 2026 23:28
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXLvouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@rynfar
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(mobile): add offline iPhone voice input - #233

Merged
rynfar merged 5 commits into
pylonfrom
upstream/2026-09-01-voice-input
Sep 1, 2026
Merged

feat(mobile): add offline iPhone voice input#233
rynfar merged 5 commits into
pylonfrom
upstream/2026-09-01-voice-input

Conversation

@rynfar

@rynfarrynfar commented Sep 1, 2026

Copy link
Copy Markdown
Collaborator

Fourth of the mobile batch. Adopted from upstream pingdotgg/t3code#8614 (352710d49). Follows #223, #224, #231.

Mobile gains on-device voice input: hold the mic in the composer, speak, and the transcript commits into the draft. Brings a voice-input controller to client-runtime, native transcription, the dictation UI, and two patched native deps.

The composer restructure, settled

This batch has cost a conflict in every cherry-pick because upstream's #8793 restructure arrives as context in each one. It turns out to be two independent changes, and treating them as one decision was the mistake:

  • ComposerSurface — animate borderRadius on a shared value, absolute glass layer, bounded collapsed radius. Touches nothing in the toolbar. Adopted (ceaae36aa); our file is now structurally identical to upstream there, so these hunks should stop conflicting.
  • The toolbar row — drop ComposerToolbarScroller for a fixed flex row. Declined, permanently. ComposerToolbarScroller is upstream's own component and they still ship it; they stopped using it because their toolbar holds four controls. Ours holds sixteen, because ControlPillMenu — which does not exist upstream at all — carries Refine, session goal, context window, agent count, the input queue, depth, resources, and reload.

Counts unchanged throughout: scroller 3, ControlPillMenu 13, QuickQuestionTrigger 2, ContextWindowIndicator 3.

Three defects I introduced, and fixed

The dictation controls had to be placed by hand rather than ported, since upstream places them inside the restructure we declined. Review found three ways that stranded the user. All 36 ported files were byte-identical to upstream — every defect was in the hand-written JSX, and none was caught by typecheck, lint, or 1042 tests.

The editor was never frozen.NewTaskDraftScreen passes readOnly={voiceInput.freezesEditor}; the thread composer did not. The keyboard stayed live during recording, and one keystroke makes resolveTranscriptCommit see a changed draft and discard the whole transcript as stale — up to five minutes of speech, silently. It also made every native read-only guard this branch adds dead code on that surface.

Send vanished on any dictation error. Gated on isVoiceInputPresented rather than voicePresentation.showsSend. Those look interchangeable but diverge in exactly one phase: error has showsSend: trueand a non-null statusLabel. So a denied permission, "no speech detected", or the stale-draft error above left no send control. Worse on entry, since dispose() no-ops in the error phase and a stale error survives navigating away and back.

Stop was unreachable for the whole dictation window. Placed inside ComposerToolbarScroller, which is the else branch of the dictation ternary — so across preparing → recording → transcribing → error there was no way to stop a running agent.

Also fixed: blocksSubmission guards on both submission entry points (canSend is derived above voiceInput and structurally cannot include it); a 4px spacer overflowing ComposerDictationToolbar's fixed 44px box and clipping the collapsed strip; the dropped pointerEvents="none" on the glass layer.

The showsSend divergence ships with a regression test — it is the one defect that is a pure predicate rather than JSX placement, and the one most likely to be reintroduced by someone reasoning that the two predicates look the same.

Pylon branding

Two leaks caught, both user-visible. app.config.ts carried upstream's "Allow T3 Code to use your microphone for voice input" — iOS renders that verbatim in the permission dialog, and the camera permission two lines below already said Pylon. Plus two T3 Code strings in docs/user/composer.md and docs/internals/voice-input.md.

No speech-recognition permission is needed, verified rather than assumed: @react-native-ai/apple uses SpeechAnalyzer/SpeechTranscriber, Apple's on-device framework, not SFSpeechRecognizer. That is what makes this offline.

Verification

Typecheck clean, lint clean, 1044 tests passing. Native build green with AppleLLM 0.12.0 and ExpoAudio.

This needs a native rebuild, not an OTA — two new native modules plus native Swift changes in t3-composer-editor/ios. runtimeVersion.policy is fingerprint, which covers new native deps and patches/, so a JS update cannot land on a stale binary.

Simulator pass verified the unavailable path only: with voice unavailable the composer renders correctly, full toolbar, no gap where the mic would be. The dictation UI itself has zero simulator coverageAppleTranscription.isAvailable() is a native check whose own error string reads "requires a supported device with iOS 26 or later", so the on-device model is absent on a simulator. Mic, status row, cancel, and the toolbar swap need a device pass before this merges.

Reviewed and integrated with Claude Opus 5 in Claude Code.


View with [code]smithAutofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

NOT READY TO MERGE. Parked as a branch commit so the conflict resolution
is not lost.
Done: all five conflicts resolved, the two new patched native deps
(@react-native-ai/apple, expo-audio) installed, the voice-input feature
directories and client-runtime module landed intact, and ThreadComposer
compiles with the controller wired (composerOwnerKey, useVoiceInputController,
resolveVoiceComposerPresentation, showsCompactDictation).
Also resolved use-composer-command-menu.ts Pylon-first: upstream's file is far
smaller than Pylon's, so every Pylon export (buildComposerCommandItems,
resolveComposerProviderSlashCommands, and the ranking helpers) is kept and only
upstream's composerSelectionAtEnd helper and owner-key ref are added.
Not done, and the reason this is parked: the dictation UI is imported but not
rendered. ComposerDictationToolbar, ComposerDictationPrimaryAction,
ComposerDictationStatus and ComposerDictationCancelAction are all unused, so the
mic never appears. Wiring them means restructuring Pylon's composer rather than
patching it — upstream wraps its toolbar row directly, while Pylon's is
ComposerToolbarRow > ComposerToolbarScroller with thirteen controls. Pylon also
declares canSend roughly 750 lines above where voiceInput can exist, so even
`canSend && !voiceInput.blocksSubmission` needs the declaration order changed.
Shipping it compiling-but-inert would look done and do nothing, so it waits.
…lbar
#8793 changed two independent things in the mobile composer, and treating
them as one decision has cost a conflict in every cherry-pick from
upstream's 2026-08-30 batch since.
The first is ComposerSurface: animate borderRadius on a shared value, put
the glass on an absolute layer, render children in their own animated view,
and bound the collapsed pill radius so the morph interpolates instead of
travelling from 999. Nothing in it touches the toolbar row.
The second is the toolbar row: drop ComposerToolbarScroller for a fixed flex
row. That one genuinely conflicts. ComposerToolbarScroller is upstream's own
component and they still ship it; they stopped using it because their
toolbar holds four controls. Pylon's holds sixteen, because ControlPillMenu
- which does not exist upstream at all - carries Refine, session goal,
context window, agent count, the input queue, depth, resources, and reload.
Those need the scroller.
So take the first, decline the second. ComposerSurface is now structurally
identical to upstream (animatedBorderRadius, AnimatedGlassSurface,
layoutTransition, animatedShapeStyle, and the bounded radius all match), and
the toolbar is untouched: scroller, 13 ControlPillMenus, QuickQuestionTrigger
and ContextWindowIndicator all at their previous counts.
Pylon's shadow wrapper survives with its comment; upstream has no equivalent,
and the radius it carries is now animated alongside the surface.
Completes the #8614 port's UI half. Upstream places the dictation controls
inside its restructured collapsed row and fixed toolbar; Pylon declined that
restructure, so they are placed into Pylon's own structure instead.
The toolbar now shows whenever isToolbarVisible rather than only when
expanded, so dictation stays reachable from the collapsed pill, and it is
wrapped in ComposerDictationToolbar. The cancel action leads the row; while
dictating, ComposerDictationStatus replaces the toolbar scroller rather than
upstream's fixed left group, so Pylon's thirteen ControlPillMenu controls
keep their scroller when not dictating. The mic sits beside send in both the
collapsed row and the toolbar, and send is hidden while dictation owns the
row.
Scroller, ControlPillMenu, QuickQuestionTrigger and ContextWindowIndicator
are all at unchanged counts.
The #8614 port carried upstream's string verbatim: "Allow T3 Code to use
your microphone for voice input." iOS shows that text in the permission
dialog, so it is product copy, not a compatibility identifier. The camera
permission two lines below already reads "Allow Pylon to access your
camera", so this was purely adoption drift.
Verified while checking permissions that no speech-recognition key is
needed: @react-native-ai/apple uses SpeechAnalyzer and SpeechTranscriber,
Apple's on-device Speech framework, rather than SFSpeechRecognizer. Only
NSMicrophoneUsageDescription applies, and it is present.
Adversarial review of the hand-placed dictation UI found three ways the
composer strands the user. All three are in code written by hand rather
than ported, and none was caught by typecheck, lint, or 1042 tests.
The editor was never frozen. NewTaskDraftScreen passes
readOnly={voiceInput.freezesEditor}; the thread composer did not, so the
keyboard stayed live during recording. One keystroke makes
resolveTranscriptCommit see a changed draft and discard the entire
transcript as stale - up to five minutes of speech, silently. It also made
every native read-only guard this branch adds dead code on this surface.
Send was gated on isVoiceInputPresented rather than
voicePresentation.showsSend. Those look equivalent but diverge in exactly
one phase: error shows a status label AND keeps send. Any dictation failure
- denied permission, no speech detected, the stale-draft error above - left
the composer with no send control until the user found the dismiss button.
Worse on entry, since dispose() no-ops in the error phase, so a stale error
survives navigating away and back.
Stop sat inside ComposerToolbarScroller, which is the else branch of the
dictation ternary, so an agent was unstoppable for the whole recording and
transcription window. Moved to the always-rendered right cluster.
Also: guard both submission entry points on blocksSubmission, since canSend
is derived above voiceInput and cannot include it; move the 4px spacer
outside ComposerDictationToolbar's fixed 44px box, where it was overflowing
and clipping the collapsed dictation strip; restore pointerEvents="none" on
the glass layer with a comment matching the new sibling structure; and
rebrand two T3 Code strings in docs.
The showsSend divergence now has a regression test. It is the one defect
here that is a pure predicate rather than JSX placement, and it is the one
most likely to be reintroduced.
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:XXL labels Sep 1, 2026
@github-actions

Copy link
Copy Markdown

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ProviderMetricMain baselineThis PRImpactPR ceiling
CodexTotal thread wire13.7 KiB13.7 KiB+8 B (+0.1%)15.1 KiB
CodexThread snapshot wire6.9 KiB6.9 KiB−3 B (−0.0%)7.3 KiB
CodexLive turn WebSocket wire6.8 KiB6.8 KiB+11 B (+0.2%)7.8 KiB
CodexLive turn WebSocket decoded58.7 KiB58.7 KiB0 B (0.0%)66.4 KiB
CodexLive turn messages11110 (0.0%)21
ClaudeTotal thread wire13.5 KiB13.8 KiB+230 B (+1.7%)15.1 KiB
ClaudeThread snapshot wire6.9 KiB6.9 KiB+2 B (+0.0%)7.3 KiB
ClaudeLive turn WebSocket wire6.6 KiB6.8 KiB+228 B (+3.4%)7.8 KiB
ClaudeLive turn WebSocket decoded58.0 KiB59.5 KiB+1.5 KiB (+2.6%)66.4 KiB
ClaudeLive turn messages911+2 (+22.2%)21

Baseline: c908bb2 · PR result: b21cb8f · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 109.5 KiB
  • Claude decoded thread snapshot: 110.2 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

@rynfar
rynfar merged commit 74377fd into pylonSep 1, 2026
19 checks passed
@rynfar
rynfar deleted the upstream/2026-09-01-voice-input branch September 1, 2026 23:28
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXLvouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@rynfar
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

feat(mobile): add offline iPhone voice input - #233

Merged
rynfar merged 5 commits into
pylonfrom
upstream/2026-09-01-voice-input
Sep 1, 2026
Merged

feat(mobile): add offline iPhone voice input#233
rynfar merged 5 commits into
pylonfrom
upstream/2026-09-01-voice-input

Conversation

@rynfar

@rynfarrynfar commented Sep 1, 2026

Copy link
Copy Markdown
Collaborator

Fourth of the mobile batch. Adopted from upstream pingdotgg/t3code#8614 (352710d49). Follows #223, #224, #231.

Mobile gains on-device voice input: hold the mic in the composer, speak, and the transcript commits into the draft. Brings a voice-input controller to client-runtime, native transcription, the dictation UI, and two patched native deps.

The composer restructure, settled

This batch has cost a conflict in every cherry-pick because upstream's #8793 restructure arrives as context in each one. It turns out to be two independent changes, and treating them as one decision was the mistake:

  • ComposerSurface — animate borderRadius on a shared value, absolute glass layer, bounded collapsed radius. Touches nothing in the toolbar. Adopted (ceaae36aa); our file is now structurally identical to upstream there, so these hunks should stop conflicting.
  • The toolbar row — drop ComposerToolbarScroller for a fixed flex row. Declined, permanently. ComposerToolbarScroller is upstream's own component and they still ship it; they stopped using it because their toolbar holds four controls. Ours holds sixteen, because ControlPillMenu — which does not exist upstream at all — carries Refine, session goal, context window, agent count, the input queue, depth, resources, and reload.

Counts unchanged throughout: scroller 3, ControlPillMenu 13, QuickQuestionTrigger 2, ContextWindowIndicator 3.

Three defects I introduced, and fixed

The dictation controls had to be placed by hand rather than ported, since upstream places them inside the restructure we declined. Review found three ways that stranded the user. All 36 ported files were byte-identical to upstream — every defect was in the hand-written JSX, and none was caught by typecheck, lint, or 1042 tests.

The editor was never frozen.NewTaskDraftScreen passes readOnly={voiceInput.freezesEditor}; the thread composer did not. The keyboard stayed live during recording, and one keystroke makes resolveTranscriptCommit see a changed draft and discard the whole transcript as stale — up to five minutes of speech, silently. It also made every native read-only guard this branch adds dead code on that surface.

Send vanished on any dictation error. Gated on isVoiceInputPresented rather than voicePresentation.showsSend. Those look interchangeable but diverge in exactly one phase: error has showsSend: trueand a non-null statusLabel. So a denied permission, "no speech detected", or the stale-draft error above left no send control. Worse on entry, since dispose() no-ops in the error phase and a stale error survives navigating away and back.

Stop was unreachable for the whole dictation window. Placed inside ComposerToolbarScroller, which is the else branch of the dictation ternary — so across preparing → recording → transcribing → error there was no way to stop a running agent.

Also fixed: blocksSubmission guards on both submission entry points (canSend is derived above voiceInput and structurally cannot include it); a 4px spacer overflowing ComposerDictationToolbar's fixed 44px box and clipping the collapsed strip; the dropped pointerEvents="none" on the glass layer.

The showsSend divergence ships with a regression test — it is the one defect that is a pure predicate rather than JSX placement, and the one most likely to be reintroduced by someone reasoning that the two predicates look the same.

Pylon branding

Two leaks caught, both user-visible. app.config.ts carried upstream's "Allow T3 Code to use your microphone for voice input" — iOS renders that verbatim in the permission dialog, and the camera permission two lines below already said Pylon. Plus two T3 Code strings in docs/user/composer.md and docs/internals/voice-input.md.

No speech-recognition permission is needed, verified rather than assumed: @react-native-ai/apple uses SpeechAnalyzer/SpeechTranscriber, Apple's on-device framework, not SFSpeechRecognizer. That is what makes this offline.

Verification

Typecheck clean, lint clean, 1044 tests passing. Native build green with AppleLLM 0.12.0 and ExpoAudio.

This needs a native rebuild, not an OTA — two new native modules plus native Swift changes in t3-composer-editor/ios. runtimeVersion.policy is fingerprint, which covers new native deps and patches/, so a JS update cannot land on a stale binary.

Simulator pass verified the unavailable path only: with voice unavailable the composer renders correctly, full toolbar, no gap where the mic would be. The dictation UI itself has zero simulator coverageAppleTranscription.isAvailable() is a native check whose own error string reads "requires a supported device with iOS 26 or later", so the on-device model is absent on a simulator. Mic, status row, cancel, and the toolbar swap need a device pass before this merges.

Reviewed and integrated with Claude Opus 5 in Claude Code.


View with [code]smithAutofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

NOT READY TO MERGE. Parked as a branch commit so the conflict resolution
is not lost.
Done: all five conflicts resolved, the two new patched native deps
(@react-native-ai/apple, expo-audio) installed, the voice-input feature
directories and client-runtime module landed intact, and ThreadComposer
compiles with the controller wired (composerOwnerKey, useVoiceInputController,
resolveVoiceComposerPresentation, showsCompactDictation).
Also resolved use-composer-command-menu.ts Pylon-first: upstream's file is far
smaller than Pylon's, so every Pylon export (buildComposerCommandItems,
resolveComposerProviderSlashCommands, and the ranking helpers) is kept and only
upstream's composerSelectionAtEnd helper and owner-key ref are added.
Not done, and the reason this is parked: the dictation UI is imported but not
rendered. ComposerDictationToolbar, ComposerDictationPrimaryAction,
ComposerDictationStatus and ComposerDictationCancelAction are all unused, so the
mic never appears. Wiring them means restructuring Pylon's composer rather than
patching it — upstream wraps its toolbar row directly, while Pylon's is
ComposerToolbarRow > ComposerToolbarScroller with thirteen controls. Pylon also
declares canSend roughly 750 lines above where voiceInput can exist, so even
`canSend && !voiceInput.blocksSubmission` needs the declaration order changed.
Shipping it compiling-but-inert would look done and do nothing, so it waits.
…lbar
#8793 changed two independent things in the mobile composer, and treating
them as one decision has cost a conflict in every cherry-pick from
upstream's 2026-08-30 batch since.
The first is ComposerSurface: animate borderRadius on a shared value, put
the glass on an absolute layer, render children in their own animated view,
and bound the collapsed pill radius so the morph interpolates instead of
travelling from 999. Nothing in it touches the toolbar row.
The second is the toolbar row: drop ComposerToolbarScroller for a fixed flex
row. That one genuinely conflicts. ComposerToolbarScroller is upstream's own
component and they still ship it; they stopped using it because their
toolbar holds four controls. Pylon's holds sixteen, because ControlPillMenu
- which does not exist upstream at all - carries Refine, session goal,
context window, agent count, the input queue, depth, resources, and reload.
Those need the scroller.
So take the first, decline the second. ComposerSurface is now structurally
identical to upstream (animatedBorderRadius, AnimatedGlassSurface,
layoutTransition, animatedShapeStyle, and the bounded radius all match), and
the toolbar is untouched: scroller, 13 ControlPillMenus, QuickQuestionTrigger
and ContextWindowIndicator all at their previous counts.
Pylon's shadow wrapper survives with its comment; upstream has no equivalent,
and the radius it carries is now animated alongside the surface.
Completes the #8614 port's UI half. Upstream places the dictation controls
inside its restructured collapsed row and fixed toolbar; Pylon declined that
restructure, so they are placed into Pylon's own structure instead.
The toolbar now shows whenever isToolbarVisible rather than only when
expanded, so dictation stays reachable from the collapsed pill, and it is
wrapped in ComposerDictationToolbar. The cancel action leads the row; while
dictating, ComposerDictationStatus replaces the toolbar scroller rather than
upstream's fixed left group, so Pylon's thirteen ControlPillMenu controls
keep their scroller when not dictating. The mic sits beside send in both the
collapsed row and the toolbar, and send is hidden while dictation owns the
row.
Scroller, ControlPillMenu, QuickQuestionTrigger and ContextWindowIndicator
are all at unchanged counts.
The #8614 port carried upstream's string verbatim: "Allow T3 Code to use
your microphone for voice input." iOS shows that text in the permission
dialog, so it is product copy, not a compatibility identifier. The camera
permission two lines below already reads "Allow Pylon to access your
camera", so this was purely adoption drift.
Verified while checking permissions that no speech-recognition key is
needed: @react-native-ai/apple uses SpeechAnalyzer and SpeechTranscriber,
Apple's on-device Speech framework, rather than SFSpeechRecognizer. Only
NSMicrophoneUsageDescription applies, and it is present.
Adversarial review of the hand-placed dictation UI found three ways the
composer strands the user. All three are in code written by hand rather
than ported, and none was caught by typecheck, lint, or 1042 tests.
The editor was never frozen. NewTaskDraftScreen passes
readOnly={voiceInput.freezesEditor}; the thread composer did not, so the
keyboard stayed live during recording. One keystroke makes
resolveTranscriptCommit see a changed draft and discard the entire
transcript as stale - up to five minutes of speech, silently. It also made
every native read-only guard this branch adds dead code on this surface.
Send was gated on isVoiceInputPresented rather than
voicePresentation.showsSend. Those look equivalent but diverge in exactly
one phase: error shows a status label AND keeps send. Any dictation failure
- denied permission, no speech detected, the stale-draft error above - left
the composer with no send control until the user found the dismiss button.
Worse on entry, since dispose() no-ops in the error phase, so a stale error
survives navigating away and back.
Stop sat inside ComposerToolbarScroller, which is the else branch of the
dictation ternary, so an agent was unstoppable for the whole recording and
transcription window. Moved to the always-rendered right cluster.
Also: guard both submission entry points on blocksSubmission, since canSend
is derived above voiceInput and cannot include it; move the 4px spacer
outside ComposerDictationToolbar's fixed 44px box, where it was overflowing
and clipping the collapsed dictation strip; restore pointerEvents="none" on
the glass layer with a comment matching the new sibling structure; and
rebrand two T3 Code strings in docs.
The showsSend divergence now has a regression test. It is the one defect
here that is a pure predicate rather than JSX placement, and it is the one
most likely to be reintroduced.
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:XXL labels Sep 1, 2026
@github-actions

Copy link
Copy Markdown

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ProviderMetricMain baselineThis PRImpactPR ceiling
CodexTotal thread wire13.7 KiB13.7 KiB+8 B (+0.1%)15.1 KiB
CodexThread snapshot wire6.9 KiB6.9 KiB−3 B (−0.0%)7.3 KiB
CodexLive turn WebSocket wire6.8 KiB6.8 KiB+11 B (+0.2%)7.8 KiB
CodexLive turn WebSocket decoded58.7 KiB58.7 KiB0 B (0.0%)66.4 KiB
CodexLive turn messages11110 (0.0%)21
ClaudeTotal thread wire13.5 KiB13.8 KiB+230 B (+1.7%)15.1 KiB
ClaudeThread snapshot wire6.9 KiB6.9 KiB+2 B (+0.0%)7.3 KiB
ClaudeLive turn WebSocket wire6.6 KiB6.8 KiB+228 B (+3.4%)7.8 KiB
ClaudeLive turn WebSocket decoded58.0 KiB59.5 KiB+1.5 KiB (+2.6%)66.4 KiB
ClaudeLive turn messages911+2 (+22.2%)21

Baseline: c908bb2 · PR result: b21cb8f · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 109.5 KiB
  • Claude decoded thread snapshot: 110.2 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

@rynfar
rynfar merged commit 74377fd into pylonSep 1, 2026
19 checks passed
@rynfar
rynfar deleted the upstream/2026-09-01-voice-input branch September 1, 2026 23:28
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXLvouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@rynfar
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(mobile): add offline iPhone voice input - #233

Merged
rynfar merged 5 commits into
pylonfrom
upstream/2026-09-01-voice-input
Sep 1, 2026
Merged

feat(mobile): add offline iPhone voice input#233
rynfar merged 5 commits into
pylonfrom
upstream/2026-09-01-voice-input

Conversation

@rynfar

@rynfarrynfar commented Sep 1, 2026

Copy link
Copy Markdown
Collaborator

Fourth of the mobile batch. Adopted from upstream pingdotgg/t3code#8614 (352710d49). Follows #223, #224, #231.

Mobile gains on-device voice input: hold the mic in the composer, speak, and the transcript commits into the draft. Brings a voice-input controller to client-runtime, native transcription, the dictation UI, and two patched native deps.

The composer restructure, settled

This batch has cost a conflict in every cherry-pick because upstream's #8793 restructure arrives as context in each one. It turns out to be two independent changes, and treating them as one decision was the mistake:

  • ComposerSurface — animate borderRadius on a shared value, absolute glass layer, bounded collapsed radius. Touches nothing in the toolbar. Adopted (ceaae36aa); our file is now structurally identical to upstream there, so these hunks should stop conflicting.
  • The toolbar row — drop ComposerToolbarScroller for a fixed flex row. Declined, permanently. ComposerToolbarScroller is upstream's own component and they still ship it; they stopped using it because their toolbar holds four controls. Ours holds sixteen, because ControlPillMenu — which does not exist upstream at all — carries Refine, session goal, context window, agent count, the input queue, depth, resources, and reload.

Counts unchanged throughout: scroller 3, ControlPillMenu 13, QuickQuestionTrigger 2, ContextWindowIndicator 3.

Three defects I introduced, and fixed

The dictation controls had to be placed by hand rather than ported, since upstream places them inside the restructure we declined. Review found three ways that stranded the user. All 36 ported files were byte-identical to upstream — every defect was in the hand-written JSX, and none was caught by typecheck, lint, or 1042 tests.

The editor was never frozen.NewTaskDraftScreen passes readOnly={voiceInput.freezesEditor}; the thread composer did not. The keyboard stayed live during recording, and one keystroke makes resolveTranscriptCommit see a changed draft and discard the whole transcript as stale — up to five minutes of speech, silently. It also made every native read-only guard this branch adds dead code on that surface.

Send vanished on any dictation error. Gated on isVoiceInputPresented rather than voicePresentation.showsSend. Those look interchangeable but diverge in exactly one phase: error has showsSend: trueand a non-null statusLabel. So a denied permission, "no speech detected", or the stale-draft error above left no send control. Worse on entry, since dispose() no-ops in the error phase and a stale error survives navigating away and back.

Stop was unreachable for the whole dictation window. Placed inside ComposerToolbarScroller, which is the else branch of the dictation ternary — so across preparing → recording → transcribing → error there was no way to stop a running agent.

Also fixed: blocksSubmission guards on both submission entry points (canSend is derived above voiceInput and structurally cannot include it); a 4px spacer overflowing ComposerDictationToolbar's fixed 44px box and clipping the collapsed strip; the dropped pointerEvents="none" on the glass layer.

The showsSend divergence ships with a regression test — it is the one defect that is a pure predicate rather than JSX placement, and the one most likely to be reintroduced by someone reasoning that the two predicates look the same.

Pylon branding

Two leaks caught, both user-visible. app.config.ts carried upstream's "Allow T3 Code to use your microphone for voice input" — iOS renders that verbatim in the permission dialog, and the camera permission two lines below already said Pylon. Plus two T3 Code strings in docs/user/composer.md and docs/internals/voice-input.md.

No speech-recognition permission is needed, verified rather than assumed: @react-native-ai/apple uses SpeechAnalyzer/SpeechTranscriber, Apple's on-device framework, not SFSpeechRecognizer. That is what makes this offline.

Verification

Typecheck clean, lint clean, 1044 tests passing. Native build green with AppleLLM 0.12.0 and ExpoAudio.

This needs a native rebuild, not an OTA — two new native modules plus native Swift changes in t3-composer-editor/ios. runtimeVersion.policy is fingerprint, which covers new native deps and patches/, so a JS update cannot land on a stale binary.

Simulator pass verified the unavailable path only: with voice unavailable the composer renders correctly, full toolbar, no gap where the mic would be. The dictation UI itself has zero simulator coverageAppleTranscription.isAvailable() is a native check whose own error string reads "requires a supported device with iOS 26 or later", so the on-device model is absent on a simulator. Mic, status row, cancel, and the toolbar swap need a device pass before this merges.

Reviewed and integrated with Claude Opus 5 in Claude Code.


View with [code]smithAutofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

NOT READY TO MERGE. Parked as a branch commit so the conflict resolution
is not lost.
Done: all five conflicts resolved, the two new patched native deps
(@react-native-ai/apple, expo-audio) installed, the voice-input feature
directories and client-runtime module landed intact, and ThreadComposer
compiles with the controller wired (composerOwnerKey, useVoiceInputController,
resolveVoiceComposerPresentation, showsCompactDictation).
Also resolved use-composer-command-menu.ts Pylon-first: upstream's file is far
smaller than Pylon's, so every Pylon export (buildComposerCommandItems,
resolveComposerProviderSlashCommands, and the ranking helpers) is kept and only
upstream's composerSelectionAtEnd helper and owner-key ref are added.
Not done, and the reason this is parked: the dictation UI is imported but not
rendered. ComposerDictationToolbar, ComposerDictationPrimaryAction,
ComposerDictationStatus and ComposerDictationCancelAction are all unused, so the
mic never appears. Wiring them means restructuring Pylon's composer rather than
patching it — upstream wraps its toolbar row directly, while Pylon's is
ComposerToolbarRow > ComposerToolbarScroller with thirteen controls. Pylon also
declares canSend roughly 750 lines above where voiceInput can exist, so even
`canSend && !voiceInput.blocksSubmission` needs the declaration order changed.
Shipping it compiling-but-inert would look done and do nothing, so it waits.
…lbar
#8793 changed two independent things in the mobile composer, and treating
them as one decision has cost a conflict in every cherry-pick from
upstream's 2026-08-30 batch since.
The first is ComposerSurface: animate borderRadius on a shared value, put
the glass on an absolute layer, render children in their own animated view,
and bound the collapsed pill radius so the morph interpolates instead of
travelling from 999. Nothing in it touches the toolbar row.
The second is the toolbar row: drop ComposerToolbarScroller for a fixed flex
row. That one genuinely conflicts. ComposerToolbarScroller is upstream's own
component and they still ship it; they stopped using it because their
toolbar holds four controls. Pylon's holds sixteen, because ControlPillMenu
- which does not exist upstream at all - carries Refine, session goal,
context window, agent count, the input queue, depth, resources, and reload.
Those need the scroller.
So take the first, decline the second. ComposerSurface is now structurally
identical to upstream (animatedBorderRadius, AnimatedGlassSurface,
layoutTransition, animatedShapeStyle, and the bounded radius all match), and
the toolbar is untouched: scroller, 13 ControlPillMenus, QuickQuestionTrigger
and ContextWindowIndicator all at their previous counts.
Pylon's shadow wrapper survives with its comment; upstream has no equivalent,
and the radius it carries is now animated alongside the surface.
Completes the #8614 port's UI half. Upstream places the dictation controls
inside its restructured collapsed row and fixed toolbar; Pylon declined that
restructure, so they are placed into Pylon's own structure instead.
The toolbar now shows whenever isToolbarVisible rather than only when
expanded, so dictation stays reachable from the collapsed pill, and it is
wrapped in ComposerDictationToolbar. The cancel action leads the row; while
dictating, ComposerDictationStatus replaces the toolbar scroller rather than
upstream's fixed left group, so Pylon's thirteen ControlPillMenu controls
keep their scroller when not dictating. The mic sits beside send in both the
collapsed row and the toolbar, and send is hidden while dictation owns the
row.
Scroller, ControlPillMenu, QuickQuestionTrigger and ContextWindowIndicator
are all at unchanged counts.
The #8614 port carried upstream's string verbatim: "Allow T3 Code to use
your microphone for voice input." iOS shows that text in the permission
dialog, so it is product copy, not a compatibility identifier. The camera
permission two lines below already reads "Allow Pylon to access your
camera", so this was purely adoption drift.
Verified while checking permissions that no speech-recognition key is
needed: @react-native-ai/apple uses SpeechAnalyzer and SpeechTranscriber,
Apple's on-device Speech framework, rather than SFSpeechRecognizer. Only
NSMicrophoneUsageDescription applies, and it is present.
Adversarial review of the hand-placed dictation UI found three ways the
composer strands the user. All three are in code written by hand rather
than ported, and none was caught by typecheck, lint, or 1042 tests.
The editor was never frozen. NewTaskDraftScreen passes
readOnly={voiceInput.freezesEditor}; the thread composer did not, so the
keyboard stayed live during recording. One keystroke makes
resolveTranscriptCommit see a changed draft and discard the entire
transcript as stale - up to five minutes of speech, silently. It also made
every native read-only guard this branch adds dead code on this surface.
Send was gated on isVoiceInputPresented rather than
voicePresentation.showsSend. Those look equivalent but diverge in exactly
one phase: error shows a status label AND keeps send. Any dictation failure
- denied permission, no speech detected, the stale-draft error above - left
the composer with no send control until the user found the dismiss button.
Worse on entry, since dispose() no-ops in the error phase, so a stale error
survives navigating away and back.
Stop sat inside ComposerToolbarScroller, which is the else branch of the
dictation ternary, so an agent was unstoppable for the whole recording and
transcription window. Moved to the always-rendered right cluster.
Also: guard both submission entry points on blocksSubmission, since canSend
is derived above voiceInput and cannot include it; move the 4px spacer
outside ComposerDictationToolbar's fixed 44px box, where it was overflowing
and clipping the collapsed dictation strip; restore pointerEvents="none" on
the glass layer with a comment matching the new sibling structure; and
rebrand two T3 Code strings in docs.
The showsSend divergence now has a regression test. It is the one defect
here that is a pure predicate rather than JSX placement, and it is the one
most likely to be reintroduced.
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:XXL labels Sep 1, 2026
@github-actions

Copy link
Copy Markdown

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ProviderMetricMain baselineThis PRImpactPR ceiling
CodexTotal thread wire13.7 KiB13.7 KiB+8 B (+0.1%)15.1 KiB
CodexThread snapshot wire6.9 KiB6.9 KiB−3 B (−0.0%)7.3 KiB
CodexLive turn WebSocket wire6.8 KiB6.8 KiB+11 B (+0.2%)7.8 KiB
CodexLive turn WebSocket decoded58.7 KiB58.7 KiB0 B (0.0%)66.4 KiB
CodexLive turn messages11110 (0.0%)21
ClaudeTotal thread wire13.5 KiB13.8 KiB+230 B (+1.7%)15.1 KiB
ClaudeThread snapshot wire6.9 KiB6.9 KiB+2 B (+0.0%)7.3 KiB
ClaudeLive turn WebSocket wire6.6 KiB6.8 KiB+228 B (+3.4%)7.8 KiB
ClaudeLive turn WebSocket decoded58.0 KiB59.5 KiB+1.5 KiB (+2.6%)66.4 KiB
ClaudeLive turn messages911+2 (+22.2%)21

Baseline: c908bb2 · PR result: b21cb8f · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 109.5 KiB
  • Claude decoded thread snapshot: 110.2 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

@rynfar
rynfar merged commit 74377fd into pylonSep 1, 2026
19 checks passed
@rynfar
rynfar deleted the upstream/2026-09-01-voice-input branch September 1, 2026 23:28
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXLvouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@rynfar
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(mobile): add offline iPhone voice input - #233

Merged
rynfar merged 5 commits into
pylonfrom
upstream/2026-09-01-voice-input
Sep 1, 2026
Merged

feat(mobile): add offline iPhone voice input#233
rynfar merged 5 commits into
pylonfrom
upstream/2026-09-01-voice-input

Conversation

@rynfar

@rynfarrynfar commented Sep 1, 2026

Copy link
Copy Markdown
Collaborator

Fourth of the mobile batch. Adopted from upstream pingdotgg/t3code#8614 (352710d49). Follows #223, #224, #231.

Mobile gains on-device voice input: hold the mic in the composer, speak, and the transcript commits into the draft. Brings a voice-input controller to client-runtime, native transcription, the dictation UI, and two patched native deps.

The composer restructure, settled

This batch has cost a conflict in every cherry-pick because upstream's #8793 restructure arrives as context in each one. It turns out to be two independent changes, and treating them as one decision was the mistake:

  • ComposerSurface — animate borderRadius on a shared value, absolute glass layer, bounded collapsed radius. Touches nothing in the toolbar. Adopted (ceaae36aa); our file is now structurally identical to upstream there, so these hunks should stop conflicting.
  • The toolbar row — drop ComposerToolbarScroller for a fixed flex row. Declined, permanently. ComposerToolbarScroller is upstream's own component and they still ship it; they stopped using it because their toolbar holds four controls. Ours holds sixteen, because ControlPillMenu — which does not exist upstream at all — carries Refine, session goal, context window, agent count, the input queue, depth, resources, and reload.

Counts unchanged throughout: scroller 3, ControlPillMenu 13, QuickQuestionTrigger 2, ContextWindowIndicator 3.

Three defects I introduced, and fixed

The dictation controls had to be placed by hand rather than ported, since upstream places them inside the restructure we declined. Review found three ways that stranded the user. All 36 ported files were byte-identical to upstream — every defect was in the hand-written JSX, and none was caught by typecheck, lint, or 1042 tests.

The editor was never frozen.NewTaskDraftScreen passes readOnly={voiceInput.freezesEditor}; the thread composer did not. The keyboard stayed live during recording, and one keystroke makes resolveTranscriptCommit see a changed draft and discard the whole transcript as stale — up to five minutes of speech, silently. It also made every native read-only guard this branch adds dead code on that surface.

Send vanished on any dictation error. Gated on isVoiceInputPresented rather than voicePresentation.showsSend. Those look interchangeable but diverge in exactly one phase: error has showsSend: trueand a non-null statusLabel. So a denied permission, "no speech detected", or the stale-draft error above left no send control. Worse on entry, since dispose() no-ops in the error phase and a stale error survives navigating away and back.

Stop was unreachable for the whole dictation window. Placed inside ComposerToolbarScroller, which is the else branch of the dictation ternary — so across preparing → recording → transcribing → error there was no way to stop a running agent.

Also fixed: blocksSubmission guards on both submission entry points (canSend is derived above voiceInput and structurally cannot include it); a 4px spacer overflowing ComposerDictationToolbar's fixed 44px box and clipping the collapsed strip; the dropped pointerEvents="none" on the glass layer.

The showsSend divergence ships with a regression test — it is the one defect that is a pure predicate rather than JSX placement, and the one most likely to be reintroduced by someone reasoning that the two predicates look the same.

Pylon branding

Two leaks caught, both user-visible. app.config.ts carried upstream's "Allow T3 Code to use your microphone for voice input" — iOS renders that verbatim in the permission dialog, and the camera permission two lines below already said Pylon. Plus two T3 Code strings in docs/user/composer.md and docs/internals/voice-input.md.

No speech-recognition permission is needed, verified rather than assumed: @react-native-ai/apple uses SpeechAnalyzer/SpeechTranscriber, Apple's on-device framework, not SFSpeechRecognizer. That is what makes this offline.

Verification

Typecheck clean, lint clean, 1044 tests passing. Native build green with AppleLLM 0.12.0 and ExpoAudio.

This needs a native rebuild, not an OTA — two new native modules plus native Swift changes in t3-composer-editor/ios. runtimeVersion.policy is fingerprint, which covers new native deps and patches/, so a JS update cannot land on a stale binary.

Simulator pass verified the unavailable path only: with voice unavailable the composer renders correctly, full toolbar, no gap where the mic would be. The dictation UI itself has zero simulator coverageAppleTranscription.isAvailable() is a native check whose own error string reads "requires a supported device with iOS 26 or later", so the on-device model is absent on a simulator. Mic, status row, cancel, and the toolbar swap need a device pass before this merges.

Reviewed and integrated with Claude Opus 5 in Claude Code.


View with [code]smithAutofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

NOT READY TO MERGE. Parked as a branch commit so the conflict resolution
is not lost.
Done: all five conflicts resolved, the two new patched native deps
(@react-native-ai/apple, expo-audio) installed, the voice-input feature
directories and client-runtime module landed intact, and ThreadComposer
compiles with the controller wired (composerOwnerKey, useVoiceInputController,
resolveVoiceComposerPresentation, showsCompactDictation).
Also resolved use-composer-command-menu.ts Pylon-first: upstream's file is far
smaller than Pylon's, so every Pylon export (buildComposerCommandItems,
resolveComposerProviderSlashCommands, and the ranking helpers) is kept and only
upstream's composerSelectionAtEnd helper and owner-key ref are added.
Not done, and the reason this is parked: the dictation UI is imported but not
rendered. ComposerDictationToolbar, ComposerDictationPrimaryAction,
ComposerDictationStatus and ComposerDictationCancelAction are all unused, so the
mic never appears. Wiring them means restructuring Pylon's composer rather than
patching it — upstream wraps its toolbar row directly, while Pylon's is
ComposerToolbarRow > ComposerToolbarScroller with thirteen controls. Pylon also
declares canSend roughly 750 lines above where voiceInput can exist, so even
`canSend && !voiceInput.blocksSubmission` needs the declaration order changed.
Shipping it compiling-but-inert would look done and do nothing, so it waits.
…lbar
#8793 changed two independent things in the mobile composer, and treating
them as one decision has cost a conflict in every cherry-pick from
upstream's 2026-08-30 batch since.
The first is ComposerSurface: animate borderRadius on a shared value, put
the glass on an absolute layer, render children in their own animated view,
and bound the collapsed pill radius so the morph interpolates instead of
travelling from 999. Nothing in it touches the toolbar row.
The second is the toolbar row: drop ComposerToolbarScroller for a fixed flex
row. That one genuinely conflicts. ComposerToolbarScroller is upstream's own
component and they still ship it; they stopped using it because their
toolbar holds four controls. Pylon's holds sixteen, because ControlPillMenu
- which does not exist upstream at all - carries Refine, session goal,
context window, agent count, the input queue, depth, resources, and reload.
Those need the scroller.
So take the first, decline the second. ComposerSurface is now structurally
identical to upstream (animatedBorderRadius, AnimatedGlassSurface,
layoutTransition, animatedShapeStyle, and the bounded radius all match), and
the toolbar is untouched: scroller, 13 ControlPillMenus, QuickQuestionTrigger
and ContextWindowIndicator all at their previous counts.
Pylon's shadow wrapper survives with its comment; upstream has no equivalent,
and the radius it carries is now animated alongside the surface.
Completes the #8614 port's UI half. Upstream places the dictation controls
inside its restructured collapsed row and fixed toolbar; Pylon declined that
restructure, so they are placed into Pylon's own structure instead.
The toolbar now shows whenever isToolbarVisible rather than only when
expanded, so dictation stays reachable from the collapsed pill, and it is
wrapped in ComposerDictationToolbar. The cancel action leads the row; while
dictating, ComposerDictationStatus replaces the toolbar scroller rather than
upstream's fixed left group, so Pylon's thirteen ControlPillMenu controls
keep their scroller when not dictating. The mic sits beside send in both the
collapsed row and the toolbar, and send is hidden while dictation owns the
row.
Scroller, ControlPillMenu, QuickQuestionTrigger and ContextWindowIndicator
are all at unchanged counts.
The #8614 port carried upstream's string verbatim: "Allow T3 Code to use
your microphone for voice input." iOS shows that text in the permission
dialog, so it is product copy, not a compatibility identifier. The camera
permission two lines below already reads "Allow Pylon to access your
camera", so this was purely adoption drift.
Verified while checking permissions that no speech-recognition key is
needed: @react-native-ai/apple uses SpeechAnalyzer and SpeechTranscriber,
Apple's on-device Speech framework, rather than SFSpeechRecognizer. Only
NSMicrophoneUsageDescription applies, and it is present.
Adversarial review of the hand-placed dictation UI found three ways the
composer strands the user. All three are in code written by hand rather
than ported, and none was caught by typecheck, lint, or 1042 tests.
The editor was never frozen. NewTaskDraftScreen passes
readOnly={voiceInput.freezesEditor}; the thread composer did not, so the
keyboard stayed live during recording. One keystroke makes
resolveTranscriptCommit see a changed draft and discard the entire
transcript as stale - up to five minutes of speech, silently. It also made
every native read-only guard this branch adds dead code on this surface.
Send was gated on isVoiceInputPresented rather than
voicePresentation.showsSend. Those look equivalent but diverge in exactly
one phase: error shows a status label AND keeps send. Any dictation failure
- denied permission, no speech detected, the stale-draft error above - left
the composer with no send control until the user found the dismiss button.
Worse on entry, since dispose() no-ops in the error phase, so a stale error
survives navigating away and back.
Stop sat inside ComposerToolbarScroller, which is the else branch of the
dictation ternary, so an agent was unstoppable for the whole recording and
transcription window. Moved to the always-rendered right cluster.
Also: guard both submission entry points on blocksSubmission, since canSend
is derived above voiceInput and cannot include it; move the 4px spacer
outside ComposerDictationToolbar's fixed 44px box, where it was overflowing
and clipping the collapsed dictation strip; restore pointerEvents="none" on
the glass layer with a comment matching the new sibling structure; and
rebrand two T3 Code strings in docs.
The showsSend divergence now has a regression test. It is the one defect
here that is a pure predicate rather than JSX placement, and it is the one
most likely to be reintroduced.
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:XXL labels Sep 1, 2026
@github-actions

Copy link
Copy Markdown

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ProviderMetricMain baselineThis PRImpactPR ceiling
CodexTotal thread wire13.7 KiB13.7 KiB+8 B (+0.1%)15.1 KiB
CodexThread snapshot wire6.9 KiB6.9 KiB−3 B (−0.0%)7.3 KiB
CodexLive turn WebSocket wire6.8 KiB6.8 KiB+11 B (+0.2%)7.8 KiB
CodexLive turn WebSocket decoded58.7 KiB58.7 KiB0 B (0.0%)66.4 KiB
CodexLive turn messages11110 (0.0%)21
ClaudeTotal thread wire13.5 KiB13.8 KiB+230 B (+1.7%)15.1 KiB
ClaudeThread snapshot wire6.9 KiB6.9 KiB+2 B (+0.0%)7.3 KiB
ClaudeLive turn WebSocket wire6.6 KiB6.8 KiB+228 B (+3.4%)7.8 KiB
ClaudeLive turn WebSocket decoded58.0 KiB59.5 KiB+1.5 KiB (+2.6%)66.4 KiB
ClaudeLive turn messages911+2 (+22.2%)21

Baseline: c908bb2 · PR result: b21cb8f · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 109.5 KiB
  • Claude decoded thread snapshot: 110.2 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

@rynfar
rynfar merged commit 74377fd into pylonSep 1, 2026
19 checks passed
@rynfar
rynfar deleted the upstream/2026-09-01-voice-input branch September 1, 2026 23:28
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXLvouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@rynfar
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

feat(mobile): add offline iPhone voice input - #233

Merged
rynfar merged 5 commits into
pylonfrom
upstream/2026-09-01-voice-input
Sep 1, 2026
Merged

feat(mobile): add offline iPhone voice input#233
rynfar merged 5 commits into
pylonfrom
upstream/2026-09-01-voice-input

Conversation

@rynfar

@rynfarrynfar commented Sep 1, 2026

Copy link
Copy Markdown
Collaborator

Fourth of the mobile batch. Adopted from upstream pingdotgg/t3code#8614 (352710d49). Follows #223, #224, #231.

Mobile gains on-device voice input: hold the mic in the composer, speak, and the transcript commits into the draft. Brings a voice-input controller to client-runtime, native transcription, the dictation UI, and two patched native deps.

The composer restructure, settled

This batch has cost a conflict in every cherry-pick because upstream's #8793 restructure arrives as context in each one. It turns out to be two independent changes, and treating them as one decision was the mistake:

  • ComposerSurface — animate borderRadius on a shared value, absolute glass layer, bounded collapsed radius. Touches nothing in the toolbar. Adopted (ceaae36aa); our file is now structurally identical to upstream there, so these hunks should stop conflicting.
  • The toolbar row — drop ComposerToolbarScroller for a fixed flex row. Declined, permanently. ComposerToolbarScroller is upstream's own component and they still ship it; they stopped using it because their toolbar holds four controls. Ours holds sixteen, because ControlPillMenu — which does not exist upstream at all — carries Refine, session goal, context window, agent count, the input queue, depth, resources, and reload.

Counts unchanged throughout: scroller 3, ControlPillMenu 13, QuickQuestionTrigger 2, ContextWindowIndicator 3.

Three defects I introduced, and fixed

The dictation controls had to be placed by hand rather than ported, since upstream places them inside the restructure we declined. Review found three ways that stranded the user. All 36 ported files were byte-identical to upstream — every defect was in the hand-written JSX, and none was caught by typecheck, lint, or 1042 tests.

The editor was never frozen.NewTaskDraftScreen passes readOnly={voiceInput.freezesEditor}; the thread composer did not. The keyboard stayed live during recording, and one keystroke makes resolveTranscriptCommit see a changed draft and discard the whole transcript as stale — up to five minutes of speech, silently. It also made every native read-only guard this branch adds dead code on that surface.

Send vanished on any dictation error. Gated on isVoiceInputPresented rather than voicePresentation.showsSend. Those look interchangeable but diverge in exactly one phase: error has showsSend: trueand a non-null statusLabel. So a denied permission, "no speech detected", or the stale-draft error above left no send control. Worse on entry, since dispose() no-ops in the error phase and a stale error survives navigating away and back.

Stop was unreachable for the whole dictation window. Placed inside ComposerToolbarScroller, which is the else branch of the dictation ternary — so across preparing → recording → transcribing → error there was no way to stop a running agent.

Also fixed: blocksSubmission guards on both submission entry points (canSend is derived above voiceInput and structurally cannot include it); a 4px spacer overflowing ComposerDictationToolbar's fixed 44px box and clipping the collapsed strip; the dropped pointerEvents="none" on the glass layer.

The showsSend divergence ships with a regression test — it is the one defect that is a pure predicate rather than JSX placement, and the one most likely to be reintroduced by someone reasoning that the two predicates look the same.

Pylon branding

Two leaks caught, both user-visible. app.config.ts carried upstream's "Allow T3 Code to use your microphone for voice input" — iOS renders that verbatim in the permission dialog, and the camera permission two lines below already said Pylon. Plus two T3 Code strings in docs/user/composer.md and docs/internals/voice-input.md.

No speech-recognition permission is needed, verified rather than assumed: @react-native-ai/apple uses SpeechAnalyzer/SpeechTranscriber, Apple's on-device framework, not SFSpeechRecognizer. That is what makes this offline.

Verification

Typecheck clean, lint clean, 1044 tests passing. Native build green with AppleLLM 0.12.0 and ExpoAudio.

This needs a native rebuild, not an OTA — two new native modules plus native Swift changes in t3-composer-editor/ios. runtimeVersion.policy is fingerprint, which covers new native deps and patches/, so a JS update cannot land on a stale binary.

Simulator pass verified the unavailable path only: with voice unavailable the composer renders correctly, full toolbar, no gap where the mic would be. The dictation UI itself has zero simulator coverageAppleTranscription.isAvailable() is a native check whose own error string reads "requires a supported device with iOS 26 or later", so the on-device model is absent on a simulator. Mic, status row, cancel, and the toolbar swap need a device pass before this merges.

Reviewed and integrated with Claude Opus 5 in Claude Code.


View with [code]smithAutofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

NOT READY TO MERGE. Parked as a branch commit so the conflict resolution
is not lost.
Done: all five conflicts resolved, the two new patched native deps
(@react-native-ai/apple, expo-audio) installed, the voice-input feature
directories and client-runtime module landed intact, and ThreadComposer
compiles with the controller wired (composerOwnerKey, useVoiceInputController,
resolveVoiceComposerPresentation, showsCompactDictation).
Also resolved use-composer-command-menu.ts Pylon-first: upstream's file is far
smaller than Pylon's, so every Pylon export (buildComposerCommandItems,
resolveComposerProviderSlashCommands, and the ranking helpers) is kept and only
upstream's composerSelectionAtEnd helper and owner-key ref are added.
Not done, and the reason this is parked: the dictation UI is imported but not
rendered. ComposerDictationToolbar, ComposerDictationPrimaryAction,
ComposerDictationStatus and ComposerDictationCancelAction are all unused, so the
mic never appears. Wiring them means restructuring Pylon's composer rather than
patching it — upstream wraps its toolbar row directly, while Pylon's is
ComposerToolbarRow > ComposerToolbarScroller with thirteen controls. Pylon also
declares canSend roughly 750 lines above where voiceInput can exist, so even
`canSend && !voiceInput.blocksSubmission` needs the declaration order changed.
Shipping it compiling-but-inert would look done and do nothing, so it waits.
…lbar
#8793 changed two independent things in the mobile composer, and treating
them as one decision has cost a conflict in every cherry-pick from
upstream's 2026-08-30 batch since.
The first is ComposerSurface: animate borderRadius on a shared value, put
the glass on an absolute layer, render children in their own animated view,
and bound the collapsed pill radius so the morph interpolates instead of
travelling from 999. Nothing in it touches the toolbar row.
The second is the toolbar row: drop ComposerToolbarScroller for a fixed flex
row. That one genuinely conflicts. ComposerToolbarScroller is upstream's own
component and they still ship it; they stopped using it because their
toolbar holds four controls. Pylon's holds sixteen, because ControlPillMenu
- which does not exist upstream at all - carries Refine, session goal,
context window, agent count, the input queue, depth, resources, and reload.
Those need the scroller.
So take the first, decline the second. ComposerSurface is now structurally
identical to upstream (animatedBorderRadius, AnimatedGlassSurface,
layoutTransition, animatedShapeStyle, and the bounded radius all match), and
the toolbar is untouched: scroller, 13 ControlPillMenus, QuickQuestionTrigger
and ContextWindowIndicator all at their previous counts.
Pylon's shadow wrapper survives with its comment; upstream has no equivalent,
and the radius it carries is now animated alongside the surface.
Completes the #8614 port's UI half. Upstream places the dictation controls
inside its restructured collapsed row and fixed toolbar; Pylon declined that
restructure, so they are placed into Pylon's own structure instead.
The toolbar now shows whenever isToolbarVisible rather than only when
expanded, so dictation stays reachable from the collapsed pill, and it is
wrapped in ComposerDictationToolbar. The cancel action leads the row; while
dictating, ComposerDictationStatus replaces the toolbar scroller rather than
upstream's fixed left group, so Pylon's thirteen ControlPillMenu controls
keep their scroller when not dictating. The mic sits beside send in both the
collapsed row and the toolbar, and send is hidden while dictation owns the
row.
Scroller, ControlPillMenu, QuickQuestionTrigger and ContextWindowIndicator
are all at unchanged counts.
The #8614 port carried upstream's string verbatim: "Allow T3 Code to use
your microphone for voice input." iOS shows that text in the permission
dialog, so it is product copy, not a compatibility identifier. The camera
permission two lines below already reads "Allow Pylon to access your
camera", so this was purely adoption drift.
Verified while checking permissions that no speech-recognition key is
needed: @react-native-ai/apple uses SpeechAnalyzer and SpeechTranscriber,
Apple's on-device Speech framework, rather than SFSpeechRecognizer. Only
NSMicrophoneUsageDescription applies, and it is present.
Adversarial review of the hand-placed dictation UI found three ways the
composer strands the user. All three are in code written by hand rather
than ported, and none was caught by typecheck, lint, or 1042 tests.
The editor was never frozen. NewTaskDraftScreen passes
readOnly={voiceInput.freezesEditor}; the thread composer did not, so the
keyboard stayed live during recording. One keystroke makes
resolveTranscriptCommit see a changed draft and discard the entire
transcript as stale - up to five minutes of speech, silently. It also made
every native read-only guard this branch adds dead code on this surface.
Send was gated on isVoiceInputPresented rather than
voicePresentation.showsSend. Those look equivalent but diverge in exactly
one phase: error shows a status label AND keeps send. Any dictation failure
- denied permission, no speech detected, the stale-draft error above - left
the composer with no send control until the user found the dismiss button.
Worse on entry, since dispose() no-ops in the error phase, so a stale error
survives navigating away and back.
Stop sat inside ComposerToolbarScroller, which is the else branch of the
dictation ternary, so an agent was unstoppable for the whole recording and
transcription window. Moved to the always-rendered right cluster.
Also: guard both submission entry points on blocksSubmission, since canSend
is derived above voiceInput and cannot include it; move the 4px spacer
outside ComposerDictationToolbar's fixed 44px box, where it was overflowing
and clipping the collapsed dictation strip; restore pointerEvents="none" on
the glass layer with a comment matching the new sibling structure; and
rebrand two T3 Code strings in docs.
The showsSend divergence now has a regression test. It is the one defect
here that is a pure predicate rather than JSX placement, and it is the one
most likely to be reintroduced.
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:XXL labels Sep 1, 2026
@github-actions

Copy link
Copy Markdown

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ProviderMetricMain baselineThis PRImpactPR ceiling
CodexTotal thread wire13.7 KiB13.7 KiB+8 B (+0.1%)15.1 KiB
CodexThread snapshot wire6.9 KiB6.9 KiB−3 B (−0.0%)7.3 KiB
CodexLive turn WebSocket wire6.8 KiB6.8 KiB+11 B (+0.2%)7.8 KiB
CodexLive turn WebSocket decoded58.7 KiB58.7 KiB0 B (0.0%)66.4 KiB
CodexLive turn messages11110 (0.0%)21
ClaudeTotal thread wire13.5 KiB13.8 KiB+230 B (+1.7%)15.1 KiB
ClaudeThread snapshot wire6.9 KiB6.9 KiB+2 B (+0.0%)7.3 KiB
ClaudeLive turn WebSocket wire6.6 KiB6.8 KiB+228 B (+3.4%)7.8 KiB
ClaudeLive turn WebSocket decoded58.0 KiB59.5 KiB+1.5 KiB (+2.6%)66.4 KiB
ClaudeLive turn messages911+2 (+22.2%)21

Baseline: c908bb2 · PR result: b21cb8f · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 109.5 KiB
  • Claude decoded thread snapshot: 110.2 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

@rynfar
rynfar merged commit 74377fd into pylonSep 1, 2026
19 checks passed
@rynfar
rynfar deleted the upstream/2026-09-01-voice-input branch September 1, 2026 23:28
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXLvouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@rynfar