Uh oh!
There was an error while loading. Please reload this page.
feat(mobile): add offline iPhone voice input - #233
Merged
Conversation
NOT READY TO MERGE. Parked as a branch commit so the conflict resolution is not lost. Done: all five conflicts resolved, the two new patched native deps (@react-native-ai/apple, expo-audio) installed, the voice-input feature directories and client-runtime module landed intact, and ThreadComposer compiles with the controller wired (composerOwnerKey, useVoiceInputController, resolveVoiceComposerPresentation, showsCompactDictation). Also resolved use-composer-command-menu.ts Pylon-first: upstream's file is far smaller than Pylon's, so every Pylon export (buildComposerCommandItems, resolveComposerProviderSlashCommands, and the ranking helpers) is kept and only upstream's composerSelectionAtEnd helper and owner-key ref are added. Not done, and the reason this is parked: the dictation UI is imported but not rendered. ComposerDictationToolbar, ComposerDictationPrimaryAction, ComposerDictationStatus and ComposerDictationCancelAction are all unused, so the mic never appears. Wiring them means restructuring Pylon's composer rather than patching it — upstream wraps its toolbar row directly, while Pylon's is ComposerToolbarRow > ComposerToolbarScroller with thirteen controls. Pylon also declares canSend roughly 750 lines above where voiceInput can exist, so even `canSend && !voiceInput.blocksSubmission` needs the declaration order changed. Shipping it compiling-but-inert would look done and do nothing, so it waits.
…lbar #8793 changed two independent things in the mobile composer, and treating them as one decision has cost a conflict in every cherry-pick from upstream's 2026-08-30 batch since. The first is ComposerSurface: animate borderRadius on a shared value, put the glass on an absolute layer, render children in their own animated view, and bound the collapsed pill radius so the morph interpolates instead of travelling from 999. Nothing in it touches the toolbar row. The second is the toolbar row: drop ComposerToolbarScroller for a fixed flex row. That one genuinely conflicts. ComposerToolbarScroller is upstream's own component and they still ship it; they stopped using it because their toolbar holds four controls. Pylon's holds sixteen, because ControlPillMenu - which does not exist upstream at all - carries Refine, session goal, context window, agent count, the input queue, depth, resources, and reload. Those need the scroller. So take the first, decline the second. ComposerSurface is now structurally identical to upstream (animatedBorderRadius, AnimatedGlassSurface, layoutTransition, animatedShapeStyle, and the bounded radius all match), and the toolbar is untouched: scroller, 13 ControlPillMenus, QuickQuestionTrigger and ContextWindowIndicator all at their previous counts. Pylon's shadow wrapper survives with its comment; upstream has no equivalent, and the radius it carries is now animated alongside the surface.
Completes the #8614 port's UI half. Upstream places the dictation controls inside its restructured collapsed row and fixed toolbar; Pylon declined that restructure, so they are placed into Pylon's own structure instead. The toolbar now shows whenever isToolbarVisible rather than only when expanded, so dictation stays reachable from the collapsed pill, and it is wrapped in ComposerDictationToolbar. The cancel action leads the row; while dictating, ComposerDictationStatus replaces the toolbar scroller rather than upstream's fixed left group, so Pylon's thirteen ControlPillMenu controls keep their scroller when not dictating. The mic sits beside send in both the collapsed row and the toolbar, and send is hidden while dictation owns the row. Scroller, ControlPillMenu, QuickQuestionTrigger and ContextWindowIndicator are all at unchanged counts.
The #8614 port carried upstream's string verbatim: "Allow T3 Code to use your microphone for voice input." iOS shows that text in the permission dialog, so it is product copy, not a compatibility identifier. The camera permission two lines below already reads "Allow Pylon to access your camera", so this was purely adoption drift. Verified while checking permissions that no speech-recognition key is needed: @react-native-ai/apple uses SpeechAnalyzer and SpeechTranscriber, Apple's on-device Speech framework, rather than SFSpeechRecognizer. Only NSMicrophoneUsageDescription applies, and it is present.
Adversarial review of the hand-placed dictation UI found three ways the
composer strands the user. All three are in code written by hand rather
than ported, and none was caught by typecheck, lint, or 1042 tests.
The editor was never frozen. NewTaskDraftScreen passes
readOnly={voiceInput.freezesEditor}; the thread composer did not, so the
keyboard stayed live during recording. One keystroke makes
resolveTranscriptCommit see a changed draft and discard the entire
transcript as stale - up to five minutes of speech, silently. It also made
every native read-only guard this branch adds dead code on this surface.
Send was gated on isVoiceInputPresented rather than
voicePresentation.showsSend. Those look equivalent but diverge in exactly
one phase: error shows a status label AND keeps send. Any dictation failure
- denied permission, no speech detected, the stale-draft error above - left
the composer with no send control until the user found the dismiss button.
Worse on entry, since dispose() no-ops in the error phase, so a stale error
survives navigating away and back.
Stop sat inside ComposerToolbarScroller, which is the else branch of the
dictation ternary, so an agent was unstoppable for the whole recording and
transcription window. Moved to the always-rendered right cluster.
Also: guard both submission entry points on blocksSubmission, since canSend
is derived above voiceInput and cannot include it; move the 4px spacer
outside ComposerDictationToolbar's fixed 44px box, where it was overflowing
and clipping the collapsed dictation strip; restore pointerEvents="none" on
the glass layer with a comment matching the new sibling structure; and
rebrand two T3 Code strings in docs.
The showsSend divergence now has a regression test. It is the one defect
here that is a pure predicate rather than JSX placement, and it is the one
most likely to be reintroduced.Thread transfer impact✅ Thread transfer remains within every enforced ceiling.
Baseline: Scenario and decoded snapshot size10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.
Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for freeto join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fourth of the mobile batch. Adopted from upstream
pingdotgg/t3code#8614(352710d49). Follows #223, #224, #231.Mobile gains on-device voice input: hold the mic in the composer, speak, and the transcript commits into the draft. Brings a voice-input controller to
client-runtime, native transcription, the dictation UI, and two patched native deps.The composer restructure, settled
This batch has cost a conflict in every cherry-pick because upstream's #8793 restructure arrives as context in each one. It turns out to be two independent changes, and treating them as one decision was the mistake:
ComposerSurface— animateborderRadiuson a shared value, absolute glass layer, bounded collapsed radius. Touches nothing in the toolbar. Adopted (ceaae36aa); our file is now structurally identical to upstream there, so these hunks should stop conflicting.ComposerToolbarScrollerfor a fixed flex row. Declined, permanently.ComposerToolbarScrolleris upstream's own component and they still ship it; they stopped using it because their toolbar holds four controls. Ours holds sixteen, becauseControlPillMenu— which does not exist upstream at all — carries Refine, session goal, context window, agent count, the input queue, depth, resources, and reload.Counts unchanged throughout: scroller 3,
ControlPillMenu13,QuickQuestionTrigger2,ContextWindowIndicator3.Three defects I introduced, and fixed
The dictation controls had to be placed by hand rather than ported, since upstream places them inside the restructure we declined. Review found three ways that stranded the user. All 36 ported files were byte-identical to upstream — every defect was in the hand-written JSX, and none was caught by typecheck, lint, or 1042 tests.
The editor was never frozen.
NewTaskDraftScreenpassesreadOnly={voiceInput.freezesEditor}; the thread composer did not. The keyboard stayed live during recording, and one keystroke makesresolveTranscriptCommitsee a changed draft and discard the whole transcript as stale — up to five minutes of speech, silently. It also made every native read-only guard this branch adds dead code on that surface.Send vanished on any dictation error. Gated on
isVoiceInputPresentedrather thanvoicePresentation.showsSend. Those look interchangeable but diverge in exactly one phase:errorhasshowsSend: trueand a non-nullstatusLabel. So a denied permission, "no speech detected", or the stale-draft error above left no send control. Worse on entry, sincedispose()no-ops in the error phase and a stale error survives navigating away and back.Stop was unreachable for the whole dictation window. Placed inside
ComposerToolbarScroller, which is theelsebranch of the dictation ternary — so acrosspreparing → recording → transcribing → errorthere was no way to stop a running agent.Also fixed:
blocksSubmissionguards on both submission entry points (canSendis derived abovevoiceInputand structurally cannot include it); a 4px spacer overflowingComposerDictationToolbar's fixed 44px box and clipping the collapsed strip; the droppedpointerEvents="none"on the glass layer.The
showsSenddivergence ships with a regression test — it is the one defect that is a pure predicate rather than JSX placement, and the one most likely to be reintroduced by someone reasoning that the two predicates look the same.Pylon branding
Two leaks caught, both user-visible.
app.config.tscarried upstream's "Allow T3 Code to use your microphone for voice input" — iOS renders that verbatim in the permission dialog, and the camera permission two lines below already said Pylon. Plus twoT3 Codestrings indocs/user/composer.mdanddocs/internals/voice-input.md.No speech-recognition permission is needed, verified rather than assumed:
@react-native-ai/appleusesSpeechAnalyzer/SpeechTranscriber, Apple's on-device framework, notSFSpeechRecognizer. That is what makes this offline.Verification
Typecheck clean, lint clean, 1044 tests passing. Native build green with
AppleLLM 0.12.0andExpoAudio.This needs a native rebuild, not an OTA — two new native modules plus native Swift changes in
t3-composer-editor/ios.runtimeVersion.policyisfingerprint, which covers new native deps andpatches/, so a JS update cannot land on a stale binary.Simulator pass verified the unavailable path only: with voice unavailable the composer renders correctly, full toolbar, no gap where the mic would be. The dictation UI itself has zero simulator coverage —
AppleTranscription.isAvailable()is a native check whose own error string reads "requires a supported device with iOS 26 or later", so the on-device model is absent on a simulator. Mic, status row, cancel, and the toolbar swap need a device pass before this merges.Reviewed and integrated with Claude Opus 5 in Claude Code.
Need help on this PR? Tag
@codesmith-botwith what you need. Autofix is disabled.