feat(voice): live dictation via macOS 26 SpeechAnalyzer - #233

Merged
Tryanks merged 7 commits into
mainfrom
feat/voice-input
Aug 20, 2026
Merged

feat(voice): live dictation via macOS 26 SpeechAnalyzer#233
Tryanks merged 7 commits into
mainfrom
feat/voice-input

Conversation

@Tryanks

Copy link
Copy Markdown
Owner

What

Voice input for the composer: a mic toggle button that streams on-device speech transcription into the message editor as you speak. macOS 26+ only — on every other platform/version the button does not exist (compile-time cfg + runtime is_supported()).

Research behind the approach: docs/voice-input-research.md. Plan: docs/plans/voice-input.md.

How

  • crates/voice (new)tcode-voice: a Swift shim (swift/shim.swift, C ABI via @_cdecl) around SpeechAnalyzer/SpeechTranscriber, compiled and linked by build.rs only when targeting macOS with an SDK ≥ 26 (TCODE_VOICE_FORCE_STUB=1 forces the stub). Swift owns the whole audio path: AVAudioEngine mic tap → AVAudioConverterAnalyzerInput stream; locale assets are installed on demand via AssetInventory. Events (Ready | Volatile | Final | Error | Ended) cross the FFI on arbitrary threads with a serialized, single-terminal contract.
  • Composer UI — mic button next to the send button (both row layouts). Click to start (spinner while assets/mic warm up), speech streams into the editor at an anchored insertion point: each volatile hypothesis rewrites the volatile tail in place, finals commit. Click again / Esc stops and keeps the text; submit, thread switch, or typing during dictation also stops it. All offsets are UTF-8 bytes (zh-CN safe).
  • Locale — follows the app locale: zh → zh_CN, otherwise en_US.
  • PackagingNSMicrophoneUsageDescription + NSSpeechRecognitionUsageDescription added to the release Info.plist.

Test plan

  • cargo run -p tcode-voice --example file_dictation -- <aiff> zh_CN — 21 volatile events, correct final, clean exit (fixture from say -v Tingting)
  • cargo check --workspace, cargo clippy — clean (only the pre-existing block v0.1.6 future-incompat note)
  • cargo test -p tcode-ui — 254 passed, incl. locale key-set test and the new transcript-edit unit test
  • TCODE_VOICE_FORCE_STUB=1 build — stub path compiles, is_supported() false; non-macOS cfg path checked warning-free under -D warnings
  • Shim symbols present in the final tcode binary; is_supported() true at runtime on macOS 26.5
  • Live mic dictation in the running app (needs a human to approve the mic TCC prompt on first use)

Out of scope (per plan)

Settings section, cloud/BYOK backends, other platforms (Windows/Linux need a local ASR engine — see research doc), and conversational voice mode.

🤖 Generated with Claude Code

Tryanksand others added 7 commits August 21, 2026 02:40
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Adds a mic toggle to the composer that streams on-device speech
transcription into the editor. macOS 26+ only: the Swift shim
(SpeechAnalyzer/SpeechTranscriber, AVAudioEngine capture) is compiled by
build.rs when the SDK allows and the button does not exist elsewhere.
Volatile hypotheses rewrite an anchored range in place; final results
commit. Esc, submit, thread switch or typing stops the session.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The mic path stamped every AnalyzerInput with an explicit bufferStartTime
and stopped via finalizeAndFinish(through: <cumulative time>). With its
own timestamps the analyzer never reported results, and the targeted
finalize waited on a timeline position that never arrived, so stop() hung
and Ended was never delivered. Feed untimestamped buffers (the analyzer
tracks the live timeline itself) and finalize through end of input, as
validated against a standalone probe. Also treat an empty buffer from the
priming rate converter as skippable rather than an error, and keep a
mic_dictation example for live testing.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The shim reports a perceptual input level (RMS over a -50 dBFS window) at
the audio callback rate as a new cosmetic Level event; the mic button
renders it as a danger-tinted glow with fast-attack/slow-decay smoothing,
so the user can see their speech is being picked up.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ze flush
Two compounding reasons dictation text never appeared:
- SpeechTranscriber with plain volatileResults buffers every result until
finalization — nothing streams while speaking. DictationTranscriber's
progressiveLongDictation preset (the dictation-purpose module; 53
locales incl. zh_CN/en_US) streams hypotheses every few hundred ms,
verified against real audio driven through the exact mic pipeline.
- The UI dropped the session and event pump immediately on stop, so even
the finalization flush was discarded. Stop is now graceful (pump lives
until the terminal event lands the flushed text); submit, destination
switches and user edits abort instead, where a late flush would land in
the wrong place.
Also keeps a stderr diagnostic for mic authorization/terminal events.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…heck
crates/voice/build.rs invokes xcrun/swiftc at build time; the
Command::new boundary exists for runtime process spawning, which build
scripts cannot route through the approved helpers.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…tub builds
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@Tryanks
Tryanks merged commit e8a5f1c into mainAug 20, 2026
3 checks passed
@Tryanks
Tryanks deleted the feat/voice-input branch August 20, 2026 19:33
@TryanksTryanks mentioned this pull request Aug 22, 2026
12 tasks
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Tryanks
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

feat(voice): live dictation via macOS 26 SpeechAnalyzer - #233

Merged
Tryanks merged 7 commits into
mainfrom
feat/voice-input
Aug 20, 2026
Merged

feat(voice): live dictation via macOS 26 SpeechAnalyzer#233
Tryanks merged 7 commits into
mainfrom
feat/voice-input

Conversation

@Tryanks

Copy link
Copy Markdown
Owner

What

Voice input for the composer: a mic toggle button that streams on-device speech transcription into the message editor as you speak. macOS 26+ only — on every other platform/version the button does not exist (compile-time cfg + runtime is_supported()).

Research behind the approach: docs/voice-input-research.md. Plan: docs/plans/voice-input.md.

How

  • crates/voice (new)tcode-voice: a Swift shim (swift/shim.swift, C ABI via @_cdecl) around SpeechAnalyzer/SpeechTranscriber, compiled and linked by build.rs only when targeting macOS with an SDK ≥ 26 (TCODE_VOICE_FORCE_STUB=1 forces the stub). Swift owns the whole audio path: AVAudioEngine mic tap → AVAudioConverterAnalyzerInput stream; locale assets are installed on demand via AssetInventory. Events (Ready | Volatile | Final | Error | Ended) cross the FFI on arbitrary threads with a serialized, single-terminal contract.
  • Composer UI — mic button next to the send button (both row layouts). Click to start (spinner while assets/mic warm up), speech streams into the editor at an anchored insertion point: each volatile hypothesis rewrites the volatile tail in place, finals commit. Click again / Esc stops and keeps the text; submit, thread switch, or typing during dictation also stops it. All offsets are UTF-8 bytes (zh-CN safe).
  • Locale — follows the app locale: zh → zh_CN, otherwise en_US.
  • PackagingNSMicrophoneUsageDescription + NSSpeechRecognitionUsageDescription added to the release Info.plist.

Test plan

  • cargo run -p tcode-voice --example file_dictation -- <aiff> zh_CN — 21 volatile events, correct final, clean exit (fixture from say -v Tingting)
  • cargo check --workspace, cargo clippy — clean (only the pre-existing block v0.1.6 future-incompat note)
  • cargo test -p tcode-ui — 254 passed, incl. locale key-set test and the new transcript-edit unit test
  • TCODE_VOICE_FORCE_STUB=1 build — stub path compiles, is_supported() false; non-macOS cfg path checked warning-free under -D warnings
  • Shim symbols present in the final tcode binary; is_supported() true at runtime on macOS 26.5
  • Live mic dictation in the running app (needs a human to approve the mic TCC prompt on first use)

Out of scope (per plan)

Settings section, cloud/BYOK backends, other platforms (Windows/Linux need a local ASR engine — see research doc), and conversational voice mode.

🤖 Generated with Claude Code

Tryanksand others added 7 commits August 21, 2026 02:40
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Adds a mic toggle to the composer that streams on-device speech
transcription into the editor. macOS 26+ only: the Swift shim
(SpeechAnalyzer/SpeechTranscriber, AVAudioEngine capture) is compiled by
build.rs when the SDK allows and the button does not exist elsewhere.
Volatile hypotheses rewrite an anchored range in place; final results
commit. Esc, submit, thread switch or typing stops the session.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The mic path stamped every AnalyzerInput with an explicit bufferStartTime
and stopped via finalizeAndFinish(through: <cumulative time>). With its
own timestamps the analyzer never reported results, and the targeted
finalize waited on a timeline position that never arrived, so stop() hung
and Ended was never delivered. Feed untimestamped buffers (the analyzer
tracks the live timeline itself) and finalize through end of input, as
validated against a standalone probe. Also treat an empty buffer from the
priming rate converter as skippable rather than an error, and keep a
mic_dictation example for live testing.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The shim reports a perceptual input level (RMS over a -50 dBFS window) at
the audio callback rate as a new cosmetic Level event; the mic button
renders it as a danger-tinted glow with fast-attack/slow-decay smoothing,
so the user can see their speech is being picked up.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ze flush
Two compounding reasons dictation text never appeared:
- SpeechTranscriber with plain volatileResults buffers every result until
finalization — nothing streams while speaking. DictationTranscriber's
progressiveLongDictation preset (the dictation-purpose module; 53
locales incl. zh_CN/en_US) streams hypotheses every few hundred ms,
verified against real audio driven through the exact mic pipeline.
- The UI dropped the session and event pump immediately on stop, so even
the finalization flush was discarded. Stop is now graceful (pump lives
until the terminal event lands the flushed text); submit, destination
switches and user edits abort instead, where a late flush would land in
the wrong place.
Also keeps a stderr diagnostic for mic authorization/terminal events.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…heck
crates/voice/build.rs invokes xcrun/swiftc at build time; the
Command::new boundary exists for runtime process spawning, which build
scripts cannot route through the approved helpers.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…tub builds
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@Tryanks
Tryanks merged commit e8a5f1c into mainAug 20, 2026
3 checks passed
@Tryanks
Tryanks deleted the feat/voice-input branch August 20, 2026 19:33
@TryanksTryanks mentioned this pull request Aug 22, 2026
12 tasks
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Tryanks
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(voice): live dictation via macOS 26 SpeechAnalyzer - #233

Merged
Tryanks merged 7 commits into
mainfrom
feat/voice-input
Aug 20, 2026
Merged

feat(voice): live dictation via macOS 26 SpeechAnalyzer#233
Tryanks merged 7 commits into
mainfrom
feat/voice-input

Conversation

@Tryanks

Copy link
Copy Markdown
Owner

What

Voice input for the composer: a mic toggle button that streams on-device speech transcription into the message editor as you speak. macOS 26+ only — on every other platform/version the button does not exist (compile-time cfg + runtime is_supported()).

Research behind the approach: docs/voice-input-research.md. Plan: docs/plans/voice-input.md.

How

  • crates/voice (new)tcode-voice: a Swift shim (swift/shim.swift, C ABI via @_cdecl) around SpeechAnalyzer/SpeechTranscriber, compiled and linked by build.rs only when targeting macOS with an SDK ≥ 26 (TCODE_VOICE_FORCE_STUB=1 forces the stub). Swift owns the whole audio path: AVAudioEngine mic tap → AVAudioConverterAnalyzerInput stream; locale assets are installed on demand via AssetInventory. Events (Ready | Volatile | Final | Error | Ended) cross the FFI on arbitrary threads with a serialized, single-terminal contract.
  • Composer UI — mic button next to the send button (both row layouts). Click to start (spinner while assets/mic warm up), speech streams into the editor at an anchored insertion point: each volatile hypothesis rewrites the volatile tail in place, finals commit. Click again / Esc stops and keeps the text; submit, thread switch, or typing during dictation also stops it. All offsets are UTF-8 bytes (zh-CN safe).
  • Locale — follows the app locale: zh → zh_CN, otherwise en_US.
  • PackagingNSMicrophoneUsageDescription + NSSpeechRecognitionUsageDescription added to the release Info.plist.

Test plan

  • cargo run -p tcode-voice --example file_dictation -- <aiff> zh_CN — 21 volatile events, correct final, clean exit (fixture from say -v Tingting)
  • cargo check --workspace, cargo clippy — clean (only the pre-existing block v0.1.6 future-incompat note)
  • cargo test -p tcode-ui — 254 passed, incl. locale key-set test and the new transcript-edit unit test
  • TCODE_VOICE_FORCE_STUB=1 build — stub path compiles, is_supported() false; non-macOS cfg path checked warning-free under -D warnings
  • Shim symbols present in the final tcode binary; is_supported() true at runtime on macOS 26.5
  • Live mic dictation in the running app (needs a human to approve the mic TCC prompt on first use)

Out of scope (per plan)

Settings section, cloud/BYOK backends, other platforms (Windows/Linux need a local ASR engine — see research doc), and conversational voice mode.

🤖 Generated with Claude Code

Tryanksand others added 7 commits August 21, 2026 02:40
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Adds a mic toggle to the composer that streams on-device speech
transcription into the editor. macOS 26+ only: the Swift shim
(SpeechAnalyzer/SpeechTranscriber, AVAudioEngine capture) is compiled by
build.rs when the SDK allows and the button does not exist elsewhere.
Volatile hypotheses rewrite an anchored range in place; final results
commit. Esc, submit, thread switch or typing stops the session.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The mic path stamped every AnalyzerInput with an explicit bufferStartTime
and stopped via finalizeAndFinish(through: <cumulative time>). With its
own timestamps the analyzer never reported results, and the targeted
finalize waited on a timeline position that never arrived, so stop() hung
and Ended was never delivered. Feed untimestamped buffers (the analyzer
tracks the live timeline itself) and finalize through end of input, as
validated against a standalone probe. Also treat an empty buffer from the
priming rate converter as skippable rather than an error, and keep a
mic_dictation example for live testing.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The shim reports a perceptual input level (RMS over a -50 dBFS window) at
the audio callback rate as a new cosmetic Level event; the mic button
renders it as a danger-tinted glow with fast-attack/slow-decay smoothing,
so the user can see their speech is being picked up.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ze flush
Two compounding reasons dictation text never appeared:
- SpeechTranscriber with plain volatileResults buffers every result until
finalization — nothing streams while speaking. DictationTranscriber's
progressiveLongDictation preset (the dictation-purpose module; 53
locales incl. zh_CN/en_US) streams hypotheses every few hundred ms,
verified against real audio driven through the exact mic pipeline.
- The UI dropped the session and event pump immediately on stop, so even
the finalization flush was discarded. Stop is now graceful (pump lives
until the terminal event lands the flushed text); submit, destination
switches and user edits abort instead, where a late flush would land in
the wrong place.
Also keeps a stderr diagnostic for mic authorization/terminal events.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…heck
crates/voice/build.rs invokes xcrun/swiftc at build time; the
Command::new boundary exists for runtime process spawning, which build
scripts cannot route through the approved helpers.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…tub builds
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@Tryanks
Tryanks merged commit e8a5f1c into mainAug 20, 2026
3 checks passed
@Tryanks
Tryanks deleted the feat/voice-input branch August 20, 2026 19:33
@TryanksTryanks mentioned this pull request Aug 22, 2026
12 tasks
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Tryanks
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(voice): live dictation via macOS 26 SpeechAnalyzer - #233

Merged
Tryanks merged 7 commits into
mainfrom
feat/voice-input
Aug 20, 2026
Merged

feat(voice): live dictation via macOS 26 SpeechAnalyzer#233
Tryanks merged 7 commits into
mainfrom
feat/voice-input

Conversation

@Tryanks

Copy link
Copy Markdown
Owner

What

Voice input for the composer: a mic toggle button that streams on-device speech transcription into the message editor as you speak. macOS 26+ only — on every other platform/version the button does not exist (compile-time cfg + runtime is_supported()).

Research behind the approach: docs/voice-input-research.md. Plan: docs/plans/voice-input.md.

How

  • crates/voice (new)tcode-voice: a Swift shim (swift/shim.swift, C ABI via @_cdecl) around SpeechAnalyzer/SpeechTranscriber, compiled and linked by build.rs only when targeting macOS with an SDK ≥ 26 (TCODE_VOICE_FORCE_STUB=1 forces the stub). Swift owns the whole audio path: AVAudioEngine mic tap → AVAudioConverterAnalyzerInput stream; locale assets are installed on demand via AssetInventory. Events (Ready | Volatile | Final | Error | Ended) cross the FFI on arbitrary threads with a serialized, single-terminal contract.
  • Composer UI — mic button next to the send button (both row layouts). Click to start (spinner while assets/mic warm up), speech streams into the editor at an anchored insertion point: each volatile hypothesis rewrites the volatile tail in place, finals commit. Click again / Esc stops and keeps the text; submit, thread switch, or typing during dictation also stops it. All offsets are UTF-8 bytes (zh-CN safe).
  • Locale — follows the app locale: zh → zh_CN, otherwise en_US.
  • PackagingNSMicrophoneUsageDescription + NSSpeechRecognitionUsageDescription added to the release Info.plist.

Test plan

  • cargo run -p tcode-voice --example file_dictation -- <aiff> zh_CN — 21 volatile events, correct final, clean exit (fixture from say -v Tingting)
  • cargo check --workspace, cargo clippy — clean (only the pre-existing block v0.1.6 future-incompat note)
  • cargo test -p tcode-ui — 254 passed, incl. locale key-set test and the new transcript-edit unit test
  • TCODE_VOICE_FORCE_STUB=1 build — stub path compiles, is_supported() false; non-macOS cfg path checked warning-free under -D warnings
  • Shim symbols present in the final tcode binary; is_supported() true at runtime on macOS 26.5
  • Live mic dictation in the running app (needs a human to approve the mic TCC prompt on first use)

Out of scope (per plan)

Settings section, cloud/BYOK backends, other platforms (Windows/Linux need a local ASR engine — see research doc), and conversational voice mode.

🤖 Generated with Claude Code

Tryanksand others added 7 commits August 21, 2026 02:40
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Adds a mic toggle to the composer that streams on-device speech
transcription into the editor. macOS 26+ only: the Swift shim
(SpeechAnalyzer/SpeechTranscriber, AVAudioEngine capture) is compiled by
build.rs when the SDK allows and the button does not exist elsewhere.
Volatile hypotheses rewrite an anchored range in place; final results
commit. Esc, submit, thread switch or typing stops the session.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The mic path stamped every AnalyzerInput with an explicit bufferStartTime
and stopped via finalizeAndFinish(through: <cumulative time>). With its
own timestamps the analyzer never reported results, and the targeted
finalize waited on a timeline position that never arrived, so stop() hung
and Ended was never delivered. Feed untimestamped buffers (the analyzer
tracks the live timeline itself) and finalize through end of input, as
validated against a standalone probe. Also treat an empty buffer from the
priming rate converter as skippable rather than an error, and keep a
mic_dictation example for live testing.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The shim reports a perceptual input level (RMS over a -50 dBFS window) at
the audio callback rate as a new cosmetic Level event; the mic button
renders it as a danger-tinted glow with fast-attack/slow-decay smoothing,
so the user can see their speech is being picked up.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ze flush
Two compounding reasons dictation text never appeared:
- SpeechTranscriber with plain volatileResults buffers every result until
finalization — nothing streams while speaking. DictationTranscriber's
progressiveLongDictation preset (the dictation-purpose module; 53
locales incl. zh_CN/en_US) streams hypotheses every few hundred ms,
verified against real audio driven through the exact mic pipeline.
- The UI dropped the session and event pump immediately on stop, so even
the finalization flush was discarded. Stop is now graceful (pump lives
until the terminal event lands the flushed text); submit, destination
switches and user edits abort instead, where a late flush would land in
the wrong place.
Also keeps a stderr diagnostic for mic authorization/terminal events.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…heck
crates/voice/build.rs invokes xcrun/swiftc at build time; the
Command::new boundary exists for runtime process spawning, which build
scripts cannot route through the approved helpers.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…tub builds
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@Tryanks
Tryanks merged commit e8a5f1c into mainAug 20, 2026
3 checks passed
@Tryanks
Tryanks deleted the feat/voice-input branch August 20, 2026 19:33
@TryanksTryanks mentioned this pull request Aug 22, 2026
12 tasks
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Tryanks
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

feat(voice): live dictation via macOS 26 SpeechAnalyzer - #233

Merged
Tryanks merged 7 commits into
mainfrom
feat/voice-input
Aug 20, 2026
Merged

feat(voice): live dictation via macOS 26 SpeechAnalyzer#233
Tryanks merged 7 commits into
mainfrom
feat/voice-input

Conversation

@Tryanks

Copy link
Copy Markdown
Owner

What

Voice input for the composer: a mic toggle button that streams on-device speech transcription into the message editor as you speak. macOS 26+ only — on every other platform/version the button does not exist (compile-time cfg + runtime is_supported()).

Research behind the approach: docs/voice-input-research.md. Plan: docs/plans/voice-input.md.

How

  • crates/voice (new)tcode-voice: a Swift shim (swift/shim.swift, C ABI via @_cdecl) around SpeechAnalyzer/SpeechTranscriber, compiled and linked by build.rs only when targeting macOS with an SDK ≥ 26 (TCODE_VOICE_FORCE_STUB=1 forces the stub). Swift owns the whole audio path: AVAudioEngine mic tap → AVAudioConverterAnalyzerInput stream; locale assets are installed on demand via AssetInventory. Events (Ready | Volatile | Final | Error | Ended) cross the FFI on arbitrary threads with a serialized, single-terminal contract.
  • Composer UI — mic button next to the send button (both row layouts). Click to start (spinner while assets/mic warm up), speech streams into the editor at an anchored insertion point: each volatile hypothesis rewrites the volatile tail in place, finals commit. Click again / Esc stops and keeps the text; submit, thread switch, or typing during dictation also stops it. All offsets are UTF-8 bytes (zh-CN safe).
  • Locale — follows the app locale: zh → zh_CN, otherwise en_US.
  • PackagingNSMicrophoneUsageDescription + NSSpeechRecognitionUsageDescription added to the release Info.plist.

Test plan

  • cargo run -p tcode-voice --example file_dictation -- <aiff> zh_CN — 21 volatile events, correct final, clean exit (fixture from say -v Tingting)
  • cargo check --workspace, cargo clippy — clean (only the pre-existing block v0.1.6 future-incompat note)
  • cargo test -p tcode-ui — 254 passed, incl. locale key-set test and the new transcript-edit unit test
  • TCODE_VOICE_FORCE_STUB=1 build — stub path compiles, is_supported() false; non-macOS cfg path checked warning-free under -D warnings
  • Shim symbols present in the final tcode binary; is_supported() true at runtime on macOS 26.5
  • Live mic dictation in the running app (needs a human to approve the mic TCC prompt on first use)

Out of scope (per plan)

Settings section, cloud/BYOK backends, other platforms (Windows/Linux need a local ASR engine — see research doc), and conversational voice mode.

🤖 Generated with Claude Code

Tryanksand others added 7 commits August 21, 2026 02:40
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Adds a mic toggle to the composer that streams on-device speech
transcription into the editor. macOS 26+ only: the Swift shim
(SpeechAnalyzer/SpeechTranscriber, AVAudioEngine capture) is compiled by
build.rs when the SDK allows and the button does not exist elsewhere.
Volatile hypotheses rewrite an anchored range in place; final results
commit. Esc, submit, thread switch or typing stops the session.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The mic path stamped every AnalyzerInput with an explicit bufferStartTime
and stopped via finalizeAndFinish(through: <cumulative time>). With its
own timestamps the analyzer never reported results, and the targeted
finalize waited on a timeline position that never arrived, so stop() hung
and Ended was never delivered. Feed untimestamped buffers (the analyzer
tracks the live timeline itself) and finalize through end of input, as
validated against a standalone probe. Also treat an empty buffer from the
priming rate converter as skippable rather than an error, and keep a
mic_dictation example for live testing.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The shim reports a perceptual input level (RMS over a -50 dBFS window) at
the audio callback rate as a new cosmetic Level event; the mic button
renders it as a danger-tinted glow with fast-attack/slow-decay smoothing,
so the user can see their speech is being picked up.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ze flush
Two compounding reasons dictation text never appeared:
- SpeechTranscriber with plain volatileResults buffers every result until
finalization — nothing streams while speaking. DictationTranscriber's
progressiveLongDictation preset (the dictation-purpose module; 53
locales incl. zh_CN/en_US) streams hypotheses every few hundred ms,
verified against real audio driven through the exact mic pipeline.
- The UI dropped the session and event pump immediately on stop, so even
the finalization flush was discarded. Stop is now graceful (pump lives
until the terminal event lands the flushed text); submit, destination
switches and user edits abort instead, where a late flush would land in
the wrong place.
Also keeps a stderr diagnostic for mic authorization/terminal events.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…heck
crates/voice/build.rs invokes xcrun/swiftc at build time; the
Command::new boundary exists for runtime process spawning, which build
scripts cannot route through the approved helpers.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…tub builds
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@Tryanks
Tryanks merged commit e8a5f1c into mainAug 20, 2026
3 checks passed
@Tryanks
Tryanks deleted the feat/voice-input branch August 20, 2026 19:33
@TryanksTryanks mentioned this pull request Aug 22, 2026
12 tasks
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Tryanks
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(voice): live dictation via macOS 26 SpeechAnalyzer - #233

Merged
Tryanks merged 7 commits into
mainfrom
feat/voice-input
Aug 20, 2026
Merged

feat(voice): live dictation via macOS 26 SpeechAnalyzer#233
Tryanks merged 7 commits into
mainfrom
feat/voice-input

Conversation

@Tryanks

Copy link
Copy Markdown
Owner

What

Voice input for the composer: a mic toggle button that streams on-device speech transcription into the message editor as you speak. macOS 26+ only — on every other platform/version the button does not exist (compile-time cfg + runtime is_supported()).

Research behind the approach: docs/voice-input-research.md. Plan: docs/plans/voice-input.md.

How

  • crates/voice (new)tcode-voice: a Swift shim (swift/shim.swift, C ABI via @_cdecl) around SpeechAnalyzer/SpeechTranscriber, compiled and linked by build.rs only when targeting macOS with an SDK ≥ 26 (TCODE_VOICE_FORCE_STUB=1 forces the stub). Swift owns the whole audio path: AVAudioEngine mic tap → AVAudioConverterAnalyzerInput stream; locale assets are installed on demand via AssetInventory. Events (Ready | Volatile | Final | Error | Ended) cross the FFI on arbitrary threads with a serialized, single-terminal contract.
  • Composer UI — mic button next to the send button (both row layouts). Click to start (spinner while assets/mic warm up), speech streams into the editor at an anchored insertion point: each volatile hypothesis rewrites the volatile tail in place, finals commit. Click again / Esc stops and keeps the text; submit, thread switch, or typing during dictation also stops it. All offsets are UTF-8 bytes (zh-CN safe).
  • Locale — follows the app locale: zh → zh_CN, otherwise en_US.
  • PackagingNSMicrophoneUsageDescription + NSSpeechRecognitionUsageDescription added to the release Info.plist.

Test plan

  • cargo run -p tcode-voice --example file_dictation -- <aiff> zh_CN — 21 volatile events, correct final, clean exit (fixture from say -v Tingting)
  • cargo check --workspace, cargo clippy — clean (only the pre-existing block v0.1.6 future-incompat note)
  • cargo test -p tcode-ui — 254 passed, incl. locale key-set test and the new transcript-edit unit test
  • TCODE_VOICE_FORCE_STUB=1 build — stub path compiles, is_supported() false; non-macOS cfg path checked warning-free under -D warnings
  • Shim symbols present in the final tcode binary; is_supported() true at runtime on macOS 26.5
  • Live mic dictation in the running app (needs a human to approve the mic TCC prompt on first use)

Out of scope (per plan)

Settings section, cloud/BYOK backends, other platforms (Windows/Linux need a local ASR engine — see research doc), and conversational voice mode.

🤖 Generated with Claude Code

Tryanksand others added 7 commits August 21, 2026 02:40
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Adds a mic toggle to the composer that streams on-device speech
transcription into the editor. macOS 26+ only: the Swift shim
(SpeechAnalyzer/SpeechTranscriber, AVAudioEngine capture) is compiled by
build.rs when the SDK allows and the button does not exist elsewhere.
Volatile hypotheses rewrite an anchored range in place; final results
commit. Esc, submit, thread switch or typing stops the session.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The mic path stamped every AnalyzerInput with an explicit bufferStartTime
and stopped via finalizeAndFinish(through: <cumulative time>). With its
own timestamps the analyzer never reported results, and the targeted
finalize waited on a timeline position that never arrived, so stop() hung
and Ended was never delivered. Feed untimestamped buffers (the analyzer
tracks the live timeline itself) and finalize through end of input, as
validated against a standalone probe. Also treat an empty buffer from the
priming rate converter as skippable rather than an error, and keep a
mic_dictation example for live testing.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The shim reports a perceptual input level (RMS over a -50 dBFS window) at
the audio callback rate as a new cosmetic Level event; the mic button
renders it as a danger-tinted glow with fast-attack/slow-decay smoothing,
so the user can see their speech is being picked up.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ze flush
Two compounding reasons dictation text never appeared:
- SpeechTranscriber with plain volatileResults buffers every result until
finalization — nothing streams while speaking. DictationTranscriber's
progressiveLongDictation preset (the dictation-purpose module; 53
locales incl. zh_CN/en_US) streams hypotheses every few hundred ms,
verified against real audio driven through the exact mic pipeline.
- The UI dropped the session and event pump immediately on stop, so even
the finalization flush was discarded. Stop is now graceful (pump lives
until the terminal event lands the flushed text); submit, destination
switches and user edits abort instead, where a late flush would land in
the wrong place.
Also keeps a stderr diagnostic for mic authorization/terminal events.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…heck
crates/voice/build.rs invokes xcrun/swiftc at build time; the
Command::new boundary exists for runtime process spawning, which build
scripts cannot route through the approved helpers.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…tub builds
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@Tryanks
Tryanks merged commit e8a5f1c into mainAug 20, 2026
3 checks passed
@Tryanks
Tryanks deleted the feat/voice-input branch August 20, 2026 19:33
@TryanksTryanks mentioned this pull request Aug 22, 2026
12 tasks
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Tryanks
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(voice): live dictation via macOS 26 SpeechAnalyzer - #233

Merged
Tryanks merged 7 commits into
mainfrom
feat/voice-input
Aug 20, 2026
Merged

feat(voice): live dictation via macOS 26 SpeechAnalyzer#233
Tryanks merged 7 commits into
mainfrom
feat/voice-input

Conversation

@Tryanks

Copy link
Copy Markdown
Owner

What

Voice input for the composer: a mic toggle button that streams on-device speech transcription into the message editor as you speak. macOS 26+ only — on every other platform/version the button does not exist (compile-time cfg + runtime is_supported()).

Research behind the approach: docs/voice-input-research.md. Plan: docs/plans/voice-input.md.

How

  • crates/voice (new)tcode-voice: a Swift shim (swift/shim.swift, C ABI via @_cdecl) around SpeechAnalyzer/SpeechTranscriber, compiled and linked by build.rs only when targeting macOS with an SDK ≥ 26 (TCODE_VOICE_FORCE_STUB=1 forces the stub). Swift owns the whole audio path: AVAudioEngine mic tap → AVAudioConverterAnalyzerInput stream; locale assets are installed on demand via AssetInventory. Events (Ready | Volatile | Final | Error | Ended) cross the FFI on arbitrary threads with a serialized, single-terminal contract.
  • Composer UI — mic button next to the send button (both row layouts). Click to start (spinner while assets/mic warm up), speech streams into the editor at an anchored insertion point: each volatile hypothesis rewrites the volatile tail in place, finals commit. Click again / Esc stops and keeps the text; submit, thread switch, or typing during dictation also stops it. All offsets are UTF-8 bytes (zh-CN safe).
  • Locale — follows the app locale: zh → zh_CN, otherwise en_US.
  • PackagingNSMicrophoneUsageDescription + NSSpeechRecognitionUsageDescription added to the release Info.plist.

Test plan

  • cargo run -p tcode-voice --example file_dictation -- <aiff> zh_CN — 21 volatile events, correct final, clean exit (fixture from say -v Tingting)
  • cargo check --workspace, cargo clippy — clean (only the pre-existing block v0.1.6 future-incompat note)
  • cargo test -p tcode-ui — 254 passed, incl. locale key-set test and the new transcript-edit unit test
  • TCODE_VOICE_FORCE_STUB=1 build — stub path compiles, is_supported() false; non-macOS cfg path checked warning-free under -D warnings
  • Shim symbols present in the final tcode binary; is_supported() true at runtime on macOS 26.5
  • Live mic dictation in the running app (needs a human to approve the mic TCC prompt on first use)

Out of scope (per plan)

Settings section, cloud/BYOK backends, other platforms (Windows/Linux need a local ASR engine — see research doc), and conversational voice mode.

🤖 Generated with Claude Code

Tryanksand others added 7 commits August 21, 2026 02:40
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Adds a mic toggle to the composer that streams on-device speech
transcription into the editor. macOS 26+ only: the Swift shim
(SpeechAnalyzer/SpeechTranscriber, AVAudioEngine capture) is compiled by
build.rs when the SDK allows and the button does not exist elsewhere.
Volatile hypotheses rewrite an anchored range in place; final results
commit. Esc, submit, thread switch or typing stops the session.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The mic path stamped every AnalyzerInput with an explicit bufferStartTime
and stopped via finalizeAndFinish(through: <cumulative time>). With its
own timestamps the analyzer never reported results, and the targeted
finalize waited on a timeline position that never arrived, so stop() hung
and Ended was never delivered. Feed untimestamped buffers (the analyzer
tracks the live timeline itself) and finalize through end of input, as
validated against a standalone probe. Also treat an empty buffer from the
priming rate converter as skippable rather than an error, and keep a
mic_dictation example for live testing.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The shim reports a perceptual input level (RMS over a -50 dBFS window) at
the audio callback rate as a new cosmetic Level event; the mic button
renders it as a danger-tinted glow with fast-attack/slow-decay smoothing,
so the user can see their speech is being picked up.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ze flush
Two compounding reasons dictation text never appeared:
- SpeechTranscriber with plain volatileResults buffers every result until
finalization — nothing streams while speaking. DictationTranscriber's
progressiveLongDictation preset (the dictation-purpose module; 53
locales incl. zh_CN/en_US) streams hypotheses every few hundred ms,
verified against real audio driven through the exact mic pipeline.
- The UI dropped the session and event pump immediately on stop, so even
the finalization flush was discarded. Stop is now graceful (pump lives
until the terminal event lands the flushed text); submit, destination
switches and user edits abort instead, where a late flush would land in
the wrong place.
Also keeps a stderr diagnostic for mic authorization/terminal events.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…heck
crates/voice/build.rs invokes xcrun/swiftc at build time; the
Command::new boundary exists for runtime process spawning, which build
scripts cannot route through the approved helpers.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…tub builds
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@Tryanks
Tryanks merged commit e8a5f1c into mainAug 20, 2026
3 checks passed
@Tryanks
Tryanks deleted the feat/voice-input branch August 20, 2026 19:33
@TryanksTryanks mentioned this pull request Aug 22, 2026
12 tasks
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Tryanks
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

feat(voice): live dictation via macOS 26 SpeechAnalyzer - #233

Merged
Tryanks merged 7 commits into
mainfrom
feat/voice-input
Aug 20, 2026
Merged

feat(voice): live dictation via macOS 26 SpeechAnalyzer#233
Tryanks merged 7 commits into
mainfrom
feat/voice-input

Conversation

@Tryanks

Copy link
Copy Markdown
Owner

What

Voice input for the composer: a mic toggle button that streams on-device speech transcription into the message editor as you speak. macOS 26+ only — on every other platform/version the button does not exist (compile-time cfg + runtime is_supported()).

Research behind the approach: docs/voice-input-research.md. Plan: docs/plans/voice-input.md.

How

  • crates/voice (new)tcode-voice: a Swift shim (swift/shim.swift, C ABI via @_cdecl) around SpeechAnalyzer/SpeechTranscriber, compiled and linked by build.rs only when targeting macOS with an SDK ≥ 26 (TCODE_VOICE_FORCE_STUB=1 forces the stub). Swift owns the whole audio path: AVAudioEngine mic tap → AVAudioConverterAnalyzerInput stream; locale assets are installed on demand via AssetInventory. Events (Ready | Volatile | Final | Error | Ended) cross the FFI on arbitrary threads with a serialized, single-terminal contract.
  • Composer UI — mic button next to the send button (both row layouts). Click to start (spinner while assets/mic warm up), speech streams into the editor at an anchored insertion point: each volatile hypothesis rewrites the volatile tail in place, finals commit. Click again / Esc stops and keeps the text; submit, thread switch, or typing during dictation also stops it. All offsets are UTF-8 bytes (zh-CN safe).
  • Locale — follows the app locale: zh → zh_CN, otherwise en_US.
  • PackagingNSMicrophoneUsageDescription + NSSpeechRecognitionUsageDescription added to the release Info.plist.

Test plan

  • cargo run -p tcode-voice --example file_dictation -- <aiff> zh_CN — 21 volatile events, correct final, clean exit (fixture from say -v Tingting)
  • cargo check --workspace, cargo clippy — clean (only the pre-existing block v0.1.6 future-incompat note)
  • cargo test -p tcode-ui — 254 passed, incl. locale key-set test and the new transcript-edit unit test
  • TCODE_VOICE_FORCE_STUB=1 build — stub path compiles, is_supported() false; non-macOS cfg path checked warning-free under -D warnings
  • Shim symbols present in the final tcode binary; is_supported() true at runtime on macOS 26.5
  • Live mic dictation in the running app (needs a human to approve the mic TCC prompt on first use)

Out of scope (per plan)

Settings section, cloud/BYOK backends, other platforms (Windows/Linux need a local ASR engine — see research doc), and conversational voice mode.

🤖 Generated with Claude Code

Tryanksand others added 7 commits August 21, 2026 02:40
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Adds a mic toggle to the composer that streams on-device speech
transcription into the editor. macOS 26+ only: the Swift shim
(SpeechAnalyzer/SpeechTranscriber, AVAudioEngine capture) is compiled by
build.rs when the SDK allows and the button does not exist elsewhere.
Volatile hypotheses rewrite an anchored range in place; final results
commit. Esc, submit, thread switch or typing stops the session.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The mic path stamped every AnalyzerInput with an explicit bufferStartTime
and stopped via finalizeAndFinish(through: <cumulative time>). With its
own timestamps the analyzer never reported results, and the targeted
finalize waited on a timeline position that never arrived, so stop() hung
and Ended was never delivered. Feed untimestamped buffers (the analyzer
tracks the live timeline itself) and finalize through end of input, as
validated against a standalone probe. Also treat an empty buffer from the
priming rate converter as skippable rather than an error, and keep a
mic_dictation example for live testing.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The shim reports a perceptual input level (RMS over a -50 dBFS window) at
the audio callback rate as a new cosmetic Level event; the mic button
renders it as a danger-tinted glow with fast-attack/slow-decay smoothing,
so the user can see their speech is being picked up.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ze flush
Two compounding reasons dictation text never appeared:
- SpeechTranscriber with plain volatileResults buffers every result until
finalization — nothing streams while speaking. DictationTranscriber's
progressiveLongDictation preset (the dictation-purpose module; 53
locales incl. zh_CN/en_US) streams hypotheses every few hundred ms,
verified against real audio driven through the exact mic pipeline.
- The UI dropped the session and event pump immediately on stop, so even
the finalization flush was discarded. Stop is now graceful (pump lives
until the terminal event lands the flushed text); submit, destination
switches and user edits abort instead, where a late flush would land in
the wrong place.
Also keeps a stderr diagnostic for mic authorization/terminal events.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…heck
crates/voice/build.rs invokes xcrun/swiftc at build time; the
Command::new boundary exists for runtime process spawning, which build
scripts cannot route through the approved helpers.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…tub builds
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@Tryanks
Tryanks merged commit e8a5f1c into mainAug 20, 2026
3 checks passed
@Tryanks
Tryanks deleted the feat/voice-input branch August 20, 2026 19:33
@TryanksTryanks mentioned this pull request Aug 22, 2026
12 tasks
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Tryanks