Repository files navigation

Utter

Local AI-powered voice input for macOS menu bar


GitHub StarsGitHub ForksGitHub Issues

PlatformSwiftApple SiliconLicense

WhisperKitMLX

Website · 中文文档


Overview

Utter is a macOS menu bar app for AI-powered voice input and dictation. It supports both fully local on-device inference and remote LLM APIs. Press a hotkey to start recording, release to transcribe, and the result is typed directly into whatever app you're using.

Three output modes are available:

  • Verbatim — raw transcription, lowest latency
  • Smart Format — transcription cleaned up by an LLM (contextual filler removal, grammar fixes, structured formatting)
  • Voice Command — speak a command and get an AI-generated response based on screen context

Features

FeatureDescription
Multiple Speech EnginesApple Speech, WhisperKit, Doubao ASR, or Qwen3-ASR
Smart Text ProcessingLocal MLX Qwen2.5/Qwen3 or remote LLM infers spoken intent — contextual cleanup, "scratch that" restarts, self-correction handling, spoken punctuation, technical terms, numbers/ranges/units, and structured formatting
LLM-Owned Spoken FormattingSpoken casing, no-space dictation, identifiers, file paths, shortcuts, emoji, Markdown tasks, dates/times, quantities, units, formulas, fractions, and digit sequences are handled by the Smart Format / Voice Command prompts instead of local hardcoded rewrite rules
Voice Edit CommandsIn Voice Command mode, an LLM classifies safe structured actions for replacing, undoing, proofreading, titling, summarizing, drafting replies, making meeting notes, extracting key points/decisions/questions/risks/deadlines/owners/action items, rewriting tone, expanding, making tables/lists, or deleting the previous Utter insertion or selected text
Verbatim & Preview BoundaryVerbatim mode, streaming HUD, integration partials, and instant-insert drafts keep ASR text close to raw output with only dictionary, whitespace, duplicate, and non-speech-artifact cleanup
Remote LLM SupportOpenAI, Claude (Anthropic format), Gemini, OpenRouter, SiliconFlow, Doubao, Bailian, MiniMax (CN & Global)
Global HotkeyConfigurable key (Fn/Ctrl/Shift/Option) with long-press, double-tap, or single-tap activation
Translation DictationUse a dedicated hotkey chord to speak in one language and insert an English, Chinese, Japanese, Korean, Spanish, French, or German translation
Screen Context OCRCaptures on-screen text via ScreenCaptureKit + Vision to help the LLM correct homophones
Voice Command ModeScreen-aware voice assistant — summarize, reply, translate based on what's on screen
Input MemoryRecent input history injected as LLM context for better continuity
Industry VocabularyChoose Medical, Legal, Finance & Accounting, or Software Technology terms for Apple Speech / Whisper biasing and output normalization; personal terms take priority
Edit RulesPersonal text replacement rules applied on every output
Language Style PresetsConcise / Formal / Casual / Custom prompt per language
Input History & StatsFull history with raw vs. processed comparison, word count stats, configurable retention
Bilingual UIChinese and English interface, independent of recognition language
Sound FeedbackAudio cues on recording start and stop
Guided OnboardingStep-by-step setup: permissions, model download, and first use

System Requirements

  • OS: macOS 26 (Tahoe) or later
  • Chip: Apple Silicon (M1 / M2 / M3 / M4)
  • Disk: ~400 MB minimum (Apple Speech + Qwen3-0.6B), up to ~4 GB with larger models

Installation

Download

Grab the latest .dmg from Releases, open it, and drag Utter.app to Applications.

"Cannot verify the developer" on first launch? The app is not notarized by Apple. Before first run, execute in Terminal:

xattr -cr /Applications/Utter.app

Or go to System Settings → Privacy & Security and click "Open Anyway".

Build from Source

# Build .app bundle + .dmg installer
bash scripts/build-app.sh
# Or for development
swift build
swift run OpenType
# Build, sign, and launch the development .app bundle
bash scripts/build-and-run.sh --verify
# Or open in Xcode
open Package.swift

Material changes follow the artifact-driven workflow in docs/sdlc/README.md. Before opening a pull request, run:

bash scripts/sdlc-checks.sh
bash scripts/ci-basic-checks.sh
swift test

The public app is Utter.app. The Swift package product remains OpenType so existing source integrations and upgrade paths continue to work.

First Run

  1. Launch Utter — it appears as a waveform icon in the menu bar
  2. The onboarding wizard guides you through permissions and model setup
  3. Grant Microphone and Accessibility permissions (required)
  4. Wait for the LLM model to download (~335 MB, one-time)
  5. Hold Fn to start dictating, release to stop and insert text

Permissions

PermissionPurposeRequired
MicrophoneAudio captureYes
AccessibilityGlobal hotkey + text injection (simulated paste)Yes
Speech RecognitionApple on-device ASR engineOnly if using Apple Speech
Screen RecordingOCR for screen context and Voice Command modeOptional
NetworkModel downloads; remote LLM API callsFirst run / remote LLM mode

Remote LLM Providers

Utter supports both OpenAI-compatible and Anthropic API formats:

ProviderAPI FormatBase URL
OpenAIOpenAIhttps://api.openai.com/v1
Anthropic ClaudeAnthropichttps://api.anthropic.com/v1
Google GeminiOpenAIhttps://generativelanguage.googleapis.com/v1beta/openai
OpenRouterOpenAIhttps://openrouter.ai/api/v1
SiliconFlowOpenAIhttps://api.siliconflow.cn/v1
Volcengine DoubaoOpenAIhttps://ark.cn-beijing.volces.com/api/v3
Alibaba BailianOpenAIhttps://dashscope.aliyuncs.com/compatible-mode/v1
MiniMax (China)OpenAIhttps://api.minimax.chat/v1
MiniMax (Global)OpenAIhttps://api.minimaxi.chat/v1

Local ASR Providers

ProviderLocal runtimeDefault model
Qwen3-ASRNative Swift + MLX on Apple Siliconmlx-community/Qwen3-ASR-1.7B-bf16

Qwen3-ASR does not call a hosted ASR API. The app downloads the selected model into the same model storage used by WhisperKit and runs inference locally through native Swift and MLX.

Project Structure

Sources/
├── App/ # Entry point, AppDelegate, AppState, VoicePipeline, AppIcon
├── Audio/ # Microphone capture (AVAudioEngine), sound playback
├── Config/ # AppSettings, ModelCatalog, RemoteModelConfig, Localization
├── Hotkey/ # Global hotkey via CGEvent tap
├── LLM/ # LLMEngine (MLX), RemoteLLMClient (OpenAI/Anthropic)
├── Output/ # Text injection (Accessibility API + clipboard paste)
├── Processing/ # TextProcessor, InputHistory, MemoryStore, personal and industry vocabulary
├── Prompts/ # PromptBuilder, prompt catalogs, style prompt presets
├── Screen/ # Screen OCR (ScreenCaptureKit + Vision)
├── Speech/ # SpeechEngine protocol, WhisperKit, Apple Speech, Doubao ASR, local ASR engines
├── UI/ # SwiftUI: MenuBar, Settings, Onboarding, Overlay, History, Models
└── Resources/ # Localization strings (en/zh-Hans), sounds, app icon
scripts/
├── build-and-run.sh # Build, sign, and launch a development .app bundle
├── build-app.sh # Build release .app bundle and .dmg installer
├── ci-basic-checks.sh # CI guardrails for linked files and resources
├── create-signing-cert.sh # Generate self-signed code signing certificate
├── generate-icon.swift # Generate AppIcon.icns from source PNG
├── test-industry-lexicons.sh # Validate vocabulary recall and non-target preservation
├── unit-test-coverage.sh # Run unit tests with coverage thresholds
└── validate-volc-asr.swift # Validate Volcengine ASR configuration manually

Tech Stack

  • WhisperKit — offline Whisper speech recognition
  • mlx-swift-lm — local LLM inference on Apple Silicon (Qwen2.5 / Qwen3)
  • SwiftUI + AppKit — native macOS UI
  • ScreenCaptureKit + Vision — screen OCR
  • AVAudioEngine — low-latency microphone capture
  • Apple Speech Framework — on-device speech recognition

License

MIT


Made with care for Apple Silicon

About

Local AI-powered voice input for macOS menu bar

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Utter

Local AI-powered voice input for macOS menu bar


GitHub StarsGitHub ForksGitHub Issues

PlatformSwiftApple SiliconLicense

WhisperKitMLX

Website · 中文文档


Overview

Utter is a macOS menu bar app for AI-powered voice input and dictation. It supports both fully local on-device inference and remote LLM APIs. Press a hotkey to start recording, release to transcribe, and the result is typed directly into whatever app you're using.

Three output modes are available:

  • Verbatim — raw transcription, lowest latency
  • Smart Format — transcription cleaned up by an LLM (contextual filler removal, grammar fixes, structured formatting)
  • Voice Command — speak a command and get an AI-generated response based on screen context

Features

FeatureDescription
Multiple Speech EnginesApple Speech, WhisperKit, Doubao ASR, or Qwen3-ASR
Smart Text ProcessingLocal MLX Qwen2.5/Qwen3 or remote LLM infers spoken intent — contextual cleanup, "scratch that" restarts, self-correction handling, spoken punctuation, technical terms, numbers/ranges/units, and structured formatting
LLM-Owned Spoken FormattingSpoken casing, no-space dictation, identifiers, file paths, shortcuts, emoji, Markdown tasks, dates/times, quantities, units, formulas, fractions, and digit sequences are handled by the Smart Format / Voice Command prompts instead of local hardcoded rewrite rules
Voice Edit CommandsIn Voice Command mode, an LLM classifies safe structured actions for replacing, undoing, proofreading, titling, summarizing, drafting replies, making meeting notes, extracting key points/decisions/questions/risks/deadlines/owners/action items, rewriting tone, expanding, making tables/lists, or deleting the previous Utter insertion or selected text
Verbatim & Preview BoundaryVerbatim mode, streaming HUD, integration partials, and instant-insert drafts keep ASR text close to raw output with only dictionary, whitespace, duplicate, and non-speech-artifact cleanup
Remote LLM SupportOpenAI, Claude (Anthropic format), Gemini, OpenRouter, SiliconFlow, Doubao, Bailian, MiniMax (CN & Global)
Global HotkeyConfigurable key (Fn/Ctrl/Shift/Option) with long-press, double-tap, or single-tap activation
Translation DictationUse a dedicated hotkey chord to speak in one language and insert an English, Chinese, Japanese, Korean, Spanish, French, or German translation
Screen Context OCRCaptures on-screen text via ScreenCaptureKit + Vision to help the LLM correct homophones
Voice Command ModeScreen-aware voice assistant — summarize, reply, translate based on what's on screen
Input MemoryRecent input history injected as LLM context for better continuity
Industry VocabularyChoose Medical, Legal, Finance & Accounting, or Software Technology terms for Apple Speech / Whisper biasing and output normalization; personal terms take priority
Edit RulesPersonal text replacement rules applied on every output
Language Style PresetsConcise / Formal / Casual / Custom prompt per language
Input History & StatsFull history with raw vs. processed comparison, word count stats, configurable retention
Bilingual UIChinese and English interface, independent of recognition language
Sound FeedbackAudio cues on recording start and stop
Guided OnboardingStep-by-step setup: permissions, model download, and first use

System Requirements

  • OS: macOS 26 (Tahoe) or later
  • Chip: Apple Silicon (M1 / M2 / M3 / M4)
  • Disk: ~400 MB minimum (Apple Speech + Qwen3-0.6B), up to ~4 GB with larger models

Installation

Download

Grab the latest .dmg from Releases, open it, and drag Utter.app to Applications.

"Cannot verify the developer" on first launch? The app is not notarized by Apple. Before first run, execute in Terminal:

xattr -cr /Applications/Utter.app

Or go to System Settings → Privacy & Security and click "Open Anyway".

Build from Source

# Build .app bundle + .dmg installer
bash scripts/build-app.sh
# Or for development
swift build
swift run OpenType
# Build, sign, and launch the development .app bundle
bash scripts/build-and-run.sh --verify
# Or open in Xcode
open Package.swift

Material changes follow the artifact-driven workflow in docs/sdlc/README.md. Before opening a pull request, run:

bash scripts/sdlc-checks.sh
bash scripts/ci-basic-checks.sh
swift test

The public app is Utter.app. The Swift package product remains OpenType so existing source integrations and upgrade paths continue to work.

First Run

  1. Launch Utter — it appears as a waveform icon in the menu bar
  2. The onboarding wizard guides you through permissions and model setup
  3. Grant Microphone and Accessibility permissions (required)
  4. Wait for the LLM model to download (~335 MB, one-time)
  5. Hold Fn to start dictating, release to stop and insert text

Permissions

PermissionPurposeRequired
MicrophoneAudio captureYes
AccessibilityGlobal hotkey + text injection (simulated paste)Yes
Speech RecognitionApple on-device ASR engineOnly if using Apple Speech
Screen RecordingOCR for screen context and Voice Command modeOptional
NetworkModel downloads; remote LLM API callsFirst run / remote LLM mode

Remote LLM Providers

Utter supports both OpenAI-compatible and Anthropic API formats:

ProviderAPI FormatBase URL
OpenAIOpenAIhttps://api.openai.com/v1
Anthropic ClaudeAnthropichttps://api.anthropic.com/v1
Google GeminiOpenAIhttps://generativelanguage.googleapis.com/v1beta/openai
OpenRouterOpenAIhttps://openrouter.ai/api/v1
SiliconFlowOpenAIhttps://api.siliconflow.cn/v1
Volcengine DoubaoOpenAIhttps://ark.cn-beijing.volces.com/api/v3
Alibaba BailianOpenAIhttps://dashscope.aliyuncs.com/compatible-mode/v1
MiniMax (China)OpenAIhttps://api.minimax.chat/v1
MiniMax (Global)OpenAIhttps://api.minimaxi.chat/v1

Local ASR Providers

ProviderLocal runtimeDefault model
Qwen3-ASRNative Swift + MLX on Apple Siliconmlx-community/Qwen3-ASR-1.7B-bf16

Qwen3-ASR does not call a hosted ASR API. The app downloads the selected model into the same model storage used by WhisperKit and runs inference locally through native Swift and MLX.

Project Structure

Sources/
├── App/ # Entry point, AppDelegate, AppState, VoicePipeline, AppIcon
├── Audio/ # Microphone capture (AVAudioEngine), sound playback
├── Config/ # AppSettings, ModelCatalog, RemoteModelConfig, Localization
├── Hotkey/ # Global hotkey via CGEvent tap
├── LLM/ # LLMEngine (MLX), RemoteLLMClient (OpenAI/Anthropic)
├── Output/ # Text injection (Accessibility API + clipboard paste)
├── Processing/ # TextProcessor, InputHistory, MemoryStore, personal and industry vocabulary
├── Prompts/ # PromptBuilder, prompt catalogs, style prompt presets
├── Screen/ # Screen OCR (ScreenCaptureKit + Vision)
├── Speech/ # SpeechEngine protocol, WhisperKit, Apple Speech, Doubao ASR, local ASR engines
├── UI/ # SwiftUI: MenuBar, Settings, Onboarding, Overlay, History, Models
└── Resources/ # Localization strings (en/zh-Hans), sounds, app icon
scripts/
├── build-and-run.sh # Build, sign, and launch a development .app bundle
├── build-app.sh # Build release .app bundle and .dmg installer
├── ci-basic-checks.sh # CI guardrails for linked files and resources
├── create-signing-cert.sh # Generate self-signed code signing certificate
├── generate-icon.swift # Generate AppIcon.icns from source PNG
├── test-industry-lexicons.sh # Validate vocabulary recall and non-target preservation
├── unit-test-coverage.sh # Run unit tests with coverage thresholds
└── validate-volc-asr.swift # Validate Volcengine ASR configuration manually

Tech Stack

  • WhisperKit — offline Whisper speech recognition
  • mlx-swift-lm — local LLM inference on Apple Silicon (Qwen2.5 / Qwen3)
  • SwiftUI + AppKit — native macOS UI
  • ScreenCaptureKit + Vision — screen OCR
  • AVAudioEngine — low-latency microphone capture
  • Apple Speech Framework — on-device speech recognition

License

MIT


Made with care for Apple Silicon

About

Local AI-powered voice input for macOS menu bar

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Utter

Local AI-powered voice input for macOS menu bar


GitHub StarsGitHub ForksGitHub Issues

PlatformSwiftApple SiliconLicense

WhisperKitMLX

Website · 中文文档


Overview

Utter is a macOS menu bar app for AI-powered voice input and dictation. It supports both fully local on-device inference and remote LLM APIs. Press a hotkey to start recording, release to transcribe, and the result is typed directly into whatever app you're using.

Three output modes are available:

  • Verbatim — raw transcription, lowest latency
  • Smart Format — transcription cleaned up by an LLM (contextual filler removal, grammar fixes, structured formatting)
  • Voice Command — speak a command and get an AI-generated response based on screen context

Features

FeatureDescription
Multiple Speech EnginesApple Speech, WhisperKit, Doubao ASR, or Qwen3-ASR
Smart Text ProcessingLocal MLX Qwen2.5/Qwen3 or remote LLM infers spoken intent — contextual cleanup, "scratch that" restarts, self-correction handling, spoken punctuation, technical terms, numbers/ranges/units, and structured formatting
LLM-Owned Spoken FormattingSpoken casing, no-space dictation, identifiers, file paths, shortcuts, emoji, Markdown tasks, dates/times, quantities, units, formulas, fractions, and digit sequences are handled by the Smart Format / Voice Command prompts instead of local hardcoded rewrite rules
Voice Edit CommandsIn Voice Command mode, an LLM classifies safe structured actions for replacing, undoing, proofreading, titling, summarizing, drafting replies, making meeting notes, extracting key points/decisions/questions/risks/deadlines/owners/action items, rewriting tone, expanding, making tables/lists, or deleting the previous Utter insertion or selected text
Verbatim & Preview BoundaryVerbatim mode, streaming HUD, integration partials, and instant-insert drafts keep ASR text close to raw output with only dictionary, whitespace, duplicate, and non-speech-artifact cleanup
Remote LLM SupportOpenAI, Claude (Anthropic format), Gemini, OpenRouter, SiliconFlow, Doubao, Bailian, MiniMax (CN & Global)
Global HotkeyConfigurable key (Fn/Ctrl/Shift/Option) with long-press, double-tap, or single-tap activation
Translation DictationUse a dedicated hotkey chord to speak in one language and insert an English, Chinese, Japanese, Korean, Spanish, French, or German translation
Screen Context OCRCaptures on-screen text via ScreenCaptureKit + Vision to help the LLM correct homophones
Voice Command ModeScreen-aware voice assistant — summarize, reply, translate based on what's on screen
Input MemoryRecent input history injected as LLM context for better continuity
Industry VocabularyChoose Medical, Legal, Finance & Accounting, or Software Technology terms for Apple Speech / Whisper biasing and output normalization; personal terms take priority
Edit RulesPersonal text replacement rules applied on every output
Language Style PresetsConcise / Formal / Casual / Custom prompt per language
Input History & StatsFull history with raw vs. processed comparison, word count stats, configurable retention
Bilingual UIChinese and English interface, independent of recognition language
Sound FeedbackAudio cues on recording start and stop
Guided OnboardingStep-by-step setup: permissions, model download, and first use

System Requirements

  • OS: macOS 26 (Tahoe) or later
  • Chip: Apple Silicon (M1 / M2 / M3 / M4)
  • Disk: ~400 MB minimum (Apple Speech + Qwen3-0.6B), up to ~4 GB with larger models

Installation

Download

Grab the latest .dmg from Releases, open it, and drag Utter.app to Applications.

"Cannot verify the developer" on first launch? The app is not notarized by Apple. Before first run, execute in Terminal:

xattr -cr /Applications/Utter.app

Or go to System Settings → Privacy & Security and click "Open Anyway".

Build from Source

# Build .app bundle + .dmg installer
bash scripts/build-app.sh
# Or for development
swift build
swift run OpenType
# Build, sign, and launch the development .app bundle
bash scripts/build-and-run.sh --verify
# Or open in Xcode
open Package.swift

Material changes follow the artifact-driven workflow in docs/sdlc/README.md. Before opening a pull request, run:

bash scripts/sdlc-checks.sh
bash scripts/ci-basic-checks.sh
swift test

The public app is Utter.app. The Swift package product remains OpenType so existing source integrations and upgrade paths continue to work.

First Run

  1. Launch Utter — it appears as a waveform icon in the menu bar
  2. The onboarding wizard guides you through permissions and model setup
  3. Grant Microphone and Accessibility permissions (required)
  4. Wait for the LLM model to download (~335 MB, one-time)
  5. Hold Fn to start dictating, release to stop and insert text

Permissions

PermissionPurposeRequired
MicrophoneAudio captureYes
AccessibilityGlobal hotkey + text injection (simulated paste)Yes
Speech RecognitionApple on-device ASR engineOnly if using Apple Speech
Screen RecordingOCR for screen context and Voice Command modeOptional
NetworkModel downloads; remote LLM API callsFirst run / remote LLM mode

Remote LLM Providers

Utter supports both OpenAI-compatible and Anthropic API formats:

ProviderAPI FormatBase URL
OpenAIOpenAIhttps://api.openai.com/v1
Anthropic ClaudeAnthropichttps://api.anthropic.com/v1
Google GeminiOpenAIhttps://generativelanguage.googleapis.com/v1beta/openai
OpenRouterOpenAIhttps://openrouter.ai/api/v1
SiliconFlowOpenAIhttps://api.siliconflow.cn/v1
Volcengine DoubaoOpenAIhttps://ark.cn-beijing.volces.com/api/v3
Alibaba BailianOpenAIhttps://dashscope.aliyuncs.com/compatible-mode/v1
MiniMax (China)OpenAIhttps://api.minimax.chat/v1
MiniMax (Global)OpenAIhttps://api.minimaxi.chat/v1

Local ASR Providers

ProviderLocal runtimeDefault model
Qwen3-ASRNative Swift + MLX on Apple Siliconmlx-community/Qwen3-ASR-1.7B-bf16

Qwen3-ASR does not call a hosted ASR API. The app downloads the selected model into the same model storage used by WhisperKit and runs inference locally through native Swift and MLX.

Project Structure

Sources/
├── App/ # Entry point, AppDelegate, AppState, VoicePipeline, AppIcon
├── Audio/ # Microphone capture (AVAudioEngine), sound playback
├── Config/ # AppSettings, ModelCatalog, RemoteModelConfig, Localization
├── Hotkey/ # Global hotkey via CGEvent tap
├── LLM/ # LLMEngine (MLX), RemoteLLMClient (OpenAI/Anthropic)
├── Output/ # Text injection (Accessibility API + clipboard paste)
├── Processing/ # TextProcessor, InputHistory, MemoryStore, personal and industry vocabulary
├── Prompts/ # PromptBuilder, prompt catalogs, style prompt presets
├── Screen/ # Screen OCR (ScreenCaptureKit + Vision)
├── Speech/ # SpeechEngine protocol, WhisperKit, Apple Speech, Doubao ASR, local ASR engines
├── UI/ # SwiftUI: MenuBar, Settings, Onboarding, Overlay, History, Models
└── Resources/ # Localization strings (en/zh-Hans), sounds, app icon
scripts/
├── build-and-run.sh # Build, sign, and launch a development .app bundle
├── build-app.sh # Build release .app bundle and .dmg installer
├── ci-basic-checks.sh # CI guardrails for linked files and resources
├── create-signing-cert.sh # Generate self-signed code signing certificate
├── generate-icon.swift # Generate AppIcon.icns from source PNG
├── test-industry-lexicons.sh # Validate vocabulary recall and non-target preservation
├── unit-test-coverage.sh # Run unit tests with coverage thresholds
└── validate-volc-asr.swift # Validate Volcengine ASR configuration manually

Tech Stack

  • WhisperKit — offline Whisper speech recognition
  • mlx-swift-lm — local LLM inference on Apple Silicon (Qwen2.5 / Qwen3)
  • SwiftUI + AppKit — native macOS UI
  • ScreenCaptureKit + Vision — screen OCR
  • AVAudioEngine — low-latency microphone capture
  • Apple Speech Framework — on-device speech recognition

License

MIT


Made with care for Apple Silicon

About

Local AI-powered voice input for macOS menu bar

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Utter

Local AI-powered voice input for macOS menu bar


GitHub StarsGitHub ForksGitHub Issues

PlatformSwiftApple SiliconLicense

WhisperKitMLX

Website · 中文文档


Overview

Utter is a macOS menu bar app for AI-powered voice input and dictation. It supports both fully local on-device inference and remote LLM APIs. Press a hotkey to start recording, release to transcribe, and the result is typed directly into whatever app you're using.

Three output modes are available:

  • Verbatim — raw transcription, lowest latency
  • Smart Format — transcription cleaned up by an LLM (contextual filler removal, grammar fixes, structured formatting)
  • Voice Command — speak a command and get an AI-generated response based on screen context

Features

FeatureDescription
Multiple Speech EnginesApple Speech, WhisperKit, Doubao ASR, or Qwen3-ASR
Smart Text ProcessingLocal MLX Qwen2.5/Qwen3 or remote LLM infers spoken intent — contextual cleanup, "scratch that" restarts, self-correction handling, spoken punctuation, technical terms, numbers/ranges/units, and structured formatting
LLM-Owned Spoken FormattingSpoken casing, no-space dictation, identifiers, file paths, shortcuts, emoji, Markdown tasks, dates/times, quantities, units, formulas, fractions, and digit sequences are handled by the Smart Format / Voice Command prompts instead of local hardcoded rewrite rules
Voice Edit CommandsIn Voice Command mode, an LLM classifies safe structured actions for replacing, undoing, proofreading, titling, summarizing, drafting replies, making meeting notes, extracting key points/decisions/questions/risks/deadlines/owners/action items, rewriting tone, expanding, making tables/lists, or deleting the previous Utter insertion or selected text
Verbatim & Preview BoundaryVerbatim mode, streaming HUD, integration partials, and instant-insert drafts keep ASR text close to raw output with only dictionary, whitespace, duplicate, and non-speech-artifact cleanup
Remote LLM SupportOpenAI, Claude (Anthropic format), Gemini, OpenRouter, SiliconFlow, Doubao, Bailian, MiniMax (CN & Global)
Global HotkeyConfigurable key (Fn/Ctrl/Shift/Option) with long-press, double-tap, or single-tap activation
Translation DictationUse a dedicated hotkey chord to speak in one language and insert an English, Chinese, Japanese, Korean, Spanish, French, or German translation
Screen Context OCRCaptures on-screen text via ScreenCaptureKit + Vision to help the LLM correct homophones
Voice Command ModeScreen-aware voice assistant — summarize, reply, translate based on what's on screen
Input MemoryRecent input history injected as LLM context for better continuity
Industry VocabularyChoose Medical, Legal, Finance & Accounting, or Software Technology terms for Apple Speech / Whisper biasing and output normalization; personal terms take priority
Edit RulesPersonal text replacement rules applied on every output
Language Style PresetsConcise / Formal / Casual / Custom prompt per language
Input History & StatsFull history with raw vs. processed comparison, word count stats, configurable retention
Bilingual UIChinese and English interface, independent of recognition language
Sound FeedbackAudio cues on recording start and stop
Guided OnboardingStep-by-step setup: permissions, model download, and first use

System Requirements

  • OS: macOS 26 (Tahoe) or later
  • Chip: Apple Silicon (M1 / M2 / M3 / M4)
  • Disk: ~400 MB minimum (Apple Speech + Qwen3-0.6B), up to ~4 GB with larger models

Installation

Download

Grab the latest .dmg from Releases, open it, and drag Utter.app to Applications.

"Cannot verify the developer" on first launch? The app is not notarized by Apple. Before first run, execute in Terminal:

xattr -cr /Applications/Utter.app

Or go to System Settings → Privacy & Security and click "Open Anyway".

Build from Source

# Build .app bundle + .dmg installer
bash scripts/build-app.sh
# Or for development
swift build
swift run OpenType
# Build, sign, and launch the development .app bundle
bash scripts/build-and-run.sh --verify
# Or open in Xcode
open Package.swift

Material changes follow the artifact-driven workflow in docs/sdlc/README.md. Before opening a pull request, run:

bash scripts/sdlc-checks.sh
bash scripts/ci-basic-checks.sh
swift test

The public app is Utter.app. The Swift package product remains OpenType so existing source integrations and upgrade paths continue to work.

First Run

  1. Launch Utter — it appears as a waveform icon in the menu bar
  2. The onboarding wizard guides you through permissions and model setup
  3. Grant Microphone and Accessibility permissions (required)
  4. Wait for the LLM model to download (~335 MB, one-time)
  5. Hold Fn to start dictating, release to stop and insert text

Permissions

PermissionPurposeRequired
MicrophoneAudio captureYes
AccessibilityGlobal hotkey + text injection (simulated paste)Yes
Speech RecognitionApple on-device ASR engineOnly if using Apple Speech
Screen RecordingOCR for screen context and Voice Command modeOptional
NetworkModel downloads; remote LLM API callsFirst run / remote LLM mode

Remote LLM Providers

Utter supports both OpenAI-compatible and Anthropic API formats:

ProviderAPI FormatBase URL
OpenAIOpenAIhttps://api.openai.com/v1
Anthropic ClaudeAnthropichttps://api.anthropic.com/v1
Google GeminiOpenAIhttps://generativelanguage.googleapis.com/v1beta/openai
OpenRouterOpenAIhttps://openrouter.ai/api/v1
SiliconFlowOpenAIhttps://api.siliconflow.cn/v1
Volcengine DoubaoOpenAIhttps://ark.cn-beijing.volces.com/api/v3
Alibaba BailianOpenAIhttps://dashscope.aliyuncs.com/compatible-mode/v1
MiniMax (China)OpenAIhttps://api.minimax.chat/v1
MiniMax (Global)OpenAIhttps://api.minimaxi.chat/v1

Local ASR Providers

ProviderLocal runtimeDefault model
Qwen3-ASRNative Swift + MLX on Apple Siliconmlx-community/Qwen3-ASR-1.7B-bf16

Qwen3-ASR does not call a hosted ASR API. The app downloads the selected model into the same model storage used by WhisperKit and runs inference locally through native Swift and MLX.

Project Structure

Sources/
├── App/ # Entry point, AppDelegate, AppState, VoicePipeline, AppIcon
├── Audio/ # Microphone capture (AVAudioEngine), sound playback
├── Config/ # AppSettings, ModelCatalog, RemoteModelConfig, Localization
├── Hotkey/ # Global hotkey via CGEvent tap
├── LLM/ # LLMEngine (MLX), RemoteLLMClient (OpenAI/Anthropic)
├── Output/ # Text injection (Accessibility API + clipboard paste)
├── Processing/ # TextProcessor, InputHistory, MemoryStore, personal and industry vocabulary
├── Prompts/ # PromptBuilder, prompt catalogs, style prompt presets
├── Screen/ # Screen OCR (ScreenCaptureKit + Vision)
├── Speech/ # SpeechEngine protocol, WhisperKit, Apple Speech, Doubao ASR, local ASR engines
├── UI/ # SwiftUI: MenuBar, Settings, Onboarding, Overlay, History, Models
└── Resources/ # Localization strings (en/zh-Hans), sounds, app icon
scripts/
├── build-and-run.sh # Build, sign, and launch a development .app bundle
├── build-app.sh # Build release .app bundle and .dmg installer
├── ci-basic-checks.sh # CI guardrails for linked files and resources
├── create-signing-cert.sh # Generate self-signed code signing certificate
├── generate-icon.swift # Generate AppIcon.icns from source PNG
├── test-industry-lexicons.sh # Validate vocabulary recall and non-target preservation
├── unit-test-coverage.sh # Run unit tests with coverage thresholds
└── validate-volc-asr.swift # Validate Volcengine ASR configuration manually

Tech Stack

  • WhisperKit — offline Whisper speech recognition
  • mlx-swift-lm — local LLM inference on Apple Silicon (Qwen2.5 / Qwen3)
  • SwiftUI + AppKit — native macOS UI
  • ScreenCaptureKit + Vision — screen OCR
  • AVAudioEngine — low-latency microphone capture
  • Apple Speech Framework — on-device speech recognition

License

MIT


Made with care for Apple Silicon

About

Local AI-powered voice input for macOS menu bar

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Utter

Local AI-powered voice input for macOS menu bar


GitHub StarsGitHub ForksGitHub Issues

PlatformSwiftApple SiliconLicense

WhisperKitMLX

Website · 中文文档


Overview

Utter is a macOS menu bar app for AI-powered voice input and dictation. It supports both fully local on-device inference and remote LLM APIs. Press a hotkey to start recording, release to transcribe, and the result is typed directly into whatever app you're using.

Three output modes are available:

  • Verbatim — raw transcription, lowest latency
  • Smart Format — transcription cleaned up by an LLM (contextual filler removal, grammar fixes, structured formatting)
  • Voice Command — speak a command and get an AI-generated response based on screen context

Features

FeatureDescription
Multiple Speech EnginesApple Speech, WhisperKit, Doubao ASR, or Qwen3-ASR
Smart Text ProcessingLocal MLX Qwen2.5/Qwen3 or remote LLM infers spoken intent — contextual cleanup, "scratch that" restarts, self-correction handling, spoken punctuation, technical terms, numbers/ranges/units, and structured formatting
LLM-Owned Spoken FormattingSpoken casing, no-space dictation, identifiers, file paths, shortcuts, emoji, Markdown tasks, dates/times, quantities, units, formulas, fractions, and digit sequences are handled by the Smart Format / Voice Command prompts instead of local hardcoded rewrite rules
Voice Edit CommandsIn Voice Command mode, an LLM classifies safe structured actions for replacing, undoing, proofreading, titling, summarizing, drafting replies, making meeting notes, extracting key points/decisions/questions/risks/deadlines/owners/action items, rewriting tone, expanding, making tables/lists, or deleting the previous Utter insertion or selected text
Verbatim & Preview BoundaryVerbatim mode, streaming HUD, integration partials, and instant-insert drafts keep ASR text close to raw output with only dictionary, whitespace, duplicate, and non-speech-artifact cleanup
Remote LLM SupportOpenAI, Claude (Anthropic format), Gemini, OpenRouter, SiliconFlow, Doubao, Bailian, MiniMax (CN & Global)
Global HotkeyConfigurable key (Fn/Ctrl/Shift/Option) with long-press, double-tap, or single-tap activation
Translation DictationUse a dedicated hotkey chord to speak in one language and insert an English, Chinese, Japanese, Korean, Spanish, French, or German translation
Screen Context OCRCaptures on-screen text via ScreenCaptureKit + Vision to help the LLM correct homophones
Voice Command ModeScreen-aware voice assistant — summarize, reply, translate based on what's on screen
Input MemoryRecent input history injected as LLM context for better continuity
Industry VocabularyChoose Medical, Legal, Finance & Accounting, or Software Technology terms for Apple Speech / Whisper biasing and output normalization; personal terms take priority
Edit RulesPersonal text replacement rules applied on every output
Language Style PresetsConcise / Formal / Casual / Custom prompt per language
Input History & StatsFull history with raw vs. processed comparison, word count stats, configurable retention
Bilingual UIChinese and English interface, independent of recognition language
Sound FeedbackAudio cues on recording start and stop
Guided OnboardingStep-by-step setup: permissions, model download, and first use

System Requirements

  • OS: macOS 26 (Tahoe) or later
  • Chip: Apple Silicon (M1 / M2 / M3 / M4)
  • Disk: ~400 MB minimum (Apple Speech + Qwen3-0.6B), up to ~4 GB with larger models

Installation

Download

Grab the latest .dmg from Releases, open it, and drag Utter.app to Applications.

"Cannot verify the developer" on first launch? The app is not notarized by Apple. Before first run, execute in Terminal:

xattr -cr /Applications/Utter.app

Or go to System Settings → Privacy & Security and click "Open Anyway".

Build from Source

# Build .app bundle + .dmg installer
bash scripts/build-app.sh
# Or for development
swift build
swift run OpenType
# Build, sign, and launch the development .app bundle
bash scripts/build-and-run.sh --verify
# Or open in Xcode
open Package.swift

Material changes follow the artifact-driven workflow in docs/sdlc/README.md. Before opening a pull request, run:

bash scripts/sdlc-checks.sh
bash scripts/ci-basic-checks.sh
swift test

The public app is Utter.app. The Swift package product remains OpenType so existing source integrations and upgrade paths continue to work.

First Run

  1. Launch Utter — it appears as a waveform icon in the menu bar
  2. The onboarding wizard guides you through permissions and model setup
  3. Grant Microphone and Accessibility permissions (required)
  4. Wait for the LLM model to download (~335 MB, one-time)
  5. Hold Fn to start dictating, release to stop and insert text

Permissions

PermissionPurposeRequired
MicrophoneAudio captureYes
AccessibilityGlobal hotkey + text injection (simulated paste)Yes
Speech RecognitionApple on-device ASR engineOnly if using Apple Speech
Screen RecordingOCR for screen context and Voice Command modeOptional
NetworkModel downloads; remote LLM API callsFirst run / remote LLM mode

Remote LLM Providers

Utter supports both OpenAI-compatible and Anthropic API formats:

ProviderAPI FormatBase URL
OpenAIOpenAIhttps://api.openai.com/v1
Anthropic ClaudeAnthropichttps://api.anthropic.com/v1
Google GeminiOpenAIhttps://generativelanguage.googleapis.com/v1beta/openai
OpenRouterOpenAIhttps://openrouter.ai/api/v1
SiliconFlowOpenAIhttps://api.siliconflow.cn/v1
Volcengine DoubaoOpenAIhttps://ark.cn-beijing.volces.com/api/v3
Alibaba BailianOpenAIhttps://dashscope.aliyuncs.com/compatible-mode/v1
MiniMax (China)OpenAIhttps://api.minimax.chat/v1
MiniMax (Global)OpenAIhttps://api.minimaxi.chat/v1

Local ASR Providers

ProviderLocal runtimeDefault model
Qwen3-ASRNative Swift + MLX on Apple Siliconmlx-community/Qwen3-ASR-1.7B-bf16

Qwen3-ASR does not call a hosted ASR API. The app downloads the selected model into the same model storage used by WhisperKit and runs inference locally through native Swift and MLX.

Project Structure

Sources/
├── App/ # Entry point, AppDelegate, AppState, VoicePipeline, AppIcon
├── Audio/ # Microphone capture (AVAudioEngine), sound playback
├── Config/ # AppSettings, ModelCatalog, RemoteModelConfig, Localization
├── Hotkey/ # Global hotkey via CGEvent tap
├── LLM/ # LLMEngine (MLX), RemoteLLMClient (OpenAI/Anthropic)
├── Output/ # Text injection (Accessibility API + clipboard paste)
├── Processing/ # TextProcessor, InputHistory, MemoryStore, personal and industry vocabulary
├── Prompts/ # PromptBuilder, prompt catalogs, style prompt presets
├── Screen/ # Screen OCR (ScreenCaptureKit + Vision)
├── Speech/ # SpeechEngine protocol, WhisperKit, Apple Speech, Doubao ASR, local ASR engines
├── UI/ # SwiftUI: MenuBar, Settings, Onboarding, Overlay, History, Models
└── Resources/ # Localization strings (en/zh-Hans), sounds, app icon
scripts/
├── build-and-run.sh # Build, sign, and launch a development .app bundle
├── build-app.sh # Build release .app bundle and .dmg installer
├── ci-basic-checks.sh # CI guardrails for linked files and resources
├── create-signing-cert.sh # Generate self-signed code signing certificate
├── generate-icon.swift # Generate AppIcon.icns from source PNG
├── test-industry-lexicons.sh # Validate vocabulary recall and non-target preservation
├── unit-test-coverage.sh # Run unit tests with coverage thresholds
└── validate-volc-asr.swift # Validate Volcengine ASR configuration manually

Tech Stack

  • WhisperKit — offline Whisper speech recognition
  • mlx-swift-lm — local LLM inference on Apple Silicon (Qwen2.5 / Qwen3)
  • SwiftUI + AppKit — native macOS UI
  • ScreenCaptureKit + Vision — screen OCR
  • AVAudioEngine — low-latency microphone capture
  • Apple Speech Framework — on-device speech recognition

License

MIT


Made with care for Apple Silicon

About

Local AI-powered voice input for macOS menu bar

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Utter

Local AI-powered voice input for macOS menu bar


GitHub StarsGitHub ForksGitHub Issues

PlatformSwiftApple SiliconLicense

WhisperKitMLX

Website · 中文文档


Overview

Utter is a macOS menu bar app for AI-powered voice input and dictation. It supports both fully local on-device inference and remote LLM APIs. Press a hotkey to start recording, release to transcribe, and the result is typed directly into whatever app you're using.

Three output modes are available:

  • Verbatim — raw transcription, lowest latency
  • Smart Format — transcription cleaned up by an LLM (contextual filler removal, grammar fixes, structured formatting)
  • Voice Command — speak a command and get an AI-generated response based on screen context

Features

FeatureDescription
Multiple Speech EnginesApple Speech, WhisperKit, Doubao ASR, or Qwen3-ASR
Smart Text ProcessingLocal MLX Qwen2.5/Qwen3 or remote LLM infers spoken intent — contextual cleanup, "scratch that" restarts, self-correction handling, spoken punctuation, technical terms, numbers/ranges/units, and structured formatting
LLM-Owned Spoken FormattingSpoken casing, no-space dictation, identifiers, file paths, shortcuts, emoji, Markdown tasks, dates/times, quantities, units, formulas, fractions, and digit sequences are handled by the Smart Format / Voice Command prompts instead of local hardcoded rewrite rules
Voice Edit CommandsIn Voice Command mode, an LLM classifies safe structured actions for replacing, undoing, proofreading, titling, summarizing, drafting replies, making meeting notes, extracting key points/decisions/questions/risks/deadlines/owners/action items, rewriting tone, expanding, making tables/lists, or deleting the previous Utter insertion or selected text
Verbatim & Preview BoundaryVerbatim mode, streaming HUD, integration partials, and instant-insert drafts keep ASR text close to raw output with only dictionary, whitespace, duplicate, and non-speech-artifact cleanup
Remote LLM SupportOpenAI, Claude (Anthropic format), Gemini, OpenRouter, SiliconFlow, Doubao, Bailian, MiniMax (CN & Global)
Global HotkeyConfigurable key (Fn/Ctrl/Shift/Option) with long-press, double-tap, or single-tap activation
Translation DictationUse a dedicated hotkey chord to speak in one language and insert an English, Chinese, Japanese, Korean, Spanish, French, or German translation
Screen Context OCRCaptures on-screen text via ScreenCaptureKit + Vision to help the LLM correct homophones
Voice Command ModeScreen-aware voice assistant — summarize, reply, translate based on what's on screen
Input MemoryRecent input history injected as LLM context for better continuity
Industry VocabularyChoose Medical, Legal, Finance & Accounting, or Software Technology terms for Apple Speech / Whisper biasing and output normalization; personal terms take priority
Edit RulesPersonal text replacement rules applied on every output
Language Style PresetsConcise / Formal / Casual / Custom prompt per language
Input History & StatsFull history with raw vs. processed comparison, word count stats, configurable retention
Bilingual UIChinese and English interface, independent of recognition language
Sound FeedbackAudio cues on recording start and stop
Guided OnboardingStep-by-step setup: permissions, model download, and first use

System Requirements

  • OS: macOS 26 (Tahoe) or later
  • Chip: Apple Silicon (M1 / M2 / M3 / M4)
  • Disk: ~400 MB minimum (Apple Speech + Qwen3-0.6B), up to ~4 GB with larger models

Installation

Download

Grab the latest .dmg from Releases, open it, and drag Utter.app to Applications.

"Cannot verify the developer" on first launch? The app is not notarized by Apple. Before first run, execute in Terminal:

xattr -cr /Applications/Utter.app

Or go to System Settings → Privacy & Security and click "Open Anyway".

Build from Source

# Build .app bundle + .dmg installer
bash scripts/build-app.sh
# Or for development
swift build
swift run OpenType
# Build, sign, and launch the development .app bundle
bash scripts/build-and-run.sh --verify
# Or open in Xcode
open Package.swift

Material changes follow the artifact-driven workflow in docs/sdlc/README.md. Before opening a pull request, run:

bash scripts/sdlc-checks.sh
bash scripts/ci-basic-checks.sh
swift test

The public app is Utter.app. The Swift package product remains OpenType so existing source integrations and upgrade paths continue to work.

First Run

  1. Launch Utter — it appears as a waveform icon in the menu bar
  2. The onboarding wizard guides you through permissions and model setup
  3. Grant Microphone and Accessibility permissions (required)
  4. Wait for the LLM model to download (~335 MB, one-time)
  5. Hold Fn to start dictating, release to stop and insert text

Permissions

PermissionPurposeRequired
MicrophoneAudio captureYes
AccessibilityGlobal hotkey + text injection (simulated paste)Yes
Speech RecognitionApple on-device ASR engineOnly if using Apple Speech
Screen RecordingOCR for screen context and Voice Command modeOptional
NetworkModel downloads; remote LLM API callsFirst run / remote LLM mode

Remote LLM Providers

Utter supports both OpenAI-compatible and Anthropic API formats:

ProviderAPI FormatBase URL
OpenAIOpenAIhttps://api.openai.com/v1
Anthropic ClaudeAnthropichttps://api.anthropic.com/v1
Google GeminiOpenAIhttps://generativelanguage.googleapis.com/v1beta/openai
OpenRouterOpenAIhttps://openrouter.ai/api/v1
SiliconFlowOpenAIhttps://api.siliconflow.cn/v1
Volcengine DoubaoOpenAIhttps://ark.cn-beijing.volces.com/api/v3
Alibaba BailianOpenAIhttps://dashscope.aliyuncs.com/compatible-mode/v1
MiniMax (China)OpenAIhttps://api.minimax.chat/v1
MiniMax (Global)OpenAIhttps://api.minimaxi.chat/v1

Local ASR Providers

ProviderLocal runtimeDefault model
Qwen3-ASRNative Swift + MLX on Apple Siliconmlx-community/Qwen3-ASR-1.7B-bf16

Qwen3-ASR does not call a hosted ASR API. The app downloads the selected model into the same model storage used by WhisperKit and runs inference locally through native Swift and MLX.

Project Structure

Sources/
├── App/ # Entry point, AppDelegate, AppState, VoicePipeline, AppIcon
├── Audio/ # Microphone capture (AVAudioEngine), sound playback
├── Config/ # AppSettings, ModelCatalog, RemoteModelConfig, Localization
├── Hotkey/ # Global hotkey via CGEvent tap
├── LLM/ # LLMEngine (MLX), RemoteLLMClient (OpenAI/Anthropic)
├── Output/ # Text injection (Accessibility API + clipboard paste)
├── Processing/ # TextProcessor, InputHistory, MemoryStore, personal and industry vocabulary
├── Prompts/ # PromptBuilder, prompt catalogs, style prompt presets
├── Screen/ # Screen OCR (ScreenCaptureKit + Vision)
├── Speech/ # SpeechEngine protocol, WhisperKit, Apple Speech, Doubao ASR, local ASR engines
├── UI/ # SwiftUI: MenuBar, Settings, Onboarding, Overlay, History, Models
└── Resources/ # Localization strings (en/zh-Hans), sounds, app icon
scripts/
├── build-and-run.sh # Build, sign, and launch a development .app bundle
├── build-app.sh # Build release .app bundle and .dmg installer
├── ci-basic-checks.sh # CI guardrails for linked files and resources
├── create-signing-cert.sh # Generate self-signed code signing certificate
├── generate-icon.swift # Generate AppIcon.icns from source PNG
├── test-industry-lexicons.sh # Validate vocabulary recall and non-target preservation
├── unit-test-coverage.sh # Run unit tests with coverage thresholds
└── validate-volc-asr.swift # Validate Volcengine ASR configuration manually

Tech Stack

  • WhisperKit — offline Whisper speech recognition
  • mlx-swift-lm — local LLM inference on Apple Silicon (Qwen2.5 / Qwen3)
  • SwiftUI + AppKit — native macOS UI
  • ScreenCaptureKit + Vision — screen OCR
  • AVAudioEngine — low-latency microphone capture
  • Apple Speech Framework — on-device speech recognition

License

MIT


Made with care for Apple Silicon

About

Local AI-powered voice input for macOS menu bar

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Utter

Local AI-powered voice input for macOS menu bar


GitHub StarsGitHub ForksGitHub Issues

PlatformSwiftApple SiliconLicense

WhisperKitMLX

Website · 中文文档


Overview

Utter is a macOS menu bar app for AI-powered voice input and dictation. It supports both fully local on-device inference and remote LLM APIs. Press a hotkey to start recording, release to transcribe, and the result is typed directly into whatever app you're using.

Three output modes are available:

  • Verbatim — raw transcription, lowest latency
  • Smart Format — transcription cleaned up by an LLM (contextual filler removal, grammar fixes, structured formatting)
  • Voice Command — speak a command and get an AI-generated response based on screen context

Features

FeatureDescription
Multiple Speech EnginesApple Speech, WhisperKit, Doubao ASR, or Qwen3-ASR
Smart Text ProcessingLocal MLX Qwen2.5/Qwen3 or remote LLM infers spoken intent — contextual cleanup, "scratch that" restarts, self-correction handling, spoken punctuation, technical terms, numbers/ranges/units, and structured formatting
LLM-Owned Spoken FormattingSpoken casing, no-space dictation, identifiers, file paths, shortcuts, emoji, Markdown tasks, dates/times, quantities, units, formulas, fractions, and digit sequences are handled by the Smart Format / Voice Command prompts instead of local hardcoded rewrite rules
Voice Edit CommandsIn Voice Command mode, an LLM classifies safe structured actions for replacing, undoing, proofreading, titling, summarizing, drafting replies, making meeting notes, extracting key points/decisions/questions/risks/deadlines/owners/action items, rewriting tone, expanding, making tables/lists, or deleting the previous Utter insertion or selected text
Verbatim & Preview BoundaryVerbatim mode, streaming HUD, integration partials, and instant-insert drafts keep ASR text close to raw output with only dictionary, whitespace, duplicate, and non-speech-artifact cleanup
Remote LLM SupportOpenAI, Claude (Anthropic format), Gemini, OpenRouter, SiliconFlow, Doubao, Bailian, MiniMax (CN & Global)
Global HotkeyConfigurable key (Fn/Ctrl/Shift/Option) with long-press, double-tap, or single-tap activation
Translation DictationUse a dedicated hotkey chord to speak in one language and insert an English, Chinese, Japanese, Korean, Spanish, French, or German translation
Screen Context OCRCaptures on-screen text via ScreenCaptureKit + Vision to help the LLM correct homophones
Voice Command ModeScreen-aware voice assistant — summarize, reply, translate based on what's on screen
Input MemoryRecent input history injected as LLM context for better continuity
Industry VocabularyChoose Medical, Legal, Finance & Accounting, or Software Technology terms for Apple Speech / Whisper biasing and output normalization; personal terms take priority
Edit RulesPersonal text replacement rules applied on every output
Language Style PresetsConcise / Formal / Casual / Custom prompt per language
Input History & StatsFull history with raw vs. processed comparison, word count stats, configurable retention
Bilingual UIChinese and English interface, independent of recognition language
Sound FeedbackAudio cues on recording start and stop
Guided OnboardingStep-by-step setup: permissions, model download, and first use

System Requirements

  • OS: macOS 26 (Tahoe) or later
  • Chip: Apple Silicon (M1 / M2 / M3 / M4)
  • Disk: ~400 MB minimum (Apple Speech + Qwen3-0.6B), up to ~4 GB with larger models

Installation

Download

Grab the latest .dmg from Releases, open it, and drag Utter.app to Applications.

"Cannot verify the developer" on first launch? The app is not notarized by Apple. Before first run, execute in Terminal:

xattr -cr /Applications/Utter.app

Or go to System Settings → Privacy & Security and click "Open Anyway".

Build from Source

# Build .app bundle + .dmg installer
bash scripts/build-app.sh
# Or for development
swift build
swift run OpenType
# Build, sign, and launch the development .app bundle
bash scripts/build-and-run.sh --verify
# Or open in Xcode
open Package.swift

Material changes follow the artifact-driven workflow in docs/sdlc/README.md. Before opening a pull request, run:

bash scripts/sdlc-checks.sh
bash scripts/ci-basic-checks.sh
swift test

The public app is Utter.app. The Swift package product remains OpenType so existing source integrations and upgrade paths continue to work.

First Run

  1. Launch Utter — it appears as a waveform icon in the menu bar
  2. The onboarding wizard guides you through permissions and model setup
  3. Grant Microphone and Accessibility permissions (required)
  4. Wait for the LLM model to download (~335 MB, one-time)
  5. Hold Fn to start dictating, release to stop and insert text

Permissions

PermissionPurposeRequired
MicrophoneAudio captureYes
AccessibilityGlobal hotkey + text injection (simulated paste)Yes
Speech RecognitionApple on-device ASR engineOnly if using Apple Speech
Screen RecordingOCR for screen context and Voice Command modeOptional
NetworkModel downloads; remote LLM API callsFirst run / remote LLM mode

Remote LLM Providers

Utter supports both OpenAI-compatible and Anthropic API formats:

ProviderAPI FormatBase URL
OpenAIOpenAIhttps://api.openai.com/v1
Anthropic ClaudeAnthropichttps://api.anthropic.com/v1
Google GeminiOpenAIhttps://generativelanguage.googleapis.com/v1beta/openai
OpenRouterOpenAIhttps://openrouter.ai/api/v1
SiliconFlowOpenAIhttps://api.siliconflow.cn/v1
Volcengine DoubaoOpenAIhttps://ark.cn-beijing.volces.com/api/v3
Alibaba BailianOpenAIhttps://dashscope.aliyuncs.com/compatible-mode/v1
MiniMax (China)OpenAIhttps://api.minimax.chat/v1
MiniMax (Global)OpenAIhttps://api.minimaxi.chat/v1

Local ASR Providers

ProviderLocal runtimeDefault model
Qwen3-ASRNative Swift + MLX on Apple Siliconmlx-community/Qwen3-ASR-1.7B-bf16

Qwen3-ASR does not call a hosted ASR API. The app downloads the selected model into the same model storage used by WhisperKit and runs inference locally through native Swift and MLX.

Project Structure

Sources/
├── App/ # Entry point, AppDelegate, AppState, VoicePipeline, AppIcon
├── Audio/ # Microphone capture (AVAudioEngine), sound playback
├── Config/ # AppSettings, ModelCatalog, RemoteModelConfig, Localization
├── Hotkey/ # Global hotkey via CGEvent tap
├── LLM/ # LLMEngine (MLX), RemoteLLMClient (OpenAI/Anthropic)
├── Output/ # Text injection (Accessibility API + clipboard paste)
├── Processing/ # TextProcessor, InputHistory, MemoryStore, personal and industry vocabulary
├── Prompts/ # PromptBuilder, prompt catalogs, style prompt presets
├── Screen/ # Screen OCR (ScreenCaptureKit + Vision)
├── Speech/ # SpeechEngine protocol, WhisperKit, Apple Speech, Doubao ASR, local ASR engines
├── UI/ # SwiftUI: MenuBar, Settings, Onboarding, Overlay, History, Models
└── Resources/ # Localization strings (en/zh-Hans), sounds, app icon
scripts/
├── build-and-run.sh # Build, sign, and launch a development .app bundle
├── build-app.sh # Build release .app bundle and .dmg installer
├── ci-basic-checks.sh # CI guardrails for linked files and resources
├── create-signing-cert.sh # Generate self-signed code signing certificate
├── generate-icon.swift # Generate AppIcon.icns from source PNG
├── test-industry-lexicons.sh # Validate vocabulary recall and non-target preservation
├── unit-test-coverage.sh # Run unit tests with coverage thresholds
└── validate-volc-asr.swift # Validate Volcengine ASR configuration manually

Tech Stack

  • WhisperKit — offline Whisper speech recognition
  • mlx-swift-lm — local LLM inference on Apple Silicon (Qwen2.5 / Qwen3)
  • SwiftUI + AppKit — native macOS UI
  • ScreenCaptureKit + Vision — screen OCR
  • AVAudioEngine — low-latency microphone capture
  • Apple Speech Framework — on-device speech recognition

License

MIT


Made with care for Apple Silicon

About

Local AI-powered voice input for macOS menu bar

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Utter

Local AI-powered voice input for macOS menu bar


GitHub StarsGitHub ForksGitHub Issues

PlatformSwiftApple SiliconLicense

WhisperKitMLX

Website · 中文文档


Overview

Utter is a macOS menu bar app for AI-powered voice input and dictation. It supports both fully local on-device inference and remote LLM APIs. Press a hotkey to start recording, release to transcribe, and the result is typed directly into whatever app you're using.

Three output modes are available:

  • Verbatim — raw transcription, lowest latency
  • Smart Format — transcription cleaned up by an LLM (contextual filler removal, grammar fixes, structured formatting)
  • Voice Command — speak a command and get an AI-generated response based on screen context

Features

FeatureDescription
Multiple Speech EnginesApple Speech, WhisperKit, Doubao ASR, or Qwen3-ASR
Smart Text ProcessingLocal MLX Qwen2.5/Qwen3 or remote LLM infers spoken intent — contextual cleanup, "scratch that" restarts, self-correction handling, spoken punctuation, technical terms, numbers/ranges/units, and structured formatting
LLM-Owned Spoken FormattingSpoken casing, no-space dictation, identifiers, file paths, shortcuts, emoji, Markdown tasks, dates/times, quantities, units, formulas, fractions, and digit sequences are handled by the Smart Format / Voice Command prompts instead of local hardcoded rewrite rules
Voice Edit CommandsIn Voice Command mode, an LLM classifies safe structured actions for replacing, undoing, proofreading, titling, summarizing, drafting replies, making meeting notes, extracting key points/decisions/questions/risks/deadlines/owners/action items, rewriting tone, expanding, making tables/lists, or deleting the previous Utter insertion or selected text
Verbatim & Preview BoundaryVerbatim mode, streaming HUD, integration partials, and instant-insert drafts keep ASR text close to raw output with only dictionary, whitespace, duplicate, and non-speech-artifact cleanup
Remote LLM SupportOpenAI, Claude (Anthropic format), Gemini, OpenRouter, SiliconFlow, Doubao, Bailian, MiniMax (CN & Global)
Global HotkeyConfigurable key (Fn/Ctrl/Shift/Option) with long-press, double-tap, or single-tap activation
Translation DictationUse a dedicated hotkey chord to speak in one language and insert an English, Chinese, Japanese, Korean, Spanish, French, or German translation
Screen Context OCRCaptures on-screen text via ScreenCaptureKit + Vision to help the LLM correct homophones
Voice Command ModeScreen-aware voice assistant — summarize, reply, translate based on what's on screen
Input MemoryRecent input history injected as LLM context for better continuity
Industry VocabularyChoose Medical, Legal, Finance & Accounting, or Software Technology terms for Apple Speech / Whisper biasing and output normalization; personal terms take priority
Edit RulesPersonal text replacement rules applied on every output
Language Style PresetsConcise / Formal / Casual / Custom prompt per language
Input History & StatsFull history with raw vs. processed comparison, word count stats, configurable retention
Bilingual UIChinese and English interface, independent of recognition language
Sound FeedbackAudio cues on recording start and stop
Guided OnboardingStep-by-step setup: permissions, model download, and first use

System Requirements

  • OS: macOS 26 (Tahoe) or later
  • Chip: Apple Silicon (M1 / M2 / M3 / M4)
  • Disk: ~400 MB minimum (Apple Speech + Qwen3-0.6B), up to ~4 GB with larger models

Installation

Download

Grab the latest .dmg from Releases, open it, and drag Utter.app to Applications.

"Cannot verify the developer" on first launch? The app is not notarized by Apple. Before first run, execute in Terminal:

xattr -cr /Applications/Utter.app

Or go to System Settings → Privacy & Security and click "Open Anyway".

Build from Source

# Build .app bundle + .dmg installer
bash scripts/build-app.sh
# Or for development
swift build
swift run OpenType
# Build, sign, and launch the development .app bundle
bash scripts/build-and-run.sh --verify
# Or open in Xcode
open Package.swift

Material changes follow the artifact-driven workflow in docs/sdlc/README.md. Before opening a pull request, run:

bash scripts/sdlc-checks.sh
bash scripts/ci-basic-checks.sh
swift test

The public app is Utter.app. The Swift package product remains OpenType so existing source integrations and upgrade paths continue to work.

First Run

  1. Launch Utter — it appears as a waveform icon in the menu bar
  2. The onboarding wizard guides you through permissions and model setup
  3. Grant Microphone and Accessibility permissions (required)
  4. Wait for the LLM model to download (~335 MB, one-time)
  5. Hold Fn to start dictating, release to stop and insert text

Permissions

PermissionPurposeRequired
MicrophoneAudio captureYes
AccessibilityGlobal hotkey + text injection (simulated paste)Yes
Speech RecognitionApple on-device ASR engineOnly if using Apple Speech
Screen RecordingOCR for screen context and Voice Command modeOptional
NetworkModel downloads; remote LLM API callsFirst run / remote LLM mode

Remote LLM Providers

Utter supports both OpenAI-compatible and Anthropic API formats:

ProviderAPI FormatBase URL
OpenAIOpenAIhttps://api.openai.com/v1
Anthropic ClaudeAnthropichttps://api.anthropic.com/v1
Google GeminiOpenAIhttps://generativelanguage.googleapis.com/v1beta/openai
OpenRouterOpenAIhttps://openrouter.ai/api/v1
SiliconFlowOpenAIhttps://api.siliconflow.cn/v1
Volcengine DoubaoOpenAIhttps://ark.cn-beijing.volces.com/api/v3
Alibaba BailianOpenAIhttps://dashscope.aliyuncs.com/compatible-mode/v1
MiniMax (China)OpenAIhttps://api.minimax.chat/v1
MiniMax (Global)OpenAIhttps://api.minimaxi.chat/v1

Local ASR Providers

ProviderLocal runtimeDefault model
Qwen3-ASRNative Swift + MLX on Apple Siliconmlx-community/Qwen3-ASR-1.7B-bf16

Qwen3-ASR does not call a hosted ASR API. The app downloads the selected model into the same model storage used by WhisperKit and runs inference locally through native Swift and MLX.

Project Structure

Sources/
├── App/ # Entry point, AppDelegate, AppState, VoicePipeline, AppIcon
├── Audio/ # Microphone capture (AVAudioEngine), sound playback
├── Config/ # AppSettings, ModelCatalog, RemoteModelConfig, Localization
├── Hotkey/ # Global hotkey via CGEvent tap
├── LLM/ # LLMEngine (MLX), RemoteLLMClient (OpenAI/Anthropic)
├── Output/ # Text injection (Accessibility API + clipboard paste)
├── Processing/ # TextProcessor, InputHistory, MemoryStore, personal and industry vocabulary
├── Prompts/ # PromptBuilder, prompt catalogs, style prompt presets
├── Screen/ # Screen OCR (ScreenCaptureKit + Vision)
├── Speech/ # SpeechEngine protocol, WhisperKit, Apple Speech, Doubao ASR, local ASR engines
├── UI/ # SwiftUI: MenuBar, Settings, Onboarding, Overlay, History, Models
└── Resources/ # Localization strings (en/zh-Hans), sounds, app icon
scripts/
├── build-and-run.sh # Build, sign, and launch a development .app bundle
├── build-app.sh # Build release .app bundle and .dmg installer
├── ci-basic-checks.sh # CI guardrails for linked files and resources
├── create-signing-cert.sh # Generate self-signed code signing certificate
├── generate-icon.swift # Generate AppIcon.icns from source PNG
├── test-industry-lexicons.sh # Validate vocabulary recall and non-target preservation
├── unit-test-coverage.sh # Run unit tests with coverage thresholds
└── validate-volc-asr.swift # Validate Volcengine ASR configuration manually

Tech Stack

  • WhisperKit — offline Whisper speech recognition
  • mlx-swift-lm — local LLM inference on Apple Silicon (Qwen2.5 / Qwen3)
  • SwiftUI + AppKit — native macOS UI
  • ScreenCaptureKit + Vision — screen OCR
  • AVAudioEngine — low-latency microphone capture
  • Apple Speech Framework — on-device speech recognition

License

MIT


Made with care for Apple Silicon

About

Local AI-powered voice input for macOS menu bar

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages