Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 4 additions & 1 deletion README.md
Original file line numberDiff line numberDiff line change
Expand Up@@ -39,7 +39,10 @@ Three output modes are available:
| Feature | Description |
|---|---|
| **Multiple Speech Engines** | Apple Speech, WhisperKit, Doubao ASR, Qwen3-ASR, or MiMo-V2.5-ASR |
| **Smart Text Processing** | Local MLX Qwen2.5/Qwen3 or remote LLM — contextual cleanup, self-correction handling, structured list formatting |
| **Smart Text Processing** | Local MLX Qwen2.5/Qwen3 or remote LLM infers spoken intent — contextual cleanup, "scratch that" restarts, self-correction handling, spoken punctuation, technical terms, numbers/ranges/units, and structured formatting |
| **LLM-Owned Spoken Formatting** | Spoken casing, no-space dictation, identifiers, file paths, shortcuts, emoji, Markdown tasks, dates/times, quantities, units, formulas, fractions, and digit sequences are handled by the Smart Format / Voice Command prompts instead of local hardcoded rewrite rules |
| **Voice Edit Commands** | In Voice Command mode, an LLM classifies safe structured actions for replacing, undoing, proofreading, titling, summarizing, drafting replies, making meeting notes, extracting key points/decisions/questions/risks/deadlines/owners/action items, rewriting tone, expanding, making tables/lists, or deleting the previous OpenType insertion or selected text |
| **Verbatim & Preview Boundary** | Verbatim mode, streaming HUD, integration partials, and instant-insert drafts keep ASR text close to raw output with only dictionary, whitespace, duplicate, and non-speech-artifact cleanup |
| **Remote LLM Support** | OpenAI, Claude (Anthropic format), Gemini, OpenRouter, SiliconFlow, Doubao, Bailian, MiniMax (CN & Global) |
| **Global Hotkey** | Configurable key (Fn/Ctrl/Shift/Option) with long-press, double-tap, or single-tap activation |
| **Screen Context OCR** | Captures on-screen text via ScreenCaptureKit + Vision to help the LLM correct homophones |
Expand Down
5 changes: 4 additions & 1 deletion README_zh.md
Original file line numberDiff line numberDiff line change
Expand Up@@ -39,7 +39,10 @@
| 功能 | 说明 |
|---|---|
| **多语音引擎** | Apple 语音识别、WhisperKit、豆包语音识别、Qwen3-ASR 或 MiMo-V2.5-ASR |
| **智能文字处理** | 本地 MLX Qwen2.5/Qwen3 或远程 LLM — 上下文感知的语气词清理、自动纠正、列表格式化 |
| **智能文字处理** | 本地 MLX Qwen2.5/Qwen3 或远程 LLM 理解口述意图 — 上下文感知的语气词清理、“算了/删掉刚才”重说处理、自动纠正、口述标点、技术词、数字/范围/单位和列表格式化 |
| **LLM 负责口述格式** | 大小写、无空格、标识符、文件路径、快捷键、表情、Markdown 任务、日期时间、数量、单位、公式、分数和数字串都由智能整理/语音指令提示词交给 LLM 判断,不在本地写死替换规则 |
| **语音编辑口令** | 在语音指令模式下,由 LLM 分类安全结构化动作,支持上一段/选区替换、撤销、校对、跨语言回复起草、接受/拒绝/追问回复、会议纪要、关键要点/结论/问题/风险/截止时间/负责人/行动项提取、标题化、摘要、语气改写、扩写、表格化、列表化、删除与改写口令 |
| **直出与预览边界** | 原文直出、流式 HUD、集成 partial 和快速插入草稿尽量保留 ASR 原文,只做词库、空白、重复转写和非语音垃圾过滤 |
| **远程 LLM** | 支持 OpenAI、Claude(Anthropic 格式)、Gemini、OpenRouter、硅基流动、豆包、百炼、MiniMax(国内/海外) |
| **全局快捷键** | 可配置按键(Fn/Ctrl/Shift/Option),支持长按、双击、单击三种触发模式 |
| **屏幕上下文 OCR** | 通过 ScreenCaptureKit + Vision 截取屏幕文字,辅助 LLM 纠正同音字 |
Expand Down
61 changes: 61 additions & 0 deletions Sources/App/VoicePipeline+EditCommandResolution.swift
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,61 @@
import AppKit
import Foundation

@MainActor
extension VoicePipeline {
func resolvedSpokenEditCommand(
raw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async -> SpokenEditCommand? {
guard VoicePipelinePolicy.shouldResolveEditCommandWithLLMFirst(outputMode: settings.outputMode) else {
return nil
}

let resolution = await resolveSpokenEditCommandWithLLM(
raw: raw,
settings: settings,
targetApp: targetApp
)
return VoicePipelinePolicy.editCommand(from: resolution)
}

func spokenEditCommandResolutionContext(
targetApp: NSRunningApplication?
) async -> SpokenEditCommandResolutionContext {
let lastInsertedText = appState.lastInsertedText.trimmingCharacters(in: .whitespacesAndNewlines)
let lastInsertion = lastInsertedText.isEmpty
? SpokenEditCommandTargetAvailability.unavailable
: .available
let selectedText = await textInserter.selectedText(targetApp: targetApp)
let selectionAvailability: SpokenEditCommandTargetAvailability
if let selectedText {
selectionAvailability = selectedText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty
? .unavailable
: .available
} else {
selectionAvailability = .unknown
}

return SpokenEditCommandResolutionContext(
lastInsertion: lastInsertion,
selectedText: selectionAvailability,
lastInsertionPreview: SpokenEditCommandResolutionContext.preview(lastInsertedText),
selectedTextPreview: SpokenEditCommandResolutionContext.preview(selectedText)
)
}

private func resolveSpokenEditCommandWithLLM(
raw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async -> SpokenEditCommandLLMResolution? {
var options = TextProcessingOptions(settings: settings)
options.llmModel = settings.llmModel
return await textProcessor.resolveSpokenEditCommandResolution(
text: raw,
options: options,
context: await spokenEditCommandResolutionContext(targetApp: targetApp)
)
}
}
282 changes: 282 additions & 0 deletions Sources/App/VoicePipeline+EditCommands.swift
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,282 @@
import AppKit
import Foundation

@MainActor
extension VoicePipeline {
func handleSpokenEditCommandIfNeeded(
raw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async -> Bool {
guard let command = await resolvedSpokenEditCommand(
raw: raw,
settings: settings,
targetApp: targetApp
) else {
return false
}

switch command {
case .replaceLast(let replacementRaw):
await replaceLastInsertion(
raw: raw,
replacementRaw: replacementRaw,
settings: settings,
targetApp: targetApp
)
return true
case .replaceSelection(let replacementRaw):
await replaceSelectedText(
raw: raw,
replacementRaw: replacementRaw,
settings: settings,
targetApp: targetApp
)
return true
case .rewriteLast(let intent):
await rewriteLastInsertion(
raw: raw,
intent: intent,
settings: settings,
targetApp: targetApp
)
return true
case .rewriteSelection(let intent):
await rewriteSelectedText(
raw: raw,
intent: intent,
settings: settings,
targetApp: targetApp
)
return true
case .deleteSelection:
await deleteSelectedText(targetApp: targetApp)
return true
case .undoLastInsertion:
await undoLastInsertion(targetApp: targetApp)
return true
}
}

private func replaceLastInsertion(
raw: String,
replacementRaw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async {
cancelScreenContextCapture()

guard !appState.lastInsertedText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty else {
showErrorHint(L("pipeline.no_previous_insert_to_replace"))
return
}

let context = replacementInputContext(settings: settings, targetApp: targetApp)
let replacementText = finalizedReplacementText(replacementRaw, settings: settings)
guard !replacementText.isEmpty else {
showNoSpeechDetected(reason: "spoken edit command has empty replacement text")
return
}

appState.processedText = replacementText
appState.phase = .inserting
appState.statusMessage = L("pipeline.replacing")

Log.sensitive("[VoicePipeline] voice edit replace \(replacementText.count) chars")
let result = await textInserter.replaceRecentInsertion(text: replacementText, targetApp: targetApp)

InputHistory.shared.addRecord(
rawText: raw,
processedText: replacementText,
wasProcessed: true,
context: context
)

appState.lastInsertedText = replacementText
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit replacement probably failed: \(reason)")
TextInserter.copyToClipboard(replacementText)
showInsertionFailedAlert(text: replacementText, reason: reason)
}
}

private func replaceSelectedText(
raw: String,
replacementRaw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async {
cancelScreenContextCapture()

let context = replacementInputContext(settings: settings, targetApp: targetApp)
let replacementText = finalizedReplacementText(replacementRaw, settings: settings)
guard !replacementText.isEmpty else {
showNoSpeechDetected(reason: "spoken edit command has empty selection replacement text")
return
}

appState.processedText = replacementText
appState.phase = .inserting
appState.statusMessage = L("pipeline.replacing")

Log.sensitive("[VoicePipeline] voice edit replace selection \(replacementText.count) chars")
let result = await textInserter.replaceSelectedText(text: replacementText, targetApp: targetApp)

InputHistory.shared.addRecord(
rawText: raw,
processedText: replacementText,
wasProcessed: true,
context: context
)

appState.lastInsertedText = replacementText
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit selection replacement probably failed: \(reason)")
TextInserter.copyToClipboard(replacementText)
showInsertionFailedAlert(text: replacementText, reason: reason)
}
}

private func replacementInputContext(
settings: AppSettings,
targetApp: NSRunningApplication?
) -> InputContext {
InputContext.capture(
targetApp: targetApp,
screenContext: "",
outputMode: .command,
inputLanguage: settings.inputLanguage,
source: .menuBar
)
}

private func finalizedReplacementText(
_ text: String,
settings: AppSettings
) -> String {
textProcessor.cleanCommandGeneratedOutput(
text,
inputLanguage: settings.inputLanguage
)
}

private func deleteSelectedText(targetApp: NSRunningApplication?) async {
cancelScreenContextCapture()

appState.processedText = ""
appState.phase = .inserting
appState.statusMessage = L("pipeline.replacing")

Log.info("[VoicePipeline] voice edit delete selection")
let result = await textInserter.deleteSelectedText(targetApp: targetApp)

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit delete selection probably failed: \(reason)")
showErrorHint(reason)
return
}

appState.lastInsertedText = ""
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()
}

private func undoLastInsertion(targetApp: NSRunningApplication?) async {
cancelScreenContextCapture()

guard !appState.lastInsertedText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty else {
showErrorHint(L("pipeline.no_previous_insert_to_replace"))
return
}

appState.processedText = ""
appState.phase = .inserting
appState.statusMessage = L("pipeline.undoing")

Log.info("[VoicePipeline] voice edit undo last insertion")
let result = await textInserter.undoRecentInsertion(targetApp: targetApp)

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit undo probably failed: \(reason)")
showErrorHint(reason)
return
}

appState.lastInsertedText = ""
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()
}

private func rewriteSelectedText(
raw: String,
intent: SelectionRewriteIntent,
settings: AppSettings,
targetApp: NSRunningApplication?
) async {
cancelScreenContextCapture()

guard let selectedText = await textInserter.selectedText(targetApp: targetApp),
!selectedText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty else {
showErrorHint(L("pipeline.no_selected_text_to_replace"))
return
}

let context = InputContext.capture(
targetApp: targetApp,
screenContext: selectedText,
outputMode: .command,
inputLanguage: settings.inputLanguage,
source: .menuBar
)
var options = TextProcessingOptions(settings: settings)
options.llmModel = settings.llmModel
let memoryContext = VoicePipelinePolicy.memoryContext(
for: .command,
settings: settings,
currentContext: context
)

appState.phase = .processing
appState.statusMessage = L("pipeline.formatting")
let rewrittenText = await textProcessor.processSelectionEdit(
selectedText: selectedText,
intent: intent,
options: options,
spokenCommand: raw,
memoryContext: memoryContext,
inputContext: context
)
guard !rewrittenText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty else {
showNoSpeechDetected(reason: "selection rewrite returned empty text")
return
}

appState.processedText = rewrittenText
appState.phase = .inserting
appState.statusMessage = L("pipeline.replacing")

let result = await textInserter.replaceSelectedText(text: rewrittenText, targetApp: targetApp)
InputHistory.shared.addRecord(rawText: raw, processedText: rewrittenText, wasProcessed: true, context: context)

appState.lastInsertedText = rewrittenText
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit selection rewrite probably failed: \(reason)")
TextInserter.copyToClipboard(rewrittenText)
showInsertionFailedAlert(text: rewrittenText, reason: reason)
}
}
}
Loading
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 4 additions & 1 deletion README.md
Original file line numberDiff line numberDiff line change
Expand Up@@ -39,7 +39,10 @@ Three output modes are available:
| Feature | Description |
|---|---|
| **Multiple Speech Engines** | Apple Speech, WhisperKit, Doubao ASR, Qwen3-ASR, or MiMo-V2.5-ASR |
| **Smart Text Processing** | Local MLX Qwen2.5/Qwen3 or remote LLM — contextual cleanup, self-correction handling, structured list formatting |
| **Smart Text Processing** | Local MLX Qwen2.5/Qwen3 or remote LLM infers spoken intent — contextual cleanup, "scratch that" restarts, self-correction handling, spoken punctuation, technical terms, numbers/ranges/units, and structured formatting |
| **LLM-Owned Spoken Formatting** | Spoken casing, no-space dictation, identifiers, file paths, shortcuts, emoji, Markdown tasks, dates/times, quantities, units, formulas, fractions, and digit sequences are handled by the Smart Format / Voice Command prompts instead of local hardcoded rewrite rules |
| **Voice Edit Commands** | In Voice Command mode, an LLM classifies safe structured actions for replacing, undoing, proofreading, titling, summarizing, drafting replies, making meeting notes, extracting key points/decisions/questions/risks/deadlines/owners/action items, rewriting tone, expanding, making tables/lists, or deleting the previous OpenType insertion or selected text |
| **Verbatim & Preview Boundary** | Verbatim mode, streaming HUD, integration partials, and instant-insert drafts keep ASR text close to raw output with only dictionary, whitespace, duplicate, and non-speech-artifact cleanup |
| **Remote LLM Support** | OpenAI, Claude (Anthropic format), Gemini, OpenRouter, SiliconFlow, Doubao, Bailian, MiniMax (CN & Global) |
| **Global Hotkey** | Configurable key (Fn/Ctrl/Shift/Option) with long-press, double-tap, or single-tap activation |
| **Screen Context OCR** | Captures on-screen text via ScreenCaptureKit + Vision to help the LLM correct homophones |
Expand Down
5 changes: 4 additions & 1 deletion README_zh.md
Original file line numberDiff line numberDiff line change
Expand Up@@ -39,7 +39,10 @@
| 功能 | 说明 |
|---|---|
| **多语音引擎** | Apple 语音识别、WhisperKit、豆包语音识别、Qwen3-ASR 或 MiMo-V2.5-ASR |
| **智能文字处理** | 本地 MLX Qwen2.5/Qwen3 或远程 LLM — 上下文感知的语气词清理、自动纠正、列表格式化 |
| **智能文字处理** | 本地 MLX Qwen2.5/Qwen3 或远程 LLM 理解口述意图 — 上下文感知的语气词清理、“算了/删掉刚才”重说处理、自动纠正、口述标点、技术词、数字/范围/单位和列表格式化 |
| **LLM 负责口述格式** | 大小写、无空格、标识符、文件路径、快捷键、表情、Markdown 任务、日期时间、数量、单位、公式、分数和数字串都由智能整理/语音指令提示词交给 LLM 判断,不在本地写死替换规则 |
| **语音编辑口令** | 在语音指令模式下,由 LLM 分类安全结构化动作,支持上一段/选区替换、撤销、校对、跨语言回复起草、接受/拒绝/追问回复、会议纪要、关键要点/结论/问题/风险/截止时间/负责人/行动项提取、标题化、摘要、语气改写、扩写、表格化、列表化、删除与改写口令 |
| **直出与预览边界** | 原文直出、流式 HUD、集成 partial 和快速插入草稿尽量保留 ASR 原文,只做词库、空白、重复转写和非语音垃圾过滤 |
| **远程 LLM** | 支持 OpenAI、Claude(Anthropic 格式)、Gemini、OpenRouter、硅基流动、豆包、百炼、MiniMax(国内/海外) |
| **全局快捷键** | 可配置按键(Fn/Ctrl/Shift/Option),支持长按、双击、单击三种触发模式 |
| **屏幕上下文 OCR** | 通过 ScreenCaptureKit + Vision 截取屏幕文字,辅助 LLM 纠正同音字 |
Expand Down
61 changes: 61 additions & 0 deletions Sources/App/VoicePipeline+EditCommandResolution.swift
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,61 @@
import AppKit
import Foundation

@MainActor
extension VoicePipeline {
func resolvedSpokenEditCommand(
raw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async -> SpokenEditCommand? {
guard VoicePipelinePolicy.shouldResolveEditCommandWithLLMFirst(outputMode: settings.outputMode) else {
return nil
}

let resolution = await resolveSpokenEditCommandWithLLM(
raw: raw,
settings: settings,
targetApp: targetApp
)
return VoicePipelinePolicy.editCommand(from: resolution)
}

func spokenEditCommandResolutionContext(
targetApp: NSRunningApplication?
) async -> SpokenEditCommandResolutionContext {
let lastInsertedText = appState.lastInsertedText.trimmingCharacters(in: .whitespacesAndNewlines)
let lastInsertion = lastInsertedText.isEmpty
? SpokenEditCommandTargetAvailability.unavailable
: .available
let selectedText = await textInserter.selectedText(targetApp: targetApp)
let selectionAvailability: SpokenEditCommandTargetAvailability
if let selectedText {
selectionAvailability = selectedText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty
? .unavailable
: .available
} else {
selectionAvailability = .unknown
}

return SpokenEditCommandResolutionContext(
lastInsertion: lastInsertion,
selectedText: selectionAvailability,
lastInsertionPreview: SpokenEditCommandResolutionContext.preview(lastInsertedText),
selectedTextPreview: SpokenEditCommandResolutionContext.preview(selectedText)
)
}

private func resolveSpokenEditCommandWithLLM(
raw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async -> SpokenEditCommandLLMResolution? {
var options = TextProcessingOptions(settings: settings)
options.llmModel = settings.llmModel
return await textProcessor.resolveSpokenEditCommandResolution(
text: raw,
options: options,
context: await spokenEditCommandResolutionContext(targetApp: targetApp)
)
}
}
282 changes: 282 additions & 0 deletions Sources/App/VoicePipeline+EditCommands.swift
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,282 @@
import AppKit
import Foundation

@MainActor
extension VoicePipeline {
func handleSpokenEditCommandIfNeeded(
raw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async -> Bool {
guard let command = await resolvedSpokenEditCommand(
raw: raw,
settings: settings,
targetApp: targetApp
) else {
return false
}

switch command {
case .replaceLast(let replacementRaw):
await replaceLastInsertion(
raw: raw,
replacementRaw: replacementRaw,
settings: settings,
targetApp: targetApp
)
return true
case .replaceSelection(let replacementRaw):
await replaceSelectedText(
raw: raw,
replacementRaw: replacementRaw,
settings: settings,
targetApp: targetApp
)
return true
case .rewriteLast(let intent):
await rewriteLastInsertion(
raw: raw,
intent: intent,
settings: settings,
targetApp: targetApp
)
return true
case .rewriteSelection(let intent):
await rewriteSelectedText(
raw: raw,
intent: intent,
settings: settings,
targetApp: targetApp
)
return true
case .deleteSelection:
await deleteSelectedText(targetApp: targetApp)
return true
case .undoLastInsertion:
await undoLastInsertion(targetApp: targetApp)
return true
}
}

private func replaceLastInsertion(
raw: String,
replacementRaw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async {
cancelScreenContextCapture()

guard !appState.lastInsertedText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty else {
showErrorHint(L("pipeline.no_previous_insert_to_replace"))
return
}

let context = replacementInputContext(settings: settings, targetApp: targetApp)
let replacementText = finalizedReplacementText(replacementRaw, settings: settings)
guard !replacementText.isEmpty else {
showNoSpeechDetected(reason: "spoken edit command has empty replacement text")
return
}

appState.processedText = replacementText
appState.phase = .inserting
appState.statusMessage = L("pipeline.replacing")

Log.sensitive("[VoicePipeline] voice edit replace \(replacementText.count) chars")
let result = await textInserter.replaceRecentInsertion(text: replacementText, targetApp: targetApp)

InputHistory.shared.addRecord(
rawText: raw,
processedText: replacementText,
wasProcessed: true,
context: context
)

appState.lastInsertedText = replacementText
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit replacement probably failed: \(reason)")
TextInserter.copyToClipboard(replacementText)
showInsertionFailedAlert(text: replacementText, reason: reason)
}
}

private func replaceSelectedText(
raw: String,
replacementRaw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async {
cancelScreenContextCapture()

let context = replacementInputContext(settings: settings, targetApp: targetApp)
let replacementText = finalizedReplacementText(replacementRaw, settings: settings)
guard !replacementText.isEmpty else {
showNoSpeechDetected(reason: "spoken edit command has empty selection replacement text")
return
}

appState.processedText = replacementText
appState.phase = .inserting
appState.statusMessage = L("pipeline.replacing")

Log.sensitive("[VoicePipeline] voice edit replace selection \(replacementText.count) chars")
let result = await textInserter.replaceSelectedText(text: replacementText, targetApp: targetApp)

InputHistory.shared.addRecord(
rawText: raw,
processedText: replacementText,
wasProcessed: true,
context: context
)

appState.lastInsertedText = replacementText
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit selection replacement probably failed: \(reason)")
TextInserter.copyToClipboard(replacementText)
showInsertionFailedAlert(text: replacementText, reason: reason)
}
}

private func replacementInputContext(
settings: AppSettings,
targetApp: NSRunningApplication?
) -> InputContext {
InputContext.capture(
targetApp: targetApp,
screenContext: "",
outputMode: .command,
inputLanguage: settings.inputLanguage,
source: .menuBar
)
}

private func finalizedReplacementText(
_ text: String,
settings: AppSettings
) -> String {
textProcessor.cleanCommandGeneratedOutput(
text,
inputLanguage: settings.inputLanguage
)
}

private func deleteSelectedText(targetApp: NSRunningApplication?) async {
cancelScreenContextCapture()

appState.processedText = ""
appState.phase = .inserting
appState.statusMessage = L("pipeline.replacing")

Log.info("[VoicePipeline] voice edit delete selection")
let result = await textInserter.deleteSelectedText(targetApp: targetApp)

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit delete selection probably failed: \(reason)")
showErrorHint(reason)
return
}

appState.lastInsertedText = ""
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()
}

private func undoLastInsertion(targetApp: NSRunningApplication?) async {
cancelScreenContextCapture()

guard !appState.lastInsertedText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty else {
showErrorHint(L("pipeline.no_previous_insert_to_replace"))
return
}

appState.processedText = ""
appState.phase = .inserting
appState.statusMessage = L("pipeline.undoing")

Log.info("[VoicePipeline] voice edit undo last insertion")
let result = await textInserter.undoRecentInsertion(targetApp: targetApp)

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit undo probably failed: \(reason)")
showErrorHint(reason)
return
}

appState.lastInsertedText = ""
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()
}

private func rewriteSelectedText(
raw: String,
intent: SelectionRewriteIntent,
settings: AppSettings,
targetApp: NSRunningApplication?
) async {
cancelScreenContextCapture()

guard let selectedText = await textInserter.selectedText(targetApp: targetApp),
!selectedText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty else {
showErrorHint(L("pipeline.no_selected_text_to_replace"))
return
}

let context = InputContext.capture(
targetApp: targetApp,
screenContext: selectedText,
outputMode: .command,
inputLanguage: settings.inputLanguage,
source: .menuBar
)
var options = TextProcessingOptions(settings: settings)
options.llmModel = settings.llmModel
let memoryContext = VoicePipelinePolicy.memoryContext(
for: .command,
settings: settings,
currentContext: context
)

appState.phase = .processing
appState.statusMessage = L("pipeline.formatting")
let rewrittenText = await textProcessor.processSelectionEdit(
selectedText: selectedText,
intent: intent,
options: options,
spokenCommand: raw,
memoryContext: memoryContext,
inputContext: context
)
guard !rewrittenText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty else {
showNoSpeechDetected(reason: "selection rewrite returned empty text")
return
}

appState.processedText = rewrittenText
appState.phase = .inserting
appState.statusMessage = L("pipeline.replacing")

let result = await textInserter.replaceSelectedText(text: rewrittenText, targetApp: targetApp)
InputHistory.shared.addRecord(rawText: raw, processedText: rewrittenText, wasProcessed: true, context: context)

appState.lastInsertedText = rewrittenText
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit selection rewrite probably failed: \(reason)")
TextInserter.copyToClipboard(rewrittenText)
showInsertionFailedAlert(text: rewrittenText, reason: reason)
}
}
}
Loading
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 4 additions & 1 deletion README.md
Original file line numberDiff line numberDiff line change
Expand Up@@ -39,7 +39,10 @@ Three output modes are available:
| Feature | Description |
|---|---|
| **Multiple Speech Engines** | Apple Speech, WhisperKit, Doubao ASR, Qwen3-ASR, or MiMo-V2.5-ASR |
| **Smart Text Processing** | Local MLX Qwen2.5/Qwen3 or remote LLM — contextual cleanup, self-correction handling, structured list formatting |
| **Smart Text Processing** | Local MLX Qwen2.5/Qwen3 or remote LLM infers spoken intent — contextual cleanup, "scratch that" restarts, self-correction handling, spoken punctuation, technical terms, numbers/ranges/units, and structured formatting |
| **LLM-Owned Spoken Formatting** | Spoken casing, no-space dictation, identifiers, file paths, shortcuts, emoji, Markdown tasks, dates/times, quantities, units, formulas, fractions, and digit sequences are handled by the Smart Format / Voice Command prompts instead of local hardcoded rewrite rules |
| **Voice Edit Commands** | In Voice Command mode, an LLM classifies safe structured actions for replacing, undoing, proofreading, titling, summarizing, drafting replies, making meeting notes, extracting key points/decisions/questions/risks/deadlines/owners/action items, rewriting tone, expanding, making tables/lists, or deleting the previous OpenType insertion or selected text |
| **Verbatim & Preview Boundary** | Verbatim mode, streaming HUD, integration partials, and instant-insert drafts keep ASR text close to raw output with only dictionary, whitespace, duplicate, and non-speech-artifact cleanup |
| **Remote LLM Support** | OpenAI, Claude (Anthropic format), Gemini, OpenRouter, SiliconFlow, Doubao, Bailian, MiniMax (CN & Global) |
| **Global Hotkey** | Configurable key (Fn/Ctrl/Shift/Option) with long-press, double-tap, or single-tap activation |
| **Screen Context OCR** | Captures on-screen text via ScreenCaptureKit + Vision to help the LLM correct homophones |
Expand Down
5 changes: 4 additions & 1 deletion README_zh.md
Original file line numberDiff line numberDiff line change
Expand Up@@ -39,7 +39,10 @@
| 功能 | 说明 |
|---|---|
| **多语音引擎** | Apple 语音识别、WhisperKit、豆包语音识别、Qwen3-ASR 或 MiMo-V2.5-ASR |
| **智能文字处理** | 本地 MLX Qwen2.5/Qwen3 或远程 LLM — 上下文感知的语气词清理、自动纠正、列表格式化 |
| **智能文字处理** | 本地 MLX Qwen2.5/Qwen3 或远程 LLM 理解口述意图 — 上下文感知的语气词清理、“算了/删掉刚才”重说处理、自动纠正、口述标点、技术词、数字/范围/单位和列表格式化 |
| **LLM 负责口述格式** | 大小写、无空格、标识符、文件路径、快捷键、表情、Markdown 任务、日期时间、数量、单位、公式、分数和数字串都由智能整理/语音指令提示词交给 LLM 判断,不在本地写死替换规则 |
| **语音编辑口令** | 在语音指令模式下,由 LLM 分类安全结构化动作,支持上一段/选区替换、撤销、校对、跨语言回复起草、接受/拒绝/追问回复、会议纪要、关键要点/结论/问题/风险/截止时间/负责人/行动项提取、标题化、摘要、语气改写、扩写、表格化、列表化、删除与改写口令 |
| **直出与预览边界** | 原文直出、流式 HUD、集成 partial 和快速插入草稿尽量保留 ASR 原文,只做词库、空白、重复转写和非语音垃圾过滤 |
| **远程 LLM** | 支持 OpenAI、Claude(Anthropic 格式)、Gemini、OpenRouter、硅基流动、豆包、百炼、MiniMax(国内/海外) |
| **全局快捷键** | 可配置按键(Fn/Ctrl/Shift/Option),支持长按、双击、单击三种触发模式 |
| **屏幕上下文 OCR** | 通过 ScreenCaptureKit + Vision 截取屏幕文字,辅助 LLM 纠正同音字 |
Expand Down
61 changes: 61 additions & 0 deletions Sources/App/VoicePipeline+EditCommandResolution.swift
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,61 @@
import AppKit
import Foundation

@MainActor
extension VoicePipeline {
func resolvedSpokenEditCommand(
raw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async -> SpokenEditCommand? {
guard VoicePipelinePolicy.shouldResolveEditCommandWithLLMFirst(outputMode: settings.outputMode) else {
return nil
}

let resolution = await resolveSpokenEditCommandWithLLM(
raw: raw,
settings: settings,
targetApp: targetApp
)
return VoicePipelinePolicy.editCommand(from: resolution)
}

func spokenEditCommandResolutionContext(
targetApp: NSRunningApplication?
) async -> SpokenEditCommandResolutionContext {
let lastInsertedText = appState.lastInsertedText.trimmingCharacters(in: .whitespacesAndNewlines)
let lastInsertion = lastInsertedText.isEmpty
? SpokenEditCommandTargetAvailability.unavailable
: .available
let selectedText = await textInserter.selectedText(targetApp: targetApp)
let selectionAvailability: SpokenEditCommandTargetAvailability
if let selectedText {
selectionAvailability = selectedText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty
? .unavailable
: .available
} else {
selectionAvailability = .unknown
}

return SpokenEditCommandResolutionContext(
lastInsertion: lastInsertion,
selectedText: selectionAvailability,
lastInsertionPreview: SpokenEditCommandResolutionContext.preview(lastInsertedText),
selectedTextPreview: SpokenEditCommandResolutionContext.preview(selectedText)
)
}

private func resolveSpokenEditCommandWithLLM(
raw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async -> SpokenEditCommandLLMResolution? {
var options = TextProcessingOptions(settings: settings)
options.llmModel = settings.llmModel
return await textProcessor.resolveSpokenEditCommandResolution(
text: raw,
options: options,
context: await spokenEditCommandResolutionContext(targetApp: targetApp)
)
}
}
282 changes: 282 additions & 0 deletions Sources/App/VoicePipeline+EditCommands.swift
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,282 @@
import AppKit
import Foundation

@MainActor
extension VoicePipeline {
func handleSpokenEditCommandIfNeeded(
raw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async -> Bool {
guard let command = await resolvedSpokenEditCommand(
raw: raw,
settings: settings,
targetApp: targetApp
) else {
return false
}

switch command {
case .replaceLast(let replacementRaw):
await replaceLastInsertion(
raw: raw,
replacementRaw: replacementRaw,
settings: settings,
targetApp: targetApp
)
return true
case .replaceSelection(let replacementRaw):
await replaceSelectedText(
raw: raw,
replacementRaw: replacementRaw,
settings: settings,
targetApp: targetApp
)
return true
case .rewriteLast(let intent):
await rewriteLastInsertion(
raw: raw,
intent: intent,
settings: settings,
targetApp: targetApp
)
return true
case .rewriteSelection(let intent):
await rewriteSelectedText(
raw: raw,
intent: intent,
settings: settings,
targetApp: targetApp
)
return true
case .deleteSelection:
await deleteSelectedText(targetApp: targetApp)
return true
case .undoLastInsertion:
await undoLastInsertion(targetApp: targetApp)
return true
}
}

private func replaceLastInsertion(
raw: String,
replacementRaw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async {
cancelScreenContextCapture()

guard !appState.lastInsertedText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty else {
showErrorHint(L("pipeline.no_previous_insert_to_replace"))
return
}

let context = replacementInputContext(settings: settings, targetApp: targetApp)
let replacementText = finalizedReplacementText(replacementRaw, settings: settings)
guard !replacementText.isEmpty else {
showNoSpeechDetected(reason: "spoken edit command has empty replacement text")
return
}

appState.processedText = replacementText
appState.phase = .inserting
appState.statusMessage = L("pipeline.replacing")

Log.sensitive("[VoicePipeline] voice edit replace \(replacementText.count) chars")
let result = await textInserter.replaceRecentInsertion(text: replacementText, targetApp: targetApp)

InputHistory.shared.addRecord(
rawText: raw,
processedText: replacementText,
wasProcessed: true,
context: context
)

appState.lastInsertedText = replacementText
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit replacement probably failed: \(reason)")
TextInserter.copyToClipboard(replacementText)
showInsertionFailedAlert(text: replacementText, reason: reason)
}
}

private func replaceSelectedText(
raw: String,
replacementRaw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async {
cancelScreenContextCapture()

let context = replacementInputContext(settings: settings, targetApp: targetApp)
let replacementText = finalizedReplacementText(replacementRaw, settings: settings)
guard !replacementText.isEmpty else {
showNoSpeechDetected(reason: "spoken edit command has empty selection replacement text")
return
}

appState.processedText = replacementText
appState.phase = .inserting
appState.statusMessage = L("pipeline.replacing")

Log.sensitive("[VoicePipeline] voice edit replace selection \(replacementText.count) chars")
let result = await textInserter.replaceSelectedText(text: replacementText, targetApp: targetApp)

InputHistory.shared.addRecord(
rawText: raw,
processedText: replacementText,
wasProcessed: true,
context: context
)

appState.lastInsertedText = replacementText
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit selection replacement probably failed: \(reason)")
TextInserter.copyToClipboard(replacementText)
showInsertionFailedAlert(text: replacementText, reason: reason)
}
}

private func replacementInputContext(
settings: AppSettings,
targetApp: NSRunningApplication?
) -> InputContext {
InputContext.capture(
targetApp: targetApp,
screenContext: "",
outputMode: .command,
inputLanguage: settings.inputLanguage,
source: .menuBar
)
}

private func finalizedReplacementText(
_ text: String,
settings: AppSettings
) -> String {
textProcessor.cleanCommandGeneratedOutput(
text,
inputLanguage: settings.inputLanguage
)
}

private func deleteSelectedText(targetApp: NSRunningApplication?) async {
cancelScreenContextCapture()

appState.processedText = ""
appState.phase = .inserting
appState.statusMessage = L("pipeline.replacing")

Log.info("[VoicePipeline] voice edit delete selection")
let result = await textInserter.deleteSelectedText(targetApp: targetApp)

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit delete selection probably failed: \(reason)")
showErrorHint(reason)
return
}

appState.lastInsertedText = ""
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()
}

private func undoLastInsertion(targetApp: NSRunningApplication?) async {
cancelScreenContextCapture()

guard !appState.lastInsertedText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty else {
showErrorHint(L("pipeline.no_previous_insert_to_replace"))
return
}

appState.processedText = ""
appState.phase = .inserting
appState.statusMessage = L("pipeline.undoing")

Log.info("[VoicePipeline] voice edit undo last insertion")
let result = await textInserter.undoRecentInsertion(targetApp: targetApp)

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit undo probably failed: \(reason)")
showErrorHint(reason)
return
}

appState.lastInsertedText = ""
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()
}

private func rewriteSelectedText(
raw: String,
intent: SelectionRewriteIntent,
settings: AppSettings,
targetApp: NSRunningApplication?
) async {
cancelScreenContextCapture()

guard let selectedText = await textInserter.selectedText(targetApp: targetApp),
!selectedText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty else {
showErrorHint(L("pipeline.no_selected_text_to_replace"))
return
}

let context = InputContext.capture(
targetApp: targetApp,
screenContext: selectedText,
outputMode: .command,
inputLanguage: settings.inputLanguage,
source: .menuBar
)
var options = TextProcessingOptions(settings: settings)
options.llmModel = settings.llmModel
let memoryContext = VoicePipelinePolicy.memoryContext(
for: .command,
settings: settings,
currentContext: context
)

appState.phase = .processing
appState.statusMessage = L("pipeline.formatting")
let rewrittenText = await textProcessor.processSelectionEdit(
selectedText: selectedText,
intent: intent,
options: options,
spokenCommand: raw,
memoryContext: memoryContext,
inputContext: context
)
guard !rewrittenText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty else {
showNoSpeechDetected(reason: "selection rewrite returned empty text")
return
}

appState.processedText = rewrittenText
appState.phase = .inserting
appState.statusMessage = L("pipeline.replacing")

let result = await textInserter.replaceSelectedText(text: rewrittenText, targetApp: targetApp)
InputHistory.shared.addRecord(rawText: raw, processedText: rewrittenText, wasProcessed: true, context: context)

appState.lastInsertedText = rewrittenText
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit selection rewrite probably failed: \(reason)")
TextInserter.copyToClipboard(rewrittenText)
showInsertionFailedAlert(text: rewrittenText, reason: reason)
}
}
}
Loading
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 4 additions & 1 deletion README.md
Original file line numberDiff line numberDiff line change
Expand Up@@ -39,7 +39,10 @@ Three output modes are available:
| Feature | Description |
|---|---|
| **Multiple Speech Engines** | Apple Speech, WhisperKit, Doubao ASR, Qwen3-ASR, or MiMo-V2.5-ASR |
| **Smart Text Processing** | Local MLX Qwen2.5/Qwen3 or remote LLM — contextual cleanup, self-correction handling, structured list formatting |
| **Smart Text Processing** | Local MLX Qwen2.5/Qwen3 or remote LLM infers spoken intent — contextual cleanup, "scratch that" restarts, self-correction handling, spoken punctuation, technical terms, numbers/ranges/units, and structured formatting |
| **LLM-Owned Spoken Formatting** | Spoken casing, no-space dictation, identifiers, file paths, shortcuts, emoji, Markdown tasks, dates/times, quantities, units, formulas, fractions, and digit sequences are handled by the Smart Format / Voice Command prompts instead of local hardcoded rewrite rules |
| **Voice Edit Commands** | In Voice Command mode, an LLM classifies safe structured actions for replacing, undoing, proofreading, titling, summarizing, drafting replies, making meeting notes, extracting key points/decisions/questions/risks/deadlines/owners/action items, rewriting tone, expanding, making tables/lists, or deleting the previous OpenType insertion or selected text |
| **Verbatim & Preview Boundary** | Verbatim mode, streaming HUD, integration partials, and instant-insert drafts keep ASR text close to raw output with only dictionary, whitespace, duplicate, and non-speech-artifact cleanup |
| **Remote LLM Support** | OpenAI, Claude (Anthropic format), Gemini, OpenRouter, SiliconFlow, Doubao, Bailian, MiniMax (CN & Global) |
| **Global Hotkey** | Configurable key (Fn/Ctrl/Shift/Option) with long-press, double-tap, or single-tap activation |
| **Screen Context OCR** | Captures on-screen text via ScreenCaptureKit + Vision to help the LLM correct homophones |
Expand Down
5 changes: 4 additions & 1 deletion README_zh.md
Original file line numberDiff line numberDiff line change
Expand Up@@ -39,7 +39,10 @@
| 功能 | 说明 |
|---|---|
| **多语音引擎** | Apple 语音识别、WhisperKit、豆包语音识别、Qwen3-ASR 或 MiMo-V2.5-ASR |
| **智能文字处理** | 本地 MLX Qwen2.5/Qwen3 或远程 LLM — 上下文感知的语气词清理、自动纠正、列表格式化 |
| **智能文字处理** | 本地 MLX Qwen2.5/Qwen3 或远程 LLM 理解口述意图 — 上下文感知的语气词清理、“算了/删掉刚才”重说处理、自动纠正、口述标点、技术词、数字/范围/单位和列表格式化 |
| **LLM 负责口述格式** | 大小写、无空格、标识符、文件路径、快捷键、表情、Markdown 任务、日期时间、数量、单位、公式、分数和数字串都由智能整理/语音指令提示词交给 LLM 判断,不在本地写死替换规则 |
| **语音编辑口令** | 在语音指令模式下,由 LLM 分类安全结构化动作,支持上一段/选区替换、撤销、校对、跨语言回复起草、接受/拒绝/追问回复、会议纪要、关键要点/结论/问题/风险/截止时间/负责人/行动项提取、标题化、摘要、语气改写、扩写、表格化、列表化、删除与改写口令 |
| **直出与预览边界** | 原文直出、流式 HUD、集成 partial 和快速插入草稿尽量保留 ASR 原文,只做词库、空白、重复转写和非语音垃圾过滤 |
| **远程 LLM** | 支持 OpenAI、Claude(Anthropic 格式)、Gemini、OpenRouter、硅基流动、豆包、百炼、MiniMax(国内/海外) |
| **全局快捷键** | 可配置按键(Fn/Ctrl/Shift/Option),支持长按、双击、单击三种触发模式 |
| **屏幕上下文 OCR** | 通过 ScreenCaptureKit + Vision 截取屏幕文字,辅助 LLM 纠正同音字 |
Expand Down
61 changes: 61 additions & 0 deletions Sources/App/VoicePipeline+EditCommandResolution.swift
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,61 @@
import AppKit
import Foundation

@MainActor
extension VoicePipeline {
func resolvedSpokenEditCommand(
raw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async -> SpokenEditCommand? {
guard VoicePipelinePolicy.shouldResolveEditCommandWithLLMFirst(outputMode: settings.outputMode) else {
return nil
}

let resolution = await resolveSpokenEditCommandWithLLM(
raw: raw,
settings: settings,
targetApp: targetApp
)
return VoicePipelinePolicy.editCommand(from: resolution)
}

func spokenEditCommandResolutionContext(
targetApp: NSRunningApplication?
) async -> SpokenEditCommandResolutionContext {
let lastInsertedText = appState.lastInsertedText.trimmingCharacters(in: .whitespacesAndNewlines)
let lastInsertion = lastInsertedText.isEmpty
? SpokenEditCommandTargetAvailability.unavailable
: .available
let selectedText = await textInserter.selectedText(targetApp: targetApp)
let selectionAvailability: SpokenEditCommandTargetAvailability
if let selectedText {
selectionAvailability = selectedText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty
? .unavailable
: .available
} else {
selectionAvailability = .unknown
}

return SpokenEditCommandResolutionContext(
lastInsertion: lastInsertion,
selectedText: selectionAvailability,
lastInsertionPreview: SpokenEditCommandResolutionContext.preview(lastInsertedText),
selectedTextPreview: SpokenEditCommandResolutionContext.preview(selectedText)
)
}

private func resolveSpokenEditCommandWithLLM(
raw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async -> SpokenEditCommandLLMResolution? {
var options = TextProcessingOptions(settings: settings)
options.llmModel = settings.llmModel
return await textProcessor.resolveSpokenEditCommandResolution(
text: raw,
options: options,
context: await spokenEditCommandResolutionContext(targetApp: targetApp)
)
}
}
282 changes: 282 additions & 0 deletions Sources/App/VoicePipeline+EditCommands.swift
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,282 @@
import AppKit
import Foundation

@MainActor
extension VoicePipeline {
func handleSpokenEditCommandIfNeeded(
raw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async -> Bool {
guard let command = await resolvedSpokenEditCommand(
raw: raw,
settings: settings,
targetApp: targetApp
) else {
return false
}

switch command {
case .replaceLast(let replacementRaw):
await replaceLastInsertion(
raw: raw,
replacementRaw: replacementRaw,
settings: settings,
targetApp: targetApp
)
return true
case .replaceSelection(let replacementRaw):
await replaceSelectedText(
raw: raw,
replacementRaw: replacementRaw,
settings: settings,
targetApp: targetApp
)
return true
case .rewriteLast(let intent):
await rewriteLastInsertion(
raw: raw,
intent: intent,
settings: settings,
targetApp: targetApp
)
return true
case .rewriteSelection(let intent):
await rewriteSelectedText(
raw: raw,
intent: intent,
settings: settings,
targetApp: targetApp
)
return true
case .deleteSelection:
await deleteSelectedText(targetApp: targetApp)
return true
case .undoLastInsertion:
await undoLastInsertion(targetApp: targetApp)
return true
}
}

private func replaceLastInsertion(
raw: String,
replacementRaw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async {
cancelScreenContextCapture()

guard !appState.lastInsertedText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty else {
showErrorHint(L("pipeline.no_previous_insert_to_replace"))
return
}

let context = replacementInputContext(settings: settings, targetApp: targetApp)
let replacementText = finalizedReplacementText(replacementRaw, settings: settings)
guard !replacementText.isEmpty else {
showNoSpeechDetected(reason: "spoken edit command has empty replacement text")
return
}

appState.processedText = replacementText
appState.phase = .inserting
appState.statusMessage = L("pipeline.replacing")

Log.sensitive("[VoicePipeline] voice edit replace \(replacementText.count) chars")
let result = await textInserter.replaceRecentInsertion(text: replacementText, targetApp: targetApp)

InputHistory.shared.addRecord(
rawText: raw,
processedText: replacementText,
wasProcessed: true,
context: context
)

appState.lastInsertedText = replacementText
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit replacement probably failed: \(reason)")
TextInserter.copyToClipboard(replacementText)
showInsertionFailedAlert(text: replacementText, reason: reason)
}
}

private func replaceSelectedText(
raw: String,
replacementRaw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async {
cancelScreenContextCapture()

let context = replacementInputContext(settings: settings, targetApp: targetApp)
let replacementText = finalizedReplacementText(replacementRaw, settings: settings)
guard !replacementText.isEmpty else {
showNoSpeechDetected(reason: "spoken edit command has empty selection replacement text")
return
}

appState.processedText = replacementText
appState.phase = .inserting
appState.statusMessage = L("pipeline.replacing")

Log.sensitive("[VoicePipeline] voice edit replace selection \(replacementText.count) chars")
let result = await textInserter.replaceSelectedText(text: replacementText, targetApp: targetApp)

InputHistory.shared.addRecord(
rawText: raw,
processedText: replacementText,
wasProcessed: true,
context: context
)

appState.lastInsertedText = replacementText
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit selection replacement probably failed: \(reason)")
TextInserter.copyToClipboard(replacementText)
showInsertionFailedAlert(text: replacementText, reason: reason)
}
}

private func replacementInputContext(
settings: AppSettings,
targetApp: NSRunningApplication?
) -> InputContext {
InputContext.capture(
targetApp: targetApp,
screenContext: "",
outputMode: .command,
inputLanguage: settings.inputLanguage,
source: .menuBar
)
}

private func finalizedReplacementText(
_ text: String,
settings: AppSettings
) -> String {
textProcessor.cleanCommandGeneratedOutput(
text,
inputLanguage: settings.inputLanguage
)
}

private func deleteSelectedText(targetApp: NSRunningApplication?) async {
cancelScreenContextCapture()

appState.processedText = ""
appState.phase = .inserting
appState.statusMessage = L("pipeline.replacing")

Log.info("[VoicePipeline] voice edit delete selection")
let result = await textInserter.deleteSelectedText(targetApp: targetApp)

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit delete selection probably failed: \(reason)")
showErrorHint(reason)
return
}

appState.lastInsertedText = ""
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()
}

private func undoLastInsertion(targetApp: NSRunningApplication?) async {
cancelScreenContextCapture()

guard !appState.lastInsertedText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty else {
showErrorHint(L("pipeline.no_previous_insert_to_replace"))
return
}

appState.processedText = ""
appState.phase = .inserting
appState.statusMessage = L("pipeline.undoing")

Log.info("[VoicePipeline] voice edit undo last insertion")
let result = await textInserter.undoRecentInsertion(targetApp: targetApp)

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit undo probably failed: \(reason)")
showErrorHint(reason)
return
}

appState.lastInsertedText = ""
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()
}

private func rewriteSelectedText(
raw: String,
intent: SelectionRewriteIntent,
settings: AppSettings,
targetApp: NSRunningApplication?
) async {
cancelScreenContextCapture()

guard let selectedText = await textInserter.selectedText(targetApp: targetApp),
!selectedText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty else {
showErrorHint(L("pipeline.no_selected_text_to_replace"))
return
}

let context = InputContext.capture(
targetApp: targetApp,
screenContext: selectedText,
outputMode: .command,
inputLanguage: settings.inputLanguage,
source: .menuBar
)
var options = TextProcessingOptions(settings: settings)
options.llmModel = settings.llmModel
let memoryContext = VoicePipelinePolicy.memoryContext(
for: .command,
settings: settings,
currentContext: context
)

appState.phase = .processing
appState.statusMessage = L("pipeline.formatting")
let rewrittenText = await textProcessor.processSelectionEdit(
selectedText: selectedText,
intent: intent,
options: options,
spokenCommand: raw,
memoryContext: memoryContext,
inputContext: context
)
guard !rewrittenText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty else {
showNoSpeechDetected(reason: "selection rewrite returned empty text")
return
}

appState.processedText = rewrittenText
appState.phase = .inserting
appState.statusMessage = L("pipeline.replacing")

let result = await textInserter.replaceSelectedText(text: rewrittenText, targetApp: targetApp)
InputHistory.shared.addRecord(rawText: raw, processedText: rewrittenText, wasProcessed: true, context: context)

appState.lastInsertedText = rewrittenText
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit selection rewrite probably failed: \(reason)")
TextInserter.copyToClipboard(rewrittenText)
showInsertionFailedAlert(text: rewrittenText, reason: reason)
}
}
}
Loading
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 4 additions & 1 deletion README.md
Original file line numberDiff line numberDiff line change
Expand Up@@ -39,7 +39,10 @@ Three output modes are available:
| Feature | Description |
|---|---|
| **Multiple Speech Engines** | Apple Speech, WhisperKit, Doubao ASR, Qwen3-ASR, or MiMo-V2.5-ASR |
| **Smart Text Processing** | Local MLX Qwen2.5/Qwen3 or remote LLM — contextual cleanup, self-correction handling, structured list formatting |
| **Smart Text Processing** | Local MLX Qwen2.5/Qwen3 or remote LLM infers spoken intent — contextual cleanup, "scratch that" restarts, self-correction handling, spoken punctuation, technical terms, numbers/ranges/units, and structured formatting |
| **LLM-Owned Spoken Formatting** | Spoken casing, no-space dictation, identifiers, file paths, shortcuts, emoji, Markdown tasks, dates/times, quantities, units, formulas, fractions, and digit sequences are handled by the Smart Format / Voice Command prompts instead of local hardcoded rewrite rules |
| **Voice Edit Commands** | In Voice Command mode, an LLM classifies safe structured actions for replacing, undoing, proofreading, titling, summarizing, drafting replies, making meeting notes, extracting key points/decisions/questions/risks/deadlines/owners/action items, rewriting tone, expanding, making tables/lists, or deleting the previous OpenType insertion or selected text |
| **Verbatim & Preview Boundary** | Verbatim mode, streaming HUD, integration partials, and instant-insert drafts keep ASR text close to raw output with only dictionary, whitespace, duplicate, and non-speech-artifact cleanup |
| **Remote LLM Support** | OpenAI, Claude (Anthropic format), Gemini, OpenRouter, SiliconFlow, Doubao, Bailian, MiniMax (CN & Global) |
| **Global Hotkey** | Configurable key (Fn/Ctrl/Shift/Option) with long-press, double-tap, or single-tap activation |
| **Screen Context OCR** | Captures on-screen text via ScreenCaptureKit + Vision to help the LLM correct homophones |
Expand Down
5 changes: 4 additions & 1 deletion README_zh.md
Original file line numberDiff line numberDiff line change
Expand Up@@ -39,7 +39,10 @@
| 功能 | 说明 |
|---|---|
| **多语音引擎** | Apple 语音识别、WhisperKit、豆包语音识别、Qwen3-ASR 或 MiMo-V2.5-ASR |
| **智能文字处理** | 本地 MLX Qwen2.5/Qwen3 或远程 LLM — 上下文感知的语气词清理、自动纠正、列表格式化 |
| **智能文字处理** | 本地 MLX Qwen2.5/Qwen3 或远程 LLM 理解口述意图 — 上下文感知的语气词清理、“算了/删掉刚才”重说处理、自动纠正、口述标点、技术词、数字/范围/单位和列表格式化 |
| **LLM 负责口述格式** | 大小写、无空格、标识符、文件路径、快捷键、表情、Markdown 任务、日期时间、数量、单位、公式、分数和数字串都由智能整理/语音指令提示词交给 LLM 判断,不在本地写死替换规则 |
| **语音编辑口令** | 在语音指令模式下,由 LLM 分类安全结构化动作,支持上一段/选区替换、撤销、校对、跨语言回复起草、接受/拒绝/追问回复、会议纪要、关键要点/结论/问题/风险/截止时间/负责人/行动项提取、标题化、摘要、语气改写、扩写、表格化、列表化、删除与改写口令 |
| **直出与预览边界** | 原文直出、流式 HUD、集成 partial 和快速插入草稿尽量保留 ASR 原文,只做词库、空白、重复转写和非语音垃圾过滤 |
| **远程 LLM** | 支持 OpenAI、Claude(Anthropic 格式)、Gemini、OpenRouter、硅基流动、豆包、百炼、MiniMax(国内/海外) |
| **全局快捷键** | 可配置按键(Fn/Ctrl/Shift/Option),支持长按、双击、单击三种触发模式 |
| **屏幕上下文 OCR** | 通过 ScreenCaptureKit + Vision 截取屏幕文字,辅助 LLM 纠正同音字 |
Expand Down
61 changes: 61 additions & 0 deletions Sources/App/VoicePipeline+EditCommandResolution.swift
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,61 @@
import AppKit
import Foundation

@MainActor
extension VoicePipeline {
func resolvedSpokenEditCommand(
raw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async -> SpokenEditCommand? {
guard VoicePipelinePolicy.shouldResolveEditCommandWithLLMFirst(outputMode: settings.outputMode) else {
return nil
}

let resolution = await resolveSpokenEditCommandWithLLM(
raw: raw,
settings: settings,
targetApp: targetApp
)
return VoicePipelinePolicy.editCommand(from: resolution)
}

func spokenEditCommandResolutionContext(
targetApp: NSRunningApplication?
) async -> SpokenEditCommandResolutionContext {
let lastInsertedText = appState.lastInsertedText.trimmingCharacters(in: .whitespacesAndNewlines)
let lastInsertion = lastInsertedText.isEmpty
? SpokenEditCommandTargetAvailability.unavailable
: .available
let selectedText = await textInserter.selectedText(targetApp: targetApp)
let selectionAvailability: SpokenEditCommandTargetAvailability
if let selectedText {
selectionAvailability = selectedText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty
? .unavailable
: .available
} else {
selectionAvailability = .unknown
}

return SpokenEditCommandResolutionContext(
lastInsertion: lastInsertion,
selectedText: selectionAvailability,
lastInsertionPreview: SpokenEditCommandResolutionContext.preview(lastInsertedText),
selectedTextPreview: SpokenEditCommandResolutionContext.preview(selectedText)
)
}

private func resolveSpokenEditCommandWithLLM(
raw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async -> SpokenEditCommandLLMResolution? {
var options = TextProcessingOptions(settings: settings)
options.llmModel = settings.llmModel
return await textProcessor.resolveSpokenEditCommandResolution(
text: raw,
options: options,
context: await spokenEditCommandResolutionContext(targetApp: targetApp)
)
}
}
282 changes: 282 additions & 0 deletions Sources/App/VoicePipeline+EditCommands.swift
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,282 @@
import AppKit
import Foundation

@MainActor
extension VoicePipeline {
func handleSpokenEditCommandIfNeeded(
raw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async -> Bool {
guard let command = await resolvedSpokenEditCommand(
raw: raw,
settings: settings,
targetApp: targetApp
) else {
return false
}

switch command {
case .replaceLast(let replacementRaw):
await replaceLastInsertion(
raw: raw,
replacementRaw: replacementRaw,
settings: settings,
targetApp: targetApp
)
return true
case .replaceSelection(let replacementRaw):
await replaceSelectedText(
raw: raw,
replacementRaw: replacementRaw,
settings: settings,
targetApp: targetApp
)
return true
case .rewriteLast(let intent):
await rewriteLastInsertion(
raw: raw,
intent: intent,
settings: settings,
targetApp: targetApp
)
return true
case .rewriteSelection(let intent):
await rewriteSelectedText(
raw: raw,
intent: intent,
settings: settings,
targetApp: targetApp
)
return true
case .deleteSelection:
await deleteSelectedText(targetApp: targetApp)
return true
case .undoLastInsertion:
await undoLastInsertion(targetApp: targetApp)
return true
}
}

private func replaceLastInsertion(
raw: String,
replacementRaw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async {
cancelScreenContextCapture()

guard !appState.lastInsertedText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty else {
showErrorHint(L("pipeline.no_previous_insert_to_replace"))
return
}

let context = replacementInputContext(settings: settings, targetApp: targetApp)
let replacementText = finalizedReplacementText(replacementRaw, settings: settings)
guard !replacementText.isEmpty else {
showNoSpeechDetected(reason: "spoken edit command has empty replacement text")
return
}

appState.processedText = replacementText
appState.phase = .inserting
appState.statusMessage = L("pipeline.replacing")

Log.sensitive("[VoicePipeline] voice edit replace \(replacementText.count) chars")
let result = await textInserter.replaceRecentInsertion(text: replacementText, targetApp: targetApp)

InputHistory.shared.addRecord(
rawText: raw,
processedText: replacementText,
wasProcessed: true,
context: context
)

appState.lastInsertedText = replacementText
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit replacement probably failed: \(reason)")
TextInserter.copyToClipboard(replacementText)
showInsertionFailedAlert(text: replacementText, reason: reason)
}
}

private func replaceSelectedText(
raw: String,
replacementRaw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async {
cancelScreenContextCapture()

let context = replacementInputContext(settings: settings, targetApp: targetApp)
let replacementText = finalizedReplacementText(replacementRaw, settings: settings)
guard !replacementText.isEmpty else {
showNoSpeechDetected(reason: "spoken edit command has empty selection replacement text")
return
}

appState.processedText = replacementText
appState.phase = .inserting
appState.statusMessage = L("pipeline.replacing")

Log.sensitive("[VoicePipeline] voice edit replace selection \(replacementText.count) chars")
let result = await textInserter.replaceSelectedText(text: replacementText, targetApp: targetApp)

InputHistory.shared.addRecord(
rawText: raw,
processedText: replacementText,
wasProcessed: true,
context: context
)

appState.lastInsertedText = replacementText
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit selection replacement probably failed: \(reason)")
TextInserter.copyToClipboard(replacementText)
showInsertionFailedAlert(text: replacementText, reason: reason)
}
}

private func replacementInputContext(
settings: AppSettings,
targetApp: NSRunningApplication?
) -> InputContext {
InputContext.capture(
targetApp: targetApp,
screenContext: "",
outputMode: .command,
inputLanguage: settings.inputLanguage,
source: .menuBar
)
}

private func finalizedReplacementText(
_ text: String,
settings: AppSettings
) -> String {
textProcessor.cleanCommandGeneratedOutput(
text,
inputLanguage: settings.inputLanguage
)
}

private func deleteSelectedText(targetApp: NSRunningApplication?) async {
cancelScreenContextCapture()

appState.processedText = ""
appState.phase = .inserting
appState.statusMessage = L("pipeline.replacing")

Log.info("[VoicePipeline] voice edit delete selection")
let result = await textInserter.deleteSelectedText(targetApp: targetApp)

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit delete selection probably failed: \(reason)")
showErrorHint(reason)
return
}

appState.lastInsertedText = ""
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()
}

private func undoLastInsertion(targetApp: NSRunningApplication?) async {
cancelScreenContextCapture()

guard !appState.lastInsertedText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty else {
showErrorHint(L("pipeline.no_previous_insert_to_replace"))
return
}

appState.processedText = ""
appState.phase = .inserting
appState.statusMessage = L("pipeline.undoing")

Log.info("[VoicePipeline] voice edit undo last insertion")
let result = await textInserter.undoRecentInsertion(targetApp: targetApp)

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit undo probably failed: \(reason)")
showErrorHint(reason)
return
}

appState.lastInsertedText = ""
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()
}

private func rewriteSelectedText(
raw: String,
intent: SelectionRewriteIntent,
settings: AppSettings,
targetApp: NSRunningApplication?
) async {
cancelScreenContextCapture()

guard let selectedText = await textInserter.selectedText(targetApp: targetApp),
!selectedText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty else {
showErrorHint(L("pipeline.no_selected_text_to_replace"))
return
}

let context = InputContext.capture(
targetApp: targetApp,
screenContext: selectedText,
outputMode: .command,
inputLanguage: settings.inputLanguage,
source: .menuBar
)
var options = TextProcessingOptions(settings: settings)
options.llmModel = settings.llmModel
let memoryContext = VoicePipelinePolicy.memoryContext(
for: .command,
settings: settings,
currentContext: context
)

appState.phase = .processing
appState.statusMessage = L("pipeline.formatting")
let rewrittenText = await textProcessor.processSelectionEdit(
selectedText: selectedText,
intent: intent,
options: options,
spokenCommand: raw,
memoryContext: memoryContext,
inputContext: context
)
guard !rewrittenText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty else {
showNoSpeechDetected(reason: "selection rewrite returned empty text")
return
}

appState.processedText = rewrittenText
appState.phase = .inserting
appState.statusMessage = L("pipeline.replacing")

let result = await textInserter.replaceSelectedText(text: rewrittenText, targetApp: targetApp)
InputHistory.shared.addRecord(rawText: raw, processedText: rewrittenText, wasProcessed: true, context: context)

appState.lastInsertedText = rewrittenText
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit selection rewrite probably failed: \(reason)")
TextInserter.copyToClipboard(rewrittenText)
showInsertionFailedAlert(text: rewrittenText, reason: reason)
}
}
}
Loading
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 4 additions & 1 deletion README.md
Original file line numberDiff line numberDiff line change
Expand Up@@ -39,7 +39,10 @@ Three output modes are available:
| Feature | Description |
|---|---|
| **Multiple Speech Engines** | Apple Speech, WhisperKit, Doubao ASR, Qwen3-ASR, or MiMo-V2.5-ASR |
| **Smart Text Processing** | Local MLX Qwen2.5/Qwen3 or remote LLM — contextual cleanup, self-correction handling, structured list formatting |
| **Smart Text Processing** | Local MLX Qwen2.5/Qwen3 or remote LLM infers spoken intent — contextual cleanup, "scratch that" restarts, self-correction handling, spoken punctuation, technical terms, numbers/ranges/units, and structured formatting |
| **LLM-Owned Spoken Formatting** | Spoken casing, no-space dictation, identifiers, file paths, shortcuts, emoji, Markdown tasks, dates/times, quantities, units, formulas, fractions, and digit sequences are handled by the Smart Format / Voice Command prompts instead of local hardcoded rewrite rules |
| **Voice Edit Commands** | In Voice Command mode, an LLM classifies safe structured actions for replacing, undoing, proofreading, titling, summarizing, drafting replies, making meeting notes, extracting key points/decisions/questions/risks/deadlines/owners/action items, rewriting tone, expanding, making tables/lists, or deleting the previous OpenType insertion or selected text |
| **Verbatim & Preview Boundary** | Verbatim mode, streaming HUD, integration partials, and instant-insert drafts keep ASR text close to raw output with only dictionary, whitespace, duplicate, and non-speech-artifact cleanup |
| **Remote LLM Support** | OpenAI, Claude (Anthropic format), Gemini, OpenRouter, SiliconFlow, Doubao, Bailian, MiniMax (CN & Global) |
| **Global Hotkey** | Configurable key (Fn/Ctrl/Shift/Option) with long-press, double-tap, or single-tap activation |
| **Screen Context OCR** | Captures on-screen text via ScreenCaptureKit + Vision to help the LLM correct homophones |
Expand Down
5 changes: 4 additions & 1 deletion README_zh.md
Original file line numberDiff line numberDiff line change
Expand Up@@ -39,7 +39,10 @@
| 功能 | 说明 |
|---|---|
| **多语音引擎** | Apple 语音识别、WhisperKit、豆包语音识别、Qwen3-ASR 或 MiMo-V2.5-ASR |
| **智能文字处理** | 本地 MLX Qwen2.5/Qwen3 或远程 LLM — 上下文感知的语气词清理、自动纠正、列表格式化 |
| **智能文字处理** | 本地 MLX Qwen2.5/Qwen3 或远程 LLM 理解口述意图 — 上下文感知的语气词清理、“算了/删掉刚才”重说处理、自动纠正、口述标点、技术词、数字/范围/单位和列表格式化 |
| **LLM 负责口述格式** | 大小写、无空格、标识符、文件路径、快捷键、表情、Markdown 任务、日期时间、数量、单位、公式、分数和数字串都由智能整理/语音指令提示词交给 LLM 判断,不在本地写死替换规则 |
| **语音编辑口令** | 在语音指令模式下,由 LLM 分类安全结构化动作,支持上一段/选区替换、撤销、校对、跨语言回复起草、接受/拒绝/追问回复、会议纪要、关键要点/结论/问题/风险/截止时间/负责人/行动项提取、标题化、摘要、语气改写、扩写、表格化、列表化、删除与改写口令 |
| **直出与预览边界** | 原文直出、流式 HUD、集成 partial 和快速插入草稿尽量保留 ASR 原文,只做词库、空白、重复转写和非语音垃圾过滤 |
| **远程 LLM** | 支持 OpenAI、Claude(Anthropic 格式)、Gemini、OpenRouter、硅基流动、豆包、百炼、MiniMax(国内/海外) |
| **全局快捷键** | 可配置按键(Fn/Ctrl/Shift/Option),支持长按、双击、单击三种触发模式 |
| **屏幕上下文 OCR** | 通过 ScreenCaptureKit + Vision 截取屏幕文字,辅助 LLM 纠正同音字 |
Expand Down
61 changes: 61 additions & 0 deletions Sources/App/VoicePipeline+EditCommandResolution.swift
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,61 @@
import AppKit
import Foundation

@MainActor
extension VoicePipeline {
func resolvedSpokenEditCommand(
raw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async -> SpokenEditCommand? {
guard VoicePipelinePolicy.shouldResolveEditCommandWithLLMFirst(outputMode: settings.outputMode) else {
return nil
}

let resolution = await resolveSpokenEditCommandWithLLM(
raw: raw,
settings: settings,
targetApp: targetApp
)
return VoicePipelinePolicy.editCommand(from: resolution)
}

func spokenEditCommandResolutionContext(
targetApp: NSRunningApplication?
) async -> SpokenEditCommandResolutionContext {
let lastInsertedText = appState.lastInsertedText.trimmingCharacters(in: .whitespacesAndNewlines)
let lastInsertion = lastInsertedText.isEmpty
? SpokenEditCommandTargetAvailability.unavailable
: .available
let selectedText = await textInserter.selectedText(targetApp: targetApp)
let selectionAvailability: SpokenEditCommandTargetAvailability
if let selectedText {
selectionAvailability = selectedText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty
? .unavailable
: .available
} else {
selectionAvailability = .unknown
}

return SpokenEditCommandResolutionContext(
lastInsertion: lastInsertion,
selectedText: selectionAvailability,
lastInsertionPreview: SpokenEditCommandResolutionContext.preview(lastInsertedText),
selectedTextPreview: SpokenEditCommandResolutionContext.preview(selectedText)
)
}

private func resolveSpokenEditCommandWithLLM(
raw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async -> SpokenEditCommandLLMResolution? {
var options = TextProcessingOptions(settings: settings)
options.llmModel = settings.llmModel
return await textProcessor.resolveSpokenEditCommandResolution(
text: raw,
options: options,
context: await spokenEditCommandResolutionContext(targetApp: targetApp)
)
}
}
282 changes: 282 additions & 0 deletions Sources/App/VoicePipeline+EditCommands.swift
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,282 @@
import AppKit
import Foundation

@MainActor
extension VoicePipeline {
func handleSpokenEditCommandIfNeeded(
raw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async -> Bool {
guard let command = await resolvedSpokenEditCommand(
raw: raw,
settings: settings,
targetApp: targetApp
) else {
return false
}

switch command {
case .replaceLast(let replacementRaw):
await replaceLastInsertion(
raw: raw,
replacementRaw: replacementRaw,
settings: settings,
targetApp: targetApp
)
return true
case .replaceSelection(let replacementRaw):
await replaceSelectedText(
raw: raw,
replacementRaw: replacementRaw,
settings: settings,
targetApp: targetApp
)
return true
case .rewriteLast(let intent):
await rewriteLastInsertion(
raw: raw,
intent: intent,
settings: settings,
targetApp: targetApp
)
return true
case .rewriteSelection(let intent):
await rewriteSelectedText(
raw: raw,
intent: intent,
settings: settings,
targetApp: targetApp
)
return true
case .deleteSelection:
await deleteSelectedText(targetApp: targetApp)
return true
case .undoLastInsertion:
await undoLastInsertion(targetApp: targetApp)
return true
}
}

private func replaceLastInsertion(
raw: String,
replacementRaw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async {
cancelScreenContextCapture()

guard !appState.lastInsertedText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty else {
showErrorHint(L("pipeline.no_previous_insert_to_replace"))
return
}

let context = replacementInputContext(settings: settings, targetApp: targetApp)
let replacementText = finalizedReplacementText(replacementRaw, settings: settings)
guard !replacementText.isEmpty else {
showNoSpeechDetected(reason: "spoken edit command has empty replacement text")
return
}

appState.processedText = replacementText
appState.phase = .inserting
appState.statusMessage = L("pipeline.replacing")

Log.sensitive("[VoicePipeline] voice edit replace \(replacementText.count) chars")
let result = await textInserter.replaceRecentInsertion(text: replacementText, targetApp: targetApp)

InputHistory.shared.addRecord(
rawText: raw,
processedText: replacementText,
wasProcessed: true,
context: context
)

appState.lastInsertedText = replacementText
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit replacement probably failed: \(reason)")
TextInserter.copyToClipboard(replacementText)
showInsertionFailedAlert(text: replacementText, reason: reason)
}
}

private func replaceSelectedText(
raw: String,
replacementRaw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async {
cancelScreenContextCapture()

let context = replacementInputContext(settings: settings, targetApp: targetApp)
let replacementText = finalizedReplacementText(replacementRaw, settings: settings)
guard !replacementText.isEmpty else {
showNoSpeechDetected(reason: "spoken edit command has empty selection replacement text")
return
}

appState.processedText = replacementText
appState.phase = .inserting
appState.statusMessage = L("pipeline.replacing")

Log.sensitive("[VoicePipeline] voice edit replace selection \(replacementText.count) chars")
let result = await textInserter.replaceSelectedText(text: replacementText, targetApp: targetApp)

InputHistory.shared.addRecord(
rawText: raw,
processedText: replacementText,
wasProcessed: true,
context: context
)

appState.lastInsertedText = replacementText
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit selection replacement probably failed: \(reason)")
TextInserter.copyToClipboard(replacementText)
showInsertionFailedAlert(text: replacementText, reason: reason)
}
}

private func replacementInputContext(
settings: AppSettings,
targetApp: NSRunningApplication?
) -> InputContext {
InputContext.capture(
targetApp: targetApp,
screenContext: "",
outputMode: .command,
inputLanguage: settings.inputLanguage,
source: .menuBar
)
}

private func finalizedReplacementText(
_ text: String,
settings: AppSettings
) -> String {
textProcessor.cleanCommandGeneratedOutput(
text,
inputLanguage: settings.inputLanguage
)
}

private func deleteSelectedText(targetApp: NSRunningApplication?) async {
cancelScreenContextCapture()

appState.processedText = ""
appState.phase = .inserting
appState.statusMessage = L("pipeline.replacing")

Log.info("[VoicePipeline] voice edit delete selection")
let result = await textInserter.deleteSelectedText(targetApp: targetApp)

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit delete selection probably failed: \(reason)")
showErrorHint(reason)
return
}

appState.lastInsertedText = ""
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()
}

private func undoLastInsertion(targetApp: NSRunningApplication?) async {
cancelScreenContextCapture()

guard !appState.lastInsertedText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty else {
showErrorHint(L("pipeline.no_previous_insert_to_replace"))
return
}

appState.processedText = ""
appState.phase = .inserting
appState.statusMessage = L("pipeline.undoing")

Log.info("[VoicePipeline] voice edit undo last insertion")
let result = await textInserter.undoRecentInsertion(targetApp: targetApp)

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit undo probably failed: \(reason)")
showErrorHint(reason)
return
}

appState.lastInsertedText = ""
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()
}

private func rewriteSelectedText(
raw: String,
intent: SelectionRewriteIntent,
settings: AppSettings,
targetApp: NSRunningApplication?
) async {
cancelScreenContextCapture()

guard let selectedText = await textInserter.selectedText(targetApp: targetApp),
!selectedText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty else {
showErrorHint(L("pipeline.no_selected_text_to_replace"))
return
}

let context = InputContext.capture(
targetApp: targetApp,
screenContext: selectedText,
outputMode: .command,
inputLanguage: settings.inputLanguage,
source: .menuBar
)
var options = TextProcessingOptions(settings: settings)
options.llmModel = settings.llmModel
let memoryContext = VoicePipelinePolicy.memoryContext(
for: .command,
settings: settings,
currentContext: context
)

appState.phase = .processing
appState.statusMessage = L("pipeline.formatting")
let rewrittenText = await textProcessor.processSelectionEdit(
selectedText: selectedText,
intent: intent,
options: options,
spokenCommand: raw,
memoryContext: memoryContext,
inputContext: context
)
guard !rewrittenText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty else {
showNoSpeechDetected(reason: "selection rewrite returned empty text")
return
}

appState.processedText = rewrittenText
appState.phase = .inserting
appState.statusMessage = L("pipeline.replacing")

let result = await textInserter.replaceSelectedText(text: rewrittenText, targetApp: targetApp)
InputHistory.shared.addRecord(rawText: raw, processedText: rewrittenText, wasProcessed: true, context: context)

appState.lastInsertedText = rewrittenText
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit selection rewrite probably failed: \(reason)")
TextInserter.copyToClipboard(rewrittenText)
showInsertionFailedAlert(text: rewrittenText, reason: reason)
}
}
}
Loading
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 4 additions & 1 deletion README.md
Original file line numberDiff line numberDiff line change
Expand Up@@ -39,7 +39,10 @@ Three output modes are available:
| Feature | Description |
|---|---|
| **Multiple Speech Engines** | Apple Speech, WhisperKit, Doubao ASR, Qwen3-ASR, or MiMo-V2.5-ASR |
| **Smart Text Processing** | Local MLX Qwen2.5/Qwen3 or remote LLM — contextual cleanup, self-correction handling, structured list formatting |
| **Smart Text Processing** | Local MLX Qwen2.5/Qwen3 or remote LLM infers spoken intent — contextual cleanup, "scratch that" restarts, self-correction handling, spoken punctuation, technical terms, numbers/ranges/units, and structured formatting |
| **LLM-Owned Spoken Formatting** | Spoken casing, no-space dictation, identifiers, file paths, shortcuts, emoji, Markdown tasks, dates/times, quantities, units, formulas, fractions, and digit sequences are handled by the Smart Format / Voice Command prompts instead of local hardcoded rewrite rules |
| **Voice Edit Commands** | In Voice Command mode, an LLM classifies safe structured actions for replacing, undoing, proofreading, titling, summarizing, drafting replies, making meeting notes, extracting key points/decisions/questions/risks/deadlines/owners/action items, rewriting tone, expanding, making tables/lists, or deleting the previous OpenType insertion or selected text |
| **Verbatim & Preview Boundary** | Verbatim mode, streaming HUD, integration partials, and instant-insert drafts keep ASR text close to raw output with only dictionary, whitespace, duplicate, and non-speech-artifact cleanup |
| **Remote LLM Support** | OpenAI, Claude (Anthropic format), Gemini, OpenRouter, SiliconFlow, Doubao, Bailian, MiniMax (CN & Global) |
| **Global Hotkey** | Configurable key (Fn/Ctrl/Shift/Option) with long-press, double-tap, or single-tap activation |
| **Screen Context OCR** | Captures on-screen text via ScreenCaptureKit + Vision to help the LLM correct homophones |
Expand Down
5 changes: 4 additions & 1 deletion README_zh.md
Original file line numberDiff line numberDiff line change
Expand Up@@ -39,7 +39,10 @@
| 功能 | 说明 |
|---|---|
| **多语音引擎** | Apple 语音识别、WhisperKit、豆包语音识别、Qwen3-ASR 或 MiMo-V2.5-ASR |
| **智能文字处理** | 本地 MLX Qwen2.5/Qwen3 或远程 LLM — 上下文感知的语气词清理、自动纠正、列表格式化 |
| **智能文字处理** | 本地 MLX Qwen2.5/Qwen3 或远程 LLM 理解口述意图 — 上下文感知的语气词清理、“算了/删掉刚才”重说处理、自动纠正、口述标点、技术词、数字/范围/单位和列表格式化 |
| **LLM 负责口述格式** | 大小写、无空格、标识符、文件路径、快捷键、表情、Markdown 任务、日期时间、数量、单位、公式、分数和数字串都由智能整理/语音指令提示词交给 LLM 判断,不在本地写死替换规则 |
| **语音编辑口令** | 在语音指令模式下,由 LLM 分类安全结构化动作,支持上一段/选区替换、撤销、校对、跨语言回复起草、接受/拒绝/追问回复、会议纪要、关键要点/结论/问题/风险/截止时间/负责人/行动项提取、标题化、摘要、语气改写、扩写、表格化、列表化、删除与改写口令 |
| **直出与预览边界** | 原文直出、流式 HUD、集成 partial 和快速插入草稿尽量保留 ASR 原文,只做词库、空白、重复转写和非语音垃圾过滤 |
| **远程 LLM** | 支持 OpenAI、Claude(Anthropic 格式)、Gemini、OpenRouter、硅基流动、豆包、百炼、MiniMax(国内/海外) |
| **全局快捷键** | 可配置按键(Fn/Ctrl/Shift/Option),支持长按、双击、单击三种触发模式 |
| **屏幕上下文 OCR** | 通过 ScreenCaptureKit + Vision 截取屏幕文字,辅助 LLM 纠正同音字 |
Expand Down
61 changes: 61 additions & 0 deletions Sources/App/VoicePipeline+EditCommandResolution.swift
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,61 @@
import AppKit
import Foundation

@MainActor
extension VoicePipeline {
func resolvedSpokenEditCommand(
raw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async -> SpokenEditCommand? {
guard VoicePipelinePolicy.shouldResolveEditCommandWithLLMFirst(outputMode: settings.outputMode) else {
return nil
}

let resolution = await resolveSpokenEditCommandWithLLM(
raw: raw,
settings: settings,
targetApp: targetApp
)
return VoicePipelinePolicy.editCommand(from: resolution)
}

func spokenEditCommandResolutionContext(
targetApp: NSRunningApplication?
) async -> SpokenEditCommandResolutionContext {
let lastInsertedText = appState.lastInsertedText.trimmingCharacters(in: .whitespacesAndNewlines)
let lastInsertion = lastInsertedText.isEmpty
? SpokenEditCommandTargetAvailability.unavailable
: .available
let selectedText = await textInserter.selectedText(targetApp: targetApp)
let selectionAvailability: SpokenEditCommandTargetAvailability
if let selectedText {
selectionAvailability = selectedText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty
? .unavailable
: .available
} else {
selectionAvailability = .unknown
}

return SpokenEditCommandResolutionContext(
lastInsertion: lastInsertion,
selectedText: selectionAvailability,
lastInsertionPreview: SpokenEditCommandResolutionContext.preview(lastInsertedText),
selectedTextPreview: SpokenEditCommandResolutionContext.preview(selectedText)
)
}

private func resolveSpokenEditCommandWithLLM(
raw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async -> SpokenEditCommandLLMResolution? {
var options = TextProcessingOptions(settings: settings)
options.llmModel = settings.llmModel
return await textProcessor.resolveSpokenEditCommandResolution(
text: raw,
options: options,
context: await spokenEditCommandResolutionContext(targetApp: targetApp)
)
}
}
282 changes: 282 additions & 0 deletions Sources/App/VoicePipeline+EditCommands.swift
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,282 @@
import AppKit
import Foundation

@MainActor
extension VoicePipeline {
func handleSpokenEditCommandIfNeeded(
raw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async -> Bool {
guard let command = await resolvedSpokenEditCommand(
raw: raw,
settings: settings,
targetApp: targetApp
) else {
return false
}

switch command {
case .replaceLast(let replacementRaw):
await replaceLastInsertion(
raw: raw,
replacementRaw: replacementRaw,
settings: settings,
targetApp: targetApp
)
return true
case .replaceSelection(let replacementRaw):
await replaceSelectedText(
raw: raw,
replacementRaw: replacementRaw,
settings: settings,
targetApp: targetApp
)
return true
case .rewriteLast(let intent):
await rewriteLastInsertion(
raw: raw,
intent: intent,
settings: settings,
targetApp: targetApp
)
return true
case .rewriteSelection(let intent):
await rewriteSelectedText(
raw: raw,
intent: intent,
settings: settings,
targetApp: targetApp
)
return true
case .deleteSelection:
await deleteSelectedText(targetApp: targetApp)
return true
case .undoLastInsertion:
await undoLastInsertion(targetApp: targetApp)
return true
}
}

private func replaceLastInsertion(
raw: String,
replacementRaw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async {
cancelScreenContextCapture()

guard !appState.lastInsertedText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty else {
showErrorHint(L("pipeline.no_previous_insert_to_replace"))
return
}

let context = replacementInputContext(settings: settings, targetApp: targetApp)
let replacementText = finalizedReplacementText(replacementRaw, settings: settings)
guard !replacementText.isEmpty else {
showNoSpeechDetected(reason: "spoken edit command has empty replacement text")
return
}

appState.processedText = replacementText
appState.phase = .inserting
appState.statusMessage = L("pipeline.replacing")

Log.sensitive("[VoicePipeline] voice edit replace \(replacementText.count) chars")
let result = await textInserter.replaceRecentInsertion(text: replacementText, targetApp: targetApp)

InputHistory.shared.addRecord(
rawText: raw,
processedText: replacementText,
wasProcessed: true,
context: context
)

appState.lastInsertedText = replacementText
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit replacement probably failed: \(reason)")
TextInserter.copyToClipboard(replacementText)
showInsertionFailedAlert(text: replacementText, reason: reason)
}
}

private func replaceSelectedText(
raw: String,
replacementRaw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async {
cancelScreenContextCapture()

let context = replacementInputContext(settings: settings, targetApp: targetApp)
let replacementText = finalizedReplacementText(replacementRaw, settings: settings)
guard !replacementText.isEmpty else {
showNoSpeechDetected(reason: "spoken edit command has empty selection replacement text")
return
}

appState.processedText = replacementText
appState.phase = .inserting
appState.statusMessage = L("pipeline.replacing")

Log.sensitive("[VoicePipeline] voice edit replace selection \(replacementText.count) chars")
let result = await textInserter.replaceSelectedText(text: replacementText, targetApp: targetApp)

InputHistory.shared.addRecord(
rawText: raw,
processedText: replacementText,
wasProcessed: true,
context: context
)

appState.lastInsertedText = replacementText
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit selection replacement probably failed: \(reason)")
TextInserter.copyToClipboard(replacementText)
showInsertionFailedAlert(text: replacementText, reason: reason)
}
}

private func replacementInputContext(
settings: AppSettings,
targetApp: NSRunningApplication?
) -> InputContext {
InputContext.capture(
targetApp: targetApp,
screenContext: "",
outputMode: .command,
inputLanguage: settings.inputLanguage,
source: .menuBar
)
}

private func finalizedReplacementText(
_ text: String,
settings: AppSettings
) -> String {
textProcessor.cleanCommandGeneratedOutput(
text,
inputLanguage: settings.inputLanguage
)
}

private func deleteSelectedText(targetApp: NSRunningApplication?) async {
cancelScreenContextCapture()

appState.processedText = ""
appState.phase = .inserting
appState.statusMessage = L("pipeline.replacing")

Log.info("[VoicePipeline] voice edit delete selection")
let result = await textInserter.deleteSelectedText(targetApp: targetApp)

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit delete selection probably failed: \(reason)")
showErrorHint(reason)
return
}

appState.lastInsertedText = ""
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()
}

private func undoLastInsertion(targetApp: NSRunningApplication?) async {
cancelScreenContextCapture()

guard !appState.lastInsertedText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty else {
showErrorHint(L("pipeline.no_previous_insert_to_replace"))
return
}

appState.processedText = ""
appState.phase = .inserting
appState.statusMessage = L("pipeline.undoing")

Log.info("[VoicePipeline] voice edit undo last insertion")
let result = await textInserter.undoRecentInsertion(targetApp: targetApp)

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit undo probably failed: \(reason)")
showErrorHint(reason)
return
}

appState.lastInsertedText = ""
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()
}

private func rewriteSelectedText(
raw: String,
intent: SelectionRewriteIntent,
settings: AppSettings,
targetApp: NSRunningApplication?
) async {
cancelScreenContextCapture()

guard let selectedText = await textInserter.selectedText(targetApp: targetApp),
!selectedText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty else {
showErrorHint(L("pipeline.no_selected_text_to_replace"))
return
}

let context = InputContext.capture(
targetApp: targetApp,
screenContext: selectedText,
outputMode: .command,
inputLanguage: settings.inputLanguage,
source: .menuBar
)
var options = TextProcessingOptions(settings: settings)
options.llmModel = settings.llmModel
let memoryContext = VoicePipelinePolicy.memoryContext(
for: .command,
settings: settings,
currentContext: context
)

appState.phase = .processing
appState.statusMessage = L("pipeline.formatting")
let rewrittenText = await textProcessor.processSelectionEdit(
selectedText: selectedText,
intent: intent,
options: options,
spokenCommand: raw,
memoryContext: memoryContext,
inputContext: context
)
guard !rewrittenText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty else {
showNoSpeechDetected(reason: "selection rewrite returned empty text")
return
}

appState.processedText = rewrittenText
appState.phase = .inserting
appState.statusMessage = L("pipeline.replacing")

let result = await textInserter.replaceSelectedText(text: rewrittenText, targetApp: targetApp)
InputHistory.shared.addRecord(rawText: raw, processedText: rewrittenText, wasProcessed: true, context: context)

appState.lastInsertedText = rewrittenText
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit selection rewrite probably failed: \(reason)")
TextInserter.copyToClipboard(rewrittenText)
showInsertionFailedAlert(text: rewrittenText, reason: reason)
}
}
}
Loading
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 4 additions & 1 deletion README.md
Original file line numberDiff line numberDiff line change
Expand Up@@ -39,7 +39,10 @@ Three output modes are available:
| Feature | Description |
|---|---|
| **Multiple Speech Engines** | Apple Speech, WhisperKit, Doubao ASR, Qwen3-ASR, or MiMo-V2.5-ASR |
| **Smart Text Processing** | Local MLX Qwen2.5/Qwen3 or remote LLM — contextual cleanup, self-correction handling, structured list formatting |
| **Smart Text Processing** | Local MLX Qwen2.5/Qwen3 or remote LLM infers spoken intent — contextual cleanup, "scratch that" restarts, self-correction handling, spoken punctuation, technical terms, numbers/ranges/units, and structured formatting |
| **LLM-Owned Spoken Formatting** | Spoken casing, no-space dictation, identifiers, file paths, shortcuts, emoji, Markdown tasks, dates/times, quantities, units, formulas, fractions, and digit sequences are handled by the Smart Format / Voice Command prompts instead of local hardcoded rewrite rules |
| **Voice Edit Commands** | In Voice Command mode, an LLM classifies safe structured actions for replacing, undoing, proofreading, titling, summarizing, drafting replies, making meeting notes, extracting key points/decisions/questions/risks/deadlines/owners/action items, rewriting tone, expanding, making tables/lists, or deleting the previous OpenType insertion or selected text |
| **Verbatim & Preview Boundary** | Verbatim mode, streaming HUD, integration partials, and instant-insert drafts keep ASR text close to raw output with only dictionary, whitespace, duplicate, and non-speech-artifact cleanup |
| **Remote LLM Support** | OpenAI, Claude (Anthropic format), Gemini, OpenRouter, SiliconFlow, Doubao, Bailian, MiniMax (CN & Global) |
| **Global Hotkey** | Configurable key (Fn/Ctrl/Shift/Option) with long-press, double-tap, or single-tap activation |
| **Screen Context OCR** | Captures on-screen text via ScreenCaptureKit + Vision to help the LLM correct homophones |
Expand Down
5 changes: 4 additions & 1 deletion README_zh.md
Original file line numberDiff line numberDiff line change
Expand Up@@ -39,7 +39,10 @@
| 功能 | 说明 |
|---|---|
| **多语音引擎** | Apple 语音识别、WhisperKit、豆包语音识别、Qwen3-ASR 或 MiMo-V2.5-ASR |
| **智能文字处理** | 本地 MLX Qwen2.5/Qwen3 或远程 LLM — 上下文感知的语气词清理、自动纠正、列表格式化 |
| **智能文字处理** | 本地 MLX Qwen2.5/Qwen3 或远程 LLM 理解口述意图 — 上下文感知的语气词清理、“算了/删掉刚才”重说处理、自动纠正、口述标点、技术词、数字/范围/单位和列表格式化 |
| **LLM 负责口述格式** | 大小写、无空格、标识符、文件路径、快捷键、表情、Markdown 任务、日期时间、数量、单位、公式、分数和数字串都由智能整理/语音指令提示词交给 LLM 判断,不在本地写死替换规则 |
| **语音编辑口令** | 在语音指令模式下,由 LLM 分类安全结构化动作,支持上一段/选区替换、撤销、校对、跨语言回复起草、接受/拒绝/追问回复、会议纪要、关键要点/结论/问题/风险/截止时间/负责人/行动项提取、标题化、摘要、语气改写、扩写、表格化、列表化、删除与改写口令 |
| **直出与预览边界** | 原文直出、流式 HUD、集成 partial 和快速插入草稿尽量保留 ASR 原文,只做词库、空白、重复转写和非语音垃圾过滤 |
| **远程 LLM** | 支持 OpenAI、Claude(Anthropic 格式)、Gemini、OpenRouter、硅基流动、豆包、百炼、MiniMax(国内/海外) |
| **全局快捷键** | 可配置按键(Fn/Ctrl/Shift/Option),支持长按、双击、单击三种触发模式 |
| **屏幕上下文 OCR** | 通过 ScreenCaptureKit + Vision 截取屏幕文字,辅助 LLM 纠正同音字 |
Expand Down
61 changes: 61 additions & 0 deletions Sources/App/VoicePipeline+EditCommandResolution.swift
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,61 @@
import AppKit
import Foundation

@MainActor
extension VoicePipeline {
func resolvedSpokenEditCommand(
raw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async -> SpokenEditCommand? {
guard VoicePipelinePolicy.shouldResolveEditCommandWithLLMFirst(outputMode: settings.outputMode) else {
return nil
}

let resolution = await resolveSpokenEditCommandWithLLM(
raw: raw,
settings: settings,
targetApp: targetApp
)
return VoicePipelinePolicy.editCommand(from: resolution)
}

func spokenEditCommandResolutionContext(
targetApp: NSRunningApplication?
) async -> SpokenEditCommandResolutionContext {
let lastInsertedText = appState.lastInsertedText.trimmingCharacters(in: .whitespacesAndNewlines)
let lastInsertion = lastInsertedText.isEmpty
? SpokenEditCommandTargetAvailability.unavailable
: .available
let selectedText = await textInserter.selectedText(targetApp: targetApp)
let selectionAvailability: SpokenEditCommandTargetAvailability
if let selectedText {
selectionAvailability = selectedText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty
? .unavailable
: .available
} else {
selectionAvailability = .unknown
}

return SpokenEditCommandResolutionContext(
lastInsertion: lastInsertion,
selectedText: selectionAvailability,
lastInsertionPreview: SpokenEditCommandResolutionContext.preview(lastInsertedText),
selectedTextPreview: SpokenEditCommandResolutionContext.preview(selectedText)
)
}

private func resolveSpokenEditCommandWithLLM(
raw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async -> SpokenEditCommandLLMResolution? {
var options = TextProcessingOptions(settings: settings)
options.llmModel = settings.llmModel
return await textProcessor.resolveSpokenEditCommandResolution(
text: raw,
options: options,
context: await spokenEditCommandResolutionContext(targetApp: targetApp)
)
}
}
282 changes: 282 additions & 0 deletions Sources/App/VoicePipeline+EditCommands.swift
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,282 @@
import AppKit
import Foundation

@MainActor
extension VoicePipeline {
func handleSpokenEditCommandIfNeeded(
raw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async -> Bool {
guard let command = await resolvedSpokenEditCommand(
raw: raw,
settings: settings,
targetApp: targetApp
) else {
return false
}

switch command {
case .replaceLast(let replacementRaw):
await replaceLastInsertion(
raw: raw,
replacementRaw: replacementRaw,
settings: settings,
targetApp: targetApp
)
return true
case .replaceSelection(let replacementRaw):
await replaceSelectedText(
raw: raw,
replacementRaw: replacementRaw,
settings: settings,
targetApp: targetApp
)
return true
case .rewriteLast(let intent):
await rewriteLastInsertion(
raw: raw,
intent: intent,
settings: settings,
targetApp: targetApp
)
return true
case .rewriteSelection(let intent):
await rewriteSelectedText(
raw: raw,
intent: intent,
settings: settings,
targetApp: targetApp
)
return true
case .deleteSelection:
await deleteSelectedText(targetApp: targetApp)
return true
case .undoLastInsertion:
await undoLastInsertion(targetApp: targetApp)
return true
}
}

private func replaceLastInsertion(
raw: String,
replacementRaw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async {
cancelScreenContextCapture()

guard !appState.lastInsertedText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty else {
showErrorHint(L("pipeline.no_previous_insert_to_replace"))
return
}

let context = replacementInputContext(settings: settings, targetApp: targetApp)
let replacementText = finalizedReplacementText(replacementRaw, settings: settings)
guard !replacementText.isEmpty else {
showNoSpeechDetected(reason: "spoken edit command has empty replacement text")
return
}

appState.processedText = replacementText
appState.phase = .inserting
appState.statusMessage = L("pipeline.replacing")

Log.sensitive("[VoicePipeline] voice edit replace \(replacementText.count) chars")
let result = await textInserter.replaceRecentInsertion(text: replacementText, targetApp: targetApp)

InputHistory.shared.addRecord(
rawText: raw,
processedText: replacementText,
wasProcessed: true,
context: context
)

appState.lastInsertedText = replacementText
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit replacement probably failed: \(reason)")
TextInserter.copyToClipboard(replacementText)
showInsertionFailedAlert(text: replacementText, reason: reason)
}
}

private func replaceSelectedText(
raw: String,
replacementRaw: String,
settings: AppSettings,
targetApp: NSRunningApplication?
) async {
cancelScreenContextCapture()

let context = replacementInputContext(settings: settings, targetApp: targetApp)
let replacementText = finalizedReplacementText(replacementRaw, settings: settings)
guard !replacementText.isEmpty else {
showNoSpeechDetected(reason: "spoken edit command has empty selection replacement text")
return
}

appState.processedText = replacementText
appState.phase = .inserting
appState.statusMessage = L("pipeline.replacing")

Log.sensitive("[VoicePipeline] voice edit replace selection \(replacementText.count) chars")
let result = await textInserter.replaceSelectedText(text: replacementText, targetApp: targetApp)

InputHistory.shared.addRecord(
rawText: raw,
processedText: replacementText,
wasProcessed: true,
context: context
)

appState.lastInsertedText = replacementText
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit selection replacement probably failed: \(reason)")
TextInserter.copyToClipboard(replacementText)
showInsertionFailedAlert(text: replacementText, reason: reason)
}
}

private func replacementInputContext(
settings: AppSettings,
targetApp: NSRunningApplication?
) -> InputContext {
InputContext.capture(
targetApp: targetApp,
screenContext: "",
outputMode: .command,
inputLanguage: settings.inputLanguage,
source: .menuBar
)
}

private func finalizedReplacementText(
_ text: String,
settings: AppSettings
) -> String {
textProcessor.cleanCommandGeneratedOutput(
text,
inputLanguage: settings.inputLanguage
)
}

private func deleteSelectedText(targetApp: NSRunningApplication?) async {
cancelScreenContextCapture()

appState.processedText = ""
appState.phase = .inserting
appState.statusMessage = L("pipeline.replacing")

Log.info("[VoicePipeline] voice edit delete selection")
let result = await textInserter.deleteSelectedText(targetApp: targetApp)

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit delete selection probably failed: \(reason)")
showErrorHint(reason)
return
}

appState.lastInsertedText = ""
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()
}

private func undoLastInsertion(targetApp: NSRunningApplication?) async {
cancelScreenContextCapture()

guard !appState.lastInsertedText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty else {
showErrorHint(L("pipeline.no_previous_insert_to_replace"))
return
}

appState.processedText = ""
appState.phase = .inserting
appState.statusMessage = L("pipeline.undoing")

Log.info("[VoicePipeline] voice edit undo last insertion")
let result = await textInserter.undoRecentInsertion(targetApp: targetApp)

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit undo probably failed: \(reason)")
showErrorHint(reason)
return
}

appState.lastInsertedText = ""
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()
}

private func rewriteSelectedText(
raw: String,
intent: SelectionRewriteIntent,
settings: AppSettings,
targetApp: NSRunningApplication?
) async {
cancelScreenContextCapture()

guard let selectedText = await textInserter.selectedText(targetApp: targetApp),
!selectedText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty else {
showErrorHint(L("pipeline.no_selected_text_to_replace"))
return
}

let context = InputContext.capture(
targetApp: targetApp,
screenContext: selectedText,
outputMode: .command,
inputLanguage: settings.inputLanguage,
source: .menuBar
)
var options = TextProcessingOptions(settings: settings)
options.llmModel = settings.llmModel
let memoryContext = VoicePipelinePolicy.memoryContext(
for: .command,
settings: settings,
currentContext: context
)

appState.phase = .processing
appState.statusMessage = L("pipeline.formatting")
let rewrittenText = await textProcessor.processSelectionEdit(
selectedText: selectedText,
intent: intent,
options: options,
spokenCommand: raw,
memoryContext: memoryContext,
inputContext: context
)
guard !rewrittenText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty else {
showNoSpeechDetected(reason: "selection rewrite returned empty text")
return
}

appState.processedText = rewrittenText
appState.phase = .inserting
appState.statusMessage = L("pipeline.replacing")

let result = await textInserter.replaceSelectedText(text: rewrittenText, targetApp: targetApp)
InputHistory.shared.addRecord(rawText: raw, processedText: rewrittenText, wasProcessed: true, context: context)

appState.lastInsertedText = rewrittenText
appState.phase = .done
appState.statusMessage = L("status.done")
hideOverlayAfterDelay()

if case .probablyFailed(let reason) = result {
Log.info("[VoicePipeline] voice edit selection rewrite probably failed: \(reason)")
TextInserter.copyToClipboard(rewrittenText)
showInsertionFailedAlert(text: rewrittenText, reason: reason)
}
}
}
Loading
Loading