Improve voice recognition and correction quality - #71
Merged
Conversation
IchenDEV
marked this pull request as ready for review
July 29, 2026 18:20
Uh oh!
There was an error while loading. Please reload this page.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for freeto join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Why
Recognition quality was being reduced by inconsistent language and dictionary context, brittle Whisper model resolution, unsuitable Qwen input audio, and overly shallow decoding fallbacks. The correction layer could also accept truncated or semantically altered LLM output, while the project lacked a repeatable way to measure raw ASR quality separately from post-processing quality.
Impact
Users should get more robust recognition across Apple Speech, Whisper, and Qwen paths, with better handling of personal vocabulary and fewer correction-stage changes to numbers, dates, units, URLs, email addresses, file paths, quoted text, polarity, and self-corrections. When correction cannot be proven faithful, OpenType now preserves the prepared recognition result instead of inserting a risky rewrite.
The included example corpus validates the evaluator pipeline only. It is not evidence of production accuracy improvement; real-device recordings are still required for comparative benchmarking.
Validation
swift test: 529 XCTest tests passed, 7 environment-dependent tests skippedbash scripts/ci-basic-checks.sh: passedResearch
See
docs/superpowers/specs/2026-07-30-voice-quality-research.mdfor source evaluation, licensing notes, phased recommendations, and the proposed real-corpus benchmark.