Expose word timestamps for Parakeet - #93

Open
ArturWierzbicki wants to merge 1 commit into
handy-computer:mainfrom
ArturWierzbicki:parakeet-timestamp-verification
Open

Expose word timestamps for Parakeet#93
ArturWierzbicki wants to merge 1 commit into
handy-computer:mainfrom
ArturWierzbicki:parakeet-timestamp-verification

Conversation

@ArturWierzbicki

@ArturWierzbickiArturWierzbicki commented Jul 19, 2026

Copy link
Copy Markdown

Summary

I have a tool that turns recorded meetings into markdown transcripts with named speakers and screenshots of what was shared on screen. Today it transcribes with Parakeet TDT via parakeet-mlx. I'd love to move to transcribe.cpp to simplify the setup.

To do so, I need word timestamps: I use them to match words to diarized speaker turns and to merge overlapping audio chunks by cutting at silences between words. transcribe.cpp already computes them, but the per-file JSON from --batch-jsonl does not expose them. This PR exposes them and verifies them: the batch JSON gains words and tokens arrays, the parakeet NeMo dumper records word timings in its reference output, and validate.py compares each word's start and end against the reference within a new 80 ms tolerance, the duration of one encoder frame. The comparison code is shared across model families; a family turns it on by dumping reference word timings and adding a timestamps entry to its tolerance file.

The check caught a real bug on its first run. A comma is its own token, and the model predicts a duration for it as it does for every token, even though nothing is spoken there.

Before this PR the decode kept those predicted durations, and word ends inherited them, drifting up to 560 ms past the speech on the validation clip. NeMo gives punctuation zero duration instead: in JFK's "And so, my fellow Americans, ...", the token " so" ends at 800 ms and the "," sits exactly there, so the word "so," ends when the audio does. The offline decode in src/arch/parakeet/model.cpp now applies the same rule; transcript text is unchanged.

Scope

  • src/arch/parakeet: punctuation tokens get zero duration in offline decode; the set of punctuation tokens is built once at model load, and vocab entries that are not valid UTF-8 skip the rewrite
  • examples/cli: words appears in batch JSONL when a run returns word or token timestamps, tokens only at token level
  • scripts: the parakeet reference dumper emits word and token timings; validate.py gains the word-timestamp comparison and unit tests
  • tests: a timestamps tolerance entry for parakeet; real-model smoke assertions for zero-duration punctuation
  • docs: tools/validate, model-family-testing, and the v2 and v3 model cards

AI Assistance

Yes, heavy Fable usage: code/docs/investigation. I directed and reviewed the work and own the change, happy to iterate more. I'm not an ASR/C++ person, so I also iterated upon a human-digestible :), zero-context explainer of the whole change for me and anyone else who wants the background: https://gist.github.com/ArturWierzbicki/3a5e21740703103e0a89a51428a862a6

Validation

  • validate.py all for parakeet-tdt-0.6b-v2: 18/18 tensors within tolerance, exact transcript, and word timestamps that deviate from the reference by 80 ms max and 3.6 ms mean.
  • validate.py all for parakeet-tdt-0.6b-v3: exact transcript, and word timestamps that match the reference exactly (0.000 ms max deviation).
  • ctest 31/31, pinned clang-format clean, validate.py unit tests pass.
  • No WER impact: transcript text is unchanged (the transcript compare stays exact). The batch JSONL change also reaches granite, gigaam, and medasr, which report word or token timestamps too; the WER harness reads only the text field, so their runs are unaffected.
  • The punctuation rule changes public token and word timings for every TDT-head variant. v2 and v3 are validated here; tdt-1.1b and the tdt_ctc hybrids are not, and their word-timestamp check turns on automatically the next time their reference dumps are regenerated.
  • The opt-in real-model smoke test: every assertion this PR adds passes; 5 pre-existing failures remain (a rejected call leaves the previous result exposed)

Verify Parakeet TDT word timestamps against the NeMo reference:
- dump reference word and token timings in the parakeet NeMo dumper
(transcript.json gains word rows; token rows gain timing)
- emit words/tokens in the CLI batch JSON, gated on the returned
timestamp kind
- compare word start/end against the reference in validate.py, within
the new "timestamps" budget in tests/tolerances/<family>.json
(80 ms = one encoder frame for parakeet)
- mirror NeMo's punctuation convention in the decoder: punctuation-only
tokens are zero-length at the previous token's end, guarded for
non-UTF-8 vocab pieces, with the punctuation set cached at model
load. The first validation run caught a 560 ms word-end divergence
this fixes; the re-run passes with max 80 ms and mean 3.6 ms on the
validation clip, and parakeet-tdt-0.6b-v3 matches the reference
exactly
- update validate, model-family-testing, and model-card docs
@cjpais

Copy link
Copy Markdown
Contributor

Thanks for this PR. I will take a look at it soon and then try to pull it. And I think this is pretty reasonable

@cjpaiscjpais closed this Jul 19, 2026
@cjpaiscjpais reopened this Jul 19, 2026
@ArturWierzbicki

Copy link
Copy Markdown
Author

hey @cjpais anything I can do to help? :) should I try resolving conflicts?

@cjpais

cjpais commented Aug 15, 2026

Copy link
Copy Markdown
Contributor

Mostly just blocked on me having time to take a look at it. So just give me a bit of time. I'm trying to handle some issues in handy right now that are taking priority and also get the release out for 0.2.0 of this library. This change will likely make it into 0.2.1 after I test

But yes rebasing would be helpful

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@ArturWierzbicki@cjpais
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Expose word timestamps for Parakeet - #93

Open
ArturWierzbicki wants to merge 1 commit into
handy-computer:mainfrom
ArturWierzbicki:parakeet-timestamp-verification
Open

Expose word timestamps for Parakeet#93
ArturWierzbicki wants to merge 1 commit into
handy-computer:mainfrom
ArturWierzbicki:parakeet-timestamp-verification

Conversation

@ArturWierzbicki

@ArturWierzbickiArturWierzbicki commented Jul 19, 2026

Copy link
Copy Markdown

Summary

I have a tool that turns recorded meetings into markdown transcripts with named speakers and screenshots of what was shared on screen. Today it transcribes with Parakeet TDT via parakeet-mlx. I'd love to move to transcribe.cpp to simplify the setup.

To do so, I need word timestamps: I use them to match words to diarized speaker turns and to merge overlapping audio chunks by cutting at silences between words. transcribe.cpp already computes them, but the per-file JSON from --batch-jsonl does not expose them. This PR exposes them and verifies them: the batch JSON gains words and tokens arrays, the parakeet NeMo dumper records word timings in its reference output, and validate.py compares each word's start and end against the reference within a new 80 ms tolerance, the duration of one encoder frame. The comparison code is shared across model families; a family turns it on by dumping reference word timings and adding a timestamps entry to its tolerance file.

The check caught a real bug on its first run. A comma is its own token, and the model predicts a duration for it as it does for every token, even though nothing is spoken there.

Before this PR the decode kept those predicted durations, and word ends inherited them, drifting up to 560 ms past the speech on the validation clip. NeMo gives punctuation zero duration instead: in JFK's "And so, my fellow Americans, ...", the token " so" ends at 800 ms and the "," sits exactly there, so the word "so," ends when the audio does. The offline decode in src/arch/parakeet/model.cpp now applies the same rule; transcript text is unchanged.

Scope

  • src/arch/parakeet: punctuation tokens get zero duration in offline decode; the set of punctuation tokens is built once at model load, and vocab entries that are not valid UTF-8 skip the rewrite
  • examples/cli: words appears in batch JSONL when a run returns word or token timestamps, tokens only at token level
  • scripts: the parakeet reference dumper emits word and token timings; validate.py gains the word-timestamp comparison and unit tests
  • tests: a timestamps tolerance entry for parakeet; real-model smoke assertions for zero-duration punctuation
  • docs: tools/validate, model-family-testing, and the v2 and v3 model cards

AI Assistance

Yes, heavy Fable usage: code/docs/investigation. I directed and reviewed the work and own the change, happy to iterate more. I'm not an ASR/C++ person, so I also iterated upon a human-digestible :), zero-context explainer of the whole change for me and anyone else who wants the background: https://gist.github.com/ArturWierzbicki/3a5e21740703103e0a89a51428a862a6

Validation

  • validate.py all for parakeet-tdt-0.6b-v2: 18/18 tensors within tolerance, exact transcript, and word timestamps that deviate from the reference by 80 ms max and 3.6 ms mean.
  • validate.py all for parakeet-tdt-0.6b-v3: exact transcript, and word timestamps that match the reference exactly (0.000 ms max deviation).
  • ctest 31/31, pinned clang-format clean, validate.py unit tests pass.
  • No WER impact: transcript text is unchanged (the transcript compare stays exact). The batch JSONL change also reaches granite, gigaam, and medasr, which report word or token timestamps too; the WER harness reads only the text field, so their runs are unaffected.
  • The punctuation rule changes public token and word timings for every TDT-head variant. v2 and v3 are validated here; tdt-1.1b and the tdt_ctc hybrids are not, and their word-timestamp check turns on automatically the next time their reference dumps are regenerated.
  • The opt-in real-model smoke test: every assertion this PR adds passes; 5 pre-existing failures remain (a rejected call leaves the previous result exposed)

Verify Parakeet TDT word timestamps against the NeMo reference:
- dump reference word and token timings in the parakeet NeMo dumper
(transcript.json gains word rows; token rows gain timing)
- emit words/tokens in the CLI batch JSON, gated on the returned
timestamp kind
- compare word start/end against the reference in validate.py, within
the new "timestamps" budget in tests/tolerances/<family>.json
(80 ms = one encoder frame for parakeet)
- mirror NeMo's punctuation convention in the decoder: punctuation-only
tokens are zero-length at the previous token's end, guarded for
non-UTF-8 vocab pieces, with the punctuation set cached at model
load. The first validation run caught a 560 ms word-end divergence
this fixes; the re-run passes with max 80 ms and mean 3.6 ms on the
validation clip, and parakeet-tdt-0.6b-v3 matches the reference
exactly
- update validate, model-family-testing, and model-card docs
@cjpais

Copy link
Copy Markdown
Contributor

Thanks for this PR. I will take a look at it soon and then try to pull it. And I think this is pretty reasonable

@cjpaiscjpais closed this Jul 19, 2026
@cjpaiscjpais reopened this Jul 19, 2026
@ArturWierzbicki

Copy link
Copy Markdown
Author

hey @cjpais anything I can do to help? :) should I try resolving conflicts?

@cjpais

cjpais commented Aug 15, 2026

Copy link
Copy Markdown
Contributor

Mostly just blocked on me having time to take a look at it. So just give me a bit of time. I'm trying to handle some issues in handy right now that are taking priority and also get the release out for 0.2.0 of this library. This change will likely make it into 0.2.1 after I test

But yes rebasing would be helpful

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@ArturWierzbicki@cjpais
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Expose word timestamps for Parakeet - #93

Open
ArturWierzbicki wants to merge 1 commit into
handy-computer:mainfrom
ArturWierzbicki:parakeet-timestamp-verification
Open

Expose word timestamps for Parakeet#93
ArturWierzbicki wants to merge 1 commit into
handy-computer:mainfrom
ArturWierzbicki:parakeet-timestamp-verification

Conversation

@ArturWierzbicki

@ArturWierzbickiArturWierzbicki commented Jul 19, 2026

Copy link
Copy Markdown

Summary

I have a tool that turns recorded meetings into markdown transcripts with named speakers and screenshots of what was shared on screen. Today it transcribes with Parakeet TDT via parakeet-mlx. I'd love to move to transcribe.cpp to simplify the setup.

To do so, I need word timestamps: I use them to match words to diarized speaker turns and to merge overlapping audio chunks by cutting at silences between words. transcribe.cpp already computes them, but the per-file JSON from --batch-jsonl does not expose them. This PR exposes them and verifies them: the batch JSON gains words and tokens arrays, the parakeet NeMo dumper records word timings in its reference output, and validate.py compares each word's start and end against the reference within a new 80 ms tolerance, the duration of one encoder frame. The comparison code is shared across model families; a family turns it on by dumping reference word timings and adding a timestamps entry to its tolerance file.

The check caught a real bug on its first run. A comma is its own token, and the model predicts a duration for it as it does for every token, even though nothing is spoken there.

Before this PR the decode kept those predicted durations, and word ends inherited them, drifting up to 560 ms past the speech on the validation clip. NeMo gives punctuation zero duration instead: in JFK's "And so, my fellow Americans, ...", the token " so" ends at 800 ms and the "," sits exactly there, so the word "so," ends when the audio does. The offline decode in src/arch/parakeet/model.cpp now applies the same rule; transcript text is unchanged.

Scope

  • src/arch/parakeet: punctuation tokens get zero duration in offline decode; the set of punctuation tokens is built once at model load, and vocab entries that are not valid UTF-8 skip the rewrite
  • examples/cli: words appears in batch JSONL when a run returns word or token timestamps, tokens only at token level
  • scripts: the parakeet reference dumper emits word and token timings; validate.py gains the word-timestamp comparison and unit tests
  • tests: a timestamps tolerance entry for parakeet; real-model smoke assertions for zero-duration punctuation
  • docs: tools/validate, model-family-testing, and the v2 and v3 model cards

AI Assistance

Yes, heavy Fable usage: code/docs/investigation. I directed and reviewed the work and own the change, happy to iterate more. I'm not an ASR/C++ person, so I also iterated upon a human-digestible :), zero-context explainer of the whole change for me and anyone else who wants the background: https://gist.github.com/ArturWierzbicki/3a5e21740703103e0a89a51428a862a6

Validation

  • validate.py all for parakeet-tdt-0.6b-v2: 18/18 tensors within tolerance, exact transcript, and word timestamps that deviate from the reference by 80 ms max and 3.6 ms mean.
  • validate.py all for parakeet-tdt-0.6b-v3: exact transcript, and word timestamps that match the reference exactly (0.000 ms max deviation).
  • ctest 31/31, pinned clang-format clean, validate.py unit tests pass.
  • No WER impact: transcript text is unchanged (the transcript compare stays exact). The batch JSONL change also reaches granite, gigaam, and medasr, which report word or token timestamps too; the WER harness reads only the text field, so their runs are unaffected.
  • The punctuation rule changes public token and word timings for every TDT-head variant. v2 and v3 are validated here; tdt-1.1b and the tdt_ctc hybrids are not, and their word-timestamp check turns on automatically the next time their reference dumps are regenerated.
  • The opt-in real-model smoke test: every assertion this PR adds passes; 5 pre-existing failures remain (a rejected call leaves the previous result exposed)

Verify Parakeet TDT word timestamps against the NeMo reference:
- dump reference word and token timings in the parakeet NeMo dumper
(transcript.json gains word rows; token rows gain timing)
- emit words/tokens in the CLI batch JSON, gated on the returned
timestamp kind
- compare word start/end against the reference in validate.py, within
the new "timestamps" budget in tests/tolerances/<family>.json
(80 ms = one encoder frame for parakeet)
- mirror NeMo's punctuation convention in the decoder: punctuation-only
tokens are zero-length at the previous token's end, guarded for
non-UTF-8 vocab pieces, with the punctuation set cached at model
load. The first validation run caught a 560 ms word-end divergence
this fixes; the re-run passes with max 80 ms and mean 3.6 ms on the
validation clip, and parakeet-tdt-0.6b-v3 matches the reference
exactly
- update validate, model-family-testing, and model-card docs
@cjpais

Copy link
Copy Markdown
Contributor

Thanks for this PR. I will take a look at it soon and then try to pull it. And I think this is pretty reasonable

@cjpaiscjpais closed this Jul 19, 2026
@cjpaiscjpais reopened this Jul 19, 2026
@ArturWierzbicki

Copy link
Copy Markdown
Author

hey @cjpais anything I can do to help? :) should I try resolving conflicts?

@cjpais

cjpais commented Aug 15, 2026

Copy link
Copy Markdown
Contributor

Mostly just blocked on me having time to take a look at it. So just give me a bit of time. I'm trying to handle some issues in handy right now that are taking priority and also get the release out for 0.2.0 of this library. This change will likely make it into 0.2.1 after I test

But yes rebasing would be helpful

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@ArturWierzbicki@cjpais
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Expose word timestamps for Parakeet - #93

Open
ArturWierzbicki wants to merge 1 commit into
handy-computer:mainfrom
ArturWierzbicki:parakeet-timestamp-verification
Open

Expose word timestamps for Parakeet#93
ArturWierzbicki wants to merge 1 commit into
handy-computer:mainfrom
ArturWierzbicki:parakeet-timestamp-verification

Conversation

@ArturWierzbicki

@ArturWierzbickiArturWierzbicki commented Jul 19, 2026

Copy link
Copy Markdown

Summary

I have a tool that turns recorded meetings into markdown transcripts with named speakers and screenshots of what was shared on screen. Today it transcribes with Parakeet TDT via parakeet-mlx. I'd love to move to transcribe.cpp to simplify the setup.

To do so, I need word timestamps: I use them to match words to diarized speaker turns and to merge overlapping audio chunks by cutting at silences between words. transcribe.cpp already computes them, but the per-file JSON from --batch-jsonl does not expose them. This PR exposes them and verifies them: the batch JSON gains words and tokens arrays, the parakeet NeMo dumper records word timings in its reference output, and validate.py compares each word's start and end against the reference within a new 80 ms tolerance, the duration of one encoder frame. The comparison code is shared across model families; a family turns it on by dumping reference word timings and adding a timestamps entry to its tolerance file.

The check caught a real bug on its first run. A comma is its own token, and the model predicts a duration for it as it does for every token, even though nothing is spoken there.

Before this PR the decode kept those predicted durations, and word ends inherited them, drifting up to 560 ms past the speech on the validation clip. NeMo gives punctuation zero duration instead: in JFK's "And so, my fellow Americans, ...", the token " so" ends at 800 ms and the "," sits exactly there, so the word "so," ends when the audio does. The offline decode in src/arch/parakeet/model.cpp now applies the same rule; transcript text is unchanged.

Scope

  • src/arch/parakeet: punctuation tokens get zero duration in offline decode; the set of punctuation tokens is built once at model load, and vocab entries that are not valid UTF-8 skip the rewrite
  • examples/cli: words appears in batch JSONL when a run returns word or token timestamps, tokens only at token level
  • scripts: the parakeet reference dumper emits word and token timings; validate.py gains the word-timestamp comparison and unit tests
  • tests: a timestamps tolerance entry for parakeet; real-model smoke assertions for zero-duration punctuation
  • docs: tools/validate, model-family-testing, and the v2 and v3 model cards

AI Assistance

Yes, heavy Fable usage: code/docs/investigation. I directed and reviewed the work and own the change, happy to iterate more. I'm not an ASR/C++ person, so I also iterated upon a human-digestible :), zero-context explainer of the whole change for me and anyone else who wants the background: https://gist.github.com/ArturWierzbicki/3a5e21740703103e0a89a51428a862a6

Validation

  • validate.py all for parakeet-tdt-0.6b-v2: 18/18 tensors within tolerance, exact transcript, and word timestamps that deviate from the reference by 80 ms max and 3.6 ms mean.
  • validate.py all for parakeet-tdt-0.6b-v3: exact transcript, and word timestamps that match the reference exactly (0.000 ms max deviation).
  • ctest 31/31, pinned clang-format clean, validate.py unit tests pass.
  • No WER impact: transcript text is unchanged (the transcript compare stays exact). The batch JSONL change also reaches granite, gigaam, and medasr, which report word or token timestamps too; the WER harness reads only the text field, so their runs are unaffected.
  • The punctuation rule changes public token and word timings for every TDT-head variant. v2 and v3 are validated here; tdt-1.1b and the tdt_ctc hybrids are not, and their word-timestamp check turns on automatically the next time their reference dumps are regenerated.
  • The opt-in real-model smoke test: every assertion this PR adds passes; 5 pre-existing failures remain (a rejected call leaves the previous result exposed)

Verify Parakeet TDT word timestamps against the NeMo reference:
- dump reference word and token timings in the parakeet NeMo dumper
(transcript.json gains word rows; token rows gain timing)
- emit words/tokens in the CLI batch JSON, gated on the returned
timestamp kind
- compare word start/end against the reference in validate.py, within
the new "timestamps" budget in tests/tolerances/<family>.json
(80 ms = one encoder frame for parakeet)
- mirror NeMo's punctuation convention in the decoder: punctuation-only
tokens are zero-length at the previous token's end, guarded for
non-UTF-8 vocab pieces, with the punctuation set cached at model
load. The first validation run caught a 560 ms word-end divergence
this fixes; the re-run passes with max 80 ms and mean 3.6 ms on the
validation clip, and parakeet-tdt-0.6b-v3 matches the reference
exactly
- update validate, model-family-testing, and model-card docs
@cjpais

Copy link
Copy Markdown
Contributor

Thanks for this PR. I will take a look at it soon and then try to pull it. And I think this is pretty reasonable

@cjpaiscjpais closed this Jul 19, 2026
@cjpaiscjpais reopened this Jul 19, 2026
@ArturWierzbicki

Copy link
Copy Markdown
Author

hey @cjpais anything I can do to help? :) should I try resolving conflicts?

@cjpais

cjpais commented Aug 15, 2026

Copy link
Copy Markdown
Contributor

Mostly just blocked on me having time to take a look at it. So just give me a bit of time. I'm trying to handle some issues in handy right now that are taking priority and also get the release out for 0.2.0 of this library. This change will likely make it into 0.2.1 after I test

But yes rebasing would be helpful

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@ArturWierzbicki@cjpais
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Expose word timestamps for Parakeet - #93

Open
ArturWierzbicki wants to merge 1 commit into
handy-computer:mainfrom
ArturWierzbicki:parakeet-timestamp-verification
Open

Expose word timestamps for Parakeet#93
ArturWierzbicki wants to merge 1 commit into
handy-computer:mainfrom
ArturWierzbicki:parakeet-timestamp-verification

Conversation

@ArturWierzbicki

@ArturWierzbickiArturWierzbicki commented Jul 19, 2026

Copy link
Copy Markdown

Summary

I have a tool that turns recorded meetings into markdown transcripts with named speakers and screenshots of what was shared on screen. Today it transcribes with Parakeet TDT via parakeet-mlx. I'd love to move to transcribe.cpp to simplify the setup.

To do so, I need word timestamps: I use them to match words to diarized speaker turns and to merge overlapping audio chunks by cutting at silences between words. transcribe.cpp already computes them, but the per-file JSON from --batch-jsonl does not expose them. This PR exposes them and verifies them: the batch JSON gains words and tokens arrays, the parakeet NeMo dumper records word timings in its reference output, and validate.py compares each word's start and end against the reference within a new 80 ms tolerance, the duration of one encoder frame. The comparison code is shared across model families; a family turns it on by dumping reference word timings and adding a timestamps entry to its tolerance file.

The check caught a real bug on its first run. A comma is its own token, and the model predicts a duration for it as it does for every token, even though nothing is spoken there.

Before this PR the decode kept those predicted durations, and word ends inherited them, drifting up to 560 ms past the speech on the validation clip. NeMo gives punctuation zero duration instead: in JFK's "And so, my fellow Americans, ...", the token " so" ends at 800 ms and the "," sits exactly there, so the word "so," ends when the audio does. The offline decode in src/arch/parakeet/model.cpp now applies the same rule; transcript text is unchanged.

Scope

  • src/arch/parakeet: punctuation tokens get zero duration in offline decode; the set of punctuation tokens is built once at model load, and vocab entries that are not valid UTF-8 skip the rewrite
  • examples/cli: words appears in batch JSONL when a run returns word or token timestamps, tokens only at token level
  • scripts: the parakeet reference dumper emits word and token timings; validate.py gains the word-timestamp comparison and unit tests
  • tests: a timestamps tolerance entry for parakeet; real-model smoke assertions for zero-duration punctuation
  • docs: tools/validate, model-family-testing, and the v2 and v3 model cards

AI Assistance

Yes, heavy Fable usage: code/docs/investigation. I directed and reviewed the work and own the change, happy to iterate more. I'm not an ASR/C++ person, so I also iterated upon a human-digestible :), zero-context explainer of the whole change for me and anyone else who wants the background: https://gist.github.com/ArturWierzbicki/3a5e21740703103e0a89a51428a862a6

Validation

  • validate.py all for parakeet-tdt-0.6b-v2: 18/18 tensors within tolerance, exact transcript, and word timestamps that deviate from the reference by 80 ms max and 3.6 ms mean.
  • validate.py all for parakeet-tdt-0.6b-v3: exact transcript, and word timestamps that match the reference exactly (0.000 ms max deviation).
  • ctest 31/31, pinned clang-format clean, validate.py unit tests pass.
  • No WER impact: transcript text is unchanged (the transcript compare stays exact). The batch JSONL change also reaches granite, gigaam, and medasr, which report word or token timestamps too; the WER harness reads only the text field, so their runs are unaffected.
  • The punctuation rule changes public token and word timings for every TDT-head variant. v2 and v3 are validated here; tdt-1.1b and the tdt_ctc hybrids are not, and their word-timestamp check turns on automatically the next time their reference dumps are regenerated.
  • The opt-in real-model smoke test: every assertion this PR adds passes; 5 pre-existing failures remain (a rejected call leaves the previous result exposed)

Verify Parakeet TDT word timestamps against the NeMo reference:
- dump reference word and token timings in the parakeet NeMo dumper
(transcript.json gains word rows; token rows gain timing)
- emit words/tokens in the CLI batch JSON, gated on the returned
timestamp kind
- compare word start/end against the reference in validate.py, within
the new "timestamps" budget in tests/tolerances/<family>.json
(80 ms = one encoder frame for parakeet)
- mirror NeMo's punctuation convention in the decoder: punctuation-only
tokens are zero-length at the previous token's end, guarded for
non-UTF-8 vocab pieces, with the punctuation set cached at model
load. The first validation run caught a 560 ms word-end divergence
this fixes; the re-run passes with max 80 ms and mean 3.6 ms on the
validation clip, and parakeet-tdt-0.6b-v3 matches the reference
exactly
- update validate, model-family-testing, and model-card docs
@cjpais

Copy link
Copy Markdown
Contributor

Thanks for this PR. I will take a look at it soon and then try to pull it. And I think this is pretty reasonable

@cjpaiscjpais closed this Jul 19, 2026
@cjpaiscjpais reopened this Jul 19, 2026
@ArturWierzbicki

Copy link
Copy Markdown
Author

hey @cjpais anything I can do to help? :) should I try resolving conflicts?

@cjpais

cjpais commented Aug 15, 2026

Copy link
Copy Markdown
Contributor

Mostly just blocked on me having time to take a look at it. So just give me a bit of time. I'm trying to handle some issues in handy right now that are taking priority and also get the release out for 0.2.0 of this library. This change will likely make it into 0.2.1 after I test

But yes rebasing would be helpful

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@ArturWierzbicki@cjpais
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Expose word timestamps for Parakeet - #93

Open
ArturWierzbicki wants to merge 1 commit into
handy-computer:mainfrom
ArturWierzbicki:parakeet-timestamp-verification
Open

Expose word timestamps for Parakeet#93
ArturWierzbicki wants to merge 1 commit into
handy-computer:mainfrom
ArturWierzbicki:parakeet-timestamp-verification

Conversation

@ArturWierzbicki

@ArturWierzbickiArturWierzbicki commented Jul 19, 2026

Copy link
Copy Markdown

Summary

I have a tool that turns recorded meetings into markdown transcripts with named speakers and screenshots of what was shared on screen. Today it transcribes with Parakeet TDT via parakeet-mlx. I'd love to move to transcribe.cpp to simplify the setup.

To do so, I need word timestamps: I use them to match words to diarized speaker turns and to merge overlapping audio chunks by cutting at silences between words. transcribe.cpp already computes them, but the per-file JSON from --batch-jsonl does not expose them. This PR exposes them and verifies them: the batch JSON gains words and tokens arrays, the parakeet NeMo dumper records word timings in its reference output, and validate.py compares each word's start and end against the reference within a new 80 ms tolerance, the duration of one encoder frame. The comparison code is shared across model families; a family turns it on by dumping reference word timings and adding a timestamps entry to its tolerance file.

The check caught a real bug on its first run. A comma is its own token, and the model predicts a duration for it as it does for every token, even though nothing is spoken there.

Before this PR the decode kept those predicted durations, and word ends inherited them, drifting up to 560 ms past the speech on the validation clip. NeMo gives punctuation zero duration instead: in JFK's "And so, my fellow Americans, ...", the token " so" ends at 800 ms and the "," sits exactly there, so the word "so," ends when the audio does. The offline decode in src/arch/parakeet/model.cpp now applies the same rule; transcript text is unchanged.

Scope

  • src/arch/parakeet: punctuation tokens get zero duration in offline decode; the set of punctuation tokens is built once at model load, and vocab entries that are not valid UTF-8 skip the rewrite
  • examples/cli: words appears in batch JSONL when a run returns word or token timestamps, tokens only at token level
  • scripts: the parakeet reference dumper emits word and token timings; validate.py gains the word-timestamp comparison and unit tests
  • tests: a timestamps tolerance entry for parakeet; real-model smoke assertions for zero-duration punctuation
  • docs: tools/validate, model-family-testing, and the v2 and v3 model cards

AI Assistance

Yes, heavy Fable usage: code/docs/investigation. I directed and reviewed the work and own the change, happy to iterate more. I'm not an ASR/C++ person, so I also iterated upon a human-digestible :), zero-context explainer of the whole change for me and anyone else who wants the background: https://gist.github.com/ArturWierzbicki/3a5e21740703103e0a89a51428a862a6

Validation

  • validate.py all for parakeet-tdt-0.6b-v2: 18/18 tensors within tolerance, exact transcript, and word timestamps that deviate from the reference by 80 ms max and 3.6 ms mean.
  • validate.py all for parakeet-tdt-0.6b-v3: exact transcript, and word timestamps that match the reference exactly (0.000 ms max deviation).
  • ctest 31/31, pinned clang-format clean, validate.py unit tests pass.
  • No WER impact: transcript text is unchanged (the transcript compare stays exact). The batch JSONL change also reaches granite, gigaam, and medasr, which report word or token timestamps too; the WER harness reads only the text field, so their runs are unaffected.
  • The punctuation rule changes public token and word timings for every TDT-head variant. v2 and v3 are validated here; tdt-1.1b and the tdt_ctc hybrids are not, and their word-timestamp check turns on automatically the next time their reference dumps are regenerated.
  • The opt-in real-model smoke test: every assertion this PR adds passes; 5 pre-existing failures remain (a rejected call leaves the previous result exposed)

Verify Parakeet TDT word timestamps against the NeMo reference:
- dump reference word and token timings in the parakeet NeMo dumper
(transcript.json gains word rows; token rows gain timing)
- emit words/tokens in the CLI batch JSON, gated on the returned
timestamp kind
- compare word start/end against the reference in validate.py, within
the new "timestamps" budget in tests/tolerances/<family>.json
(80 ms = one encoder frame for parakeet)
- mirror NeMo's punctuation convention in the decoder: punctuation-only
tokens are zero-length at the previous token's end, guarded for
non-UTF-8 vocab pieces, with the punctuation set cached at model
load. The first validation run caught a 560 ms word-end divergence
this fixes; the re-run passes with max 80 ms and mean 3.6 ms on the
validation clip, and parakeet-tdt-0.6b-v3 matches the reference
exactly
- update validate, model-family-testing, and model-card docs
@cjpais

Copy link
Copy Markdown
Contributor

Thanks for this PR. I will take a look at it soon and then try to pull it. And I think this is pretty reasonable

@cjpaiscjpais closed this Jul 19, 2026
@cjpaiscjpais reopened this Jul 19, 2026
@ArturWierzbicki

Copy link
Copy Markdown
Author

hey @cjpais anything I can do to help? :) should I try resolving conflicts?

@cjpais

cjpais commented Aug 15, 2026

Copy link
Copy Markdown
Contributor

Mostly just blocked on me having time to take a look at it. So just give me a bit of time. I'm trying to handle some issues in handy right now that are taking priority and also get the release out for 0.2.0 of this library. This change will likely make it into 0.2.1 after I test

But yes rebasing would be helpful

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@ArturWierzbicki@cjpais
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Expose word timestamps for Parakeet - #93

Open
ArturWierzbicki wants to merge 1 commit into
handy-computer:mainfrom
ArturWierzbicki:parakeet-timestamp-verification
Open

Expose word timestamps for Parakeet#93
ArturWierzbicki wants to merge 1 commit into
handy-computer:mainfrom
ArturWierzbicki:parakeet-timestamp-verification

Conversation

@ArturWierzbicki

@ArturWierzbickiArturWierzbicki commented Jul 19, 2026

Copy link
Copy Markdown

Summary

I have a tool that turns recorded meetings into markdown transcripts with named speakers and screenshots of what was shared on screen. Today it transcribes with Parakeet TDT via parakeet-mlx. I'd love to move to transcribe.cpp to simplify the setup.

To do so, I need word timestamps: I use them to match words to diarized speaker turns and to merge overlapping audio chunks by cutting at silences between words. transcribe.cpp already computes them, but the per-file JSON from --batch-jsonl does not expose them. This PR exposes them and verifies them: the batch JSON gains words and tokens arrays, the parakeet NeMo dumper records word timings in its reference output, and validate.py compares each word's start and end against the reference within a new 80 ms tolerance, the duration of one encoder frame. The comparison code is shared across model families; a family turns it on by dumping reference word timings and adding a timestamps entry to its tolerance file.

The check caught a real bug on its first run. A comma is its own token, and the model predicts a duration for it as it does for every token, even though nothing is spoken there.

Before this PR the decode kept those predicted durations, and word ends inherited them, drifting up to 560 ms past the speech on the validation clip. NeMo gives punctuation zero duration instead: in JFK's "And so, my fellow Americans, ...", the token " so" ends at 800 ms and the "," sits exactly there, so the word "so," ends when the audio does. The offline decode in src/arch/parakeet/model.cpp now applies the same rule; transcript text is unchanged.

Scope

  • src/arch/parakeet: punctuation tokens get zero duration in offline decode; the set of punctuation tokens is built once at model load, and vocab entries that are not valid UTF-8 skip the rewrite
  • examples/cli: words appears in batch JSONL when a run returns word or token timestamps, tokens only at token level
  • scripts: the parakeet reference dumper emits word and token timings; validate.py gains the word-timestamp comparison and unit tests
  • tests: a timestamps tolerance entry for parakeet; real-model smoke assertions for zero-duration punctuation
  • docs: tools/validate, model-family-testing, and the v2 and v3 model cards

AI Assistance

Yes, heavy Fable usage: code/docs/investigation. I directed and reviewed the work and own the change, happy to iterate more. I'm not an ASR/C++ person, so I also iterated upon a human-digestible :), zero-context explainer of the whole change for me and anyone else who wants the background: https://gist.github.com/ArturWierzbicki/3a5e21740703103e0a89a51428a862a6

Validation

  • validate.py all for parakeet-tdt-0.6b-v2: 18/18 tensors within tolerance, exact transcript, and word timestamps that deviate from the reference by 80 ms max and 3.6 ms mean.
  • validate.py all for parakeet-tdt-0.6b-v3: exact transcript, and word timestamps that match the reference exactly (0.000 ms max deviation).
  • ctest 31/31, pinned clang-format clean, validate.py unit tests pass.
  • No WER impact: transcript text is unchanged (the transcript compare stays exact). The batch JSONL change also reaches granite, gigaam, and medasr, which report word or token timestamps too; the WER harness reads only the text field, so their runs are unaffected.
  • The punctuation rule changes public token and word timings for every TDT-head variant. v2 and v3 are validated here; tdt-1.1b and the tdt_ctc hybrids are not, and their word-timestamp check turns on automatically the next time their reference dumps are regenerated.
  • The opt-in real-model smoke test: every assertion this PR adds passes; 5 pre-existing failures remain (a rejected call leaves the previous result exposed)

Verify Parakeet TDT word timestamps against the NeMo reference:
- dump reference word and token timings in the parakeet NeMo dumper
(transcript.json gains word rows; token rows gain timing)
- emit words/tokens in the CLI batch JSON, gated on the returned
timestamp kind
- compare word start/end against the reference in validate.py, within
the new "timestamps" budget in tests/tolerances/<family>.json
(80 ms = one encoder frame for parakeet)
- mirror NeMo's punctuation convention in the decoder: punctuation-only
tokens are zero-length at the previous token's end, guarded for
non-UTF-8 vocab pieces, with the punctuation set cached at model
load. The first validation run caught a 560 ms word-end divergence
this fixes; the re-run passes with max 80 ms and mean 3.6 ms on the
validation clip, and parakeet-tdt-0.6b-v3 matches the reference
exactly
- update validate, model-family-testing, and model-card docs
@cjpais

Copy link
Copy Markdown
Contributor

Thanks for this PR. I will take a look at it soon and then try to pull it. And I think this is pretty reasonable

@cjpaiscjpais closed this Jul 19, 2026
@cjpaiscjpais reopened this Jul 19, 2026
@ArturWierzbicki

Copy link
Copy Markdown
Author

hey @cjpais anything I can do to help? :) should I try resolving conflicts?

@cjpais

cjpais commented Aug 15, 2026

Copy link
Copy Markdown
Contributor

Mostly just blocked on me having time to take a look at it. So just give me a bit of time. I'm trying to handle some issues in handy right now that are taking priority and also get the release out for 0.2.0 of this library. This change will likely make it into 0.2.1 after I test

But yes rebasing would be helpful

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@ArturWierzbicki@cjpais
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Expose word timestamps for Parakeet - #93

Open
ArturWierzbicki wants to merge 1 commit into
handy-computer:mainfrom
ArturWierzbicki:parakeet-timestamp-verification
Open

Expose word timestamps for Parakeet#93
ArturWierzbicki wants to merge 1 commit into
handy-computer:mainfrom
ArturWierzbicki:parakeet-timestamp-verification

Conversation

@ArturWierzbicki

@ArturWierzbickiArturWierzbicki commented Jul 19, 2026

Copy link
Copy Markdown

Summary

I have a tool that turns recorded meetings into markdown transcripts with named speakers and screenshots of what was shared on screen. Today it transcribes with Parakeet TDT via parakeet-mlx. I'd love to move to transcribe.cpp to simplify the setup.

To do so, I need word timestamps: I use them to match words to diarized speaker turns and to merge overlapping audio chunks by cutting at silences between words. transcribe.cpp already computes them, but the per-file JSON from --batch-jsonl does not expose them. This PR exposes them and verifies them: the batch JSON gains words and tokens arrays, the parakeet NeMo dumper records word timings in its reference output, and validate.py compares each word's start and end against the reference within a new 80 ms tolerance, the duration of one encoder frame. The comparison code is shared across model families; a family turns it on by dumping reference word timings and adding a timestamps entry to its tolerance file.

The check caught a real bug on its first run. A comma is its own token, and the model predicts a duration for it as it does for every token, even though nothing is spoken there.

Before this PR the decode kept those predicted durations, and word ends inherited them, drifting up to 560 ms past the speech on the validation clip. NeMo gives punctuation zero duration instead: in JFK's "And so, my fellow Americans, ...", the token " so" ends at 800 ms and the "," sits exactly there, so the word "so," ends when the audio does. The offline decode in src/arch/parakeet/model.cpp now applies the same rule; transcript text is unchanged.

Scope

  • src/arch/parakeet: punctuation tokens get zero duration in offline decode; the set of punctuation tokens is built once at model load, and vocab entries that are not valid UTF-8 skip the rewrite
  • examples/cli: words appears in batch JSONL when a run returns word or token timestamps, tokens only at token level
  • scripts: the parakeet reference dumper emits word and token timings; validate.py gains the word-timestamp comparison and unit tests
  • tests: a timestamps tolerance entry for parakeet; real-model smoke assertions for zero-duration punctuation
  • docs: tools/validate, model-family-testing, and the v2 and v3 model cards

AI Assistance

Yes, heavy Fable usage: code/docs/investigation. I directed and reviewed the work and own the change, happy to iterate more. I'm not an ASR/C++ person, so I also iterated upon a human-digestible :), zero-context explainer of the whole change for me and anyone else who wants the background: https://gist.github.com/ArturWierzbicki/3a5e21740703103e0a89a51428a862a6

Validation

  • validate.py all for parakeet-tdt-0.6b-v2: 18/18 tensors within tolerance, exact transcript, and word timestamps that deviate from the reference by 80 ms max and 3.6 ms mean.
  • validate.py all for parakeet-tdt-0.6b-v3: exact transcript, and word timestamps that match the reference exactly (0.000 ms max deviation).
  • ctest 31/31, pinned clang-format clean, validate.py unit tests pass.
  • No WER impact: transcript text is unchanged (the transcript compare stays exact). The batch JSONL change also reaches granite, gigaam, and medasr, which report word or token timestamps too; the WER harness reads only the text field, so their runs are unaffected.
  • The punctuation rule changes public token and word timings for every TDT-head variant. v2 and v3 are validated here; tdt-1.1b and the tdt_ctc hybrids are not, and their word-timestamp check turns on automatically the next time their reference dumps are regenerated.
  • The opt-in real-model smoke test: every assertion this PR adds passes; 5 pre-existing failures remain (a rejected call leaves the previous result exposed)

Verify Parakeet TDT word timestamps against the NeMo reference:
- dump reference word and token timings in the parakeet NeMo dumper
(transcript.json gains word rows; token rows gain timing)
- emit words/tokens in the CLI batch JSON, gated on the returned
timestamp kind
- compare word start/end against the reference in validate.py, within
the new "timestamps" budget in tests/tolerances/<family>.json
(80 ms = one encoder frame for parakeet)
- mirror NeMo's punctuation convention in the decoder: punctuation-only
tokens are zero-length at the previous token's end, guarded for
non-UTF-8 vocab pieces, with the punctuation set cached at model
load. The first validation run caught a 560 ms word-end divergence
this fixes; the re-run passes with max 80 ms and mean 3.6 ms on the
validation clip, and parakeet-tdt-0.6b-v3 matches the reference
exactly
- update validate, model-family-testing, and model-card docs
@cjpais

Copy link
Copy Markdown
Contributor

Thanks for this PR. I will take a look at it soon and then try to pull it. And I think this is pretty reasonable

@cjpaiscjpais closed this Jul 19, 2026
@cjpaiscjpais reopened this Jul 19, 2026
@ArturWierzbicki

Copy link
Copy Markdown
Author

hey @cjpais anything I can do to help? :) should I try resolving conflicts?

@cjpais

cjpais commented Aug 15, 2026

Copy link
Copy Markdown
Contributor

Mostly just blocked on me having time to take a look at it. So just give me a bit of time. I'm trying to handle some issues in handy right now that are taking priority and also get the release out for 0.2.0 of this library. This change will likely make it into 0.2.1 after I test

But yes rebasing would be helpful

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@ArturWierzbicki@cjpais