canary: add long-form chunking - #112

Draft
brauliobo wants to merge 2 commits into
handy-computer:mainfrom
brauliobo:canary-v2-long-form
Draft

canary: add long-form chunking#112
brauliobo wants to merge 2 commits into
handy-computer:mainfrom
brauliobo:canary-v2-long-form

Conversation

@brauliobo

Copy link
Copy Markdown

Dependency

This is a stacked PR on top of #111 (canary: add v2 word timestamps). The branch intentionally includes that commit; please review this PR after #111, then rebase it onto main once #111 lands.

Summary

  • add offline long-form transcription for Canary 1B v2 using 30-40 second windows, low-energy acoustic boundaries, and 1 second of overlap
  • rebase timestamped chunk output onto the source timeline and assign overlap words by midpoint ownership
  • reconcile text-only overlap against the previous hypothesis suffix so repeated phrases are not mistaken for duplicated context
  • report Canary 1B v2 audio capacity as unbounded while leaving older Canary variants at their existing positional-table limit
  • document that long-form input remains memory-resident and batch rows are processed serially

Motivation

Canary 1B v2 has a 40 second decoder window. Running longer audio as one window can silently omit substantial interior speech even when inference reaches the final phrase. On the 84.381 second samples/whole-earth.wav, the pre-long-form path returned 141 normalized words with a 35.040 second internal transcript gap. The chunked path returned 217 normalized words with a maximum 1.840 second inter-word gap and continuous coverage through the final phrase.

The first fixed-duration prototype also lost repeated words at chunk seams. Selecting low-energy boundaries and restricting text deduplication to a true suffix fixed that failure rather than masking it with a broader overlap heuristic.

Correctness

Adversarial repeated-speech tests passed without missing or duplicated repetitions:

AudioWordsResult
44 s88exact 4x repetition
55 s110exact 5x repetition
88 s176exact 8x repetition
121 s242exact 11x repetition

The 39 s, 40 s, 40.001 s, and 41 s threshold cases also passed. Every benchmark returned monotonic positive word intervals within the audio duration, with normalized full text matching the word sequence.

Performance

Warm Q8_0 measurements on a shared RTX 4060 Ti:

Mode55 s121 s
timestamps2,784.51 ms (19.8x realtime)6,366.06 ms (19.0x)
text only2,042.28 ms (26.9x realtime)4,376.20 ms (27.6x)

F32 timestamp validation at 55 seconds passed on an RTX 3090 in 899.36 ms (61.2x realtime). The shared RTX 4060 Ti did not have enough free memory for the F32 model's 6.1 GiB CUDA allocation, so that cell was moved to the 3090 rather than changing running services.

Verification

  • 52/52 configured CTest tests passed; 18 unrelated fixture-gated tests remained skipped by default
  • gated timestamped and text-only Canary Q8_0 real-model smoke passed manually
  • pinned clang-format check passed
  • git diff --check passed

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@brauliobo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

canary: add long-form chunking - #112

Draft
brauliobo wants to merge 2 commits into
handy-computer:mainfrom
brauliobo:canary-v2-long-form
Draft

canary: add long-form chunking#112
brauliobo wants to merge 2 commits into
handy-computer:mainfrom
brauliobo:canary-v2-long-form

Conversation

@brauliobo

Copy link
Copy Markdown

Dependency

This is a stacked PR on top of #111 (canary: add v2 word timestamps). The branch intentionally includes that commit; please review this PR after #111, then rebase it onto main once #111 lands.

Summary

  • add offline long-form transcription for Canary 1B v2 using 30-40 second windows, low-energy acoustic boundaries, and 1 second of overlap
  • rebase timestamped chunk output onto the source timeline and assign overlap words by midpoint ownership
  • reconcile text-only overlap against the previous hypothesis suffix so repeated phrases are not mistaken for duplicated context
  • report Canary 1B v2 audio capacity as unbounded while leaving older Canary variants at their existing positional-table limit
  • document that long-form input remains memory-resident and batch rows are processed serially

Motivation

Canary 1B v2 has a 40 second decoder window. Running longer audio as one window can silently omit substantial interior speech even when inference reaches the final phrase. On the 84.381 second samples/whole-earth.wav, the pre-long-form path returned 141 normalized words with a 35.040 second internal transcript gap. The chunked path returned 217 normalized words with a maximum 1.840 second inter-word gap and continuous coverage through the final phrase.

The first fixed-duration prototype also lost repeated words at chunk seams. Selecting low-energy boundaries and restricting text deduplication to a true suffix fixed that failure rather than masking it with a broader overlap heuristic.

Correctness

Adversarial repeated-speech tests passed without missing or duplicated repetitions:

AudioWordsResult
44 s88exact 4x repetition
55 s110exact 5x repetition
88 s176exact 8x repetition
121 s242exact 11x repetition

The 39 s, 40 s, 40.001 s, and 41 s threshold cases also passed. Every benchmark returned monotonic positive word intervals within the audio duration, with normalized full text matching the word sequence.

Performance

Warm Q8_0 measurements on a shared RTX 4060 Ti:

Mode55 s121 s
timestamps2,784.51 ms (19.8x realtime)6,366.06 ms (19.0x)
text only2,042.28 ms (26.9x realtime)4,376.20 ms (27.6x)

F32 timestamp validation at 55 seconds passed on an RTX 3090 in 899.36 ms (61.2x realtime). The shared RTX 4060 Ti did not have enough free memory for the F32 model's 6.1 GiB CUDA allocation, so that cell was moved to the 3090 rather than changing running services.

Verification

  • 52/52 configured CTest tests passed; 18 unrelated fixture-gated tests remained skipped by default
  • gated timestamped and text-only Canary Q8_0 real-model smoke passed manually
  • pinned clang-format check passed
  • git diff --check passed

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@brauliobo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

canary: add long-form chunking - #112

Draft
brauliobo wants to merge 2 commits into
handy-computer:mainfrom
brauliobo:canary-v2-long-form
Draft

canary: add long-form chunking#112
brauliobo wants to merge 2 commits into
handy-computer:mainfrom
brauliobo:canary-v2-long-form

Conversation

@brauliobo

Copy link
Copy Markdown

Dependency

This is a stacked PR on top of #111 (canary: add v2 word timestamps). The branch intentionally includes that commit; please review this PR after #111, then rebase it onto main once #111 lands.

Summary

  • add offline long-form transcription for Canary 1B v2 using 30-40 second windows, low-energy acoustic boundaries, and 1 second of overlap
  • rebase timestamped chunk output onto the source timeline and assign overlap words by midpoint ownership
  • reconcile text-only overlap against the previous hypothesis suffix so repeated phrases are not mistaken for duplicated context
  • report Canary 1B v2 audio capacity as unbounded while leaving older Canary variants at their existing positional-table limit
  • document that long-form input remains memory-resident and batch rows are processed serially

Motivation

Canary 1B v2 has a 40 second decoder window. Running longer audio as one window can silently omit substantial interior speech even when inference reaches the final phrase. On the 84.381 second samples/whole-earth.wav, the pre-long-form path returned 141 normalized words with a 35.040 second internal transcript gap. The chunked path returned 217 normalized words with a maximum 1.840 second inter-word gap and continuous coverage through the final phrase.

The first fixed-duration prototype also lost repeated words at chunk seams. Selecting low-energy boundaries and restricting text deduplication to a true suffix fixed that failure rather than masking it with a broader overlap heuristic.

Correctness

Adversarial repeated-speech tests passed without missing or duplicated repetitions:

AudioWordsResult
44 s88exact 4x repetition
55 s110exact 5x repetition
88 s176exact 8x repetition
121 s242exact 11x repetition

The 39 s, 40 s, 40.001 s, and 41 s threshold cases also passed. Every benchmark returned monotonic positive word intervals within the audio duration, with normalized full text matching the word sequence.

Performance

Warm Q8_0 measurements on a shared RTX 4060 Ti:

Mode55 s121 s
timestamps2,784.51 ms (19.8x realtime)6,366.06 ms (19.0x)
text only2,042.28 ms (26.9x realtime)4,376.20 ms (27.6x)

F32 timestamp validation at 55 seconds passed on an RTX 3090 in 899.36 ms (61.2x realtime). The shared RTX 4060 Ti did not have enough free memory for the F32 model's 6.1 GiB CUDA allocation, so that cell was moved to the 3090 rather than changing running services.

Verification

  • 52/52 configured CTest tests passed; 18 unrelated fixture-gated tests remained skipped by default
  • gated timestamped and text-only Canary Q8_0 real-model smoke passed manually
  • pinned clang-format check passed
  • git diff --check passed

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@brauliobo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

canary: add long-form chunking - #112

Draft
brauliobo wants to merge 2 commits into
handy-computer:mainfrom
brauliobo:canary-v2-long-form
Draft

canary: add long-form chunking#112
brauliobo wants to merge 2 commits into
handy-computer:mainfrom
brauliobo:canary-v2-long-form

Conversation

@brauliobo

Copy link
Copy Markdown

Dependency

This is a stacked PR on top of #111 (canary: add v2 word timestamps). The branch intentionally includes that commit; please review this PR after #111, then rebase it onto main once #111 lands.

Summary

  • add offline long-form transcription for Canary 1B v2 using 30-40 second windows, low-energy acoustic boundaries, and 1 second of overlap
  • rebase timestamped chunk output onto the source timeline and assign overlap words by midpoint ownership
  • reconcile text-only overlap against the previous hypothesis suffix so repeated phrases are not mistaken for duplicated context
  • report Canary 1B v2 audio capacity as unbounded while leaving older Canary variants at their existing positional-table limit
  • document that long-form input remains memory-resident and batch rows are processed serially

Motivation

Canary 1B v2 has a 40 second decoder window. Running longer audio as one window can silently omit substantial interior speech even when inference reaches the final phrase. On the 84.381 second samples/whole-earth.wav, the pre-long-form path returned 141 normalized words with a 35.040 second internal transcript gap. The chunked path returned 217 normalized words with a maximum 1.840 second inter-word gap and continuous coverage through the final phrase.

The first fixed-duration prototype also lost repeated words at chunk seams. Selecting low-energy boundaries and restricting text deduplication to a true suffix fixed that failure rather than masking it with a broader overlap heuristic.

Correctness

Adversarial repeated-speech tests passed without missing or duplicated repetitions:

AudioWordsResult
44 s88exact 4x repetition
55 s110exact 5x repetition
88 s176exact 8x repetition
121 s242exact 11x repetition

The 39 s, 40 s, 40.001 s, and 41 s threshold cases also passed. Every benchmark returned monotonic positive word intervals within the audio duration, with normalized full text matching the word sequence.

Performance

Warm Q8_0 measurements on a shared RTX 4060 Ti:

Mode55 s121 s
timestamps2,784.51 ms (19.8x realtime)6,366.06 ms (19.0x)
text only2,042.28 ms (26.9x realtime)4,376.20 ms (27.6x)

F32 timestamp validation at 55 seconds passed on an RTX 3090 in 899.36 ms (61.2x realtime). The shared RTX 4060 Ti did not have enough free memory for the F32 model's 6.1 GiB CUDA allocation, so that cell was moved to the 3090 rather than changing running services.

Verification

  • 52/52 configured CTest tests passed; 18 unrelated fixture-gated tests remained skipped by default
  • gated timestamped and text-only Canary Q8_0 real-model smoke passed manually
  • pinned clang-format check passed
  • git diff --check passed

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@brauliobo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

canary: add long-form chunking - #112

Draft
brauliobo wants to merge 2 commits into
handy-computer:mainfrom
brauliobo:canary-v2-long-form
Draft

canary: add long-form chunking#112
brauliobo wants to merge 2 commits into
handy-computer:mainfrom
brauliobo:canary-v2-long-form

Conversation

@brauliobo

Copy link
Copy Markdown

Dependency

This is a stacked PR on top of #111 (canary: add v2 word timestamps). The branch intentionally includes that commit; please review this PR after #111, then rebase it onto main once #111 lands.

Summary

  • add offline long-form transcription for Canary 1B v2 using 30-40 second windows, low-energy acoustic boundaries, and 1 second of overlap
  • rebase timestamped chunk output onto the source timeline and assign overlap words by midpoint ownership
  • reconcile text-only overlap against the previous hypothesis suffix so repeated phrases are not mistaken for duplicated context
  • report Canary 1B v2 audio capacity as unbounded while leaving older Canary variants at their existing positional-table limit
  • document that long-form input remains memory-resident and batch rows are processed serially

Motivation

Canary 1B v2 has a 40 second decoder window. Running longer audio as one window can silently omit substantial interior speech even when inference reaches the final phrase. On the 84.381 second samples/whole-earth.wav, the pre-long-form path returned 141 normalized words with a 35.040 second internal transcript gap. The chunked path returned 217 normalized words with a maximum 1.840 second inter-word gap and continuous coverage through the final phrase.

The first fixed-duration prototype also lost repeated words at chunk seams. Selecting low-energy boundaries and restricting text deduplication to a true suffix fixed that failure rather than masking it with a broader overlap heuristic.

Correctness

Adversarial repeated-speech tests passed without missing or duplicated repetitions:

AudioWordsResult
44 s88exact 4x repetition
55 s110exact 5x repetition
88 s176exact 8x repetition
121 s242exact 11x repetition

The 39 s, 40 s, 40.001 s, and 41 s threshold cases also passed. Every benchmark returned monotonic positive word intervals within the audio duration, with normalized full text matching the word sequence.

Performance

Warm Q8_0 measurements on a shared RTX 4060 Ti:

Mode55 s121 s
timestamps2,784.51 ms (19.8x realtime)6,366.06 ms (19.0x)
text only2,042.28 ms (26.9x realtime)4,376.20 ms (27.6x)

F32 timestamp validation at 55 seconds passed on an RTX 3090 in 899.36 ms (61.2x realtime). The shared RTX 4060 Ti did not have enough free memory for the F32 model's 6.1 GiB CUDA allocation, so that cell was moved to the 3090 rather than changing running services.

Verification

  • 52/52 configured CTest tests passed; 18 unrelated fixture-gated tests remained skipped by default
  • gated timestamped and text-only Canary Q8_0 real-model smoke passed manually
  • pinned clang-format check passed
  • git diff --check passed

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@brauliobo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

canary: add long-form chunking - #112

Draft
brauliobo wants to merge 2 commits into
handy-computer:mainfrom
brauliobo:canary-v2-long-form
Draft

canary: add long-form chunking#112
brauliobo wants to merge 2 commits into
handy-computer:mainfrom
brauliobo:canary-v2-long-form

Conversation

@brauliobo

Copy link
Copy Markdown

Dependency

This is a stacked PR on top of #111 (canary: add v2 word timestamps). The branch intentionally includes that commit; please review this PR after #111, then rebase it onto main once #111 lands.

Summary

  • add offline long-form transcription for Canary 1B v2 using 30-40 second windows, low-energy acoustic boundaries, and 1 second of overlap
  • rebase timestamped chunk output onto the source timeline and assign overlap words by midpoint ownership
  • reconcile text-only overlap against the previous hypothesis suffix so repeated phrases are not mistaken for duplicated context
  • report Canary 1B v2 audio capacity as unbounded while leaving older Canary variants at their existing positional-table limit
  • document that long-form input remains memory-resident and batch rows are processed serially

Motivation

Canary 1B v2 has a 40 second decoder window. Running longer audio as one window can silently omit substantial interior speech even when inference reaches the final phrase. On the 84.381 second samples/whole-earth.wav, the pre-long-form path returned 141 normalized words with a 35.040 second internal transcript gap. The chunked path returned 217 normalized words with a maximum 1.840 second inter-word gap and continuous coverage through the final phrase.

The first fixed-duration prototype also lost repeated words at chunk seams. Selecting low-energy boundaries and restricting text deduplication to a true suffix fixed that failure rather than masking it with a broader overlap heuristic.

Correctness

Adversarial repeated-speech tests passed without missing or duplicated repetitions:

AudioWordsResult
44 s88exact 4x repetition
55 s110exact 5x repetition
88 s176exact 8x repetition
121 s242exact 11x repetition

The 39 s, 40 s, 40.001 s, and 41 s threshold cases also passed. Every benchmark returned monotonic positive word intervals within the audio duration, with normalized full text matching the word sequence.

Performance

Warm Q8_0 measurements on a shared RTX 4060 Ti:

Mode55 s121 s
timestamps2,784.51 ms (19.8x realtime)6,366.06 ms (19.0x)
text only2,042.28 ms (26.9x realtime)4,376.20 ms (27.6x)

F32 timestamp validation at 55 seconds passed on an RTX 3090 in 899.36 ms (61.2x realtime). The shared RTX 4060 Ti did not have enough free memory for the F32 model's 6.1 GiB CUDA allocation, so that cell was moved to the 3090 rather than changing running services.

Verification

  • 52/52 configured CTest tests passed; 18 unrelated fixture-gated tests remained skipped by default
  • gated timestamped and text-only Canary Q8_0 real-model smoke passed manually
  • pinned clang-format check passed
  • git diff --check passed

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@brauliobo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

canary: add long-form chunking - #112

Draft
brauliobo wants to merge 2 commits into
handy-computer:mainfrom
brauliobo:canary-v2-long-form
Draft

canary: add long-form chunking#112
brauliobo wants to merge 2 commits into
handy-computer:mainfrom
brauliobo:canary-v2-long-form

Conversation

@brauliobo

Copy link
Copy Markdown

Dependency

This is a stacked PR on top of #111 (canary: add v2 word timestamps). The branch intentionally includes that commit; please review this PR after #111, then rebase it onto main once #111 lands.

Summary

  • add offline long-form transcription for Canary 1B v2 using 30-40 second windows, low-energy acoustic boundaries, and 1 second of overlap
  • rebase timestamped chunk output onto the source timeline and assign overlap words by midpoint ownership
  • reconcile text-only overlap against the previous hypothesis suffix so repeated phrases are not mistaken for duplicated context
  • report Canary 1B v2 audio capacity as unbounded while leaving older Canary variants at their existing positional-table limit
  • document that long-form input remains memory-resident and batch rows are processed serially

Motivation

Canary 1B v2 has a 40 second decoder window. Running longer audio as one window can silently omit substantial interior speech even when inference reaches the final phrase. On the 84.381 second samples/whole-earth.wav, the pre-long-form path returned 141 normalized words with a 35.040 second internal transcript gap. The chunked path returned 217 normalized words with a maximum 1.840 second inter-word gap and continuous coverage through the final phrase.

The first fixed-duration prototype also lost repeated words at chunk seams. Selecting low-energy boundaries and restricting text deduplication to a true suffix fixed that failure rather than masking it with a broader overlap heuristic.

Correctness

Adversarial repeated-speech tests passed without missing or duplicated repetitions:

AudioWordsResult
44 s88exact 4x repetition
55 s110exact 5x repetition
88 s176exact 8x repetition
121 s242exact 11x repetition

The 39 s, 40 s, 40.001 s, and 41 s threshold cases also passed. Every benchmark returned monotonic positive word intervals within the audio duration, with normalized full text matching the word sequence.

Performance

Warm Q8_0 measurements on a shared RTX 4060 Ti:

Mode55 s121 s
timestamps2,784.51 ms (19.8x realtime)6,366.06 ms (19.0x)
text only2,042.28 ms (26.9x realtime)4,376.20 ms (27.6x)

F32 timestamp validation at 55 seconds passed on an RTX 3090 in 899.36 ms (61.2x realtime). The shared RTX 4060 Ti did not have enough free memory for the F32 model's 6.1 GiB CUDA allocation, so that cell was moved to the 3090 rather than changing running services.

Verification

  • 52/52 configured CTest tests passed; 18 unrelated fixture-gated tests remained skipped by default
  • gated timestamped and text-only Canary Q8_0 real-model smoke passed manually
  • pinned clang-format check passed
  • git diff --check passed

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@brauliobo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

canary: add long-form chunking - #112

Draft
brauliobo wants to merge 2 commits into
handy-computer:mainfrom
brauliobo:canary-v2-long-form
Draft

canary: add long-form chunking#112
brauliobo wants to merge 2 commits into
handy-computer:mainfrom
brauliobo:canary-v2-long-form

Conversation

@brauliobo

Copy link
Copy Markdown

Dependency

This is a stacked PR on top of #111 (canary: add v2 word timestamps). The branch intentionally includes that commit; please review this PR after #111, then rebase it onto main once #111 lands.

Summary

  • add offline long-form transcription for Canary 1B v2 using 30-40 second windows, low-energy acoustic boundaries, and 1 second of overlap
  • rebase timestamped chunk output onto the source timeline and assign overlap words by midpoint ownership
  • reconcile text-only overlap against the previous hypothesis suffix so repeated phrases are not mistaken for duplicated context
  • report Canary 1B v2 audio capacity as unbounded while leaving older Canary variants at their existing positional-table limit
  • document that long-form input remains memory-resident and batch rows are processed serially

Motivation

Canary 1B v2 has a 40 second decoder window. Running longer audio as one window can silently omit substantial interior speech even when inference reaches the final phrase. On the 84.381 second samples/whole-earth.wav, the pre-long-form path returned 141 normalized words with a 35.040 second internal transcript gap. The chunked path returned 217 normalized words with a maximum 1.840 second inter-word gap and continuous coverage through the final phrase.

The first fixed-duration prototype also lost repeated words at chunk seams. Selecting low-energy boundaries and restricting text deduplication to a true suffix fixed that failure rather than masking it with a broader overlap heuristic.

Correctness

Adversarial repeated-speech tests passed without missing or duplicated repetitions:

AudioWordsResult
44 s88exact 4x repetition
55 s110exact 5x repetition
88 s176exact 8x repetition
121 s242exact 11x repetition

The 39 s, 40 s, 40.001 s, and 41 s threshold cases also passed. Every benchmark returned monotonic positive word intervals within the audio duration, with normalized full text matching the word sequence.

Performance

Warm Q8_0 measurements on a shared RTX 4060 Ti:

Mode55 s121 s
timestamps2,784.51 ms (19.8x realtime)6,366.06 ms (19.0x)
text only2,042.28 ms (26.9x realtime)4,376.20 ms (27.6x)

F32 timestamp validation at 55 seconds passed on an RTX 3090 in 899.36 ms (61.2x realtime). The shared RTX 4060 Ti did not have enough free memory for the F32 model's 6.1 GiB CUDA allocation, so that cell was moved to the 3090 rather than changing running services.

Verification

  • 52/52 configured CTest tests passed; 18 unrelated fixture-gated tests remained skipped by default
  • gated timestamped and text-only Canary Q8_0 real-model smoke passed manually
  • pinned clang-format check passed
  • git diff --check passed

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@brauliobo