stream: cut per-chunk overhead in web streams pipelines - #26

Open
anonrig wants to merge 3 commits into
mainfrom
cursor/webstreams-pipeline-perf-93a7
Open

stream: cut per-chunk overhead in web streams pipelines#26
anonrig wants to merge 3 commits into
mainfrom
cursor/webstreams-pipeline-perf-93a7

Conversation

@anonrig

Copy link
Copy Markdown
Owner

What

Reduces per-chunk overhead across common web streams pipelines (fetch body decompression/upload, TextDecoderStream/TextEncoderStream transcoding, Readable.toWeb/Writable.toWeb piping), targeting workloads that move large payloads in small (e.g. 4 KiB) chunks.

Three changes:

  1. CompressionStream/DecompressionStream: process chunks without threadpool round trips.
    Previously these classes wrapped a zlib stream.Duplex in the web streams adapters, so every written chunk was dispatched to the threadpool and its completion observed on a later event loop turn. For a 64 MiB body in 4 KiB chunks that is 16384 threadpool round trips, which dominates the cost of the stream. The classes are now built on a TransformStream that drives the raw zlib/brotli handle synchronously (writeSync), processing inputs larger than 64 KiB in slices with an event-loop turn in between so huge chunks cannot block the loop for their full duration. Output is emitted in up to 64 KiB chunks, either zero-copy or right-sized copies.

  2. TextEncoderStream: fast-path chunk encoding.
    The encode-and-enqueue algorithm was a literal transcription of the spec's per-code-unit loop (single-character string extraction, charCodeAt, string accumulator append for every code unit). Its only observable effects are the surrogate hand-off at chunk boundaries and U+FFFD replacement, which TextEncoder.encode() already performs, so the loop is replaced with explicit chunk-boundary handling plus a single encode() call.

  3. Readable.toWeb(): avoid copying buffers that are the sole view of their ArrayBuffer.
    Chunks were always copied before being enqueued to avoid exposing Buffer pool slices. Buffers that own their entire ArrayBuffer (as fs streams produce for every read) cannot alias pooled memory and are now exposed zero-copy.

Why

Pipelines like fetch() → DecompressionStream → TextDecoderStream → for await, fs.createReadStream() → CompressionStream → fetch POST, and fs → TextDecoderStream → TextEncoderStream → fs spend the majority of their time in per-chunk machinery rather than in the codecs doing the actual work.

(Benchmark results from an isolated machine to be added after the full validation run completes; preliminary measurements show ~13x on the TextEncoderStream stage and ~3x on the DecompressionStream stage for 4 KiB chunks.)

Testing

  • New/existing: test/parallel/test-whatwg-webstreams-compression.js, test-webstreams-compression-bad-chunks.js, test-webstreams-decompression-reject-trailing.js, test-webstreams-compression-buffer-source.js, test-compression-decompression-stream.js, test-whatwg-webstreams-encoding.js, adapters tests.
  • WPT: compression, encoding, streams suites.
  • Adds benchmark/webstreams/compression.js and benchmark/webstreams/encoding.js.
Open in WebOpen in Cursor

@cursor
cursorBotforce-pushed the cursor/webstreams-pipeline-perf-93a7 branch from 796edc7 to 8019292CompareAugust 20, 2026 16:30
setImmediate ??= require('timers').setImmediate;
await new Promise(setImmediate);
}
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Large chunks yield without copying input

High Severity

Inputs larger than kInputSliceSize are processed in slices with an await between them, but each slice is only a view of the original chunk. After transform yields, other tasks can detach or mutate that ArrayBuffer before later slices run, so remaining data is compressed from invalid or changed memory. The WPT detach cases (compression-with-detach, decompression-with-detach) expect the write to keep working after the source is detached.

Additional Locations (1)
Fix in CursorFix in Web

Reviewed by Cursor Bugbot for commit 8019292. Configure here.

Rewrite CompressionStream and DecompressionStream on top of a
TransformStream that drives the zlib (or brotli) handle synchronously,
instead of wrapping a zlib stream.Duplex in the web streams adapters.
The previous implementation dispatched every written chunk to the
threadpool and waited for the event loop to observe its completion,
which dominates the cost of streaming small chunks: a 64 MiB body
written in 4 KiB chunks paid for 16384 threadpool round trips plus the
Transform and adapter machinery around them. Processing the chunks
inline removes that latency entirely while performing the same work.
Inputs larger than 64 KiB are processed in slices with a turn of the
event loop in between so that huge chunks cannot block the loop for
their full duration.
Output is emitted in up to 64 KiB chunks, either as zero-copy views or
as right-sized copies (so small outputs do not retain large buffers),
which also reduces the per-chunk overhead imposed on the rest of the
pipeline downstream.
Assisted-by: Cursor
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
The encode-and-enqueue algorithm was implemented as a literal
transcription of the spec's per-code-unit loop: it extracted a
single-character string, called charCodeAt(), and appended to an
accumulator string for every code unit of every chunk, allocating
millions of temporary strings for large payloads.
The only observable effects of that loop are that a high surrogate at
the end of a chunk is held back to pair with a low surrogate starting
the next chunk, and that unpaired surrogates encode as U+FFFD, which
TextEncoder already does. Handle the chunk boundary explicitly and
encode the rest of the chunk with a single encode() call.
Assisted-by: Cursor
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
@cursor
cursorBotforce-pushed the cursor/webstreams-pipeline-perf-93a7 branch 3 times, most recently from eea0fe9 to 936186fCompareAugust 20, 2026 16:45

@cursorcursorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

There are 2 total unresolved issues (including 1 from previous review).

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, have a team admin enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 936186f. Configure here.

Comment threadlib/internal/webstreams/adapters.js Outdated
The kValidateChunk and kDestroyOnSyncError hooks existed only for the
previous stream.Duplex-based CompressionStream implementation, which no
longer uses the adapters.
Assisted-by: Cursor
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
@cursor
cursorBotforce-pushed the cursor/webstreams-pipeline-perf-93a7 branch from 936186f to e3f6bd2CompareAugust 20, 2026 17:11
@anonrig
anonrig marked this pull request as ready for review August 20, 2026 17:12
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@anonrig
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

stream: cut per-chunk overhead in web streams pipelines - #26

Open
anonrig wants to merge 3 commits into
mainfrom
cursor/webstreams-pipeline-perf-93a7
Open

stream: cut per-chunk overhead in web streams pipelines#26
anonrig wants to merge 3 commits into
mainfrom
cursor/webstreams-pipeline-perf-93a7

Conversation

@anonrig

Copy link
Copy Markdown
Owner

What

Reduces per-chunk overhead across common web streams pipelines (fetch body decompression/upload, TextDecoderStream/TextEncoderStream transcoding, Readable.toWeb/Writable.toWeb piping), targeting workloads that move large payloads in small (e.g. 4 KiB) chunks.

Three changes:

  1. CompressionStream/DecompressionStream: process chunks without threadpool round trips.
    Previously these classes wrapped a zlib stream.Duplex in the web streams adapters, so every written chunk was dispatched to the threadpool and its completion observed on a later event loop turn. For a 64 MiB body in 4 KiB chunks that is 16384 threadpool round trips, which dominates the cost of the stream. The classes are now built on a TransformStream that drives the raw zlib/brotli handle synchronously (writeSync), processing inputs larger than 64 KiB in slices with an event-loop turn in between so huge chunks cannot block the loop for their full duration. Output is emitted in up to 64 KiB chunks, either zero-copy or right-sized copies.

  2. TextEncoderStream: fast-path chunk encoding.
    The encode-and-enqueue algorithm was a literal transcription of the spec's per-code-unit loop (single-character string extraction, charCodeAt, string accumulator append for every code unit). Its only observable effects are the surrogate hand-off at chunk boundaries and U+FFFD replacement, which TextEncoder.encode() already performs, so the loop is replaced with explicit chunk-boundary handling plus a single encode() call.

  3. Readable.toWeb(): avoid copying buffers that are the sole view of their ArrayBuffer.
    Chunks were always copied before being enqueued to avoid exposing Buffer pool slices. Buffers that own their entire ArrayBuffer (as fs streams produce for every read) cannot alias pooled memory and are now exposed zero-copy.

Why

Pipelines like fetch() → DecompressionStream → TextDecoderStream → for await, fs.createReadStream() → CompressionStream → fetch POST, and fs → TextDecoderStream → TextEncoderStream → fs spend the majority of their time in per-chunk machinery rather than in the codecs doing the actual work.

(Benchmark results from an isolated machine to be added after the full validation run completes; preliminary measurements show ~13x on the TextEncoderStream stage and ~3x on the DecompressionStream stage for 4 KiB chunks.)

Testing

  • New/existing: test/parallel/test-whatwg-webstreams-compression.js, test-webstreams-compression-bad-chunks.js, test-webstreams-decompression-reject-trailing.js, test-webstreams-compression-buffer-source.js, test-compression-decompression-stream.js, test-whatwg-webstreams-encoding.js, adapters tests.
  • WPT: compression, encoding, streams suites.
  • Adds benchmark/webstreams/compression.js and benchmark/webstreams/encoding.js.
Open in WebOpen in Cursor

@cursor
cursorBotforce-pushed the cursor/webstreams-pipeline-perf-93a7 branch from 796edc7 to 8019292CompareAugust 20, 2026 16:30
setImmediate ??= require('timers').setImmediate;
await new Promise(setImmediate);
}
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Large chunks yield without copying input

High Severity

Inputs larger than kInputSliceSize are processed in slices with an await between them, but each slice is only a view of the original chunk. After transform yields, other tasks can detach or mutate that ArrayBuffer before later slices run, so remaining data is compressed from invalid or changed memory. The WPT detach cases (compression-with-detach, decompression-with-detach) expect the write to keep working after the source is detached.

Additional Locations (1)
Fix in CursorFix in Web

Reviewed by Cursor Bugbot for commit 8019292. Configure here.

Rewrite CompressionStream and DecompressionStream on top of a
TransformStream that drives the zlib (or brotli) handle synchronously,
instead of wrapping a zlib stream.Duplex in the web streams adapters.
The previous implementation dispatched every written chunk to the
threadpool and waited for the event loop to observe its completion,
which dominates the cost of streaming small chunks: a 64 MiB body
written in 4 KiB chunks paid for 16384 threadpool round trips plus the
Transform and adapter machinery around them. Processing the chunks
inline removes that latency entirely while performing the same work.
Inputs larger than 64 KiB are processed in slices with a turn of the
event loop in between so that huge chunks cannot block the loop for
their full duration.
Output is emitted in up to 64 KiB chunks, either as zero-copy views or
as right-sized copies (so small outputs do not retain large buffers),
which also reduces the per-chunk overhead imposed on the rest of the
pipeline downstream.
Assisted-by: Cursor
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
The encode-and-enqueue algorithm was implemented as a literal
transcription of the spec's per-code-unit loop: it extracted a
single-character string, called charCodeAt(), and appended to an
accumulator string for every code unit of every chunk, allocating
millions of temporary strings for large payloads.
The only observable effects of that loop are that a high surrogate at
the end of a chunk is held back to pair with a low surrogate starting
the next chunk, and that unpaired surrogates encode as U+FFFD, which
TextEncoder already does. Handle the chunk boundary explicitly and
encode the rest of the chunk with a single encode() call.
Assisted-by: Cursor
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
@cursor
cursorBotforce-pushed the cursor/webstreams-pipeline-perf-93a7 branch 3 times, most recently from eea0fe9 to 936186fCompareAugust 20, 2026 16:45

@cursorcursorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

There are 2 total unresolved issues (including 1 from previous review).

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, have a team admin enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 936186f. Configure here.

Comment threadlib/internal/webstreams/adapters.js Outdated
The kValidateChunk and kDestroyOnSyncError hooks existed only for the
previous stream.Duplex-based CompressionStream implementation, which no
longer uses the adapters.
Assisted-by: Cursor
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
@cursor
cursorBotforce-pushed the cursor/webstreams-pipeline-perf-93a7 branch from 936186f to e3f6bd2CompareAugust 20, 2026 17:11
@anonrig
anonrig marked this pull request as ready for review August 20, 2026 17:12
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@anonrig
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

stream: cut per-chunk overhead in web streams pipelines - #26

Open
anonrig wants to merge 3 commits into
mainfrom
cursor/webstreams-pipeline-perf-93a7
Open

stream: cut per-chunk overhead in web streams pipelines#26
anonrig wants to merge 3 commits into
mainfrom
cursor/webstreams-pipeline-perf-93a7

Conversation

@anonrig

Copy link
Copy Markdown
Owner

What

Reduces per-chunk overhead across common web streams pipelines (fetch body decompression/upload, TextDecoderStream/TextEncoderStream transcoding, Readable.toWeb/Writable.toWeb piping), targeting workloads that move large payloads in small (e.g. 4 KiB) chunks.

Three changes:

  1. CompressionStream/DecompressionStream: process chunks without threadpool round trips.
    Previously these classes wrapped a zlib stream.Duplex in the web streams adapters, so every written chunk was dispatched to the threadpool and its completion observed on a later event loop turn. For a 64 MiB body in 4 KiB chunks that is 16384 threadpool round trips, which dominates the cost of the stream. The classes are now built on a TransformStream that drives the raw zlib/brotli handle synchronously (writeSync), processing inputs larger than 64 KiB in slices with an event-loop turn in between so huge chunks cannot block the loop for their full duration. Output is emitted in up to 64 KiB chunks, either zero-copy or right-sized copies.

  2. TextEncoderStream: fast-path chunk encoding.
    The encode-and-enqueue algorithm was a literal transcription of the spec's per-code-unit loop (single-character string extraction, charCodeAt, string accumulator append for every code unit). Its only observable effects are the surrogate hand-off at chunk boundaries and U+FFFD replacement, which TextEncoder.encode() already performs, so the loop is replaced with explicit chunk-boundary handling plus a single encode() call.

  3. Readable.toWeb(): avoid copying buffers that are the sole view of their ArrayBuffer.
    Chunks were always copied before being enqueued to avoid exposing Buffer pool slices. Buffers that own their entire ArrayBuffer (as fs streams produce for every read) cannot alias pooled memory and are now exposed zero-copy.

Why

Pipelines like fetch() → DecompressionStream → TextDecoderStream → for await, fs.createReadStream() → CompressionStream → fetch POST, and fs → TextDecoderStream → TextEncoderStream → fs spend the majority of their time in per-chunk machinery rather than in the codecs doing the actual work.

(Benchmark results from an isolated machine to be added after the full validation run completes; preliminary measurements show ~13x on the TextEncoderStream stage and ~3x on the DecompressionStream stage for 4 KiB chunks.)

Testing

  • New/existing: test/parallel/test-whatwg-webstreams-compression.js, test-webstreams-compression-bad-chunks.js, test-webstreams-decompression-reject-trailing.js, test-webstreams-compression-buffer-source.js, test-compression-decompression-stream.js, test-whatwg-webstreams-encoding.js, adapters tests.
  • WPT: compression, encoding, streams suites.
  • Adds benchmark/webstreams/compression.js and benchmark/webstreams/encoding.js.
Open in WebOpen in Cursor

@cursor
cursorBotforce-pushed the cursor/webstreams-pipeline-perf-93a7 branch from 796edc7 to 8019292CompareAugust 20, 2026 16:30
setImmediate ??= require('timers').setImmediate;
await new Promise(setImmediate);
}
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Large chunks yield without copying input

High Severity

Inputs larger than kInputSliceSize are processed in slices with an await between them, but each slice is only a view of the original chunk. After transform yields, other tasks can detach or mutate that ArrayBuffer before later slices run, so remaining data is compressed from invalid or changed memory. The WPT detach cases (compression-with-detach, decompression-with-detach) expect the write to keep working after the source is detached.

Additional Locations (1)
Fix in CursorFix in Web

Reviewed by Cursor Bugbot for commit 8019292. Configure here.

Rewrite CompressionStream and DecompressionStream on top of a
TransformStream that drives the zlib (or brotli) handle synchronously,
instead of wrapping a zlib stream.Duplex in the web streams adapters.
The previous implementation dispatched every written chunk to the
threadpool and waited for the event loop to observe its completion,
which dominates the cost of streaming small chunks: a 64 MiB body
written in 4 KiB chunks paid for 16384 threadpool round trips plus the
Transform and adapter machinery around them. Processing the chunks
inline removes that latency entirely while performing the same work.
Inputs larger than 64 KiB are processed in slices with a turn of the
event loop in between so that huge chunks cannot block the loop for
their full duration.
Output is emitted in up to 64 KiB chunks, either as zero-copy views or
as right-sized copies (so small outputs do not retain large buffers),
which also reduces the per-chunk overhead imposed on the rest of the
pipeline downstream.
Assisted-by: Cursor
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
The encode-and-enqueue algorithm was implemented as a literal
transcription of the spec's per-code-unit loop: it extracted a
single-character string, called charCodeAt(), and appended to an
accumulator string for every code unit of every chunk, allocating
millions of temporary strings for large payloads.
The only observable effects of that loop are that a high surrogate at
the end of a chunk is held back to pair with a low surrogate starting
the next chunk, and that unpaired surrogates encode as U+FFFD, which
TextEncoder already does. Handle the chunk boundary explicitly and
encode the rest of the chunk with a single encode() call.
Assisted-by: Cursor
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
@cursor
cursorBotforce-pushed the cursor/webstreams-pipeline-perf-93a7 branch 3 times, most recently from eea0fe9 to 936186fCompareAugust 20, 2026 16:45

@cursorcursorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

There are 2 total unresolved issues (including 1 from previous review).

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, have a team admin enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 936186f. Configure here.

Comment threadlib/internal/webstreams/adapters.js Outdated
The kValidateChunk and kDestroyOnSyncError hooks existed only for the
previous stream.Duplex-based CompressionStream implementation, which no
longer uses the adapters.
Assisted-by: Cursor
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
@cursor
cursorBotforce-pushed the cursor/webstreams-pipeline-perf-93a7 branch from 936186f to e3f6bd2CompareAugust 20, 2026 17:11
@anonrig
anonrig marked this pull request as ready for review August 20, 2026 17:12
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@anonrig
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

stream: cut per-chunk overhead in web streams pipelines - #26

Open
anonrig wants to merge 3 commits into
mainfrom
cursor/webstreams-pipeline-perf-93a7
Open

stream: cut per-chunk overhead in web streams pipelines#26
anonrig wants to merge 3 commits into
mainfrom
cursor/webstreams-pipeline-perf-93a7

Conversation

@anonrig

Copy link
Copy Markdown
Owner

What

Reduces per-chunk overhead across common web streams pipelines (fetch body decompression/upload, TextDecoderStream/TextEncoderStream transcoding, Readable.toWeb/Writable.toWeb piping), targeting workloads that move large payloads in small (e.g. 4 KiB) chunks.

Three changes:

  1. CompressionStream/DecompressionStream: process chunks without threadpool round trips.
    Previously these classes wrapped a zlib stream.Duplex in the web streams adapters, so every written chunk was dispatched to the threadpool and its completion observed on a later event loop turn. For a 64 MiB body in 4 KiB chunks that is 16384 threadpool round trips, which dominates the cost of the stream. The classes are now built on a TransformStream that drives the raw zlib/brotli handle synchronously (writeSync), processing inputs larger than 64 KiB in slices with an event-loop turn in between so huge chunks cannot block the loop for their full duration. Output is emitted in up to 64 KiB chunks, either zero-copy or right-sized copies.

  2. TextEncoderStream: fast-path chunk encoding.
    The encode-and-enqueue algorithm was a literal transcription of the spec's per-code-unit loop (single-character string extraction, charCodeAt, string accumulator append for every code unit). Its only observable effects are the surrogate hand-off at chunk boundaries and U+FFFD replacement, which TextEncoder.encode() already performs, so the loop is replaced with explicit chunk-boundary handling plus a single encode() call.

  3. Readable.toWeb(): avoid copying buffers that are the sole view of their ArrayBuffer.
    Chunks were always copied before being enqueued to avoid exposing Buffer pool slices. Buffers that own their entire ArrayBuffer (as fs streams produce for every read) cannot alias pooled memory and are now exposed zero-copy.

Why

Pipelines like fetch() → DecompressionStream → TextDecoderStream → for await, fs.createReadStream() → CompressionStream → fetch POST, and fs → TextDecoderStream → TextEncoderStream → fs spend the majority of their time in per-chunk machinery rather than in the codecs doing the actual work.

(Benchmark results from an isolated machine to be added after the full validation run completes; preliminary measurements show ~13x on the TextEncoderStream stage and ~3x on the DecompressionStream stage for 4 KiB chunks.)

Testing

  • New/existing: test/parallel/test-whatwg-webstreams-compression.js, test-webstreams-compression-bad-chunks.js, test-webstreams-decompression-reject-trailing.js, test-webstreams-compression-buffer-source.js, test-compression-decompression-stream.js, test-whatwg-webstreams-encoding.js, adapters tests.
  • WPT: compression, encoding, streams suites.
  • Adds benchmark/webstreams/compression.js and benchmark/webstreams/encoding.js.
Open in WebOpen in Cursor

@cursor
cursorBotforce-pushed the cursor/webstreams-pipeline-perf-93a7 branch from 796edc7 to 8019292CompareAugust 20, 2026 16:30
setImmediate ??= require('timers').setImmediate;
await new Promise(setImmediate);
}
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Large chunks yield without copying input

High Severity

Inputs larger than kInputSliceSize are processed in slices with an await between them, but each slice is only a view of the original chunk. After transform yields, other tasks can detach or mutate that ArrayBuffer before later slices run, so remaining data is compressed from invalid or changed memory. The WPT detach cases (compression-with-detach, decompression-with-detach) expect the write to keep working after the source is detached.

Additional Locations (1)
Fix in CursorFix in Web

Reviewed by Cursor Bugbot for commit 8019292. Configure here.

Rewrite CompressionStream and DecompressionStream on top of a
TransformStream that drives the zlib (or brotli) handle synchronously,
instead of wrapping a zlib stream.Duplex in the web streams adapters.
The previous implementation dispatched every written chunk to the
threadpool and waited for the event loop to observe its completion,
which dominates the cost of streaming small chunks: a 64 MiB body
written in 4 KiB chunks paid for 16384 threadpool round trips plus the
Transform and adapter machinery around them. Processing the chunks
inline removes that latency entirely while performing the same work.
Inputs larger than 64 KiB are processed in slices with a turn of the
event loop in between so that huge chunks cannot block the loop for
their full duration.
Output is emitted in up to 64 KiB chunks, either as zero-copy views or
as right-sized copies (so small outputs do not retain large buffers),
which also reduces the per-chunk overhead imposed on the rest of the
pipeline downstream.
Assisted-by: Cursor
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
The encode-and-enqueue algorithm was implemented as a literal
transcription of the spec's per-code-unit loop: it extracted a
single-character string, called charCodeAt(), and appended to an
accumulator string for every code unit of every chunk, allocating
millions of temporary strings for large payloads.
The only observable effects of that loop are that a high surrogate at
the end of a chunk is held back to pair with a low surrogate starting
the next chunk, and that unpaired surrogates encode as U+FFFD, which
TextEncoder already does. Handle the chunk boundary explicitly and
encode the rest of the chunk with a single encode() call.
Assisted-by: Cursor
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
@cursor
cursorBotforce-pushed the cursor/webstreams-pipeline-perf-93a7 branch 3 times, most recently from eea0fe9 to 936186fCompareAugust 20, 2026 16:45

@cursorcursorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

There are 2 total unresolved issues (including 1 from previous review).

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, have a team admin enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 936186f. Configure here.

Comment threadlib/internal/webstreams/adapters.js Outdated
The kValidateChunk and kDestroyOnSyncError hooks existed only for the
previous stream.Duplex-based CompressionStream implementation, which no
longer uses the adapters.
Assisted-by: Cursor
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
@cursor
cursorBotforce-pushed the cursor/webstreams-pipeline-perf-93a7 branch from 936186f to e3f6bd2CompareAugust 20, 2026 17:11
@anonrig
anonrig marked this pull request as ready for review August 20, 2026 17:12
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@anonrig
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

stream: cut per-chunk overhead in web streams pipelines - #26

Open
anonrig wants to merge 3 commits into
mainfrom
cursor/webstreams-pipeline-perf-93a7
Open

stream: cut per-chunk overhead in web streams pipelines#26
anonrig wants to merge 3 commits into
mainfrom
cursor/webstreams-pipeline-perf-93a7

Conversation

@anonrig

Copy link
Copy Markdown
Owner

What

Reduces per-chunk overhead across common web streams pipelines (fetch body decompression/upload, TextDecoderStream/TextEncoderStream transcoding, Readable.toWeb/Writable.toWeb piping), targeting workloads that move large payloads in small (e.g. 4 KiB) chunks.

Three changes:

  1. CompressionStream/DecompressionStream: process chunks without threadpool round trips.
    Previously these classes wrapped a zlib stream.Duplex in the web streams adapters, so every written chunk was dispatched to the threadpool and its completion observed on a later event loop turn. For a 64 MiB body in 4 KiB chunks that is 16384 threadpool round trips, which dominates the cost of the stream. The classes are now built on a TransformStream that drives the raw zlib/brotli handle synchronously (writeSync), processing inputs larger than 64 KiB in slices with an event-loop turn in between so huge chunks cannot block the loop for their full duration. Output is emitted in up to 64 KiB chunks, either zero-copy or right-sized copies.

  2. TextEncoderStream: fast-path chunk encoding.
    The encode-and-enqueue algorithm was a literal transcription of the spec's per-code-unit loop (single-character string extraction, charCodeAt, string accumulator append for every code unit). Its only observable effects are the surrogate hand-off at chunk boundaries and U+FFFD replacement, which TextEncoder.encode() already performs, so the loop is replaced with explicit chunk-boundary handling plus a single encode() call.

  3. Readable.toWeb(): avoid copying buffers that are the sole view of their ArrayBuffer.
    Chunks were always copied before being enqueued to avoid exposing Buffer pool slices. Buffers that own their entire ArrayBuffer (as fs streams produce for every read) cannot alias pooled memory and are now exposed zero-copy.

Why

Pipelines like fetch() → DecompressionStream → TextDecoderStream → for await, fs.createReadStream() → CompressionStream → fetch POST, and fs → TextDecoderStream → TextEncoderStream → fs spend the majority of their time in per-chunk machinery rather than in the codecs doing the actual work.

(Benchmark results from an isolated machine to be added after the full validation run completes; preliminary measurements show ~13x on the TextEncoderStream stage and ~3x on the DecompressionStream stage for 4 KiB chunks.)

Testing

  • New/existing: test/parallel/test-whatwg-webstreams-compression.js, test-webstreams-compression-bad-chunks.js, test-webstreams-decompression-reject-trailing.js, test-webstreams-compression-buffer-source.js, test-compression-decompression-stream.js, test-whatwg-webstreams-encoding.js, adapters tests.
  • WPT: compression, encoding, streams suites.
  • Adds benchmark/webstreams/compression.js and benchmark/webstreams/encoding.js.
Open in WebOpen in Cursor

@cursor
cursorBotforce-pushed the cursor/webstreams-pipeline-perf-93a7 branch from 796edc7 to 8019292CompareAugust 20, 2026 16:30
setImmediate ??= require('timers').setImmediate;
await new Promise(setImmediate);
}
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Large chunks yield without copying input

High Severity

Inputs larger than kInputSliceSize are processed in slices with an await between them, but each slice is only a view of the original chunk. After transform yields, other tasks can detach or mutate that ArrayBuffer before later slices run, so remaining data is compressed from invalid or changed memory. The WPT detach cases (compression-with-detach, decompression-with-detach) expect the write to keep working after the source is detached.

Additional Locations (1)
Fix in CursorFix in Web

Reviewed by Cursor Bugbot for commit 8019292. Configure here.

Rewrite CompressionStream and DecompressionStream on top of a
TransformStream that drives the zlib (or brotli) handle synchronously,
instead of wrapping a zlib stream.Duplex in the web streams adapters.
The previous implementation dispatched every written chunk to the
threadpool and waited for the event loop to observe its completion,
which dominates the cost of streaming small chunks: a 64 MiB body
written in 4 KiB chunks paid for 16384 threadpool round trips plus the
Transform and adapter machinery around them. Processing the chunks
inline removes that latency entirely while performing the same work.
Inputs larger than 64 KiB are processed in slices with a turn of the
event loop in between so that huge chunks cannot block the loop for
their full duration.
Output is emitted in up to 64 KiB chunks, either as zero-copy views or
as right-sized copies (so small outputs do not retain large buffers),
which also reduces the per-chunk overhead imposed on the rest of the
pipeline downstream.
Assisted-by: Cursor
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
The encode-and-enqueue algorithm was implemented as a literal
transcription of the spec's per-code-unit loop: it extracted a
single-character string, called charCodeAt(), and appended to an
accumulator string for every code unit of every chunk, allocating
millions of temporary strings for large payloads.
The only observable effects of that loop are that a high surrogate at
the end of a chunk is held back to pair with a low surrogate starting
the next chunk, and that unpaired surrogates encode as U+FFFD, which
TextEncoder already does. Handle the chunk boundary explicitly and
encode the rest of the chunk with a single encode() call.
Assisted-by: Cursor
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
@cursor
cursorBotforce-pushed the cursor/webstreams-pipeline-perf-93a7 branch 3 times, most recently from eea0fe9 to 936186fCompareAugust 20, 2026 16:45

@cursorcursorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

There are 2 total unresolved issues (including 1 from previous review).

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, have a team admin enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 936186f. Configure here.

Comment threadlib/internal/webstreams/adapters.js Outdated
The kValidateChunk and kDestroyOnSyncError hooks existed only for the
previous stream.Duplex-based CompressionStream implementation, which no
longer uses the adapters.
Assisted-by: Cursor
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
@cursor
cursorBotforce-pushed the cursor/webstreams-pipeline-perf-93a7 branch from 936186f to e3f6bd2CompareAugust 20, 2026 17:11
@anonrig
anonrig marked this pull request as ready for review August 20, 2026 17:12
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@anonrig
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

stream: cut per-chunk overhead in web streams pipelines - #26

Open
anonrig wants to merge 3 commits into
mainfrom
cursor/webstreams-pipeline-perf-93a7
Open

stream: cut per-chunk overhead in web streams pipelines#26
anonrig wants to merge 3 commits into
mainfrom
cursor/webstreams-pipeline-perf-93a7

Conversation

@anonrig

Copy link
Copy Markdown
Owner

What

Reduces per-chunk overhead across common web streams pipelines (fetch body decompression/upload, TextDecoderStream/TextEncoderStream transcoding, Readable.toWeb/Writable.toWeb piping), targeting workloads that move large payloads in small (e.g. 4 KiB) chunks.

Three changes:

  1. CompressionStream/DecompressionStream: process chunks without threadpool round trips.
    Previously these classes wrapped a zlib stream.Duplex in the web streams adapters, so every written chunk was dispatched to the threadpool and its completion observed on a later event loop turn. For a 64 MiB body in 4 KiB chunks that is 16384 threadpool round trips, which dominates the cost of the stream. The classes are now built on a TransformStream that drives the raw zlib/brotli handle synchronously (writeSync), processing inputs larger than 64 KiB in slices with an event-loop turn in between so huge chunks cannot block the loop for their full duration. Output is emitted in up to 64 KiB chunks, either zero-copy or right-sized copies.

  2. TextEncoderStream: fast-path chunk encoding.
    The encode-and-enqueue algorithm was a literal transcription of the spec's per-code-unit loop (single-character string extraction, charCodeAt, string accumulator append for every code unit). Its only observable effects are the surrogate hand-off at chunk boundaries and U+FFFD replacement, which TextEncoder.encode() already performs, so the loop is replaced with explicit chunk-boundary handling plus a single encode() call.

  3. Readable.toWeb(): avoid copying buffers that are the sole view of their ArrayBuffer.
    Chunks were always copied before being enqueued to avoid exposing Buffer pool slices. Buffers that own their entire ArrayBuffer (as fs streams produce for every read) cannot alias pooled memory and are now exposed zero-copy.

Why

Pipelines like fetch() → DecompressionStream → TextDecoderStream → for await, fs.createReadStream() → CompressionStream → fetch POST, and fs → TextDecoderStream → TextEncoderStream → fs spend the majority of their time in per-chunk machinery rather than in the codecs doing the actual work.

(Benchmark results from an isolated machine to be added after the full validation run completes; preliminary measurements show ~13x on the TextEncoderStream stage and ~3x on the DecompressionStream stage for 4 KiB chunks.)

Testing

  • New/existing: test/parallel/test-whatwg-webstreams-compression.js, test-webstreams-compression-bad-chunks.js, test-webstreams-decompression-reject-trailing.js, test-webstreams-compression-buffer-source.js, test-compression-decompression-stream.js, test-whatwg-webstreams-encoding.js, adapters tests.
  • WPT: compression, encoding, streams suites.
  • Adds benchmark/webstreams/compression.js and benchmark/webstreams/encoding.js.
Open in WebOpen in Cursor

@cursor
cursorBotforce-pushed the cursor/webstreams-pipeline-perf-93a7 branch from 796edc7 to 8019292CompareAugust 20, 2026 16:30
setImmediate ??= require('timers').setImmediate;
await new Promise(setImmediate);
}
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Large chunks yield without copying input

High Severity

Inputs larger than kInputSliceSize are processed in slices with an await between them, but each slice is only a view of the original chunk. After transform yields, other tasks can detach or mutate that ArrayBuffer before later slices run, so remaining data is compressed from invalid or changed memory. The WPT detach cases (compression-with-detach, decompression-with-detach) expect the write to keep working after the source is detached.

Additional Locations (1)
Fix in CursorFix in Web

Reviewed by Cursor Bugbot for commit 8019292. Configure here.

Rewrite CompressionStream and DecompressionStream on top of a
TransformStream that drives the zlib (or brotli) handle synchronously,
instead of wrapping a zlib stream.Duplex in the web streams adapters.
The previous implementation dispatched every written chunk to the
threadpool and waited for the event loop to observe its completion,
which dominates the cost of streaming small chunks: a 64 MiB body
written in 4 KiB chunks paid for 16384 threadpool round trips plus the
Transform and adapter machinery around them. Processing the chunks
inline removes that latency entirely while performing the same work.
Inputs larger than 64 KiB are processed in slices with a turn of the
event loop in between so that huge chunks cannot block the loop for
their full duration.
Output is emitted in up to 64 KiB chunks, either as zero-copy views or
as right-sized copies (so small outputs do not retain large buffers),
which also reduces the per-chunk overhead imposed on the rest of the
pipeline downstream.
Assisted-by: Cursor
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
The encode-and-enqueue algorithm was implemented as a literal
transcription of the spec's per-code-unit loop: it extracted a
single-character string, called charCodeAt(), and appended to an
accumulator string for every code unit of every chunk, allocating
millions of temporary strings for large payloads.
The only observable effects of that loop are that a high surrogate at
the end of a chunk is held back to pair with a low surrogate starting
the next chunk, and that unpaired surrogates encode as U+FFFD, which
TextEncoder already does. Handle the chunk boundary explicitly and
encode the rest of the chunk with a single encode() call.
Assisted-by: Cursor
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
@cursor
cursorBotforce-pushed the cursor/webstreams-pipeline-perf-93a7 branch 3 times, most recently from eea0fe9 to 936186fCompareAugust 20, 2026 16:45

@cursorcursorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

There are 2 total unresolved issues (including 1 from previous review).

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, have a team admin enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 936186f. Configure here.

Comment threadlib/internal/webstreams/adapters.js Outdated
The kValidateChunk and kDestroyOnSyncError hooks existed only for the
previous stream.Duplex-based CompressionStream implementation, which no
longer uses the adapters.
Assisted-by: Cursor
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
@cursor
cursorBotforce-pushed the cursor/webstreams-pipeline-perf-93a7 branch from 936186f to e3f6bd2CompareAugust 20, 2026 17:11
@anonrig
anonrig marked this pull request as ready for review August 20, 2026 17:12
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@anonrig
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

stream: cut per-chunk overhead in web streams pipelines - #26

Open
anonrig wants to merge 3 commits into
mainfrom
cursor/webstreams-pipeline-perf-93a7
Open

stream: cut per-chunk overhead in web streams pipelines#26
anonrig wants to merge 3 commits into
mainfrom
cursor/webstreams-pipeline-perf-93a7

Conversation

@anonrig

Copy link
Copy Markdown
Owner

What

Reduces per-chunk overhead across common web streams pipelines (fetch body decompression/upload, TextDecoderStream/TextEncoderStream transcoding, Readable.toWeb/Writable.toWeb piping), targeting workloads that move large payloads in small (e.g. 4 KiB) chunks.

Three changes:

  1. CompressionStream/DecompressionStream: process chunks without threadpool round trips.
    Previously these classes wrapped a zlib stream.Duplex in the web streams adapters, so every written chunk was dispatched to the threadpool and its completion observed on a later event loop turn. For a 64 MiB body in 4 KiB chunks that is 16384 threadpool round trips, which dominates the cost of the stream. The classes are now built on a TransformStream that drives the raw zlib/brotli handle synchronously (writeSync), processing inputs larger than 64 KiB in slices with an event-loop turn in between so huge chunks cannot block the loop for their full duration. Output is emitted in up to 64 KiB chunks, either zero-copy or right-sized copies.

  2. TextEncoderStream: fast-path chunk encoding.
    The encode-and-enqueue algorithm was a literal transcription of the spec's per-code-unit loop (single-character string extraction, charCodeAt, string accumulator append for every code unit). Its only observable effects are the surrogate hand-off at chunk boundaries and U+FFFD replacement, which TextEncoder.encode() already performs, so the loop is replaced with explicit chunk-boundary handling plus a single encode() call.

  3. Readable.toWeb(): avoid copying buffers that are the sole view of their ArrayBuffer.
    Chunks were always copied before being enqueued to avoid exposing Buffer pool slices. Buffers that own their entire ArrayBuffer (as fs streams produce for every read) cannot alias pooled memory and are now exposed zero-copy.

Why

Pipelines like fetch() → DecompressionStream → TextDecoderStream → for await, fs.createReadStream() → CompressionStream → fetch POST, and fs → TextDecoderStream → TextEncoderStream → fs spend the majority of their time in per-chunk machinery rather than in the codecs doing the actual work.

(Benchmark results from an isolated machine to be added after the full validation run completes; preliminary measurements show ~13x on the TextEncoderStream stage and ~3x on the DecompressionStream stage for 4 KiB chunks.)

Testing

  • New/existing: test/parallel/test-whatwg-webstreams-compression.js, test-webstreams-compression-bad-chunks.js, test-webstreams-decompression-reject-trailing.js, test-webstreams-compression-buffer-source.js, test-compression-decompression-stream.js, test-whatwg-webstreams-encoding.js, adapters tests.
  • WPT: compression, encoding, streams suites.
  • Adds benchmark/webstreams/compression.js and benchmark/webstreams/encoding.js.
Open in WebOpen in Cursor

@cursor
cursorBotforce-pushed the cursor/webstreams-pipeline-perf-93a7 branch from 796edc7 to 8019292CompareAugust 20, 2026 16:30
setImmediate ??= require('timers').setImmediate;
await new Promise(setImmediate);
}
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Large chunks yield without copying input

High Severity

Inputs larger than kInputSliceSize are processed in slices with an await between them, but each slice is only a view of the original chunk. After transform yields, other tasks can detach or mutate that ArrayBuffer before later slices run, so remaining data is compressed from invalid or changed memory. The WPT detach cases (compression-with-detach, decompression-with-detach) expect the write to keep working after the source is detached.

Additional Locations (1)
Fix in CursorFix in Web

Reviewed by Cursor Bugbot for commit 8019292. Configure here.

Rewrite CompressionStream and DecompressionStream on top of a
TransformStream that drives the zlib (or brotli) handle synchronously,
instead of wrapping a zlib stream.Duplex in the web streams adapters.
The previous implementation dispatched every written chunk to the
threadpool and waited for the event loop to observe its completion,
which dominates the cost of streaming small chunks: a 64 MiB body
written in 4 KiB chunks paid for 16384 threadpool round trips plus the
Transform and adapter machinery around them. Processing the chunks
inline removes that latency entirely while performing the same work.
Inputs larger than 64 KiB are processed in slices with a turn of the
event loop in between so that huge chunks cannot block the loop for
their full duration.
Output is emitted in up to 64 KiB chunks, either as zero-copy views or
as right-sized copies (so small outputs do not retain large buffers),
which also reduces the per-chunk overhead imposed on the rest of the
pipeline downstream.
Assisted-by: Cursor
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
The encode-and-enqueue algorithm was implemented as a literal
transcription of the spec's per-code-unit loop: it extracted a
single-character string, called charCodeAt(), and appended to an
accumulator string for every code unit of every chunk, allocating
millions of temporary strings for large payloads.
The only observable effects of that loop are that a high surrogate at
the end of a chunk is held back to pair with a low surrogate starting
the next chunk, and that unpaired surrogates encode as U+FFFD, which
TextEncoder already does. Handle the chunk boundary explicitly and
encode the rest of the chunk with a single encode() call.
Assisted-by: Cursor
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
@cursor
cursorBotforce-pushed the cursor/webstreams-pipeline-perf-93a7 branch 3 times, most recently from eea0fe9 to 936186fCompareAugust 20, 2026 16:45

@cursorcursorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

There are 2 total unresolved issues (including 1 from previous review).

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, have a team admin enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 936186f. Configure here.

Comment threadlib/internal/webstreams/adapters.js Outdated
The kValidateChunk and kDestroyOnSyncError hooks existed only for the
previous stream.Duplex-based CompressionStream implementation, which no
longer uses the adapters.
Assisted-by: Cursor
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
@cursor
cursorBotforce-pushed the cursor/webstreams-pipeline-perf-93a7 branch from 936186f to e3f6bd2CompareAugust 20, 2026 17:11
@anonrig
anonrig marked this pull request as ready for review August 20, 2026 17:12
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@anonrig
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

stream: cut per-chunk overhead in web streams pipelines - #26

Open
anonrig wants to merge 3 commits into
mainfrom
cursor/webstreams-pipeline-perf-93a7
Open

stream: cut per-chunk overhead in web streams pipelines#26
anonrig wants to merge 3 commits into
mainfrom
cursor/webstreams-pipeline-perf-93a7

Conversation

@anonrig

Copy link
Copy Markdown
Owner

What

Reduces per-chunk overhead across common web streams pipelines (fetch body decompression/upload, TextDecoderStream/TextEncoderStream transcoding, Readable.toWeb/Writable.toWeb piping), targeting workloads that move large payloads in small (e.g. 4 KiB) chunks.

Three changes:

  1. CompressionStream/DecompressionStream: process chunks without threadpool round trips.
    Previously these classes wrapped a zlib stream.Duplex in the web streams adapters, so every written chunk was dispatched to the threadpool and its completion observed on a later event loop turn. For a 64 MiB body in 4 KiB chunks that is 16384 threadpool round trips, which dominates the cost of the stream. The classes are now built on a TransformStream that drives the raw zlib/brotli handle synchronously (writeSync), processing inputs larger than 64 KiB in slices with an event-loop turn in between so huge chunks cannot block the loop for their full duration. Output is emitted in up to 64 KiB chunks, either zero-copy or right-sized copies.

  2. TextEncoderStream: fast-path chunk encoding.
    The encode-and-enqueue algorithm was a literal transcription of the spec's per-code-unit loop (single-character string extraction, charCodeAt, string accumulator append for every code unit). Its only observable effects are the surrogate hand-off at chunk boundaries and U+FFFD replacement, which TextEncoder.encode() already performs, so the loop is replaced with explicit chunk-boundary handling plus a single encode() call.

  3. Readable.toWeb(): avoid copying buffers that are the sole view of their ArrayBuffer.
    Chunks were always copied before being enqueued to avoid exposing Buffer pool slices. Buffers that own their entire ArrayBuffer (as fs streams produce for every read) cannot alias pooled memory and are now exposed zero-copy.

Why

Pipelines like fetch() → DecompressionStream → TextDecoderStream → for await, fs.createReadStream() → CompressionStream → fetch POST, and fs → TextDecoderStream → TextEncoderStream → fs spend the majority of their time in per-chunk machinery rather than in the codecs doing the actual work.

(Benchmark results from an isolated machine to be added after the full validation run completes; preliminary measurements show ~13x on the TextEncoderStream stage and ~3x on the DecompressionStream stage for 4 KiB chunks.)

Testing

  • New/existing: test/parallel/test-whatwg-webstreams-compression.js, test-webstreams-compression-bad-chunks.js, test-webstreams-decompression-reject-trailing.js, test-webstreams-compression-buffer-source.js, test-compression-decompression-stream.js, test-whatwg-webstreams-encoding.js, adapters tests.
  • WPT: compression, encoding, streams suites.
  • Adds benchmark/webstreams/compression.js and benchmark/webstreams/encoding.js.
Open in WebOpen in Cursor

@cursor
cursorBotforce-pushed the cursor/webstreams-pipeline-perf-93a7 branch from 796edc7 to 8019292CompareAugust 20, 2026 16:30
setImmediate ??= require('timers').setImmediate;
await new Promise(setImmediate);
}
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Large chunks yield without copying input

High Severity

Inputs larger than kInputSliceSize are processed in slices with an await between them, but each slice is only a view of the original chunk. After transform yields, other tasks can detach or mutate that ArrayBuffer before later slices run, so remaining data is compressed from invalid or changed memory. The WPT detach cases (compression-with-detach, decompression-with-detach) expect the write to keep working after the source is detached.

Additional Locations (1)
Fix in CursorFix in Web

Reviewed by Cursor Bugbot for commit 8019292. Configure here.

Rewrite CompressionStream and DecompressionStream on top of a
TransformStream that drives the zlib (or brotli) handle synchronously,
instead of wrapping a zlib stream.Duplex in the web streams adapters.
The previous implementation dispatched every written chunk to the
threadpool and waited for the event loop to observe its completion,
which dominates the cost of streaming small chunks: a 64 MiB body
written in 4 KiB chunks paid for 16384 threadpool round trips plus the
Transform and adapter machinery around them. Processing the chunks
inline removes that latency entirely while performing the same work.
Inputs larger than 64 KiB are processed in slices with a turn of the
event loop in between so that huge chunks cannot block the loop for
their full duration.
Output is emitted in up to 64 KiB chunks, either as zero-copy views or
as right-sized copies (so small outputs do not retain large buffers),
which also reduces the per-chunk overhead imposed on the rest of the
pipeline downstream.
Assisted-by: Cursor
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
The encode-and-enqueue algorithm was implemented as a literal
transcription of the spec's per-code-unit loop: it extracted a
single-character string, called charCodeAt(), and appended to an
accumulator string for every code unit of every chunk, allocating
millions of temporary strings for large payloads.
The only observable effects of that loop are that a high surrogate at
the end of a chunk is held back to pair with a low surrogate starting
the next chunk, and that unpaired surrogates encode as U+FFFD, which
TextEncoder already does. Handle the chunk boundary explicitly and
encode the rest of the chunk with a single encode() call.
Assisted-by: Cursor
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
@cursor
cursorBotforce-pushed the cursor/webstreams-pipeline-perf-93a7 branch 3 times, most recently from eea0fe9 to 936186fCompareAugust 20, 2026 16:45

@cursorcursorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

There are 2 total unresolved issues (including 1 from previous review).

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, have a team admin enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 936186f. Configure here.

Comment threadlib/internal/webstreams/adapters.js Outdated
The kValidateChunk and kDestroyOnSyncError hooks existed only for the
previous stream.Duplex-based CompressionStream implementation, which no
longer uses the adapters.
Assisted-by: Cursor
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
@cursor
cursorBotforce-pushed the cursor/webstreams-pipeline-perf-93a7 branch from 936186f to e3f6bd2CompareAugust 20, 2026 17:11
@anonrig
anonrig marked this pull request as ready for review August 20, 2026 17:12
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@anonrig