src: use simdutf for UTF-8 ⇄ UTF-16 transcoding in StringDecoder and StringBytes::Write - #65324

Closed
codebytere wants to merge 2 commits into
nodejs:mainfrom
codebytere:perf/src-simdutf-utf8-transcoding
Closed

src: use simdutf for UTF-8 ⇄ UTF-16 transcoding in StringDecoder and StringBytes::Write#65324
codebytere wants to merge 2 commits into
nodejs:mainfrom
codebytere:perf/src-simdutf-utf8-transcoding

Conversation

@codebytere

@codebyterecodebytere commented Aug 16, 2026

Copy link
Copy Markdown
Member

StringDecoder gets 3–12× faster on UTF-8, and writing non-ASCII strings as UTF-8 (Buffer#write, Buffer.from(string), every stream/socket/fs string write) gets 4–5× faster. A 64 KiB JSON-RPC round trip over a child's stdio (stringify → write → readline → parse, non-ASCII payload) drops from 1.28 ms to 0.90 ms with only the parent patched, 0.67 ms with both ends.

string_decoder/string-decoder.js encoding='utf8' inLen=1024 chunkLen=1024 *** 1125.82 % ±8.25%
string_decoder/string-decoder.js encoding='utf8' inLen=128 chunkLen=1024 *** 322.79 % ±2.28%
string_decoder/string-decoder.js encoding='utf8' inLen=32 chunkLen=1024 *** 98.53 % ±1.37%
string_decoder/string-decoder.js encoding='utf8' inLen=1024 chunkLen=16 *** 5.15 % ±0.58%
buffers/buffer-write-string-utf8.js (new) len=65536 chars='two-byte' *** 391.17 %
buffers/buffer-write-string-utf8.js len=65536 chars='two-byte-lone-surrogate' *** 314.76 %
buffers/buffer-write-string-utf8.js len=256 chars='two-byte' *** 289.28 %
buffers/buffer-write-string-utf8.js len=256 chars='two-byte-astral' *** 228.97 %
one-byte strings / other encodings ~0 % n.s.

(Linux x64, benchmark/compare.js, 30 runs, significance as in compare.R.)

Two commits:

string_decoder: decode UTF-8 via StringBytes::Encode - the decoder built strings with String::NewFromUtf8(); buffer.toString() already goes through StringBytes::Encode(), which uses simdutf. MakeString() now calls the same function (the kMaxLengthERR_STRING_TOO_LONG check stays in front), so the decoder returns exactly what toString() returns for the same bytes. The partial-character bookkeeping is untouched.

src: use simdutf for two-byte strings in UTF-8 writes - StringBytes::Write(UTF8) used String::WriteUtf8V2() for two-byte strings. For strings longer than 32 code units it now validates with simdutf::validate_utf16() (lone surrogates are repaired into a scratch buffer with to_well_formed_utf16(), so the output stays byte-identical to V8's kReplaceInvalidUtf8) and transcodes with convert_utf16_to_utf8() - but only when the destination is known to fit (buflen >= 3 × length or >= utf8_length_from_utf16()); otherwise it falls back to V8, so partial writes truncate at the same character boundary as before. Shorter strings and one-byte strings keep the V8 path.

Tests:test-string-decoder-utf8-large.js (new: large inputs, chunk boundaries inside multi-byte sequences, invalid sequences vs toString()); test-buffer-write-utf8-two-byte.js (new: independent reference encoder; BMP/astral/lone surrogates at start/middle/end; exact-fit, one-short and 3× destinations; partial writes). Existing string_decoder, buffer, stream, net, http, child_process and readline suites pass.


Disclosure: the code, test, measurements and this description were written by Claude Code, directed and reviewed by @codebytere.

@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

Review requested:

  • @nodejs/performance

@nodejs-github-botnodejs-github-bot added buffer Issues and PRs related to the buffer subsystem. c++ Issues and PRs that require attention from people who are familiar with C++. needs-ci PRs that need a full CI run. labels Aug 16, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.
Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).
buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
@codebytere
codebytereforce-pushed the perf/src-simdutf-utf8-transcoding branch from ec59ddb to 0c89308CompareAugust 16, 2026 15:51
@codebyterecodebytere added request-ci Add this label to start a Jenkins CI on a PR. and removed needs-ci PRs that need a full CI run. labels Aug 16, 2026
@codecov

codecovBot commented Aug 16, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 90.11%. Comparing base (30bff4a) to head (0c89308).
⚠️ Report is 38 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #65324 +/- ##
==========================================
- Coverage 90.13% 90.11% -0.03% 
==========================================
Files 752 752 Lines 251568 251578 +10 Branches 47270 47273 +3 ==========================================
- Hits 226759 226706 -53 - Misses 16168 16216 +48 - Partials 8641 8656 +15 
Files with missing linesCoverage Δ
src/string_bytes.cc75.23% <100.00%> (+1.49%)⬆️
src/string_decoder.cc92.17% <100.00%> (-0.34%)⬇️

... and 37 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@github-actionsgithub-actionsBot removed the request-ci Add this label to start a Jenkins CI on a PR. label Aug 16, 2026
@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

Comment threadsrc/string_bytes.cc
// 3 * length is what StorageSize() hands most callers; only compute
// the exact length when the buffer is smaller than that.
if (buflen >= 3 * length ||
buflen >= simdutf::utf8_length_from_utf16(data, length)) {

@lemirelemireAug 16, 2026

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We have simdutf::utf8_length_from_utf16_with_replacement that would be appropriate in this function.

Note that we added convert_utf16_to_utf8_with_replacement in arelease this year. (But this code is still quite fine.)

@lemire

Copy link
Copy Markdown
Member

@codebytere I might send you a PR later today. Hold on a bit.

@lemire

Copy link
Copy Markdown
Member

@codebytere The simdutf version might not be ready yet. So I still recommend merging this PR. It is good work. We might be able to improve it later, but this should not stop this PR.

@github-actions

github-actionsBot commented Aug 17, 2026

Copy link
Copy Markdown
Contributor

@codebyterecodebytere added the commit-queue PRs queued for automated landing through the Commit Queue. label Aug 17, 2026
@codebytere

Copy link
Copy Markdown
MemberAuthor

@jasnell is commit queue broken?

@jasnell

Copy link
Copy Markdown
Member

Entirely possible

@nodejs-github-botnodejs-github-bot added commit-queue-failed PRs whose Commit Queue landing failed and need manual intervention before retrying. and removed commit-queue PRs queued for automated landing through the Commit Queue. labels Aug 18, 2026
@nodejs-github-bot

Copy link
Copy Markdown
Collaborator
Commit Queue failed
- Loading data for nodejs/node/pull/65324
✔ Done loading data for nodejs/node/pull/65324
----------------------------------- PR info ------------------------------------
Title src: use simdutf for UTF-8 ⇄ UTF-16 transcoding in StringDecoder and StringBytes::Write (#65324)
Author Shelley Vohr <shelley.vohr@gmail.com> (@codebytere)
Branch codebytere:perf/src-simdutf-utf8-transcoding -> nodejs:main
Labels buffer, c++, commit-queue
Commits 2
- string_decoder: decode UTF-8 via StringBytes::Encode
- src: use simdutf for two-byte strings in UTF-8 writes
Committers 1
- Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: https://github.com/nodejs/node/pull/65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
------------------------------ Generated metadata ------------------------------
PR-URL: https://github.com/nodejs/node/pull/65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
--------------------------------------------------------------------------------
ℹ This PR was created on Sun, 16 Aug 2026 15:37:44 GMT
✔ Approvals: 4
✔ - Yagiz Nizipli (@anonrig) (TSC): https://github.com/nodejs/node/pull/65324#pullrequestreview-4947047263
✔ - Gürgün Dayıoğlu (@gurgunday): https://github.com/nodejs/node/pull/65324#pullrequestreview-4947063127
✔ - Daniel Lemire (@lemire): https://github.com/nodejs/node/pull/65324#pullrequestreview-4947151736
✔ - James M Snell (@jasnell) (TSC): https://github.com/nodejs/node/pull/65324#pullrequestreview-4948248917
✔ Last GitHub CI successful
ℹ Last Full PR CI on 2026-08-16T18:58:59Z: https://ci.nodejs.org/job/node-test-pull-request/75896/
- Querying data for job/node-test-pull-request/75896/
✔ Build data downloaded
✔ Last Jenkins CI successful
--------------------------------------------------------------------------------
✔ No git cherry-pick in progress
✔ No git am in progress
✔ No git rebase in progress
--------------------------------------------------------------------------------
- Bringing origin/main up to date...
From https://github.com/nodejs/node
* branch main -> FETCH_HEAD
✔ origin/main is now up-to-date
- Downloading patch for 65324
From https://github.com/nodejs/node
* branch refs/pull/65324/merge -> FETCH_HEAD
✔ Fetched commits as 0d84f29fa77a..0c89308bba0c
--------------------------------------------------------------------------------
[main 308a46181c] string_decoder: decode UTF-8 via StringBytes::Encode
Author: Shelley Vohr <shelley.vohr@gmail.com>
Date: Sat Aug 15 22:09:21 2026 +0000
2 files changed, 94 insertions(+), 15 deletions(-)
create mode 100644 test/parallel/test-string-decoder-utf8-large.js
[main 619cb0dd4c] src: use simdutf for two-byte strings in UTF-8 writes
Author: Shelley Vohr <shelley.vohr@gmail.com>
Date: Sat Aug 15 22:34:28 2026 +0000
3 files changed, 193 insertions(+), 1 deletion(-)
create mode 100644 benchmark/buffers/buffer-write-string-utf8.js
create mode 100644 test/parallel/test-buffer-write-utf8-two-byte.js
✔ Patches applied
There are 2 commits in the PR. Attempting autorebase.
(node:381) [DEP0190] DeprecationWarning: Passing args to a child process with shell option true can lead to security vulnerabilities, as the arguments are not escaped, only concatenated.
(Use `node --trace-deprecation ...` to show where the warning was created)
Rebasing (2/4)
Executing: git node land --amend --yes
--------------------------------- New Message ----------------------------------
string_decoder: decode UTF-8 via StringBytes::Encode

StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.

benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.

Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>

[detached HEAD 1055900345] string_decoder: decode UTF-8 via StringBytes::Encode
Author: Shelley Vohr <shelley.vohr@gmail.com>
Date: Sat Aug 15 22:09:21 2026 +0000
2 files changed, 94 insertions(+), 15 deletions(-)
create mode 100644 test/parallel/test-string-decoder-utf8-large.js
Rebasing (3/4)
Rebasing (4/4)
Executing: git node land --amend --yes
--------------------------------- New Message ----------------------------------
src: use simdutf for two-byte strings in UTF-8 writes

StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.

Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).

buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.

Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>

[detached HEAD 55c27c9b17] src: use simdutf for two-byte strings in UTF-8 writes
Author: Shelley Vohr <shelley.vohr@gmail.com>
Date: Sat Aug 15 22:34:28 2026 +0000
3 files changed, 193 insertions(+), 1 deletion(-)
create mode 100644 benchmark/buffers/buffer-write-string-utf8.js
create mode 100644 test/parallel/test-buffer-write-utf8-two-byte.js
Successfully rebased and updated refs/heads/main.

ℹ Add commit-queue-squash label to land the PR as one commit, or commit-queue-rebase to land as separate commits.

https://github.com/nodejs/node/actions/runs/32156233387

@aduh95aduh95 added commit-queue PRs queued for automated landing through the Commit Queue. commit-queue-rebase PRs the Commit Queue should land as multiple self-contained commits. and removed commit-queue-failed PRs whose Commit Queue landing failed and need manual intervention before retrying. labels Aug 18, 2026
@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

Landed in 29a083c...55e4ca3

nodejs-github-bot pushed a commit that referenced this pull request Aug 18, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
nodejs-github-bot pushed a commit that referenced this pull request Aug 18, 2026
StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.
Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).
buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
@nodejs-github-botnodejs-github-bot removed the commit-queue PRs queued for automated landing through the Commit Queue. label Aug 18, 2026
@codebytere
codebytere deleted the perf/src-simdutf-utf8-transcoding branch August 18, 2026 19:22
aduh95 pushed a commit that referenced this pull request Aug 25, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
aduh95 pushed a commit that referenced this pull request Aug 25, 2026
StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.
Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).
buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
aduh95 pushed a commit that referenced this pull request Aug 25, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
aduh95 pushed a commit that referenced this pull request Aug 25, 2026
StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.
Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).
buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
@aduh95aduh95 added dont-land-on-v22.x PRs that should not land on the v22.x-staging branch and should not be released in v22.x. dont-land-on-v24.x PRs that should not land on the v24.x-staging branch and should not be released in v24.x. labels Aug 27, 2026
aduh95 pushed a commit that referenced this pull request Aug 27, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bufferIssues and PRs related to the buffer subsystem.c++Issues and PRs that require attention from people who are familiar with C++.commit-queue-rebasePRs the Commit Queue should land as multiple self-contained commits.dont-land-on-v22.xPRs that should not land on the v22.x-staging branch and should not be released in v22.x.dont-land-on-v24.xPRs that should not land on the v24.x-staging branch and should not be released in v24.x.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants

@codebytere@nodejs-github-bot@lemire@jasnell@anonrig@mertcanaltin@gurgunday@aduh95
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

src: use simdutf for UTF-8 ⇄ UTF-16 transcoding in StringDecoder and StringBytes::Write - #65324

Closed
codebytere wants to merge 2 commits into
nodejs:mainfrom
codebytere:perf/src-simdutf-utf8-transcoding
Closed

src: use simdutf for UTF-8 ⇄ UTF-16 transcoding in StringDecoder and StringBytes::Write#65324
codebytere wants to merge 2 commits into
nodejs:mainfrom
codebytere:perf/src-simdutf-utf8-transcoding

Conversation

@codebytere

@codebyterecodebytere commented Aug 16, 2026

Copy link
Copy Markdown
Member

StringDecoder gets 3–12× faster on UTF-8, and writing non-ASCII strings as UTF-8 (Buffer#write, Buffer.from(string), every stream/socket/fs string write) gets 4–5× faster. A 64 KiB JSON-RPC round trip over a child's stdio (stringify → write → readline → parse, non-ASCII payload) drops from 1.28 ms to 0.90 ms with only the parent patched, 0.67 ms with both ends.

string_decoder/string-decoder.js encoding='utf8' inLen=1024 chunkLen=1024 *** 1125.82 % ±8.25%
string_decoder/string-decoder.js encoding='utf8' inLen=128 chunkLen=1024 *** 322.79 % ±2.28%
string_decoder/string-decoder.js encoding='utf8' inLen=32 chunkLen=1024 *** 98.53 % ±1.37%
string_decoder/string-decoder.js encoding='utf8' inLen=1024 chunkLen=16 *** 5.15 % ±0.58%
buffers/buffer-write-string-utf8.js (new) len=65536 chars='two-byte' *** 391.17 %
buffers/buffer-write-string-utf8.js len=65536 chars='two-byte-lone-surrogate' *** 314.76 %
buffers/buffer-write-string-utf8.js len=256 chars='two-byte' *** 289.28 %
buffers/buffer-write-string-utf8.js len=256 chars='two-byte-astral' *** 228.97 %
one-byte strings / other encodings ~0 % n.s.

(Linux x64, benchmark/compare.js, 30 runs, significance as in compare.R.)

Two commits:

string_decoder: decode UTF-8 via StringBytes::Encode - the decoder built strings with String::NewFromUtf8(); buffer.toString() already goes through StringBytes::Encode(), which uses simdutf. MakeString() now calls the same function (the kMaxLengthERR_STRING_TOO_LONG check stays in front), so the decoder returns exactly what toString() returns for the same bytes. The partial-character bookkeeping is untouched.

src: use simdutf for two-byte strings in UTF-8 writes - StringBytes::Write(UTF8) used String::WriteUtf8V2() for two-byte strings. For strings longer than 32 code units it now validates with simdutf::validate_utf16() (lone surrogates are repaired into a scratch buffer with to_well_formed_utf16(), so the output stays byte-identical to V8's kReplaceInvalidUtf8) and transcodes with convert_utf16_to_utf8() - but only when the destination is known to fit (buflen >= 3 × length or >= utf8_length_from_utf16()); otherwise it falls back to V8, so partial writes truncate at the same character boundary as before. Shorter strings and one-byte strings keep the V8 path.

Tests:test-string-decoder-utf8-large.js (new: large inputs, chunk boundaries inside multi-byte sequences, invalid sequences vs toString()); test-buffer-write-utf8-two-byte.js (new: independent reference encoder; BMP/astral/lone surrogates at start/middle/end; exact-fit, one-short and 3× destinations; partial writes). Existing string_decoder, buffer, stream, net, http, child_process and readline suites pass.


Disclosure: the code, test, measurements and this description were written by Claude Code, directed and reviewed by @codebytere.

@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

Review requested:

  • @nodejs/performance

@nodejs-github-botnodejs-github-bot added buffer Issues and PRs related to the buffer subsystem. c++ Issues and PRs that require attention from people who are familiar with C++. needs-ci PRs that need a full CI run. labels Aug 16, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.
Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).
buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
@codebytere
codebytereforce-pushed the perf/src-simdutf-utf8-transcoding branch from ec59ddb to 0c89308CompareAugust 16, 2026 15:51
@codebyterecodebytere added request-ci Add this label to start a Jenkins CI on a PR. and removed needs-ci PRs that need a full CI run. labels Aug 16, 2026
@codecov

codecovBot commented Aug 16, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 90.11%. Comparing base (30bff4a) to head (0c89308).
⚠️ Report is 38 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #65324 +/- ##
==========================================
- Coverage 90.13% 90.11% -0.03% 
==========================================
Files 752 752 Lines 251568 251578 +10 Branches 47270 47273 +3 ==========================================
- Hits 226759 226706 -53 - Misses 16168 16216 +48 - Partials 8641 8656 +15 
Files with missing linesCoverage Δ
src/string_bytes.cc75.23% <100.00%> (+1.49%)⬆️
src/string_decoder.cc92.17% <100.00%> (-0.34%)⬇️

... and 37 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@github-actionsgithub-actionsBot removed the request-ci Add this label to start a Jenkins CI on a PR. label Aug 16, 2026
@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

Comment threadsrc/string_bytes.cc
// 3 * length is what StorageSize() hands most callers; only compute
// the exact length when the buffer is smaller than that.
if (buflen >= 3 * length ||
buflen >= simdutf::utf8_length_from_utf16(data, length)) {

@lemirelemireAug 16, 2026

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We have simdutf::utf8_length_from_utf16_with_replacement that would be appropriate in this function.

Note that we added convert_utf16_to_utf8_with_replacement in arelease this year. (But this code is still quite fine.)

@lemire

Copy link
Copy Markdown
Member

@codebytere I might send you a PR later today. Hold on a bit.

@lemire

Copy link
Copy Markdown
Member

@codebytere The simdutf version might not be ready yet. So I still recommend merging this PR. It is good work. We might be able to improve it later, but this should not stop this PR.

@github-actions

github-actionsBot commented Aug 17, 2026

Copy link
Copy Markdown
Contributor

@codebyterecodebytere added the commit-queue PRs queued for automated landing through the Commit Queue. label Aug 17, 2026
@codebytere

Copy link
Copy Markdown
MemberAuthor

@jasnell is commit queue broken?

@jasnell

Copy link
Copy Markdown
Member

Entirely possible

@nodejs-github-botnodejs-github-bot added commit-queue-failed PRs whose Commit Queue landing failed and need manual intervention before retrying. and removed commit-queue PRs queued for automated landing through the Commit Queue. labels Aug 18, 2026
@nodejs-github-bot

Copy link
Copy Markdown
Collaborator
Commit Queue failed
- Loading data for nodejs/node/pull/65324
✔ Done loading data for nodejs/node/pull/65324
----------------------------------- PR info ------------------------------------
Title src: use simdutf for UTF-8 ⇄ UTF-16 transcoding in StringDecoder and StringBytes::Write (#65324)
Author Shelley Vohr <shelley.vohr@gmail.com> (@codebytere)
Branch codebytere:perf/src-simdutf-utf8-transcoding -> nodejs:main
Labels buffer, c++, commit-queue
Commits 2
- string_decoder: decode UTF-8 via StringBytes::Encode
- src: use simdutf for two-byte strings in UTF-8 writes
Committers 1
- Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: https://github.com/nodejs/node/pull/65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
------------------------------ Generated metadata ------------------------------
PR-URL: https://github.com/nodejs/node/pull/65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
--------------------------------------------------------------------------------
ℹ This PR was created on Sun, 16 Aug 2026 15:37:44 GMT
✔ Approvals: 4
✔ - Yagiz Nizipli (@anonrig) (TSC): https://github.com/nodejs/node/pull/65324#pullrequestreview-4947047263
✔ - Gürgün Dayıoğlu (@gurgunday): https://github.com/nodejs/node/pull/65324#pullrequestreview-4947063127
✔ - Daniel Lemire (@lemire): https://github.com/nodejs/node/pull/65324#pullrequestreview-4947151736
✔ - James M Snell (@jasnell) (TSC): https://github.com/nodejs/node/pull/65324#pullrequestreview-4948248917
✔ Last GitHub CI successful
ℹ Last Full PR CI on 2026-08-16T18:58:59Z: https://ci.nodejs.org/job/node-test-pull-request/75896/
- Querying data for job/node-test-pull-request/75896/
✔ Build data downloaded
✔ Last Jenkins CI successful
--------------------------------------------------------------------------------
✔ No git cherry-pick in progress
✔ No git am in progress
✔ No git rebase in progress
--------------------------------------------------------------------------------
- Bringing origin/main up to date...
From https://github.com/nodejs/node
* branch main -> FETCH_HEAD
✔ origin/main is now up-to-date
- Downloading patch for 65324
From https://github.com/nodejs/node
* branch refs/pull/65324/merge -> FETCH_HEAD
✔ Fetched commits as 0d84f29fa77a..0c89308bba0c
--------------------------------------------------------------------------------
[main 308a46181c] string_decoder: decode UTF-8 via StringBytes::Encode
Author: Shelley Vohr <shelley.vohr@gmail.com>
Date: Sat Aug 15 22:09:21 2026 +0000
2 files changed, 94 insertions(+), 15 deletions(-)
create mode 100644 test/parallel/test-string-decoder-utf8-large.js
[main 619cb0dd4c] src: use simdutf for two-byte strings in UTF-8 writes
Author: Shelley Vohr <shelley.vohr@gmail.com>
Date: Sat Aug 15 22:34:28 2026 +0000
3 files changed, 193 insertions(+), 1 deletion(-)
create mode 100644 benchmark/buffers/buffer-write-string-utf8.js
create mode 100644 test/parallel/test-buffer-write-utf8-two-byte.js
✔ Patches applied
There are 2 commits in the PR. Attempting autorebase.
(node:381) [DEP0190] DeprecationWarning: Passing args to a child process with shell option true can lead to security vulnerabilities, as the arguments are not escaped, only concatenated.
(Use `node --trace-deprecation ...` to show where the warning was created)
Rebasing (2/4)
Executing: git node land --amend --yes
--------------------------------- New Message ----------------------------------
string_decoder: decode UTF-8 via StringBytes::Encode

StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.

benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.

Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>

[detached HEAD 1055900345] string_decoder: decode UTF-8 via StringBytes::Encode
Author: Shelley Vohr <shelley.vohr@gmail.com>
Date: Sat Aug 15 22:09:21 2026 +0000
2 files changed, 94 insertions(+), 15 deletions(-)
create mode 100644 test/parallel/test-string-decoder-utf8-large.js
Rebasing (3/4)
Rebasing (4/4)
Executing: git node land --amend --yes
--------------------------------- New Message ----------------------------------
src: use simdutf for two-byte strings in UTF-8 writes

StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.

Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).

buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.

Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>

[detached HEAD 55c27c9b17] src: use simdutf for two-byte strings in UTF-8 writes
Author: Shelley Vohr <shelley.vohr@gmail.com>
Date: Sat Aug 15 22:34:28 2026 +0000
3 files changed, 193 insertions(+), 1 deletion(-)
create mode 100644 benchmark/buffers/buffer-write-string-utf8.js
create mode 100644 test/parallel/test-buffer-write-utf8-two-byte.js
Successfully rebased and updated refs/heads/main.

ℹ Add commit-queue-squash label to land the PR as one commit, or commit-queue-rebase to land as separate commits.

https://github.com/nodejs/node/actions/runs/32156233387

@aduh95aduh95 added commit-queue PRs queued for automated landing through the Commit Queue. commit-queue-rebase PRs the Commit Queue should land as multiple self-contained commits. and removed commit-queue-failed PRs whose Commit Queue landing failed and need manual intervention before retrying. labels Aug 18, 2026
@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

Landed in 29a083c...55e4ca3

nodejs-github-bot pushed a commit that referenced this pull request Aug 18, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
nodejs-github-bot pushed a commit that referenced this pull request Aug 18, 2026
StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.
Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).
buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
@nodejs-github-botnodejs-github-bot removed the commit-queue PRs queued for automated landing through the Commit Queue. label Aug 18, 2026
@codebytere
codebytere deleted the perf/src-simdutf-utf8-transcoding branch August 18, 2026 19:22
aduh95 pushed a commit that referenced this pull request Aug 25, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
aduh95 pushed a commit that referenced this pull request Aug 25, 2026
StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.
Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).
buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
aduh95 pushed a commit that referenced this pull request Aug 25, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
aduh95 pushed a commit that referenced this pull request Aug 25, 2026
StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.
Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).
buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
@aduh95aduh95 added dont-land-on-v22.x PRs that should not land on the v22.x-staging branch and should not be released in v22.x. dont-land-on-v24.x PRs that should not land on the v24.x-staging branch and should not be released in v24.x. labels Aug 27, 2026
aduh95 pushed a commit that referenced this pull request Aug 27, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bufferIssues and PRs related to the buffer subsystem.c++Issues and PRs that require attention from people who are familiar with C++.commit-queue-rebasePRs the Commit Queue should land as multiple self-contained commits.dont-land-on-v22.xPRs that should not land on the v22.x-staging branch and should not be released in v22.x.dont-land-on-v24.xPRs that should not land on the v24.x-staging branch and should not be released in v24.x.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants

@codebytere@nodejs-github-bot@lemire@jasnell@anonrig@mertcanaltin@gurgunday@aduh95
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

src: use simdutf for UTF-8 ⇄ UTF-16 transcoding in StringDecoder and StringBytes::Write - #65324

Closed
codebytere wants to merge 2 commits into
nodejs:mainfrom
codebytere:perf/src-simdutf-utf8-transcoding
Closed

src: use simdutf for UTF-8 ⇄ UTF-16 transcoding in StringDecoder and StringBytes::Write#65324
codebytere wants to merge 2 commits into
nodejs:mainfrom
codebytere:perf/src-simdutf-utf8-transcoding

Conversation

@codebytere

@codebyterecodebytere commented Aug 16, 2026

Copy link
Copy Markdown
Member

StringDecoder gets 3–12× faster on UTF-8, and writing non-ASCII strings as UTF-8 (Buffer#write, Buffer.from(string), every stream/socket/fs string write) gets 4–5× faster. A 64 KiB JSON-RPC round trip over a child's stdio (stringify → write → readline → parse, non-ASCII payload) drops from 1.28 ms to 0.90 ms with only the parent patched, 0.67 ms with both ends.

string_decoder/string-decoder.js encoding='utf8' inLen=1024 chunkLen=1024 *** 1125.82 % ±8.25%
string_decoder/string-decoder.js encoding='utf8' inLen=128 chunkLen=1024 *** 322.79 % ±2.28%
string_decoder/string-decoder.js encoding='utf8' inLen=32 chunkLen=1024 *** 98.53 % ±1.37%
string_decoder/string-decoder.js encoding='utf8' inLen=1024 chunkLen=16 *** 5.15 % ±0.58%
buffers/buffer-write-string-utf8.js (new) len=65536 chars='two-byte' *** 391.17 %
buffers/buffer-write-string-utf8.js len=65536 chars='two-byte-lone-surrogate' *** 314.76 %
buffers/buffer-write-string-utf8.js len=256 chars='two-byte' *** 289.28 %
buffers/buffer-write-string-utf8.js len=256 chars='two-byte-astral' *** 228.97 %
one-byte strings / other encodings ~0 % n.s.

(Linux x64, benchmark/compare.js, 30 runs, significance as in compare.R.)

Two commits:

string_decoder: decode UTF-8 via StringBytes::Encode - the decoder built strings with String::NewFromUtf8(); buffer.toString() already goes through StringBytes::Encode(), which uses simdutf. MakeString() now calls the same function (the kMaxLengthERR_STRING_TOO_LONG check stays in front), so the decoder returns exactly what toString() returns for the same bytes. The partial-character bookkeeping is untouched.

src: use simdutf for two-byte strings in UTF-8 writes - StringBytes::Write(UTF8) used String::WriteUtf8V2() for two-byte strings. For strings longer than 32 code units it now validates with simdutf::validate_utf16() (lone surrogates are repaired into a scratch buffer with to_well_formed_utf16(), so the output stays byte-identical to V8's kReplaceInvalidUtf8) and transcodes with convert_utf16_to_utf8() - but only when the destination is known to fit (buflen >= 3 × length or >= utf8_length_from_utf16()); otherwise it falls back to V8, so partial writes truncate at the same character boundary as before. Shorter strings and one-byte strings keep the V8 path.

Tests:test-string-decoder-utf8-large.js (new: large inputs, chunk boundaries inside multi-byte sequences, invalid sequences vs toString()); test-buffer-write-utf8-two-byte.js (new: independent reference encoder; BMP/astral/lone surrogates at start/middle/end; exact-fit, one-short and 3× destinations; partial writes). Existing string_decoder, buffer, stream, net, http, child_process and readline suites pass.


Disclosure: the code, test, measurements and this description were written by Claude Code, directed and reviewed by @codebytere.

@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

Review requested:

  • @nodejs/performance

@nodejs-github-botnodejs-github-bot added buffer Issues and PRs related to the buffer subsystem. c++ Issues and PRs that require attention from people who are familiar with C++. needs-ci PRs that need a full CI run. labels Aug 16, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.
Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).
buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
@codebytere
codebytereforce-pushed the perf/src-simdutf-utf8-transcoding branch from ec59ddb to 0c89308CompareAugust 16, 2026 15:51
@codebyterecodebytere added request-ci Add this label to start a Jenkins CI on a PR. and removed needs-ci PRs that need a full CI run. labels Aug 16, 2026
@codecov

codecovBot commented Aug 16, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 90.11%. Comparing base (30bff4a) to head (0c89308).
⚠️ Report is 38 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #65324 +/- ##
==========================================
- Coverage 90.13% 90.11% -0.03% 
==========================================
Files 752 752 Lines 251568 251578 +10 Branches 47270 47273 +3 ==========================================
- Hits 226759 226706 -53 - Misses 16168 16216 +48 - Partials 8641 8656 +15 
Files with missing linesCoverage Δ
src/string_bytes.cc75.23% <100.00%> (+1.49%)⬆️
src/string_decoder.cc92.17% <100.00%> (-0.34%)⬇️

... and 37 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@github-actionsgithub-actionsBot removed the request-ci Add this label to start a Jenkins CI on a PR. label Aug 16, 2026
@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

Comment threadsrc/string_bytes.cc
// 3 * length is what StorageSize() hands most callers; only compute
// the exact length when the buffer is smaller than that.
if (buflen >= 3 * length ||
buflen >= simdutf::utf8_length_from_utf16(data, length)) {

@lemirelemireAug 16, 2026

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We have simdutf::utf8_length_from_utf16_with_replacement that would be appropriate in this function.

Note that we added convert_utf16_to_utf8_with_replacement in arelease this year. (But this code is still quite fine.)

@lemire

Copy link
Copy Markdown
Member

@codebytere I might send you a PR later today. Hold on a bit.

@lemire

Copy link
Copy Markdown
Member

@codebytere The simdutf version might not be ready yet. So I still recommend merging this PR. It is good work. We might be able to improve it later, but this should not stop this PR.

@github-actions

github-actionsBot commented Aug 17, 2026

Copy link
Copy Markdown
Contributor

@codebyterecodebytere added the commit-queue PRs queued for automated landing through the Commit Queue. label Aug 17, 2026
@codebytere

Copy link
Copy Markdown
MemberAuthor

@jasnell is commit queue broken?

@jasnell

Copy link
Copy Markdown
Member

Entirely possible

@nodejs-github-botnodejs-github-bot added commit-queue-failed PRs whose Commit Queue landing failed and need manual intervention before retrying. and removed commit-queue PRs queued for automated landing through the Commit Queue. labels Aug 18, 2026
@nodejs-github-bot

Copy link
Copy Markdown
Collaborator
Commit Queue failed
- Loading data for nodejs/node/pull/65324
✔ Done loading data for nodejs/node/pull/65324
----------------------------------- PR info ------------------------------------
Title src: use simdutf for UTF-8 ⇄ UTF-16 transcoding in StringDecoder and StringBytes::Write (#65324)
Author Shelley Vohr <shelley.vohr@gmail.com> (@codebytere)
Branch codebytere:perf/src-simdutf-utf8-transcoding -> nodejs:main
Labels buffer, c++, commit-queue
Commits 2
- string_decoder: decode UTF-8 via StringBytes::Encode
- src: use simdutf for two-byte strings in UTF-8 writes
Committers 1
- Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: https://github.com/nodejs/node/pull/65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
------------------------------ Generated metadata ------------------------------
PR-URL: https://github.com/nodejs/node/pull/65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
--------------------------------------------------------------------------------
ℹ This PR was created on Sun, 16 Aug 2026 15:37:44 GMT
✔ Approvals: 4
✔ - Yagiz Nizipli (@anonrig) (TSC): https://github.com/nodejs/node/pull/65324#pullrequestreview-4947047263
✔ - Gürgün Dayıoğlu (@gurgunday): https://github.com/nodejs/node/pull/65324#pullrequestreview-4947063127
✔ - Daniel Lemire (@lemire): https://github.com/nodejs/node/pull/65324#pullrequestreview-4947151736
✔ - James M Snell (@jasnell) (TSC): https://github.com/nodejs/node/pull/65324#pullrequestreview-4948248917
✔ Last GitHub CI successful
ℹ Last Full PR CI on 2026-08-16T18:58:59Z: https://ci.nodejs.org/job/node-test-pull-request/75896/
- Querying data for job/node-test-pull-request/75896/
✔ Build data downloaded
✔ Last Jenkins CI successful
--------------------------------------------------------------------------------
✔ No git cherry-pick in progress
✔ No git am in progress
✔ No git rebase in progress
--------------------------------------------------------------------------------
- Bringing origin/main up to date...
From https://github.com/nodejs/node
* branch main -> FETCH_HEAD
✔ origin/main is now up-to-date
- Downloading patch for 65324
From https://github.com/nodejs/node
* branch refs/pull/65324/merge -> FETCH_HEAD
✔ Fetched commits as 0d84f29fa77a..0c89308bba0c
--------------------------------------------------------------------------------
[main 308a46181c] string_decoder: decode UTF-8 via StringBytes::Encode
Author: Shelley Vohr <shelley.vohr@gmail.com>
Date: Sat Aug 15 22:09:21 2026 +0000
2 files changed, 94 insertions(+), 15 deletions(-)
create mode 100644 test/parallel/test-string-decoder-utf8-large.js
[main 619cb0dd4c] src: use simdutf for two-byte strings in UTF-8 writes
Author: Shelley Vohr <shelley.vohr@gmail.com>
Date: Sat Aug 15 22:34:28 2026 +0000
3 files changed, 193 insertions(+), 1 deletion(-)
create mode 100644 benchmark/buffers/buffer-write-string-utf8.js
create mode 100644 test/parallel/test-buffer-write-utf8-two-byte.js
✔ Patches applied
There are 2 commits in the PR. Attempting autorebase.
(node:381) [DEP0190] DeprecationWarning: Passing args to a child process with shell option true can lead to security vulnerabilities, as the arguments are not escaped, only concatenated.
(Use `node --trace-deprecation ...` to show where the warning was created)
Rebasing (2/4)
Executing: git node land --amend --yes
--------------------------------- New Message ----------------------------------
string_decoder: decode UTF-8 via StringBytes::Encode

StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.

benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.

Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>

[detached HEAD 1055900345] string_decoder: decode UTF-8 via StringBytes::Encode
Author: Shelley Vohr <shelley.vohr@gmail.com>
Date: Sat Aug 15 22:09:21 2026 +0000
2 files changed, 94 insertions(+), 15 deletions(-)
create mode 100644 test/parallel/test-string-decoder-utf8-large.js
Rebasing (3/4)
Rebasing (4/4)
Executing: git node land --amend --yes
--------------------------------- New Message ----------------------------------
src: use simdutf for two-byte strings in UTF-8 writes

StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.

Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).

buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.

Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>

[detached HEAD 55c27c9b17] src: use simdutf for two-byte strings in UTF-8 writes
Author: Shelley Vohr <shelley.vohr@gmail.com>
Date: Sat Aug 15 22:34:28 2026 +0000
3 files changed, 193 insertions(+), 1 deletion(-)
create mode 100644 benchmark/buffers/buffer-write-string-utf8.js
create mode 100644 test/parallel/test-buffer-write-utf8-two-byte.js
Successfully rebased and updated refs/heads/main.

ℹ Add commit-queue-squash label to land the PR as one commit, or commit-queue-rebase to land as separate commits.

https://github.com/nodejs/node/actions/runs/32156233387

@aduh95aduh95 added commit-queue PRs queued for automated landing through the Commit Queue. commit-queue-rebase PRs the Commit Queue should land as multiple self-contained commits. and removed commit-queue-failed PRs whose Commit Queue landing failed and need manual intervention before retrying. labels Aug 18, 2026
@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

Landed in 29a083c...55e4ca3

nodejs-github-bot pushed a commit that referenced this pull request Aug 18, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
nodejs-github-bot pushed a commit that referenced this pull request Aug 18, 2026
StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.
Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).
buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
@nodejs-github-botnodejs-github-bot removed the commit-queue PRs queued for automated landing through the Commit Queue. label Aug 18, 2026
@codebytere
codebytere deleted the perf/src-simdutf-utf8-transcoding branch August 18, 2026 19:22
aduh95 pushed a commit that referenced this pull request Aug 25, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
aduh95 pushed a commit that referenced this pull request Aug 25, 2026
StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.
Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).
buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
aduh95 pushed a commit that referenced this pull request Aug 25, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
aduh95 pushed a commit that referenced this pull request Aug 25, 2026
StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.
Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).
buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
@aduh95aduh95 added dont-land-on-v22.x PRs that should not land on the v22.x-staging branch and should not be released in v22.x. dont-land-on-v24.x PRs that should not land on the v24.x-staging branch and should not be released in v24.x. labels Aug 27, 2026
aduh95 pushed a commit that referenced this pull request Aug 27, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bufferIssues and PRs related to the buffer subsystem.c++Issues and PRs that require attention from people who are familiar with C++.commit-queue-rebasePRs the Commit Queue should land as multiple self-contained commits.dont-land-on-v22.xPRs that should not land on the v22.x-staging branch and should not be released in v22.x.dont-land-on-v24.xPRs that should not land on the v24.x-staging branch and should not be released in v24.x.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants

@codebytere@nodejs-github-bot@lemire@jasnell@anonrig@mertcanaltin@gurgunday@aduh95
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

src: use simdutf for UTF-8 ⇄ UTF-16 transcoding in StringDecoder and StringBytes::Write - #65324

Closed
codebytere wants to merge 2 commits into
nodejs:mainfrom
codebytere:perf/src-simdutf-utf8-transcoding
Closed

src: use simdutf for UTF-8 ⇄ UTF-16 transcoding in StringDecoder and StringBytes::Write#65324
codebytere wants to merge 2 commits into
nodejs:mainfrom
codebytere:perf/src-simdutf-utf8-transcoding

Conversation

@codebytere

@codebyterecodebytere commented Aug 16, 2026

Copy link
Copy Markdown
Member

StringDecoder gets 3–12× faster on UTF-8, and writing non-ASCII strings as UTF-8 (Buffer#write, Buffer.from(string), every stream/socket/fs string write) gets 4–5× faster. A 64 KiB JSON-RPC round trip over a child's stdio (stringify → write → readline → parse, non-ASCII payload) drops from 1.28 ms to 0.90 ms with only the parent patched, 0.67 ms with both ends.

string_decoder/string-decoder.js encoding='utf8' inLen=1024 chunkLen=1024 *** 1125.82 % ±8.25%
string_decoder/string-decoder.js encoding='utf8' inLen=128 chunkLen=1024 *** 322.79 % ±2.28%
string_decoder/string-decoder.js encoding='utf8' inLen=32 chunkLen=1024 *** 98.53 % ±1.37%
string_decoder/string-decoder.js encoding='utf8' inLen=1024 chunkLen=16 *** 5.15 % ±0.58%
buffers/buffer-write-string-utf8.js (new) len=65536 chars='two-byte' *** 391.17 %
buffers/buffer-write-string-utf8.js len=65536 chars='two-byte-lone-surrogate' *** 314.76 %
buffers/buffer-write-string-utf8.js len=256 chars='two-byte' *** 289.28 %
buffers/buffer-write-string-utf8.js len=256 chars='two-byte-astral' *** 228.97 %
one-byte strings / other encodings ~0 % n.s.

(Linux x64, benchmark/compare.js, 30 runs, significance as in compare.R.)

Two commits:

string_decoder: decode UTF-8 via StringBytes::Encode - the decoder built strings with String::NewFromUtf8(); buffer.toString() already goes through StringBytes::Encode(), which uses simdutf. MakeString() now calls the same function (the kMaxLengthERR_STRING_TOO_LONG check stays in front), so the decoder returns exactly what toString() returns for the same bytes. The partial-character bookkeeping is untouched.

src: use simdutf for two-byte strings in UTF-8 writes - StringBytes::Write(UTF8) used String::WriteUtf8V2() for two-byte strings. For strings longer than 32 code units it now validates with simdutf::validate_utf16() (lone surrogates are repaired into a scratch buffer with to_well_formed_utf16(), so the output stays byte-identical to V8's kReplaceInvalidUtf8) and transcodes with convert_utf16_to_utf8() - but only when the destination is known to fit (buflen >= 3 × length or >= utf8_length_from_utf16()); otherwise it falls back to V8, so partial writes truncate at the same character boundary as before. Shorter strings and one-byte strings keep the V8 path.

Tests:test-string-decoder-utf8-large.js (new: large inputs, chunk boundaries inside multi-byte sequences, invalid sequences vs toString()); test-buffer-write-utf8-two-byte.js (new: independent reference encoder; BMP/astral/lone surrogates at start/middle/end; exact-fit, one-short and 3× destinations; partial writes). Existing string_decoder, buffer, stream, net, http, child_process and readline suites pass.


Disclosure: the code, test, measurements and this description were written by Claude Code, directed and reviewed by @codebytere.

@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

Review requested:

  • @nodejs/performance

@nodejs-github-botnodejs-github-bot added buffer Issues and PRs related to the buffer subsystem. c++ Issues and PRs that require attention from people who are familiar with C++. needs-ci PRs that need a full CI run. labels Aug 16, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.
Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).
buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
@codebytere
codebytereforce-pushed the perf/src-simdutf-utf8-transcoding branch from ec59ddb to 0c89308CompareAugust 16, 2026 15:51
@codebyterecodebytere added request-ci Add this label to start a Jenkins CI on a PR. and removed needs-ci PRs that need a full CI run. labels Aug 16, 2026
@codecov

codecovBot commented Aug 16, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 90.11%. Comparing base (30bff4a) to head (0c89308).
⚠️ Report is 38 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #65324 +/- ##
==========================================
- Coverage 90.13% 90.11% -0.03% 
==========================================
Files 752 752 Lines 251568 251578 +10 Branches 47270 47273 +3 ==========================================
- Hits 226759 226706 -53 - Misses 16168 16216 +48 - Partials 8641 8656 +15 
Files with missing linesCoverage Δ
src/string_bytes.cc75.23% <100.00%> (+1.49%)⬆️
src/string_decoder.cc92.17% <100.00%> (-0.34%)⬇️

... and 37 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@github-actionsgithub-actionsBot removed the request-ci Add this label to start a Jenkins CI on a PR. label Aug 16, 2026
@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

Comment threadsrc/string_bytes.cc
// 3 * length is what StorageSize() hands most callers; only compute
// the exact length when the buffer is smaller than that.
if (buflen >= 3 * length ||
buflen >= simdutf::utf8_length_from_utf16(data, length)) {

@lemirelemireAug 16, 2026

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We have simdutf::utf8_length_from_utf16_with_replacement that would be appropriate in this function.

Note that we added convert_utf16_to_utf8_with_replacement in arelease this year. (But this code is still quite fine.)

@lemire

Copy link
Copy Markdown
Member

@codebytere I might send you a PR later today. Hold on a bit.

@lemire

Copy link
Copy Markdown
Member

@codebytere The simdutf version might not be ready yet. So I still recommend merging this PR. It is good work. We might be able to improve it later, but this should not stop this PR.

@github-actions

github-actionsBot commented Aug 17, 2026

Copy link
Copy Markdown
Contributor

@codebyterecodebytere added the commit-queue PRs queued for automated landing through the Commit Queue. label Aug 17, 2026
@codebytere

Copy link
Copy Markdown
MemberAuthor

@jasnell is commit queue broken?

@jasnell

Copy link
Copy Markdown
Member

Entirely possible

@nodejs-github-botnodejs-github-bot added commit-queue-failed PRs whose Commit Queue landing failed and need manual intervention before retrying. and removed commit-queue PRs queued for automated landing through the Commit Queue. labels Aug 18, 2026
@nodejs-github-bot

Copy link
Copy Markdown
Collaborator
Commit Queue failed
- Loading data for nodejs/node/pull/65324
✔ Done loading data for nodejs/node/pull/65324
----------------------------------- PR info ------------------------------------
Title src: use simdutf for UTF-8 ⇄ UTF-16 transcoding in StringDecoder and StringBytes::Write (#65324)
Author Shelley Vohr <shelley.vohr@gmail.com> (@codebytere)
Branch codebytere:perf/src-simdutf-utf8-transcoding -> nodejs:main
Labels buffer, c++, commit-queue
Commits 2
- string_decoder: decode UTF-8 via StringBytes::Encode
- src: use simdutf for two-byte strings in UTF-8 writes
Committers 1
- Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: https://github.com/nodejs/node/pull/65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
------------------------------ Generated metadata ------------------------------
PR-URL: https://github.com/nodejs/node/pull/65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
--------------------------------------------------------------------------------
ℹ This PR was created on Sun, 16 Aug 2026 15:37:44 GMT
✔ Approvals: 4
✔ - Yagiz Nizipli (@anonrig) (TSC): https://github.com/nodejs/node/pull/65324#pullrequestreview-4947047263
✔ - Gürgün Dayıoğlu (@gurgunday): https://github.com/nodejs/node/pull/65324#pullrequestreview-4947063127
✔ - Daniel Lemire (@lemire): https://github.com/nodejs/node/pull/65324#pullrequestreview-4947151736
✔ - James M Snell (@jasnell) (TSC): https://github.com/nodejs/node/pull/65324#pullrequestreview-4948248917
✔ Last GitHub CI successful
ℹ Last Full PR CI on 2026-08-16T18:58:59Z: https://ci.nodejs.org/job/node-test-pull-request/75896/
- Querying data for job/node-test-pull-request/75896/
✔ Build data downloaded
✔ Last Jenkins CI successful
--------------------------------------------------------------------------------
✔ No git cherry-pick in progress
✔ No git am in progress
✔ No git rebase in progress
--------------------------------------------------------------------------------
- Bringing origin/main up to date...
From https://github.com/nodejs/node
* branch main -> FETCH_HEAD
✔ origin/main is now up-to-date
- Downloading patch for 65324
From https://github.com/nodejs/node
* branch refs/pull/65324/merge -> FETCH_HEAD
✔ Fetched commits as 0d84f29fa77a..0c89308bba0c
--------------------------------------------------------------------------------
[main 308a46181c] string_decoder: decode UTF-8 via StringBytes::Encode
Author: Shelley Vohr <shelley.vohr@gmail.com>
Date: Sat Aug 15 22:09:21 2026 +0000
2 files changed, 94 insertions(+), 15 deletions(-)
create mode 100644 test/parallel/test-string-decoder-utf8-large.js
[main 619cb0dd4c] src: use simdutf for two-byte strings in UTF-8 writes
Author: Shelley Vohr <shelley.vohr@gmail.com>
Date: Sat Aug 15 22:34:28 2026 +0000
3 files changed, 193 insertions(+), 1 deletion(-)
create mode 100644 benchmark/buffers/buffer-write-string-utf8.js
create mode 100644 test/parallel/test-buffer-write-utf8-two-byte.js
✔ Patches applied
There are 2 commits in the PR. Attempting autorebase.
(node:381) [DEP0190] DeprecationWarning: Passing args to a child process with shell option true can lead to security vulnerabilities, as the arguments are not escaped, only concatenated.
(Use `node --trace-deprecation ...` to show where the warning was created)
Rebasing (2/4)
Executing: git node land --amend --yes
--------------------------------- New Message ----------------------------------
string_decoder: decode UTF-8 via StringBytes::Encode

StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.

benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.

Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>

[detached HEAD 1055900345] string_decoder: decode UTF-8 via StringBytes::Encode
Author: Shelley Vohr <shelley.vohr@gmail.com>
Date: Sat Aug 15 22:09:21 2026 +0000
2 files changed, 94 insertions(+), 15 deletions(-)
create mode 100644 test/parallel/test-string-decoder-utf8-large.js
Rebasing (3/4)
Rebasing (4/4)
Executing: git node land --amend --yes
--------------------------------- New Message ----------------------------------
src: use simdutf for two-byte strings in UTF-8 writes

StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.

Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).

buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.

Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>

[detached HEAD 55c27c9b17] src: use simdutf for two-byte strings in UTF-8 writes
Author: Shelley Vohr <shelley.vohr@gmail.com>
Date: Sat Aug 15 22:34:28 2026 +0000
3 files changed, 193 insertions(+), 1 deletion(-)
create mode 100644 benchmark/buffers/buffer-write-string-utf8.js
create mode 100644 test/parallel/test-buffer-write-utf8-two-byte.js
Successfully rebased and updated refs/heads/main.

ℹ Add commit-queue-squash label to land the PR as one commit, or commit-queue-rebase to land as separate commits.

https://github.com/nodejs/node/actions/runs/32156233387

@aduh95aduh95 added commit-queue PRs queued for automated landing through the Commit Queue. commit-queue-rebase PRs the Commit Queue should land as multiple self-contained commits. and removed commit-queue-failed PRs whose Commit Queue landing failed and need manual intervention before retrying. labels Aug 18, 2026
@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

Landed in 29a083c...55e4ca3

nodejs-github-bot pushed a commit that referenced this pull request Aug 18, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
nodejs-github-bot pushed a commit that referenced this pull request Aug 18, 2026
StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.
Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).
buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
@nodejs-github-botnodejs-github-bot removed the commit-queue PRs queued for automated landing through the Commit Queue. label Aug 18, 2026
@codebytere
codebytere deleted the perf/src-simdutf-utf8-transcoding branch August 18, 2026 19:22
aduh95 pushed a commit that referenced this pull request Aug 25, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
aduh95 pushed a commit that referenced this pull request Aug 25, 2026
StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.
Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).
buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
aduh95 pushed a commit that referenced this pull request Aug 25, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
aduh95 pushed a commit that referenced this pull request Aug 25, 2026
StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.
Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).
buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
@aduh95aduh95 added dont-land-on-v22.x PRs that should not land on the v22.x-staging branch and should not be released in v22.x. dont-land-on-v24.x PRs that should not land on the v24.x-staging branch and should not be released in v24.x. labels Aug 27, 2026
aduh95 pushed a commit that referenced this pull request Aug 27, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bufferIssues and PRs related to the buffer subsystem.c++Issues and PRs that require attention from people who are familiar with C++.commit-queue-rebasePRs the Commit Queue should land as multiple self-contained commits.dont-land-on-v22.xPRs that should not land on the v22.x-staging branch and should not be released in v22.x.dont-land-on-v24.xPRs that should not land on the v24.x-staging branch and should not be released in v24.x.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants

@codebytere@nodejs-github-bot@lemire@jasnell@anonrig@mertcanaltin@gurgunday@aduh95
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

src: use simdutf for UTF-8 ⇄ UTF-16 transcoding in StringDecoder and StringBytes::Write - #65324

Closed
codebytere wants to merge 2 commits into
nodejs:mainfrom
codebytere:perf/src-simdutf-utf8-transcoding
Closed

src: use simdutf for UTF-8 ⇄ UTF-16 transcoding in StringDecoder and StringBytes::Write#65324
codebytere wants to merge 2 commits into
nodejs:mainfrom
codebytere:perf/src-simdutf-utf8-transcoding

Conversation

@codebytere

@codebyterecodebytere commented Aug 16, 2026

Copy link
Copy Markdown
Member

StringDecoder gets 3–12× faster on UTF-8, and writing non-ASCII strings as UTF-8 (Buffer#write, Buffer.from(string), every stream/socket/fs string write) gets 4–5× faster. A 64 KiB JSON-RPC round trip over a child's stdio (stringify → write → readline → parse, non-ASCII payload) drops from 1.28 ms to 0.90 ms with only the parent patched, 0.67 ms with both ends.

string_decoder/string-decoder.js encoding='utf8' inLen=1024 chunkLen=1024 *** 1125.82 % ±8.25%
string_decoder/string-decoder.js encoding='utf8' inLen=128 chunkLen=1024 *** 322.79 % ±2.28%
string_decoder/string-decoder.js encoding='utf8' inLen=32 chunkLen=1024 *** 98.53 % ±1.37%
string_decoder/string-decoder.js encoding='utf8' inLen=1024 chunkLen=16 *** 5.15 % ±0.58%
buffers/buffer-write-string-utf8.js (new) len=65536 chars='two-byte' *** 391.17 %
buffers/buffer-write-string-utf8.js len=65536 chars='two-byte-lone-surrogate' *** 314.76 %
buffers/buffer-write-string-utf8.js len=256 chars='two-byte' *** 289.28 %
buffers/buffer-write-string-utf8.js len=256 chars='two-byte-astral' *** 228.97 %
one-byte strings / other encodings ~0 % n.s.

(Linux x64, benchmark/compare.js, 30 runs, significance as in compare.R.)

Two commits:

string_decoder: decode UTF-8 via StringBytes::Encode - the decoder built strings with String::NewFromUtf8(); buffer.toString() already goes through StringBytes::Encode(), which uses simdutf. MakeString() now calls the same function (the kMaxLengthERR_STRING_TOO_LONG check stays in front), so the decoder returns exactly what toString() returns for the same bytes. The partial-character bookkeeping is untouched.

src: use simdutf for two-byte strings in UTF-8 writes - StringBytes::Write(UTF8) used String::WriteUtf8V2() for two-byte strings. For strings longer than 32 code units it now validates with simdutf::validate_utf16() (lone surrogates are repaired into a scratch buffer with to_well_formed_utf16(), so the output stays byte-identical to V8's kReplaceInvalidUtf8) and transcodes with convert_utf16_to_utf8() - but only when the destination is known to fit (buflen >= 3 × length or >= utf8_length_from_utf16()); otherwise it falls back to V8, so partial writes truncate at the same character boundary as before. Shorter strings and one-byte strings keep the V8 path.

Tests:test-string-decoder-utf8-large.js (new: large inputs, chunk boundaries inside multi-byte sequences, invalid sequences vs toString()); test-buffer-write-utf8-two-byte.js (new: independent reference encoder; BMP/astral/lone surrogates at start/middle/end; exact-fit, one-short and 3× destinations; partial writes). Existing string_decoder, buffer, stream, net, http, child_process and readline suites pass.


Disclosure: the code, test, measurements and this description were written by Claude Code, directed and reviewed by @codebytere.

@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

Review requested:

  • @nodejs/performance

@nodejs-github-botnodejs-github-bot added buffer Issues and PRs related to the buffer subsystem. c++ Issues and PRs that require attention from people who are familiar with C++. needs-ci PRs that need a full CI run. labels Aug 16, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.
Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).
buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
@codebytere
codebytereforce-pushed the perf/src-simdutf-utf8-transcoding branch from ec59ddb to 0c89308CompareAugust 16, 2026 15:51
@codebyterecodebytere added request-ci Add this label to start a Jenkins CI on a PR. and removed needs-ci PRs that need a full CI run. labels Aug 16, 2026
@codecov

codecovBot commented Aug 16, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 90.11%. Comparing base (30bff4a) to head (0c89308).
⚠️ Report is 38 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #65324 +/- ##
==========================================
- Coverage 90.13% 90.11% -0.03% 
==========================================
Files 752 752 Lines 251568 251578 +10 Branches 47270 47273 +3 ==========================================
- Hits 226759 226706 -53 - Misses 16168 16216 +48 - Partials 8641 8656 +15 
Files with missing linesCoverage Δ
src/string_bytes.cc75.23% <100.00%> (+1.49%)⬆️
src/string_decoder.cc92.17% <100.00%> (-0.34%)⬇️

... and 37 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@github-actionsgithub-actionsBot removed the request-ci Add this label to start a Jenkins CI on a PR. label Aug 16, 2026
@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

Comment threadsrc/string_bytes.cc
// 3 * length is what StorageSize() hands most callers; only compute
// the exact length when the buffer is smaller than that.
if (buflen >= 3 * length ||
buflen >= simdutf::utf8_length_from_utf16(data, length)) {

@lemirelemireAug 16, 2026

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We have simdutf::utf8_length_from_utf16_with_replacement that would be appropriate in this function.

Note that we added convert_utf16_to_utf8_with_replacement in arelease this year. (But this code is still quite fine.)

@lemire

Copy link
Copy Markdown
Member

@codebytere I might send you a PR later today. Hold on a bit.

@lemire

Copy link
Copy Markdown
Member

@codebytere The simdutf version might not be ready yet. So I still recommend merging this PR. It is good work. We might be able to improve it later, but this should not stop this PR.

@github-actions

github-actionsBot commented Aug 17, 2026

Copy link
Copy Markdown
Contributor

@codebyterecodebytere added the commit-queue PRs queued for automated landing through the Commit Queue. label Aug 17, 2026
@codebytere

Copy link
Copy Markdown
MemberAuthor

@jasnell is commit queue broken?

@jasnell

Copy link
Copy Markdown
Member

Entirely possible

@nodejs-github-botnodejs-github-bot added commit-queue-failed PRs whose Commit Queue landing failed and need manual intervention before retrying. and removed commit-queue PRs queued for automated landing through the Commit Queue. labels Aug 18, 2026
@nodejs-github-bot

Copy link
Copy Markdown
Collaborator
Commit Queue failed
- Loading data for nodejs/node/pull/65324
✔ Done loading data for nodejs/node/pull/65324
----------------------------------- PR info ------------------------------------
Title src: use simdutf for UTF-8 ⇄ UTF-16 transcoding in StringDecoder and StringBytes::Write (#65324)
Author Shelley Vohr <shelley.vohr@gmail.com> (@codebytere)
Branch codebytere:perf/src-simdutf-utf8-transcoding -> nodejs:main
Labels buffer, c++, commit-queue
Commits 2
- string_decoder: decode UTF-8 via StringBytes::Encode
- src: use simdutf for two-byte strings in UTF-8 writes
Committers 1
- Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: https://github.com/nodejs/node/pull/65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
------------------------------ Generated metadata ------------------------------
PR-URL: https://github.com/nodejs/node/pull/65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
--------------------------------------------------------------------------------
ℹ This PR was created on Sun, 16 Aug 2026 15:37:44 GMT
✔ Approvals: 4
✔ - Yagiz Nizipli (@anonrig) (TSC): https://github.com/nodejs/node/pull/65324#pullrequestreview-4947047263
✔ - Gürgün Dayıoğlu (@gurgunday): https://github.com/nodejs/node/pull/65324#pullrequestreview-4947063127
✔ - Daniel Lemire (@lemire): https://github.com/nodejs/node/pull/65324#pullrequestreview-4947151736
✔ - James M Snell (@jasnell) (TSC): https://github.com/nodejs/node/pull/65324#pullrequestreview-4948248917
✔ Last GitHub CI successful
ℹ Last Full PR CI on 2026-08-16T18:58:59Z: https://ci.nodejs.org/job/node-test-pull-request/75896/
- Querying data for job/node-test-pull-request/75896/
✔ Build data downloaded
✔ Last Jenkins CI successful
--------------------------------------------------------------------------------
✔ No git cherry-pick in progress
✔ No git am in progress
✔ No git rebase in progress
--------------------------------------------------------------------------------
- Bringing origin/main up to date...
From https://github.com/nodejs/node
* branch main -> FETCH_HEAD
✔ origin/main is now up-to-date
- Downloading patch for 65324
From https://github.com/nodejs/node
* branch refs/pull/65324/merge -> FETCH_HEAD
✔ Fetched commits as 0d84f29fa77a..0c89308bba0c
--------------------------------------------------------------------------------
[main 308a46181c] string_decoder: decode UTF-8 via StringBytes::Encode
Author: Shelley Vohr <shelley.vohr@gmail.com>
Date: Sat Aug 15 22:09:21 2026 +0000
2 files changed, 94 insertions(+), 15 deletions(-)
create mode 100644 test/parallel/test-string-decoder-utf8-large.js
[main 619cb0dd4c] src: use simdutf for two-byte strings in UTF-8 writes
Author: Shelley Vohr <shelley.vohr@gmail.com>
Date: Sat Aug 15 22:34:28 2026 +0000
3 files changed, 193 insertions(+), 1 deletion(-)
create mode 100644 benchmark/buffers/buffer-write-string-utf8.js
create mode 100644 test/parallel/test-buffer-write-utf8-two-byte.js
✔ Patches applied
There are 2 commits in the PR. Attempting autorebase.
(node:381) [DEP0190] DeprecationWarning: Passing args to a child process with shell option true can lead to security vulnerabilities, as the arguments are not escaped, only concatenated.
(Use `node --trace-deprecation ...` to show where the warning was created)
Rebasing (2/4)
Executing: git node land --amend --yes
--------------------------------- New Message ----------------------------------
string_decoder: decode UTF-8 via StringBytes::Encode

StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.

benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.

Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>

[detached HEAD 1055900345] string_decoder: decode UTF-8 via StringBytes::Encode
Author: Shelley Vohr <shelley.vohr@gmail.com>
Date: Sat Aug 15 22:09:21 2026 +0000
2 files changed, 94 insertions(+), 15 deletions(-)
create mode 100644 test/parallel/test-string-decoder-utf8-large.js
Rebasing (3/4)
Rebasing (4/4)
Executing: git node land --amend --yes
--------------------------------- New Message ----------------------------------
src: use simdutf for two-byte strings in UTF-8 writes

StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.

Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).

buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.

Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>

[detached HEAD 55c27c9b17] src: use simdutf for two-byte strings in UTF-8 writes
Author: Shelley Vohr <shelley.vohr@gmail.com>
Date: Sat Aug 15 22:34:28 2026 +0000
3 files changed, 193 insertions(+), 1 deletion(-)
create mode 100644 benchmark/buffers/buffer-write-string-utf8.js
create mode 100644 test/parallel/test-buffer-write-utf8-two-byte.js
Successfully rebased and updated refs/heads/main.

ℹ Add commit-queue-squash label to land the PR as one commit, or commit-queue-rebase to land as separate commits.

https://github.com/nodejs/node/actions/runs/32156233387

@aduh95aduh95 added commit-queue PRs queued for automated landing through the Commit Queue. commit-queue-rebase PRs the Commit Queue should land as multiple self-contained commits. and removed commit-queue-failed PRs whose Commit Queue landing failed and need manual intervention before retrying. labels Aug 18, 2026
@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

Landed in 29a083c...55e4ca3

nodejs-github-bot pushed a commit that referenced this pull request Aug 18, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
nodejs-github-bot pushed a commit that referenced this pull request Aug 18, 2026
StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.
Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).
buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
@nodejs-github-botnodejs-github-bot removed the commit-queue PRs queued for automated landing through the Commit Queue. label Aug 18, 2026
@codebytere
codebytere deleted the perf/src-simdutf-utf8-transcoding branch August 18, 2026 19:22
aduh95 pushed a commit that referenced this pull request Aug 25, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
aduh95 pushed a commit that referenced this pull request Aug 25, 2026
StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.
Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).
buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
aduh95 pushed a commit that referenced this pull request Aug 25, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
aduh95 pushed a commit that referenced this pull request Aug 25, 2026
StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.
Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).
buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
@aduh95aduh95 added dont-land-on-v22.x PRs that should not land on the v22.x-staging branch and should not be released in v22.x. dont-land-on-v24.x PRs that should not land on the v24.x-staging branch and should not be released in v24.x. labels Aug 27, 2026
aduh95 pushed a commit that referenced this pull request Aug 27, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bufferIssues and PRs related to the buffer subsystem.c++Issues and PRs that require attention from people who are familiar with C++.commit-queue-rebasePRs the Commit Queue should land as multiple self-contained commits.dont-land-on-v22.xPRs that should not land on the v22.x-staging branch and should not be released in v22.x.dont-land-on-v24.xPRs that should not land on the v24.x-staging branch and should not be released in v24.x.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants

@codebytere@nodejs-github-bot@lemire@jasnell@anonrig@mertcanaltin@gurgunday@aduh95
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

src: use simdutf for UTF-8 ⇄ UTF-16 transcoding in StringDecoder and StringBytes::Write - #65324

Closed
codebytere wants to merge 2 commits into
nodejs:mainfrom
codebytere:perf/src-simdutf-utf8-transcoding
Closed

src: use simdutf for UTF-8 ⇄ UTF-16 transcoding in StringDecoder and StringBytes::Write#65324
codebytere wants to merge 2 commits into
nodejs:mainfrom
codebytere:perf/src-simdutf-utf8-transcoding

Conversation

@codebytere

@codebyterecodebytere commented Aug 16, 2026

Copy link
Copy Markdown
Member

StringDecoder gets 3–12× faster on UTF-8, and writing non-ASCII strings as UTF-8 (Buffer#write, Buffer.from(string), every stream/socket/fs string write) gets 4–5× faster. A 64 KiB JSON-RPC round trip over a child's stdio (stringify → write → readline → parse, non-ASCII payload) drops from 1.28 ms to 0.90 ms with only the parent patched, 0.67 ms with both ends.

string_decoder/string-decoder.js encoding='utf8' inLen=1024 chunkLen=1024 *** 1125.82 % ±8.25%
string_decoder/string-decoder.js encoding='utf8' inLen=128 chunkLen=1024 *** 322.79 % ±2.28%
string_decoder/string-decoder.js encoding='utf8' inLen=32 chunkLen=1024 *** 98.53 % ±1.37%
string_decoder/string-decoder.js encoding='utf8' inLen=1024 chunkLen=16 *** 5.15 % ±0.58%
buffers/buffer-write-string-utf8.js (new) len=65536 chars='two-byte' *** 391.17 %
buffers/buffer-write-string-utf8.js len=65536 chars='two-byte-lone-surrogate' *** 314.76 %
buffers/buffer-write-string-utf8.js len=256 chars='two-byte' *** 289.28 %
buffers/buffer-write-string-utf8.js len=256 chars='two-byte-astral' *** 228.97 %
one-byte strings / other encodings ~0 % n.s.

(Linux x64, benchmark/compare.js, 30 runs, significance as in compare.R.)

Two commits:

string_decoder: decode UTF-8 via StringBytes::Encode - the decoder built strings with String::NewFromUtf8(); buffer.toString() already goes through StringBytes::Encode(), which uses simdutf. MakeString() now calls the same function (the kMaxLengthERR_STRING_TOO_LONG check stays in front), so the decoder returns exactly what toString() returns for the same bytes. The partial-character bookkeeping is untouched.

src: use simdutf for two-byte strings in UTF-8 writes - StringBytes::Write(UTF8) used String::WriteUtf8V2() for two-byte strings. For strings longer than 32 code units it now validates with simdutf::validate_utf16() (lone surrogates are repaired into a scratch buffer with to_well_formed_utf16(), so the output stays byte-identical to V8's kReplaceInvalidUtf8) and transcodes with convert_utf16_to_utf8() - but only when the destination is known to fit (buflen >= 3 × length or >= utf8_length_from_utf16()); otherwise it falls back to V8, so partial writes truncate at the same character boundary as before. Shorter strings and one-byte strings keep the V8 path.

Tests:test-string-decoder-utf8-large.js (new: large inputs, chunk boundaries inside multi-byte sequences, invalid sequences vs toString()); test-buffer-write-utf8-two-byte.js (new: independent reference encoder; BMP/astral/lone surrogates at start/middle/end; exact-fit, one-short and 3× destinations; partial writes). Existing string_decoder, buffer, stream, net, http, child_process and readline suites pass.


Disclosure: the code, test, measurements and this description were written by Claude Code, directed and reviewed by @codebytere.

@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

Review requested:

  • @nodejs/performance

@nodejs-github-botnodejs-github-bot added buffer Issues and PRs related to the buffer subsystem. c++ Issues and PRs that require attention from people who are familiar with C++. needs-ci PRs that need a full CI run. labels Aug 16, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.
Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).
buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
@codebytere
codebytereforce-pushed the perf/src-simdutf-utf8-transcoding branch from ec59ddb to 0c89308CompareAugust 16, 2026 15:51
@codebyterecodebytere added request-ci Add this label to start a Jenkins CI on a PR. and removed needs-ci PRs that need a full CI run. labels Aug 16, 2026
@codecov

codecovBot commented Aug 16, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 90.11%. Comparing base (30bff4a) to head (0c89308).
⚠️ Report is 38 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #65324 +/- ##
==========================================
- Coverage 90.13% 90.11% -0.03% 
==========================================
Files 752 752 Lines 251568 251578 +10 Branches 47270 47273 +3 ==========================================
- Hits 226759 226706 -53 - Misses 16168 16216 +48 - Partials 8641 8656 +15 
Files with missing linesCoverage Δ
src/string_bytes.cc75.23% <100.00%> (+1.49%)⬆️
src/string_decoder.cc92.17% <100.00%> (-0.34%)⬇️

... and 37 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@github-actionsgithub-actionsBot removed the request-ci Add this label to start a Jenkins CI on a PR. label Aug 16, 2026
@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

Comment threadsrc/string_bytes.cc
// 3 * length is what StorageSize() hands most callers; only compute
// the exact length when the buffer is smaller than that.
if (buflen >= 3 * length ||
buflen >= simdutf::utf8_length_from_utf16(data, length)) {

@lemirelemireAug 16, 2026

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We have simdutf::utf8_length_from_utf16_with_replacement that would be appropriate in this function.

Note that we added convert_utf16_to_utf8_with_replacement in arelease this year. (But this code is still quite fine.)

@lemire

Copy link
Copy Markdown
Member

@codebytere I might send you a PR later today. Hold on a bit.

@lemire

Copy link
Copy Markdown
Member

@codebytere The simdutf version might not be ready yet. So I still recommend merging this PR. It is good work. We might be able to improve it later, but this should not stop this PR.

@github-actions

github-actionsBot commented Aug 17, 2026

Copy link
Copy Markdown
Contributor

@codebyterecodebytere added the commit-queue PRs queued for automated landing through the Commit Queue. label Aug 17, 2026
@codebytere

Copy link
Copy Markdown
MemberAuthor

@jasnell is commit queue broken?

@jasnell

Copy link
Copy Markdown
Member

Entirely possible

@nodejs-github-botnodejs-github-bot added commit-queue-failed PRs whose Commit Queue landing failed and need manual intervention before retrying. and removed commit-queue PRs queued for automated landing through the Commit Queue. labels Aug 18, 2026
@nodejs-github-bot

Copy link
Copy Markdown
Collaborator
Commit Queue failed
- Loading data for nodejs/node/pull/65324
✔ Done loading data for nodejs/node/pull/65324
----------------------------------- PR info ------------------------------------
Title src: use simdutf for UTF-8 ⇄ UTF-16 transcoding in StringDecoder and StringBytes::Write (#65324)
Author Shelley Vohr <shelley.vohr@gmail.com> (@codebytere)
Branch codebytere:perf/src-simdutf-utf8-transcoding -> nodejs:main
Labels buffer, c++, commit-queue
Commits 2
- string_decoder: decode UTF-8 via StringBytes::Encode
- src: use simdutf for two-byte strings in UTF-8 writes
Committers 1
- Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: https://github.com/nodejs/node/pull/65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
------------------------------ Generated metadata ------------------------------
PR-URL: https://github.com/nodejs/node/pull/65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
--------------------------------------------------------------------------------
ℹ This PR was created on Sun, 16 Aug 2026 15:37:44 GMT
✔ Approvals: 4
✔ - Yagiz Nizipli (@anonrig) (TSC): https://github.com/nodejs/node/pull/65324#pullrequestreview-4947047263
✔ - Gürgün Dayıoğlu (@gurgunday): https://github.com/nodejs/node/pull/65324#pullrequestreview-4947063127
✔ - Daniel Lemire (@lemire): https://github.com/nodejs/node/pull/65324#pullrequestreview-4947151736
✔ - James M Snell (@jasnell) (TSC): https://github.com/nodejs/node/pull/65324#pullrequestreview-4948248917
✔ Last GitHub CI successful
ℹ Last Full PR CI on 2026-08-16T18:58:59Z: https://ci.nodejs.org/job/node-test-pull-request/75896/
- Querying data for job/node-test-pull-request/75896/
✔ Build data downloaded
✔ Last Jenkins CI successful
--------------------------------------------------------------------------------
✔ No git cherry-pick in progress
✔ No git am in progress
✔ No git rebase in progress
--------------------------------------------------------------------------------
- Bringing origin/main up to date...
From https://github.com/nodejs/node
* branch main -> FETCH_HEAD
✔ origin/main is now up-to-date
- Downloading patch for 65324
From https://github.com/nodejs/node
* branch refs/pull/65324/merge -> FETCH_HEAD
✔ Fetched commits as 0d84f29fa77a..0c89308bba0c
--------------------------------------------------------------------------------
[main 308a46181c] string_decoder: decode UTF-8 via StringBytes::Encode
Author: Shelley Vohr <shelley.vohr@gmail.com>
Date: Sat Aug 15 22:09:21 2026 +0000
2 files changed, 94 insertions(+), 15 deletions(-)
create mode 100644 test/parallel/test-string-decoder-utf8-large.js
[main 619cb0dd4c] src: use simdutf for two-byte strings in UTF-8 writes
Author: Shelley Vohr <shelley.vohr@gmail.com>
Date: Sat Aug 15 22:34:28 2026 +0000
3 files changed, 193 insertions(+), 1 deletion(-)
create mode 100644 benchmark/buffers/buffer-write-string-utf8.js
create mode 100644 test/parallel/test-buffer-write-utf8-two-byte.js
✔ Patches applied
There are 2 commits in the PR. Attempting autorebase.
(node:381) [DEP0190] DeprecationWarning: Passing args to a child process with shell option true can lead to security vulnerabilities, as the arguments are not escaped, only concatenated.
(Use `node --trace-deprecation ...` to show where the warning was created)
Rebasing (2/4)
Executing: git node land --amend --yes
--------------------------------- New Message ----------------------------------
string_decoder: decode UTF-8 via StringBytes::Encode

StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.

benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.

Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>

[detached HEAD 1055900345] string_decoder: decode UTF-8 via StringBytes::Encode
Author: Shelley Vohr <shelley.vohr@gmail.com>
Date: Sat Aug 15 22:09:21 2026 +0000
2 files changed, 94 insertions(+), 15 deletions(-)
create mode 100644 test/parallel/test-string-decoder-utf8-large.js
Rebasing (3/4)
Rebasing (4/4)
Executing: git node land --amend --yes
--------------------------------- New Message ----------------------------------
src: use simdutf for two-byte strings in UTF-8 writes

StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.

Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).

buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.

Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>

[detached HEAD 55c27c9b17] src: use simdutf for two-byte strings in UTF-8 writes
Author: Shelley Vohr <shelley.vohr@gmail.com>
Date: Sat Aug 15 22:34:28 2026 +0000
3 files changed, 193 insertions(+), 1 deletion(-)
create mode 100644 benchmark/buffers/buffer-write-string-utf8.js
create mode 100644 test/parallel/test-buffer-write-utf8-two-byte.js
Successfully rebased and updated refs/heads/main.

ℹ Add commit-queue-squash label to land the PR as one commit, or commit-queue-rebase to land as separate commits.

https://github.com/nodejs/node/actions/runs/32156233387

@aduh95aduh95 added commit-queue PRs queued for automated landing through the Commit Queue. commit-queue-rebase PRs the Commit Queue should land as multiple self-contained commits. and removed commit-queue-failed PRs whose Commit Queue landing failed and need manual intervention before retrying. labels Aug 18, 2026
@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

Landed in 29a083c...55e4ca3

nodejs-github-bot pushed a commit that referenced this pull request Aug 18, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
nodejs-github-bot pushed a commit that referenced this pull request Aug 18, 2026
StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.
Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).
buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
@nodejs-github-botnodejs-github-bot removed the commit-queue PRs queued for automated landing through the Commit Queue. label Aug 18, 2026
@codebytere
codebytere deleted the perf/src-simdutf-utf8-transcoding branch August 18, 2026 19:22
aduh95 pushed a commit that referenced this pull request Aug 25, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
aduh95 pushed a commit that referenced this pull request Aug 25, 2026
StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.
Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).
buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
aduh95 pushed a commit that referenced this pull request Aug 25, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
aduh95 pushed a commit that referenced this pull request Aug 25, 2026
StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.
Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).
buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
@aduh95aduh95 added dont-land-on-v22.x PRs that should not land on the v22.x-staging branch and should not be released in v22.x. dont-land-on-v24.x PRs that should not land on the v24.x-staging branch and should not be released in v24.x. labels Aug 27, 2026
aduh95 pushed a commit that referenced this pull request Aug 27, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bufferIssues and PRs related to the buffer subsystem.c++Issues and PRs that require attention from people who are familiar with C++.commit-queue-rebasePRs the Commit Queue should land as multiple self-contained commits.dont-land-on-v22.xPRs that should not land on the v22.x-staging branch and should not be released in v22.x.dont-land-on-v24.xPRs that should not land on the v24.x-staging branch and should not be released in v24.x.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants

@codebytere@nodejs-github-bot@lemire@jasnell@anonrig@mertcanaltin@gurgunday@aduh95
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

src: use simdutf for UTF-8 ⇄ UTF-16 transcoding in StringDecoder and StringBytes::Write - #65324

Closed
codebytere wants to merge 2 commits into
nodejs:mainfrom
codebytere:perf/src-simdutf-utf8-transcoding
Closed

src: use simdutf for UTF-8 ⇄ UTF-16 transcoding in StringDecoder and StringBytes::Write#65324
codebytere wants to merge 2 commits into
nodejs:mainfrom
codebytere:perf/src-simdutf-utf8-transcoding

Conversation

@codebytere

@codebyterecodebytere commented Aug 16, 2026

Copy link
Copy Markdown
Member

StringDecoder gets 3–12× faster on UTF-8, and writing non-ASCII strings as UTF-8 (Buffer#write, Buffer.from(string), every stream/socket/fs string write) gets 4–5× faster. A 64 KiB JSON-RPC round trip over a child's stdio (stringify → write → readline → parse, non-ASCII payload) drops from 1.28 ms to 0.90 ms with only the parent patched, 0.67 ms with both ends.

string_decoder/string-decoder.js encoding='utf8' inLen=1024 chunkLen=1024 *** 1125.82 % ±8.25%
string_decoder/string-decoder.js encoding='utf8' inLen=128 chunkLen=1024 *** 322.79 % ±2.28%
string_decoder/string-decoder.js encoding='utf8' inLen=32 chunkLen=1024 *** 98.53 % ±1.37%
string_decoder/string-decoder.js encoding='utf8' inLen=1024 chunkLen=16 *** 5.15 % ±0.58%
buffers/buffer-write-string-utf8.js (new) len=65536 chars='two-byte' *** 391.17 %
buffers/buffer-write-string-utf8.js len=65536 chars='two-byte-lone-surrogate' *** 314.76 %
buffers/buffer-write-string-utf8.js len=256 chars='two-byte' *** 289.28 %
buffers/buffer-write-string-utf8.js len=256 chars='two-byte-astral' *** 228.97 %
one-byte strings / other encodings ~0 % n.s.

(Linux x64, benchmark/compare.js, 30 runs, significance as in compare.R.)

Two commits:

string_decoder: decode UTF-8 via StringBytes::Encode - the decoder built strings with String::NewFromUtf8(); buffer.toString() already goes through StringBytes::Encode(), which uses simdutf. MakeString() now calls the same function (the kMaxLengthERR_STRING_TOO_LONG check stays in front), so the decoder returns exactly what toString() returns for the same bytes. The partial-character bookkeeping is untouched.

src: use simdutf for two-byte strings in UTF-8 writes - StringBytes::Write(UTF8) used String::WriteUtf8V2() for two-byte strings. For strings longer than 32 code units it now validates with simdutf::validate_utf16() (lone surrogates are repaired into a scratch buffer with to_well_formed_utf16(), so the output stays byte-identical to V8's kReplaceInvalidUtf8) and transcodes with convert_utf16_to_utf8() - but only when the destination is known to fit (buflen >= 3 × length or >= utf8_length_from_utf16()); otherwise it falls back to V8, so partial writes truncate at the same character boundary as before. Shorter strings and one-byte strings keep the V8 path.

Tests:test-string-decoder-utf8-large.js (new: large inputs, chunk boundaries inside multi-byte sequences, invalid sequences vs toString()); test-buffer-write-utf8-two-byte.js (new: independent reference encoder; BMP/astral/lone surrogates at start/middle/end; exact-fit, one-short and 3× destinations; partial writes). Existing string_decoder, buffer, stream, net, http, child_process and readline suites pass.


Disclosure: the code, test, measurements and this description were written by Claude Code, directed and reviewed by @codebytere.

@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

Review requested:

  • @nodejs/performance

@nodejs-github-botnodejs-github-bot added buffer Issues and PRs related to the buffer subsystem. c++ Issues and PRs that require attention from people who are familiar with C++. needs-ci PRs that need a full CI run. labels Aug 16, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.
Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).
buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
@codebytere
codebytereforce-pushed the perf/src-simdutf-utf8-transcoding branch from ec59ddb to 0c89308CompareAugust 16, 2026 15:51
@codebyterecodebytere added request-ci Add this label to start a Jenkins CI on a PR. and removed needs-ci PRs that need a full CI run. labels Aug 16, 2026
@codecov

codecovBot commented Aug 16, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 90.11%. Comparing base (30bff4a) to head (0c89308).
⚠️ Report is 38 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #65324 +/- ##
==========================================
- Coverage 90.13% 90.11% -0.03% 
==========================================
Files 752 752 Lines 251568 251578 +10 Branches 47270 47273 +3 ==========================================
- Hits 226759 226706 -53 - Misses 16168 16216 +48 - Partials 8641 8656 +15 
Files with missing linesCoverage Δ
src/string_bytes.cc75.23% <100.00%> (+1.49%)⬆️
src/string_decoder.cc92.17% <100.00%> (-0.34%)⬇️

... and 37 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@github-actionsgithub-actionsBot removed the request-ci Add this label to start a Jenkins CI on a PR. label Aug 16, 2026
@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

Comment threadsrc/string_bytes.cc
// 3 * length is what StorageSize() hands most callers; only compute
// the exact length when the buffer is smaller than that.
if (buflen >= 3 * length ||
buflen >= simdutf::utf8_length_from_utf16(data, length)) {

@lemirelemireAug 16, 2026

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We have simdutf::utf8_length_from_utf16_with_replacement that would be appropriate in this function.

Note that we added convert_utf16_to_utf8_with_replacement in arelease this year. (But this code is still quite fine.)

@lemire

Copy link
Copy Markdown
Member

@codebytere I might send you a PR later today. Hold on a bit.

@lemire

Copy link
Copy Markdown
Member

@codebytere The simdutf version might not be ready yet. So I still recommend merging this PR. It is good work. We might be able to improve it later, but this should not stop this PR.

@github-actions

github-actionsBot commented Aug 17, 2026

Copy link
Copy Markdown
Contributor

@codebyterecodebytere added the commit-queue PRs queued for automated landing through the Commit Queue. label Aug 17, 2026
@codebytere

Copy link
Copy Markdown
MemberAuthor

@jasnell is commit queue broken?

@jasnell

Copy link
Copy Markdown
Member

Entirely possible

@nodejs-github-botnodejs-github-bot added commit-queue-failed PRs whose Commit Queue landing failed and need manual intervention before retrying. and removed commit-queue PRs queued for automated landing through the Commit Queue. labels Aug 18, 2026
@nodejs-github-bot

Copy link
Copy Markdown
Collaborator
Commit Queue failed
- Loading data for nodejs/node/pull/65324
✔ Done loading data for nodejs/node/pull/65324
----------------------------------- PR info ------------------------------------
Title src: use simdutf for UTF-8 ⇄ UTF-16 transcoding in StringDecoder and StringBytes::Write (#65324)
Author Shelley Vohr <shelley.vohr@gmail.com> (@codebytere)
Branch codebytere:perf/src-simdutf-utf8-transcoding -> nodejs:main
Labels buffer, c++, commit-queue
Commits 2
- string_decoder: decode UTF-8 via StringBytes::Encode
- src: use simdutf for two-byte strings in UTF-8 writes
Committers 1
- Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: https://github.com/nodejs/node/pull/65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
------------------------------ Generated metadata ------------------------------
PR-URL: https://github.com/nodejs/node/pull/65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
--------------------------------------------------------------------------------
ℹ This PR was created on Sun, 16 Aug 2026 15:37:44 GMT
✔ Approvals: 4
✔ - Yagiz Nizipli (@anonrig) (TSC): https://github.com/nodejs/node/pull/65324#pullrequestreview-4947047263
✔ - Gürgün Dayıoğlu (@gurgunday): https://github.com/nodejs/node/pull/65324#pullrequestreview-4947063127
✔ - Daniel Lemire (@lemire): https://github.com/nodejs/node/pull/65324#pullrequestreview-4947151736
✔ - James M Snell (@jasnell) (TSC): https://github.com/nodejs/node/pull/65324#pullrequestreview-4948248917
✔ Last GitHub CI successful
ℹ Last Full PR CI on 2026-08-16T18:58:59Z: https://ci.nodejs.org/job/node-test-pull-request/75896/
- Querying data for job/node-test-pull-request/75896/
✔ Build data downloaded
✔ Last Jenkins CI successful
--------------------------------------------------------------------------------
✔ No git cherry-pick in progress
✔ No git am in progress
✔ No git rebase in progress
--------------------------------------------------------------------------------
- Bringing origin/main up to date...
From https://github.com/nodejs/node
* branch main -> FETCH_HEAD
✔ origin/main is now up-to-date
- Downloading patch for 65324
From https://github.com/nodejs/node
* branch refs/pull/65324/merge -> FETCH_HEAD
✔ Fetched commits as 0d84f29fa77a..0c89308bba0c
--------------------------------------------------------------------------------
[main 308a46181c] string_decoder: decode UTF-8 via StringBytes::Encode
Author: Shelley Vohr <shelley.vohr@gmail.com>
Date: Sat Aug 15 22:09:21 2026 +0000
2 files changed, 94 insertions(+), 15 deletions(-)
create mode 100644 test/parallel/test-string-decoder-utf8-large.js
[main 619cb0dd4c] src: use simdutf for two-byte strings in UTF-8 writes
Author: Shelley Vohr <shelley.vohr@gmail.com>
Date: Sat Aug 15 22:34:28 2026 +0000
3 files changed, 193 insertions(+), 1 deletion(-)
create mode 100644 benchmark/buffers/buffer-write-string-utf8.js
create mode 100644 test/parallel/test-buffer-write-utf8-two-byte.js
✔ Patches applied
There are 2 commits in the PR. Attempting autorebase.
(node:381) [DEP0190] DeprecationWarning: Passing args to a child process with shell option true can lead to security vulnerabilities, as the arguments are not escaped, only concatenated.
(Use `node --trace-deprecation ...` to show where the warning was created)
Rebasing (2/4)
Executing: git node land --amend --yes
--------------------------------- New Message ----------------------------------
string_decoder: decode UTF-8 via StringBytes::Encode

StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.

benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.

Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>

[detached HEAD 1055900345] string_decoder: decode UTF-8 via StringBytes::Encode
Author: Shelley Vohr <shelley.vohr@gmail.com>
Date: Sat Aug 15 22:09:21 2026 +0000
2 files changed, 94 insertions(+), 15 deletions(-)
create mode 100644 test/parallel/test-string-decoder-utf8-large.js
Rebasing (3/4)
Rebasing (4/4)
Executing: git node land --amend --yes
--------------------------------- New Message ----------------------------------
src: use simdutf for two-byte strings in UTF-8 writes

StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.

Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).

buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.

Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>

[detached HEAD 55c27c9b17] src: use simdutf for two-byte strings in UTF-8 writes
Author: Shelley Vohr <shelley.vohr@gmail.com>
Date: Sat Aug 15 22:34:28 2026 +0000
3 files changed, 193 insertions(+), 1 deletion(-)
create mode 100644 benchmark/buffers/buffer-write-string-utf8.js
create mode 100644 test/parallel/test-buffer-write-utf8-two-byte.js
Successfully rebased and updated refs/heads/main.

ℹ Add commit-queue-squash label to land the PR as one commit, or commit-queue-rebase to land as separate commits.

https://github.com/nodejs/node/actions/runs/32156233387

@aduh95aduh95 added commit-queue PRs queued for automated landing through the Commit Queue. commit-queue-rebase PRs the Commit Queue should land as multiple self-contained commits. and removed commit-queue-failed PRs whose Commit Queue landing failed and need manual intervention before retrying. labels Aug 18, 2026
@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

Landed in 29a083c...55e4ca3

nodejs-github-bot pushed a commit that referenced this pull request Aug 18, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
nodejs-github-bot pushed a commit that referenced this pull request Aug 18, 2026
StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.
Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).
buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
@nodejs-github-botnodejs-github-bot removed the commit-queue PRs queued for automated landing through the Commit Queue. label Aug 18, 2026
@codebytere
codebytere deleted the perf/src-simdutf-utf8-transcoding branch August 18, 2026 19:22
aduh95 pushed a commit that referenced this pull request Aug 25, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
aduh95 pushed a commit that referenced this pull request Aug 25, 2026
StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.
Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).
buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
aduh95 pushed a commit that referenced this pull request Aug 25, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
aduh95 pushed a commit that referenced this pull request Aug 25, 2026
StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.
Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).
buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
@aduh95aduh95 added dont-land-on-v22.x PRs that should not land on the v22.x-staging branch and should not be released in v22.x. dont-land-on-v24.x PRs that should not land on the v24.x-staging branch and should not be released in v24.x. labels Aug 27, 2026
aduh95 pushed a commit that referenced this pull request Aug 27, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bufferIssues and PRs related to the buffer subsystem.c++Issues and PRs that require attention from people who are familiar with C++.commit-queue-rebasePRs the Commit Queue should land as multiple self-contained commits.dont-land-on-v22.xPRs that should not land on the v22.x-staging branch and should not be released in v22.x.dont-land-on-v24.xPRs that should not land on the v24.x-staging branch and should not be released in v24.x.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants

@codebytere@nodejs-github-bot@lemire@jasnell@anonrig@mertcanaltin@gurgunday@aduh95
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

src: use simdutf for UTF-8 ⇄ UTF-16 transcoding in StringDecoder and StringBytes::Write - #65324

Closed
codebytere wants to merge 2 commits into
nodejs:mainfrom
codebytere:perf/src-simdutf-utf8-transcoding
Closed

src: use simdutf for UTF-8 ⇄ UTF-16 transcoding in StringDecoder and StringBytes::Write#65324
codebytere wants to merge 2 commits into
nodejs:mainfrom
codebytere:perf/src-simdutf-utf8-transcoding

Conversation

@codebytere

@codebyterecodebytere commented Aug 16, 2026

Copy link
Copy Markdown
Member

StringDecoder gets 3–12× faster on UTF-8, and writing non-ASCII strings as UTF-8 (Buffer#write, Buffer.from(string), every stream/socket/fs string write) gets 4–5× faster. A 64 KiB JSON-RPC round trip over a child's stdio (stringify → write → readline → parse, non-ASCII payload) drops from 1.28 ms to 0.90 ms with only the parent patched, 0.67 ms with both ends.

string_decoder/string-decoder.js encoding='utf8' inLen=1024 chunkLen=1024 *** 1125.82 % ±8.25%
string_decoder/string-decoder.js encoding='utf8' inLen=128 chunkLen=1024 *** 322.79 % ±2.28%
string_decoder/string-decoder.js encoding='utf8' inLen=32 chunkLen=1024 *** 98.53 % ±1.37%
string_decoder/string-decoder.js encoding='utf8' inLen=1024 chunkLen=16 *** 5.15 % ±0.58%
buffers/buffer-write-string-utf8.js (new) len=65536 chars='two-byte' *** 391.17 %
buffers/buffer-write-string-utf8.js len=65536 chars='two-byte-lone-surrogate' *** 314.76 %
buffers/buffer-write-string-utf8.js len=256 chars='two-byte' *** 289.28 %
buffers/buffer-write-string-utf8.js len=256 chars='two-byte-astral' *** 228.97 %
one-byte strings / other encodings ~0 % n.s.

(Linux x64, benchmark/compare.js, 30 runs, significance as in compare.R.)

Two commits:

string_decoder: decode UTF-8 via StringBytes::Encode - the decoder built strings with String::NewFromUtf8(); buffer.toString() already goes through StringBytes::Encode(), which uses simdutf. MakeString() now calls the same function (the kMaxLengthERR_STRING_TOO_LONG check stays in front), so the decoder returns exactly what toString() returns for the same bytes. The partial-character bookkeeping is untouched.

src: use simdutf for two-byte strings in UTF-8 writes - StringBytes::Write(UTF8) used String::WriteUtf8V2() for two-byte strings. For strings longer than 32 code units it now validates with simdutf::validate_utf16() (lone surrogates are repaired into a scratch buffer with to_well_formed_utf16(), so the output stays byte-identical to V8's kReplaceInvalidUtf8) and transcodes with convert_utf16_to_utf8() - but only when the destination is known to fit (buflen >= 3 × length or >= utf8_length_from_utf16()); otherwise it falls back to V8, so partial writes truncate at the same character boundary as before. Shorter strings and one-byte strings keep the V8 path.

Tests:test-string-decoder-utf8-large.js (new: large inputs, chunk boundaries inside multi-byte sequences, invalid sequences vs toString()); test-buffer-write-utf8-two-byte.js (new: independent reference encoder; BMP/astral/lone surrogates at start/middle/end; exact-fit, one-short and 3× destinations; partial writes). Existing string_decoder, buffer, stream, net, http, child_process and readline suites pass.


Disclosure: the code, test, measurements and this description were written by Claude Code, directed and reviewed by @codebytere.

@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

Review requested:

  • @nodejs/performance

@nodejs-github-botnodejs-github-bot added buffer Issues and PRs related to the buffer subsystem. c++ Issues and PRs that require attention from people who are familiar with C++. needs-ci PRs that need a full CI run. labels Aug 16, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.
Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).
buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
@codebytere
codebytereforce-pushed the perf/src-simdutf-utf8-transcoding branch from ec59ddb to 0c89308CompareAugust 16, 2026 15:51
@codebyterecodebytere added request-ci Add this label to start a Jenkins CI on a PR. and removed needs-ci PRs that need a full CI run. labels Aug 16, 2026
@codecov

codecovBot commented Aug 16, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 90.11%. Comparing base (30bff4a) to head (0c89308).
⚠️ Report is 38 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #65324 +/- ##
==========================================
- Coverage 90.13% 90.11% -0.03% 
==========================================
Files 752 752 Lines 251568 251578 +10 Branches 47270 47273 +3 ==========================================
- Hits 226759 226706 -53 - Misses 16168 16216 +48 - Partials 8641 8656 +15 
Files with missing linesCoverage Δ
src/string_bytes.cc75.23% <100.00%> (+1.49%)⬆️
src/string_decoder.cc92.17% <100.00%> (-0.34%)⬇️

... and 37 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@github-actionsgithub-actionsBot removed the request-ci Add this label to start a Jenkins CI on a PR. label Aug 16, 2026
@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

Comment threadsrc/string_bytes.cc
// 3 * length is what StorageSize() hands most callers; only compute
// the exact length when the buffer is smaller than that.
if (buflen >= 3 * length ||
buflen >= simdutf::utf8_length_from_utf16(data, length)) {

@lemirelemireAug 16, 2026

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We have simdutf::utf8_length_from_utf16_with_replacement that would be appropriate in this function.

Note that we added convert_utf16_to_utf8_with_replacement in arelease this year. (But this code is still quite fine.)

@lemire

Copy link
Copy Markdown
Member

@codebytere I might send you a PR later today. Hold on a bit.

@lemire

Copy link
Copy Markdown
Member

@codebytere The simdutf version might not be ready yet. So I still recommend merging this PR. It is good work. We might be able to improve it later, but this should not stop this PR.

@github-actions

github-actionsBot commented Aug 17, 2026

Copy link
Copy Markdown
Contributor

@codebyterecodebytere added the commit-queue PRs queued for automated landing through the Commit Queue. label Aug 17, 2026
@codebytere

Copy link
Copy Markdown
MemberAuthor

@jasnell is commit queue broken?

@jasnell

Copy link
Copy Markdown
Member

Entirely possible

@nodejs-github-botnodejs-github-bot added commit-queue-failed PRs whose Commit Queue landing failed and need manual intervention before retrying. and removed commit-queue PRs queued for automated landing through the Commit Queue. labels Aug 18, 2026
@nodejs-github-bot

Copy link
Copy Markdown
Collaborator
Commit Queue failed
- Loading data for nodejs/node/pull/65324
✔ Done loading data for nodejs/node/pull/65324
----------------------------------- PR info ------------------------------------
Title src: use simdutf for UTF-8 ⇄ UTF-16 transcoding in StringDecoder and StringBytes::Write (#65324)
Author Shelley Vohr <shelley.vohr@gmail.com> (@codebytere)
Branch codebytere:perf/src-simdutf-utf8-transcoding -> nodejs:main
Labels buffer, c++, commit-queue
Commits 2
- string_decoder: decode UTF-8 via StringBytes::Encode
- src: use simdutf for two-byte strings in UTF-8 writes
Committers 1
- Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: https://github.com/nodejs/node/pull/65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
------------------------------ Generated metadata ------------------------------
PR-URL: https://github.com/nodejs/node/pull/65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
--------------------------------------------------------------------------------
ℹ This PR was created on Sun, 16 Aug 2026 15:37:44 GMT
✔ Approvals: 4
✔ - Yagiz Nizipli (@anonrig) (TSC): https://github.com/nodejs/node/pull/65324#pullrequestreview-4947047263
✔ - Gürgün Dayıoğlu (@gurgunday): https://github.com/nodejs/node/pull/65324#pullrequestreview-4947063127
✔ - Daniel Lemire (@lemire): https://github.com/nodejs/node/pull/65324#pullrequestreview-4947151736
✔ - James M Snell (@jasnell) (TSC): https://github.com/nodejs/node/pull/65324#pullrequestreview-4948248917
✔ Last GitHub CI successful
ℹ Last Full PR CI on 2026-08-16T18:58:59Z: https://ci.nodejs.org/job/node-test-pull-request/75896/
- Querying data for job/node-test-pull-request/75896/
✔ Build data downloaded
✔ Last Jenkins CI successful
--------------------------------------------------------------------------------
✔ No git cherry-pick in progress
✔ No git am in progress
✔ No git rebase in progress
--------------------------------------------------------------------------------
- Bringing origin/main up to date...
From https://github.com/nodejs/node
* branch main -> FETCH_HEAD
✔ origin/main is now up-to-date
- Downloading patch for 65324
From https://github.com/nodejs/node
* branch refs/pull/65324/merge -> FETCH_HEAD
✔ Fetched commits as 0d84f29fa77a..0c89308bba0c
--------------------------------------------------------------------------------
[main 308a46181c] string_decoder: decode UTF-8 via StringBytes::Encode
Author: Shelley Vohr <shelley.vohr@gmail.com>
Date: Sat Aug 15 22:09:21 2026 +0000
2 files changed, 94 insertions(+), 15 deletions(-)
create mode 100644 test/parallel/test-string-decoder-utf8-large.js
[main 619cb0dd4c] src: use simdutf for two-byte strings in UTF-8 writes
Author: Shelley Vohr <shelley.vohr@gmail.com>
Date: Sat Aug 15 22:34:28 2026 +0000
3 files changed, 193 insertions(+), 1 deletion(-)
create mode 100644 benchmark/buffers/buffer-write-string-utf8.js
create mode 100644 test/parallel/test-buffer-write-utf8-two-byte.js
✔ Patches applied
There are 2 commits in the PR. Attempting autorebase.
(node:381) [DEP0190] DeprecationWarning: Passing args to a child process with shell option true can lead to security vulnerabilities, as the arguments are not escaped, only concatenated.
(Use `node --trace-deprecation ...` to show where the warning was created)
Rebasing (2/4)
Executing: git node land --amend --yes
--------------------------------- New Message ----------------------------------
string_decoder: decode UTF-8 via StringBytes::Encode

StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.

benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.

Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>

[detached HEAD 1055900345] string_decoder: decode UTF-8 via StringBytes::Encode
Author: Shelley Vohr <shelley.vohr@gmail.com>
Date: Sat Aug 15 22:09:21 2026 +0000
2 files changed, 94 insertions(+), 15 deletions(-)
create mode 100644 test/parallel/test-string-decoder-utf8-large.js
Rebasing (3/4)
Rebasing (4/4)
Executing: git node land --amend --yes
--------------------------------- New Message ----------------------------------
src: use simdutf for two-byte strings in UTF-8 writes

StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.

Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).

buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.

Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>

[detached HEAD 55c27c9b17] src: use simdutf for two-byte strings in UTF-8 writes
Author: Shelley Vohr <shelley.vohr@gmail.com>
Date: Sat Aug 15 22:34:28 2026 +0000
3 files changed, 193 insertions(+), 1 deletion(-)
create mode 100644 benchmark/buffers/buffer-write-string-utf8.js
create mode 100644 test/parallel/test-buffer-write-utf8-two-byte.js
Successfully rebased and updated refs/heads/main.

ℹ Add commit-queue-squash label to land the PR as one commit, or commit-queue-rebase to land as separate commits.

https://github.com/nodejs/node/actions/runs/32156233387

@aduh95aduh95 added commit-queue PRs queued for automated landing through the Commit Queue. commit-queue-rebase PRs the Commit Queue should land as multiple self-contained commits. and removed commit-queue-failed PRs whose Commit Queue landing failed and need manual intervention before retrying. labels Aug 18, 2026
@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

Landed in 29a083c...55e4ca3

nodejs-github-bot pushed a commit that referenced this pull request Aug 18, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
nodejs-github-bot pushed a commit that referenced this pull request Aug 18, 2026
StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.
Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).
buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
@nodejs-github-botnodejs-github-bot removed the commit-queue PRs queued for automated landing through the Commit Queue. label Aug 18, 2026
@codebytere
codebytere deleted the perf/src-simdutf-utf8-transcoding branch August 18, 2026 19:22
aduh95 pushed a commit that referenced this pull request Aug 25, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
aduh95 pushed a commit that referenced this pull request Aug 25, 2026
StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.
Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).
buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
aduh95 pushed a commit that referenced this pull request Aug 25, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
aduh95 pushed a commit that referenced this pull request Aug 25, 2026
StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.
Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).
buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
@aduh95aduh95 added dont-land-on-v22.x PRs that should not land on the v22.x-staging branch and should not be released in v22.x. dont-land-on-v24.x PRs that should not land on the v24.x-staging branch and should not be released in v24.x. labels Aug 27, 2026
aduh95 pushed a commit that referenced this pull request Aug 27, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.
benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.
Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bufferIssues and PRs related to the buffer subsystem.c++Issues and PRs that require attention from people who are familiar with C++.commit-queue-rebasePRs the Commit Queue should land as multiple self-contained commits.dont-land-on-v22.xPRs that should not land on the v22.x-staging branch and should not be released in v22.x.dont-land-on-v24.xPRs that should not land on the v24.x-staging branch and should not be released in v24.x.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants

@codebytere@nodejs-github-bot@lemire@jasnell@anonrig@mertcanaltin@gurgunday@aduh95