Skip to content

src: use simdutf for UTF-8 ⇄ UTF-16 transcoding in StringDecoder and StringBytes::Write - #65324

Open
codebytere wants to merge 2 commits into
nodejs:mainfrom
codebytere:perf/src-simdutf-utf8-transcoding
Open

src: use simdutf for UTF-8 ⇄ UTF-16 transcoding in StringDecoder and StringBytes::Write#65324
codebytere wants to merge 2 commits into
nodejs:mainfrom
codebytere:perf/src-simdutf-utf8-transcoding

Conversation

@codebytere

@codebytere codebytere commented Aug 16, 2026

Copy link
Copy Markdown
Member

StringDecoder gets 3–12× faster on UTF-8, and writing non-ASCII strings as UTF-8 (Buffer#write, Buffer.from(string),
every stream/socket/fs string write) gets 4–5× faster. A 64 KiB JSON-RPC round trip over a child's stdio (stringify → write
→ readline → parse, non-ASCII payload) drops from 1.28 ms to 0.90 ms with only the parent patched, 0.67 ms with both ends.

string_decoder/string-decoder.js encoding='utf8' inLen=1024 chunkLen=1024      ***  1125.82 %  ±8.25%
string_decoder/string-decoder.js encoding='utf8' inLen=128  chunkLen=1024      ***   322.79 %  ±2.28%
string_decoder/string-decoder.js encoding='utf8' inLen=32   chunkLen=1024      ***    98.53 %  ±1.37%
string_decoder/string-decoder.js encoding='utf8' inLen=1024 chunkLen=16        ***     5.15 %  ±0.58%
buffers/buffer-write-string-utf8.js (new)  len=65536 chars='two-byte'                 ***   391.17 %
buffers/buffer-write-string-utf8.js        len=65536 chars='two-byte-lone-surrogate'  ***   314.76 %
buffers/buffer-write-string-utf8.js        len=256   chars='two-byte'                 ***   289.28 %
buffers/buffer-write-string-utf8.js        len=256   chars='two-byte-astral'          ***   228.97 %
one-byte strings / other encodings                                                             ~0 %  n.s.

(Linux x64, benchmark/compare.js, 30 runs, significance as in compare.R.)

Two commits:

string_decoder: decode UTF-8 via StringBytes::Encode - the decoder built strings with String::NewFromUtf8();
buffer.toString() already goes through StringBytes::Encode(), which uses simdutf. MakeString() now calls the same
function (the kMaxLengthERR_STRING_TOO_LONG check stays in front), so the decoder returns exactly what toString()
returns for the same bytes. The partial-character bookkeeping is untouched.

src: use simdutf for two-byte strings in UTF-8 writes - StringBytes::Write(UTF8) used String::WriteUtf8V2() for
two-byte strings. For strings longer than 32 code units it now validates with simdutf::validate_utf16() (lone surrogates
are repaired into a scratch buffer with to_well_formed_utf16(), so the output stays byte-identical to V8's
kReplaceInvalidUtf8) and transcodes with convert_utf16_to_utf8() - but only when the destination is known to fit
(buflen >= 3 × length or >= utf8_length_from_utf16()); otherwise it falls back to V8, so partial writes truncate at the
same character boundary as before. Shorter strings and one-byte strings keep the V8 path.

Tests: test-string-decoder-utf8-large.js (new: large inputs, chunk boundaries inside multi-byte sequences, invalid
sequences vs toString()); test-buffer-write-utf8-two-byte.js (new: independent reference encoder; BMP/astral/lone
surrogates at start/middle/end; exact-fit, one-short and 3× destinations; partial writes). Existing string_decoder, buffer,
stream, net, http, child_process and readline suites pass.


Disclosure: the code, test, measurements and this description were written by Claude Code, directed and reviewed by @codebytere.

@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

Review requested:

  • @nodejs/performance

@nodejs-github-bot nodejs-github-bot added buffer Issues and PRs related to the buffer subsystem. c++ Issues and PRs that require attention from people who are familiar with C++. needs-ci PRs that need a full CI run. labels Aug 16, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.

benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.

Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.

Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).

buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.

Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
@codebytere
codebytere force-pushed the perf/src-simdutf-utf8-transcoding branch from ec59ddb to 0c89308 Compare August 16, 2026 15:51
@codebytere codebytere added request-ci Add this label to start a Jenkins CI on a PR. and removed needs-ci PRs that need a full CI run. labels Aug 16, 2026
@codecov

codecov Bot commented Aug 16, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 90.11%. Comparing base (30bff4a) to head (0c89308).
⚠️ Report is 6 commits behind head on main.

Additional details and impacted files
@@            Coverage Diff             @@
##             main   #65324      +/-   ##
==========================================
- Coverage   90.13%   90.11%   -0.03%     
==========================================
  Files         752      752              
  Lines      251568   251578      +10     
  Branches    47270    47273       +3     
==========================================
- Hits       226759   226706      -53     
- Misses      16168    16216      +48     
- Partials     8641     8656      +15     
Files with missing lines Coverage Δ
src/string_bytes.cc 75.23% <100.00%> (+1.49%) ⬆️
src/string_decoder.cc 92.17% <100.00%> (-0.34%) ⬇️

... and 37 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@codebytere
codebytere requested review from addaleax and anonrig August 16, 2026 18:37
@github-actions github-actions Bot removed the request-ci Add this label to start a Jenkins CI on a PR. label Aug 16, 2026
@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

Comment thread src/string_bytes.cc
// 3 * length is what StorageSize() hands most callers; only compute
// the exact length when the buffer is smaller than that.
if (buflen >= 3 * length ||
buflen >= simdutf::utf8_length_from_utf16(data, length)) {

@lemire lemire Aug 16, 2026

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We have simdutf::utf8_length_from_utf16_with_replacement that would be appropriate in this function.

Note that we added convert_utf16_to_utf8_with_replacement in arelease this year. (But this code is still quite fine.)

@lemire

lemire commented Aug 16, 2026

Copy link
Copy Markdown
Member

@codebytere I might send you a PR later today. Hold on a bit.

@lemire

lemire commented Aug 16, 2026

Copy link
Copy Markdown
Member

@codebytere The simdutf version might not be ready yet. So I still recommend merging this PR. It is good work. We might be able to improve it later, but this should not stop this PR.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

buffer Issues and PRs related to the buffer subsystem. c++ Issues and PRs that require attention from people who are familiar with C++.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants