src: improve compile cache performance and size - #63861

Open
anonrig wants to merge 3 commits into
nodejs:mainfrom
anonrig:compile-cache-perf
Open

src: improve compile cache performance and size#63861
anonrig wants to merge 3 commits into
nodejs:mainfrom
anonrig:compile-cache-perf

Conversation

@anonrig

Copy link
Copy Markdown
Member

Improves the on-disk compile cache (NODE_COMPILE_CACHE / module.enableCompileCache()):

  • Read path: read cache files with a single exactly-sized read (using the file size from fstat) instead of an exponentially growing buffer, which previously cost O(log N) syscalls/allocations and ~2N bytes of copying per file.
  • Size: compress the cache content on disk with zstd (level 1, prioritizing speed since persistence happens at shutdown), falling back to raw storage when not compressible. Shrinks cache directories ~2-4x and makes the crc32 integrity check cheaper since it now runs over the compressed bytes. The magic number is bumped so files in the old format are discarded as cache misses and overwritten in place.
  • Consume path: hand the cache to V8 through a non-owning CachedData wrapper (BufferNotOwned) instead of copying the entire buffer on every cache hit. The underlying buffer is owned by the cache entry, which outlives the synchronous compilation (same pattern as the vm cached-data path in node_contextify.cc).

Corrupted cache files keep degrading to silent cache misses and are regenerated; a corrupted size header can no longer cause an oversized allocation since the zstd frame content size is cross-checked first. Added test/parallel/test-compile-cache-corrupted.js covering bad magic, truncation, content bit-flips, and header corruption.

No public API or documented behavior changes; the file format is private to src/compile_cache.cc. Benchmark numbers (cache size and warm-startup timings) to follow in a comment.

This change was developed with AI assistance (see Co-authored-by trailer).

@nodejs-github-botnodejs-github-bot added c++ Issues and PRs that require attention from people who are familiar with C++. lib / src Issues and PRs involving general changes in the lib/ or src/ directories. needs-ci PRs that need a full CI run. labels Jun 11, 2026
@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

Review requested:

  • @nodejs/loaders
  • @nodejs/vm

@anonrig

Copy link
Copy Markdown
MemberAuthor

Verification results (macOS arm64, release build, vs. baseline at the merge-base):

Tests: all 23 parallel/test-compile-cache* pass (22 existing unchanged + the new corruption test).

On-disk size

ScenarioBaselineThis PRReduction
10 MB snapshot/typescript.js fixture (single big CJS file)1,818,380 B752,924 B2.42x
300 small/medium CJS modules274,316 B179,915 B1.52x

Warm-start, self-controlled (same binary, cache on vs. off, interleaved 20 runs, trimmed means — avoids binary-layout skew between builds):

ScenarioCache benefit (baseline)Cache benefit (this PR)
big-file+26.4 ms+25.0 ms
many-modules+1.7 ms+1.5 ms

Warm-start with a hot page cache is neutral within noise (~1 ms on the pathological single-10 MB-file case, which is dominated by one large zstd decompress; with cold page cache or slower storage the 2.4x smaller read wins). Cold-start adds one-time compression at persist (~35 ms for the 1.8 MB blob at level 1, proportionally less for typical files).

The second commit reuses zstd contexts (one ZSTD_DCtx on the handler, one ZSTD_CCtx across Persist()), which removed most of the per-file decompression overhead observed with one-shot contexts on the many-modules scenario.

Improve the compile cache by:
- Reading cache files with a single exactly-sized read using the file
size from fstat instead of reading into an exponentially growing
buffer, which previously cost O(log N) syscalls and allocations and
about 2N bytes of copying per file.
- Compressing the cache content on disk with zstd at level 1, falling
back to raw storage when the data is not compressible. This shrinks
cache directories by about 2-4x. The magic number is bumped so that
files in the old format are discarded as cache misses and then
overwritten in place.
- Handing the cache to V8 through a non-owning CachedData wrapper
instead of copying the whole buffer on every cache hit.
Corrupted cache files keep degrading to silent cache misses and are
regenerated, now covered by a regression test.
Co-authored-by: Grok <grok@x.ai>
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
@anonrig
anonrigforce-pushed the compile-cache-perf branch from 816f4ef to 0cf8818CompareJune 11, 2026 22:38
Creating and freeing a zstd context for every cache file costs more
than the (de)compression itself for small caches. Lazily create one
decompression context on the handler and reuse it across reads, and
share one compression context across all entries in Persist().
Co-authored-by: Grok <grok@x.ai>
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
@anonrig
anonrigforce-pushed the compile-cache-perf branch from 0cf8818 to f427fb3CompareJune 11, 2026 22:40
@joyeecheung

joyeecheung commented Jun 12, 2026

Copy link
Copy Markdown
Member

It seems the performance gains are within noise and it's mostly only the compression changes size? Can you split them into different PRs and measure them individually?

  1. I am not sure if the read changes actually gives any wins for small files - in the happy path where the cache is small, it's better to just read once and resize rather than stat and read (which are two file system calls instead of just one). Also there is a TOCTOU risk in doing stat, and fstat is not realiable across platforms, so the loop condition should not be gated on the fstat result but we must always read until EOF is reached in case the size is not accurate.
  2. It doesn't appear that compression alone does much to performance or it might actually hurt but just got compensated by other changes. In that case it's better to make that configurable and let users choose instead.
  3. Avoiding the copy would be better and there were precedents in builtin caches that it actually helped. I suspect this was the only one that actually helps performance while the other two may not. Hence it's better to split and measure individually.

lemire added a commit to lemire/node that referenced this pull request Jun 12, 2026
Builds on the zstd compression in nodejs#63861 by embedding a small zstd
dictionary trained on a diverse corpus of real modules, so each
small/medium compile-cache entry compresses better. Per entry we keep
the smaller of the plain and dictionary-assisted frame, so the
dictionary only ever helps.
- Add src/compile_cache_zstd.dict (16 KiB). It is trained on V8 code
caches harvested (via vm.compileFunction, the same shape the CJS
loader produces) from a diverse corpus: bundled npm packages, lib/,
tools/ and a few deps.
- Add tools/generate_compile_cache_dict.py and a node.gyp action that
generates compile_cache_zstd_dict.h into SHARED_INTERMEDIATE_DIR at
build time; no generated header is checked in. libnode include_dirs
updated to pick it up.
- Prepare the CDict/DDict once per process (shared across all handlers
and Workers, matching the lazy-context approach from nodejs#63861) and use
them in Persist() and ReadCacheFile(). Persist() compresses the plain
and dict frames into separate buffers and selects the smaller, so the
written bytes and recorded size always agree. The dictionary is only
tried for entries up to 256 KiB; larger blobs never benefit, so the
second compression is skipped to avoid wasted work. Falls back to
plain zstd if dictionary preparation fails.
- The dictionary is embedded in the binary because the compile cache
must be usable early, portably, and without extra filesystem state.
- No on-disk format change: dict-assisted frames carry the dictID, plain
frames carry none, and a single DDict decompresses both.
- Size, measured on data held out from training (per-entry min policy):
diverse modules go from ~1.87x (plain zstd) to ~2.44x with the
dictionary (~24% smaller on disk); on test/parallel, which is not in
the training corpus at all, ~1.74x -> ~2.22x (~22% smaller). A real
end-to-end run (npm --version, ~70 modules) is ~15% smaller. Read
time is unchanged and the extra write-time work is negligible.
- Add a multi-module write/read roundtrip test and a startup benchmark
(standard createBenchmark harness).
Builds on the zstd compression in nodejs#63861 by embedding a small zstd
dictionary trained on a diverse corpus of real modules, so each
small/medium compile-cache entry compresses better. Per entry we keep
the smaller of the plain and dictionary-assisted frame, so the
dictionary only ever helps.
- Add src/compile_cache_zstd.dict (16 KiB). It is trained on V8 code
caches harvested (via vm.compileFunction, the same shape the CJS
loader produces) from a diverse corpus: bundled npm packages, lib/,
tools/ and a few deps.
- Add tools/generate_compile_cache_dict.py and a node.gyp action that
generates compile_cache_zstd_dict.h into SHARED_INTERMEDIATE_DIR at
build time; no generated header is checked in. libnode include_dirs
updated to pick it up.
- Prepare the CDict/DDict once per process (shared across all handlers
and Workers, matching the lazy-context approach from nodejs#63861) and use
them in Persist() and ReadCacheFile(). Persist() compresses the plain
and dict frames into separate buffers and selects the smaller, so the
written bytes and recorded size always agree. The dictionary is only
tried for entries up to 256 KiB; larger blobs never benefit, so the
second compression is skipped to avoid wasted work. Falls back to
plain zstd if dictionary preparation fails.
- The dictionary is embedded in the binary because the compile cache
must be usable early, portably, and without extra filesystem state.
- No on-disk format change: dict-assisted frames carry the dictID, plain
frames carry none, and a single DDict decompresses both.
- Size, measured on data held out from training (per-entry min policy):
diverse modules go from ~1.87x (plain zstd) to ~2.44x with the
dictionary (~24% smaller on disk); on test/parallel, which is not in
the training corpus at all, ~1.74x -> ~2.22x (~22% smaller). A real
end-to-end run (npm --version, ~70 modules) is ~15% smaller. Read
time is unchanged and the extra write-time work is negligible.
- Add a multi-module write/read roundtrip test and a startup benchmark
(standard createBenchmark harness).

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If we do include dictionary files in our source tree, I'd say we should either include instructions for how to (at least approximately) re-create this file, or even generate it at build time entirely. strings src/compile_cache_zstd.dict results in a surprising amount of Node.js-core specificity, which seems surprising given that the commit that adds it claims it was trained on public npm packages.

(strings src/compile_cache_zstd.dict output in the fold)
wUUu2
_HOFJW
[Z)S
nding
options
Paz
types
PaN
process
versions
openssl
internal/process/pre_execution
PdR`
prepareMainThreadExecution
Pc6Wl
markBootstrapComplete
internal/options
"<";
__createBinding
__setModuleDefault
PaZY4J
birname
note
signatures
ObjectSetPrototypeOf
PromisePrototypeThen
PromiseWithResolvers
RegExpPrototypeExec
RegExpPrototypeSymbolReplace
StringPrototypeSplit
StringPrototypeToLowerCase
_createBlob
_createBlobFromFilePath
getDataObject
kMaxLength
TextDecoder
markTransferMode
Pbvyo
isAnyArrayBuffer
PcBS
isArrayBufferView
PaZ
require
PaN"
util
InvalidArgumentError
ConnectTimeoutError
SessionCache
maybeNormalizeConnectError
PromiseWithResolvers
nonOpStart
Pf&
writableStreamMarkFirstWriteRequestInFlight
createPromiseCallback
Pb:6k
nonOpWrite
$Pg>
writableStreamDefaultWriterEnsureReadyPromiseRejected
writableStreamDefaultControllerClose
DOMException
Pc^v
extractHighWaterMark
extractSizeAlgorithm
Pb~:
kEmptyObject
wgIgAygCACIAIAQgAWtqIQUgASAAa0ECaiEGAkADQCABLQAAIABButUAai0AAEcNAyAAQQJGDQEgAEEBaiEAIAQgAUEBaiIBRw0ACyADIAU2AgAMwwILIAMoAgQhACADQgA3AwAgAyAAIAZBAWoiARAuIgBFBEBB4wEhAgyqAgsgA0H1ATYCHCADIAE2AhQgAyAANgIMQQAhAgzCAgtB9AEhAiABIARGDcECIAMoAgAiACAEIAFraiEFIAEgAGtBAWohBgJAA0AgAS0AACAAQbjVAGotAABHDQIgAEEBRg0BIABBAWohACAEIAFBAWoiAUcNAAsgAyAFNgIADMICCyADQYEEOwEoIAMoAgQhACADQgA3AwAgAyAAIAZBAWoiARAuIgANAwwCCyADQQA2AgALQQAhAiADQQA2AhwgAyABNgIUIANB5R82AhAgA0EINgIMDL8CC0HVASECDKUCCyADQfMBNgIcIAMgATYCFCADIAA2AgxBACECDL0CC0EAIQACQCADKAI4IgJFDQAgAigCQ
memoMethod
perf
calculatedSizcludes
kTypes
uvErrmapGet
ArrayPrototypeIndexOf
classRegExp
MathMax
ArrayIsArray
isWindows
ObjectIsExtensible
isPermissionModelError
getExpectedArgumentLength
ArrayPrototypeJoin
kIsNodeError
get source
get ports
CloseEvent
get listening
header.js.map
@[@b
node:path
./large-numbers.js
"<";
escapeHTML
PcB_C+
convertChangesToXML
#noProxy
#ProxyAgent
#getProxy
#timeoutConnection
Pc~}
#drainPendingRequests
connect
Pbv2&_
addRequest
createSocket
.get
.get
getMilestoneTimestamp
getTimeOriginTimestamp
PbjK
internalBinding
performance
constants
ERR_INVALID_ARG_TYPE
ERR_INVALID_ARG_VALUE
Pb:e
ERR_OUT_OF_RANGE
isErrorStackTraceLimitWritable
Buffer
Pa&
inspect
validateBoolean
validateFunction
validateNumber
validateString
validateOneOf
PbbH
validateObject
validateInteger
isReadableStream
isWritableStream
isNodeStream
utilColors
PbJ%
lazyUtilColors
PbNn6
__importStar
DQCABLQAAQSBHDUggBCABQQFqIgFHDQALQSUhAwxpC0ElIQMMaAsgAi0ALUEBcQRAQcMBIQMMTwsgAigCBCEAQQAhAyACQQA2AgQgAiAAIAEQKSIABEAgAkEmNgIcIAIgADYCDCACIAFBAWo2AhQMaAsgAUEBaiEBDFwLIAFBAWohASACLwEwIgBBgAFxBEBBACEAAkAgAigCOCIDRQ0AIAMoAlQiA0UNACACIAMRAAAhAAsgAEUNBiAAQRVHDR8gAkEFNgIcIAIgATYCFCACQfkXNgIQIAJBFTYCDEEAIQMMZwsCQCAAQaAEcUGgBEcNACACLQAtQQJxDQBBACEDIAJBADYCHCACIAE2AhQgAkGWEzYCECACQQQ2AgwMZwsgAgJ/IAIvATBBFHFBFEYEQEEBIAItAChBAUYNARogAi8BMkHlAEYMAQsgAi0AKUEFRgs6AC5BACEAAkAgAigCOCIDRQ0AIAMoAiQiA0UNACACIAMRAAAhAAsCQAJAAkACQAJAIAAOFgIBAAQEBAQEBAQE
get ok
get statusText
get headers
get body
get bodyUsed
clone
parsePullArgs
PcNs
validateBackpressure
primordials
PcJ=
internal/encoding
TextEncoder
internal/errors
Pav
codes
internal/util
internal/util/types
internal/validators
onComplete
E`w PaN
onError
node:assert
../core/symbols
../core/errors
../core/util
Pab
beep
ArrayFrom
ArrayPrototypeFilter
ArrayPrototypeIncludes
ArrayPrototypeMap
ArrayPrototypePush
PcB,t
ArrayPrototypePushApply
ArrayPrototypeSlice
ObjectDefineProperty
Pb>c
ObjectKeys
PdRc
ObjectPrototypeHasOwnProperty
Pbf)
ReflectGet
SafeMap
SafeSet
StringPrototypeSlice
Error
PbFOT
_flushFlag
__esModule
.desc.get
stop
destroyer
primordials
internal/errors
Pav
codes
internal/streams/utils
__esModule
fs/promises
PaF-M
path
PaZ
require
Pa"/
module
__filename
__dirname

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@anonrig I'm going to mark this as un-resolved again, since there was no response here as far as I can tell

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, I didn't see this up until now and didn't pressed the "resolve" button at all. Will address your concerns.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

didn't pressed the "resolve" button at all

Github begged to differ:

Screenshot 2026-06-14 005643

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AI "babysit" mode might be the root cause

Comment threadsrc/compile_cache.h
compiler_cache_store_;
// Lazily created zstd decompression context, reused across cache reads
// to avoid recreating its workspace for every file.
ZSTD_DCtx_s* zstd_dctx_ = nullptr;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can this be held in a unique_ptr instead of a raw pointer?

Comment threadsrc/compile_cache.cc
if (cctx != nullptr) {
ZSTD_freeCCtx(cctx);
}
});

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

unique_ptr for this as well

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

c++Issues and PRs that require attention from people who are familiar with C++.lib / srcIssues and PRs involving general changes in the lib/ or src/ directories.needs-ciPRs that need a full CI run.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@anonrig@nodejs-github-bot@joyeecheung@lemire@jasnell@addaleax
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

src: improve compile cache performance and size - #63861

Open
anonrig wants to merge 3 commits into
nodejs:mainfrom
anonrig:compile-cache-perf
Open

src: improve compile cache performance and size#63861
anonrig wants to merge 3 commits into
nodejs:mainfrom
anonrig:compile-cache-perf

Conversation

@anonrig

Copy link
Copy Markdown
Member

Improves the on-disk compile cache (NODE_COMPILE_CACHE / module.enableCompileCache()):

  • Read path: read cache files with a single exactly-sized read (using the file size from fstat) instead of an exponentially growing buffer, which previously cost O(log N) syscalls/allocations and ~2N bytes of copying per file.
  • Size: compress the cache content on disk with zstd (level 1, prioritizing speed since persistence happens at shutdown), falling back to raw storage when not compressible. Shrinks cache directories ~2-4x and makes the crc32 integrity check cheaper since it now runs over the compressed bytes. The magic number is bumped so files in the old format are discarded as cache misses and overwritten in place.
  • Consume path: hand the cache to V8 through a non-owning CachedData wrapper (BufferNotOwned) instead of copying the entire buffer on every cache hit. The underlying buffer is owned by the cache entry, which outlives the synchronous compilation (same pattern as the vm cached-data path in node_contextify.cc).

Corrupted cache files keep degrading to silent cache misses and are regenerated; a corrupted size header can no longer cause an oversized allocation since the zstd frame content size is cross-checked first. Added test/parallel/test-compile-cache-corrupted.js covering bad magic, truncation, content bit-flips, and header corruption.

No public API or documented behavior changes; the file format is private to src/compile_cache.cc. Benchmark numbers (cache size and warm-startup timings) to follow in a comment.

This change was developed with AI assistance (see Co-authored-by trailer).

@nodejs-github-botnodejs-github-bot added c++ Issues and PRs that require attention from people who are familiar with C++. lib / src Issues and PRs involving general changes in the lib/ or src/ directories. needs-ci PRs that need a full CI run. labels Jun 11, 2026
@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

Review requested:

  • @nodejs/loaders
  • @nodejs/vm

@anonrig

Copy link
Copy Markdown
MemberAuthor

Verification results (macOS arm64, release build, vs. baseline at the merge-base):

Tests: all 23 parallel/test-compile-cache* pass (22 existing unchanged + the new corruption test).

On-disk size

ScenarioBaselineThis PRReduction
10 MB snapshot/typescript.js fixture (single big CJS file)1,818,380 B752,924 B2.42x
300 small/medium CJS modules274,316 B179,915 B1.52x

Warm-start, self-controlled (same binary, cache on vs. off, interleaved 20 runs, trimmed means — avoids binary-layout skew between builds):

ScenarioCache benefit (baseline)Cache benefit (this PR)
big-file+26.4 ms+25.0 ms
many-modules+1.7 ms+1.5 ms

Warm-start with a hot page cache is neutral within noise (~1 ms on the pathological single-10 MB-file case, which is dominated by one large zstd decompress; with cold page cache or slower storage the 2.4x smaller read wins). Cold-start adds one-time compression at persist (~35 ms for the 1.8 MB blob at level 1, proportionally less for typical files).

The second commit reuses zstd contexts (one ZSTD_DCtx on the handler, one ZSTD_CCtx across Persist()), which removed most of the per-file decompression overhead observed with one-shot contexts on the many-modules scenario.

Improve the compile cache by:
- Reading cache files with a single exactly-sized read using the file
size from fstat instead of reading into an exponentially growing
buffer, which previously cost O(log N) syscalls and allocations and
about 2N bytes of copying per file.
- Compressing the cache content on disk with zstd at level 1, falling
back to raw storage when the data is not compressible. This shrinks
cache directories by about 2-4x. The magic number is bumped so that
files in the old format are discarded as cache misses and then
overwritten in place.
- Handing the cache to V8 through a non-owning CachedData wrapper
instead of copying the whole buffer on every cache hit.
Corrupted cache files keep degrading to silent cache misses and are
regenerated, now covered by a regression test.
Co-authored-by: Grok <grok@x.ai>
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
@anonrig
anonrigforce-pushed the compile-cache-perf branch from 816f4ef to 0cf8818CompareJune 11, 2026 22:38
Creating and freeing a zstd context for every cache file costs more
than the (de)compression itself for small caches. Lazily create one
decompression context on the handler and reuse it across reads, and
share one compression context across all entries in Persist().
Co-authored-by: Grok <grok@x.ai>
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
@anonrig
anonrigforce-pushed the compile-cache-perf branch from 0cf8818 to f427fb3CompareJune 11, 2026 22:40
@joyeecheung

joyeecheung commented Jun 12, 2026

Copy link
Copy Markdown
Member

It seems the performance gains are within noise and it's mostly only the compression changes size? Can you split them into different PRs and measure them individually?

  1. I am not sure if the read changes actually gives any wins for small files - in the happy path where the cache is small, it's better to just read once and resize rather than stat and read (which are two file system calls instead of just one). Also there is a TOCTOU risk in doing stat, and fstat is not realiable across platforms, so the loop condition should not be gated on the fstat result but we must always read until EOF is reached in case the size is not accurate.
  2. It doesn't appear that compression alone does much to performance or it might actually hurt but just got compensated by other changes. In that case it's better to make that configurable and let users choose instead.
  3. Avoiding the copy would be better and there were precedents in builtin caches that it actually helped. I suspect this was the only one that actually helps performance while the other two may not. Hence it's better to split and measure individually.

lemire added a commit to lemire/node that referenced this pull request Jun 12, 2026
Builds on the zstd compression in nodejs#63861 by embedding a small zstd
dictionary trained on a diverse corpus of real modules, so each
small/medium compile-cache entry compresses better. Per entry we keep
the smaller of the plain and dictionary-assisted frame, so the
dictionary only ever helps.
- Add src/compile_cache_zstd.dict (16 KiB). It is trained on V8 code
caches harvested (via vm.compileFunction, the same shape the CJS
loader produces) from a diverse corpus: bundled npm packages, lib/,
tools/ and a few deps.
- Add tools/generate_compile_cache_dict.py and a node.gyp action that
generates compile_cache_zstd_dict.h into SHARED_INTERMEDIATE_DIR at
build time; no generated header is checked in. libnode include_dirs
updated to pick it up.
- Prepare the CDict/DDict once per process (shared across all handlers
and Workers, matching the lazy-context approach from nodejs#63861) and use
them in Persist() and ReadCacheFile(). Persist() compresses the plain
and dict frames into separate buffers and selects the smaller, so the
written bytes and recorded size always agree. The dictionary is only
tried for entries up to 256 KiB; larger blobs never benefit, so the
second compression is skipped to avoid wasted work. Falls back to
plain zstd if dictionary preparation fails.
- The dictionary is embedded in the binary because the compile cache
must be usable early, portably, and without extra filesystem state.
- No on-disk format change: dict-assisted frames carry the dictID, plain
frames carry none, and a single DDict decompresses both.
- Size, measured on data held out from training (per-entry min policy):
diverse modules go from ~1.87x (plain zstd) to ~2.44x with the
dictionary (~24% smaller on disk); on test/parallel, which is not in
the training corpus at all, ~1.74x -> ~2.22x (~22% smaller). A real
end-to-end run (npm --version, ~70 modules) is ~15% smaller. Read
time is unchanged and the extra write-time work is negligible.
- Add a multi-module write/read roundtrip test and a startup benchmark
(standard createBenchmark harness).
Builds on the zstd compression in nodejs#63861 by embedding a small zstd
dictionary trained on a diverse corpus of real modules, so each
small/medium compile-cache entry compresses better. Per entry we keep
the smaller of the plain and dictionary-assisted frame, so the
dictionary only ever helps.
- Add src/compile_cache_zstd.dict (16 KiB). It is trained on V8 code
caches harvested (via vm.compileFunction, the same shape the CJS
loader produces) from a diverse corpus: bundled npm packages, lib/,
tools/ and a few deps.
- Add tools/generate_compile_cache_dict.py and a node.gyp action that
generates compile_cache_zstd_dict.h into SHARED_INTERMEDIATE_DIR at
build time; no generated header is checked in. libnode include_dirs
updated to pick it up.
- Prepare the CDict/DDict once per process (shared across all handlers
and Workers, matching the lazy-context approach from nodejs#63861) and use
them in Persist() and ReadCacheFile(). Persist() compresses the plain
and dict frames into separate buffers and selects the smaller, so the
written bytes and recorded size always agree. The dictionary is only
tried for entries up to 256 KiB; larger blobs never benefit, so the
second compression is skipped to avoid wasted work. Falls back to
plain zstd if dictionary preparation fails.
- The dictionary is embedded in the binary because the compile cache
must be usable early, portably, and without extra filesystem state.
- No on-disk format change: dict-assisted frames carry the dictID, plain
frames carry none, and a single DDict decompresses both.
- Size, measured on data held out from training (per-entry min policy):
diverse modules go from ~1.87x (plain zstd) to ~2.44x with the
dictionary (~24% smaller on disk); on test/parallel, which is not in
the training corpus at all, ~1.74x -> ~2.22x (~22% smaller). A real
end-to-end run (npm --version, ~70 modules) is ~15% smaller. Read
time is unchanged and the extra write-time work is negligible.
- Add a multi-module write/read roundtrip test and a startup benchmark
(standard createBenchmark harness).

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If we do include dictionary files in our source tree, I'd say we should either include instructions for how to (at least approximately) re-create this file, or even generate it at build time entirely. strings src/compile_cache_zstd.dict results in a surprising amount of Node.js-core specificity, which seems surprising given that the commit that adds it claims it was trained on public npm packages.

(strings src/compile_cache_zstd.dict output in the fold)
wUUu2
_HOFJW
[Z)S
nding
options
Paz
types
PaN
process
versions
openssl
internal/process/pre_execution
PdR`
prepareMainThreadExecution
Pc6Wl
markBootstrapComplete
internal/options
"<";
__createBinding
__setModuleDefault
PaZY4J
birname
note
signatures
ObjectSetPrototypeOf
PromisePrototypeThen
PromiseWithResolvers
RegExpPrototypeExec
RegExpPrototypeSymbolReplace
StringPrototypeSplit
StringPrototypeToLowerCase
_createBlob
_createBlobFromFilePath
getDataObject
kMaxLength
TextDecoder
markTransferMode
Pbvyo
isAnyArrayBuffer
PcBS
isArrayBufferView
PaZ
require
PaN"
util
InvalidArgumentError
ConnectTimeoutError
SessionCache
maybeNormalizeConnectError
PromiseWithResolvers
nonOpStart
Pf&
writableStreamMarkFirstWriteRequestInFlight
createPromiseCallback
Pb:6k
nonOpWrite
$Pg>
writableStreamDefaultWriterEnsureReadyPromiseRejected
writableStreamDefaultControllerClose
DOMException
Pc^v
extractHighWaterMark
extractSizeAlgorithm
Pb~:
kEmptyObject
wgIgAygCACIAIAQgAWtqIQUgASAAa0ECaiEGAkADQCABLQAAIABButUAai0AAEcNAyAAQQJGDQEgAEEBaiEAIAQgAUEBaiIBRw0ACyADIAU2AgAMwwILIAMoAgQhACADQgA3AwAgAyAAIAZBAWoiARAuIgBFBEBB4wEhAgyqAgsgA0H1ATYCHCADIAE2AhQgAyAANgIMQQAhAgzCAgtB9AEhAiABIARGDcECIAMoAgAiACAEIAFraiEFIAEgAGtBAWohBgJAA0AgAS0AACAAQbjVAGotAABHDQIgAEEBRg0BIABBAWohACAEIAFBAWoiAUcNAAsgAyAFNgIADMICCyADQYEEOwEoIAMoAgQhACADQgA3AwAgAyAAIAZBAWoiARAuIgANAwwCCyADQQA2AgALQQAhAiADQQA2AhwgAyABNgIUIANB5R82AhAgA0EINgIMDL8CC0HVASECDKUCCyADQfMBNgIcIAMgATYCFCADIAA2AgxBACECDL0CC0EAIQACQCADKAI4IgJFDQAgAigCQ
memoMethod
perf
calculatedSizcludes
kTypes
uvErrmapGet
ArrayPrototypeIndexOf
classRegExp
MathMax
ArrayIsArray
isWindows
ObjectIsExtensible
isPermissionModelError
getExpectedArgumentLength
ArrayPrototypeJoin
kIsNodeError
get source
get ports
CloseEvent
get listening
header.js.map
@[@b
node:path
./large-numbers.js
"<";
escapeHTML
PcB_C+
convertChangesToXML
#noProxy
#ProxyAgent
#getProxy
#timeoutConnection
Pc~}
#drainPendingRequests
connect
Pbv2&_
addRequest
createSocket
.get
.get
getMilestoneTimestamp
getTimeOriginTimestamp
PbjK
internalBinding
performance
constants
ERR_INVALID_ARG_TYPE
ERR_INVALID_ARG_VALUE
Pb:e
ERR_OUT_OF_RANGE
isErrorStackTraceLimitWritable
Buffer
Pa&
inspect
validateBoolean
validateFunction
validateNumber
validateString
validateOneOf
PbbH
validateObject
validateInteger
isReadableStream
isWritableStream
isNodeStream
utilColors
PbJ%
lazyUtilColors
PbNn6
__importStar
DQCABLQAAQSBHDUggBCABQQFqIgFHDQALQSUhAwxpC0ElIQMMaAsgAi0ALUEBcQRAQcMBIQMMTwsgAigCBCEAQQAhAyACQQA2AgQgAiAAIAEQKSIABEAgAkEmNgIcIAIgADYCDCACIAFBAWo2AhQMaAsgAUEBaiEBDFwLIAFBAWohASACLwEwIgBBgAFxBEBBACEAAkAgAigCOCIDRQ0AIAMoAlQiA0UNACACIAMRAAAhAAsgAEUNBiAAQRVHDR8gAkEFNgIcIAIgATYCFCACQfkXNgIQIAJBFTYCDEEAIQMMZwsCQCAAQaAEcUGgBEcNACACLQAtQQJxDQBBACEDIAJBADYCHCACIAE2AhQgAkGWEzYCECACQQQ2AgwMZwsgAgJ/IAIvATBBFHFBFEYEQEEBIAItAChBAUYNARogAi8BMkHlAEYMAQsgAi0AKUEFRgs6AC5BACEAAkAgAigCOCIDRQ0AIAMoAiQiA0UNACACIAMRAAAhAAsCQAJAAkACQAJAIAAOFgIBAAQEBAQEBAQE
get ok
get statusText
get headers
get body
get bodyUsed
clone
parsePullArgs
PcNs
validateBackpressure
primordials
PcJ=
internal/encoding
TextEncoder
internal/errors
Pav
codes
internal/util
internal/util/types
internal/validators
onComplete
E`w PaN
onError
node:assert
../core/symbols
../core/errors
../core/util
Pab
beep
ArrayFrom
ArrayPrototypeFilter
ArrayPrototypeIncludes
ArrayPrototypeMap
ArrayPrototypePush
PcB,t
ArrayPrototypePushApply
ArrayPrototypeSlice
ObjectDefineProperty
Pb>c
ObjectKeys
PdRc
ObjectPrototypeHasOwnProperty
Pbf)
ReflectGet
SafeMap
SafeSet
StringPrototypeSlice
Error
PbFOT
_flushFlag
__esModule
.desc.get
stop
destroyer
primordials
internal/errors
Pav
codes
internal/streams/utils
__esModule
fs/promises
PaF-M
path
PaZ
require
Pa"/
module
__filename
__dirname

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@anonrig I'm going to mark this as un-resolved again, since there was no response here as far as I can tell

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, I didn't see this up until now and didn't pressed the "resolve" button at all. Will address your concerns.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

didn't pressed the "resolve" button at all

Github begged to differ:

Screenshot 2026-06-14 005643

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AI "babysit" mode might be the root cause

Comment threadsrc/compile_cache.h
compiler_cache_store_;
// Lazily created zstd decompression context, reused across cache reads
// to avoid recreating its workspace for every file.
ZSTD_DCtx_s* zstd_dctx_ = nullptr;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can this be held in a unique_ptr instead of a raw pointer?

Comment threadsrc/compile_cache.cc
if (cctx != nullptr) {
ZSTD_freeCCtx(cctx);
}
});

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

unique_ptr for this as well

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

c++Issues and PRs that require attention from people who are familiar with C++.lib / srcIssues and PRs involving general changes in the lib/ or src/ directories.needs-ciPRs that need a full CI run.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@anonrig@nodejs-github-bot@joyeecheung@lemire@jasnell@addaleax
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

src: improve compile cache performance and size - #63861

Open
anonrig wants to merge 3 commits into
nodejs:mainfrom
anonrig:compile-cache-perf
Open

src: improve compile cache performance and size#63861
anonrig wants to merge 3 commits into
nodejs:mainfrom
anonrig:compile-cache-perf

Conversation

@anonrig

Copy link
Copy Markdown
Member

Improves the on-disk compile cache (NODE_COMPILE_CACHE / module.enableCompileCache()):

  • Read path: read cache files with a single exactly-sized read (using the file size from fstat) instead of an exponentially growing buffer, which previously cost O(log N) syscalls/allocations and ~2N bytes of copying per file.
  • Size: compress the cache content on disk with zstd (level 1, prioritizing speed since persistence happens at shutdown), falling back to raw storage when not compressible. Shrinks cache directories ~2-4x and makes the crc32 integrity check cheaper since it now runs over the compressed bytes. The magic number is bumped so files in the old format are discarded as cache misses and overwritten in place.
  • Consume path: hand the cache to V8 through a non-owning CachedData wrapper (BufferNotOwned) instead of copying the entire buffer on every cache hit. The underlying buffer is owned by the cache entry, which outlives the synchronous compilation (same pattern as the vm cached-data path in node_contextify.cc).

Corrupted cache files keep degrading to silent cache misses and are regenerated; a corrupted size header can no longer cause an oversized allocation since the zstd frame content size is cross-checked first. Added test/parallel/test-compile-cache-corrupted.js covering bad magic, truncation, content bit-flips, and header corruption.

No public API or documented behavior changes; the file format is private to src/compile_cache.cc. Benchmark numbers (cache size and warm-startup timings) to follow in a comment.

This change was developed with AI assistance (see Co-authored-by trailer).

@nodejs-github-botnodejs-github-bot added c++ Issues and PRs that require attention from people who are familiar with C++. lib / src Issues and PRs involving general changes in the lib/ or src/ directories. needs-ci PRs that need a full CI run. labels Jun 11, 2026
@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

Review requested:

  • @nodejs/loaders
  • @nodejs/vm

@anonrig

Copy link
Copy Markdown
MemberAuthor

Verification results (macOS arm64, release build, vs. baseline at the merge-base):

Tests: all 23 parallel/test-compile-cache* pass (22 existing unchanged + the new corruption test).

On-disk size

ScenarioBaselineThis PRReduction
10 MB snapshot/typescript.js fixture (single big CJS file)1,818,380 B752,924 B2.42x
300 small/medium CJS modules274,316 B179,915 B1.52x

Warm-start, self-controlled (same binary, cache on vs. off, interleaved 20 runs, trimmed means — avoids binary-layout skew between builds):

ScenarioCache benefit (baseline)Cache benefit (this PR)
big-file+26.4 ms+25.0 ms
many-modules+1.7 ms+1.5 ms

Warm-start with a hot page cache is neutral within noise (~1 ms on the pathological single-10 MB-file case, which is dominated by one large zstd decompress; with cold page cache or slower storage the 2.4x smaller read wins). Cold-start adds one-time compression at persist (~35 ms for the 1.8 MB blob at level 1, proportionally less for typical files).

The second commit reuses zstd contexts (one ZSTD_DCtx on the handler, one ZSTD_CCtx across Persist()), which removed most of the per-file decompression overhead observed with one-shot contexts on the many-modules scenario.

Improve the compile cache by:
- Reading cache files with a single exactly-sized read using the file
size from fstat instead of reading into an exponentially growing
buffer, which previously cost O(log N) syscalls and allocations and
about 2N bytes of copying per file.
- Compressing the cache content on disk with zstd at level 1, falling
back to raw storage when the data is not compressible. This shrinks
cache directories by about 2-4x. The magic number is bumped so that
files in the old format are discarded as cache misses and then
overwritten in place.
- Handing the cache to V8 through a non-owning CachedData wrapper
instead of copying the whole buffer on every cache hit.
Corrupted cache files keep degrading to silent cache misses and are
regenerated, now covered by a regression test.
Co-authored-by: Grok <grok@x.ai>
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
@anonrig
anonrigforce-pushed the compile-cache-perf branch from 816f4ef to 0cf8818CompareJune 11, 2026 22:38
Creating and freeing a zstd context for every cache file costs more
than the (de)compression itself for small caches. Lazily create one
decompression context on the handler and reuse it across reads, and
share one compression context across all entries in Persist().
Co-authored-by: Grok <grok@x.ai>
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
@anonrig
anonrigforce-pushed the compile-cache-perf branch from 0cf8818 to f427fb3CompareJune 11, 2026 22:40
@joyeecheung

joyeecheung commented Jun 12, 2026

Copy link
Copy Markdown
Member

It seems the performance gains are within noise and it's mostly only the compression changes size? Can you split them into different PRs and measure them individually?

  1. I am not sure if the read changes actually gives any wins for small files - in the happy path where the cache is small, it's better to just read once and resize rather than stat and read (which are two file system calls instead of just one). Also there is a TOCTOU risk in doing stat, and fstat is not realiable across platforms, so the loop condition should not be gated on the fstat result but we must always read until EOF is reached in case the size is not accurate.
  2. It doesn't appear that compression alone does much to performance or it might actually hurt but just got compensated by other changes. In that case it's better to make that configurable and let users choose instead.
  3. Avoiding the copy would be better and there were precedents in builtin caches that it actually helped. I suspect this was the only one that actually helps performance while the other two may not. Hence it's better to split and measure individually.

lemire added a commit to lemire/node that referenced this pull request Jun 12, 2026
Builds on the zstd compression in nodejs#63861 by embedding a small zstd
dictionary trained on a diverse corpus of real modules, so each
small/medium compile-cache entry compresses better. Per entry we keep
the smaller of the plain and dictionary-assisted frame, so the
dictionary only ever helps.
- Add src/compile_cache_zstd.dict (16 KiB). It is trained on V8 code
caches harvested (via vm.compileFunction, the same shape the CJS
loader produces) from a diverse corpus: bundled npm packages, lib/,
tools/ and a few deps.
- Add tools/generate_compile_cache_dict.py and a node.gyp action that
generates compile_cache_zstd_dict.h into SHARED_INTERMEDIATE_DIR at
build time; no generated header is checked in. libnode include_dirs
updated to pick it up.
- Prepare the CDict/DDict once per process (shared across all handlers
and Workers, matching the lazy-context approach from nodejs#63861) and use
them in Persist() and ReadCacheFile(). Persist() compresses the plain
and dict frames into separate buffers and selects the smaller, so the
written bytes and recorded size always agree. The dictionary is only
tried for entries up to 256 KiB; larger blobs never benefit, so the
second compression is skipped to avoid wasted work. Falls back to
plain zstd if dictionary preparation fails.
- The dictionary is embedded in the binary because the compile cache
must be usable early, portably, and without extra filesystem state.
- No on-disk format change: dict-assisted frames carry the dictID, plain
frames carry none, and a single DDict decompresses both.
- Size, measured on data held out from training (per-entry min policy):
diverse modules go from ~1.87x (plain zstd) to ~2.44x with the
dictionary (~24% smaller on disk); on test/parallel, which is not in
the training corpus at all, ~1.74x -> ~2.22x (~22% smaller). A real
end-to-end run (npm --version, ~70 modules) is ~15% smaller. Read
time is unchanged and the extra write-time work is negligible.
- Add a multi-module write/read roundtrip test and a startup benchmark
(standard createBenchmark harness).
Builds on the zstd compression in nodejs#63861 by embedding a small zstd
dictionary trained on a diverse corpus of real modules, so each
small/medium compile-cache entry compresses better. Per entry we keep
the smaller of the plain and dictionary-assisted frame, so the
dictionary only ever helps.
- Add src/compile_cache_zstd.dict (16 KiB). It is trained on V8 code
caches harvested (via vm.compileFunction, the same shape the CJS
loader produces) from a diverse corpus: bundled npm packages, lib/,
tools/ and a few deps.
- Add tools/generate_compile_cache_dict.py and a node.gyp action that
generates compile_cache_zstd_dict.h into SHARED_INTERMEDIATE_DIR at
build time; no generated header is checked in. libnode include_dirs
updated to pick it up.
- Prepare the CDict/DDict once per process (shared across all handlers
and Workers, matching the lazy-context approach from nodejs#63861) and use
them in Persist() and ReadCacheFile(). Persist() compresses the plain
and dict frames into separate buffers and selects the smaller, so the
written bytes and recorded size always agree. The dictionary is only
tried for entries up to 256 KiB; larger blobs never benefit, so the
second compression is skipped to avoid wasted work. Falls back to
plain zstd if dictionary preparation fails.
- The dictionary is embedded in the binary because the compile cache
must be usable early, portably, and without extra filesystem state.
- No on-disk format change: dict-assisted frames carry the dictID, plain
frames carry none, and a single DDict decompresses both.
- Size, measured on data held out from training (per-entry min policy):
diverse modules go from ~1.87x (plain zstd) to ~2.44x with the
dictionary (~24% smaller on disk); on test/parallel, which is not in
the training corpus at all, ~1.74x -> ~2.22x (~22% smaller). A real
end-to-end run (npm --version, ~70 modules) is ~15% smaller. Read
time is unchanged and the extra write-time work is negligible.
- Add a multi-module write/read roundtrip test and a startup benchmark
(standard createBenchmark harness).

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If we do include dictionary files in our source tree, I'd say we should either include instructions for how to (at least approximately) re-create this file, or even generate it at build time entirely. strings src/compile_cache_zstd.dict results in a surprising amount of Node.js-core specificity, which seems surprising given that the commit that adds it claims it was trained on public npm packages.

(strings src/compile_cache_zstd.dict output in the fold)
wUUu2
_HOFJW
[Z)S
nding
options
Paz
types
PaN
process
versions
openssl
internal/process/pre_execution
PdR`
prepareMainThreadExecution
Pc6Wl
markBootstrapComplete
internal/options
"<";
__createBinding
__setModuleDefault
PaZY4J
birname
note
signatures
ObjectSetPrototypeOf
PromisePrototypeThen
PromiseWithResolvers
RegExpPrototypeExec
RegExpPrototypeSymbolReplace
StringPrototypeSplit
StringPrototypeToLowerCase
_createBlob
_createBlobFromFilePath
getDataObject
kMaxLength
TextDecoder
markTransferMode
Pbvyo
isAnyArrayBuffer
PcBS
isArrayBufferView
PaZ
require
PaN"
util
InvalidArgumentError
ConnectTimeoutError
SessionCache
maybeNormalizeConnectError
PromiseWithResolvers
nonOpStart
Pf&
writableStreamMarkFirstWriteRequestInFlight
createPromiseCallback
Pb:6k
nonOpWrite
$Pg>
writableStreamDefaultWriterEnsureReadyPromiseRejected
writableStreamDefaultControllerClose
DOMException
Pc^v
extractHighWaterMark
extractSizeAlgorithm
Pb~:
kEmptyObject
wgIgAygCACIAIAQgAWtqIQUgASAAa0ECaiEGAkADQCABLQAAIABButUAai0AAEcNAyAAQQJGDQEgAEEBaiEAIAQgAUEBaiIBRw0ACyADIAU2AgAMwwILIAMoAgQhACADQgA3AwAgAyAAIAZBAWoiARAuIgBFBEBB4wEhAgyqAgsgA0H1ATYCHCADIAE2AhQgAyAANgIMQQAhAgzCAgtB9AEhAiABIARGDcECIAMoAgAiACAEIAFraiEFIAEgAGtBAWohBgJAA0AgAS0AACAAQbjVAGotAABHDQIgAEEBRg0BIABBAWohACAEIAFBAWoiAUcNAAsgAyAFNgIADMICCyADQYEEOwEoIAMoAgQhACADQgA3AwAgAyAAIAZBAWoiARAuIgANAwwCCyADQQA2AgALQQAhAiADQQA2AhwgAyABNgIUIANB5R82AhAgA0EINgIMDL8CC0HVASECDKUCCyADQfMBNgIcIAMgATYCFCADIAA2AgxBACECDL0CC0EAIQACQCADKAI4IgJFDQAgAigCQ
memoMethod
perf
calculatedSizcludes
kTypes
uvErrmapGet
ArrayPrototypeIndexOf
classRegExp
MathMax
ArrayIsArray
isWindows
ObjectIsExtensible
isPermissionModelError
getExpectedArgumentLength
ArrayPrototypeJoin
kIsNodeError
get source
get ports
CloseEvent
get listening
header.js.map
@[@b
node:path
./large-numbers.js
"<";
escapeHTML
PcB_C+
convertChangesToXML
#noProxy
#ProxyAgent
#getProxy
#timeoutConnection
Pc~}
#drainPendingRequests
connect
Pbv2&_
addRequest
createSocket
.get
.get
getMilestoneTimestamp
getTimeOriginTimestamp
PbjK
internalBinding
performance
constants
ERR_INVALID_ARG_TYPE
ERR_INVALID_ARG_VALUE
Pb:e
ERR_OUT_OF_RANGE
isErrorStackTraceLimitWritable
Buffer
Pa&
inspect
validateBoolean
validateFunction
validateNumber
validateString
validateOneOf
PbbH
validateObject
validateInteger
isReadableStream
isWritableStream
isNodeStream
utilColors
PbJ%
lazyUtilColors
PbNn6
__importStar
DQCABLQAAQSBHDUggBCABQQFqIgFHDQALQSUhAwxpC0ElIQMMaAsgAi0ALUEBcQRAQcMBIQMMTwsgAigCBCEAQQAhAyACQQA2AgQgAiAAIAEQKSIABEAgAkEmNgIcIAIgADYCDCACIAFBAWo2AhQMaAsgAUEBaiEBDFwLIAFBAWohASACLwEwIgBBgAFxBEBBACEAAkAgAigCOCIDRQ0AIAMoAlQiA0UNACACIAMRAAAhAAsgAEUNBiAAQRVHDR8gAkEFNgIcIAIgATYCFCACQfkXNgIQIAJBFTYCDEEAIQMMZwsCQCAAQaAEcUGgBEcNACACLQAtQQJxDQBBACEDIAJBADYCHCACIAE2AhQgAkGWEzYCECACQQQ2AgwMZwsgAgJ/IAIvATBBFHFBFEYEQEEBIAItAChBAUYNARogAi8BMkHlAEYMAQsgAi0AKUEFRgs6AC5BACEAAkAgAigCOCIDRQ0AIAMoAiQiA0UNACACIAMRAAAhAAsCQAJAAkACQAJAIAAOFgIBAAQEBAQEBAQE
get ok
get statusText
get headers
get body
get bodyUsed
clone
parsePullArgs
PcNs
validateBackpressure
primordials
PcJ=
internal/encoding
TextEncoder
internal/errors
Pav
codes
internal/util
internal/util/types
internal/validators
onComplete
E`w PaN
onError
node:assert
../core/symbols
../core/errors
../core/util
Pab
beep
ArrayFrom
ArrayPrototypeFilter
ArrayPrototypeIncludes
ArrayPrototypeMap
ArrayPrototypePush
PcB,t
ArrayPrototypePushApply
ArrayPrototypeSlice
ObjectDefineProperty
Pb>c
ObjectKeys
PdRc
ObjectPrototypeHasOwnProperty
Pbf)
ReflectGet
SafeMap
SafeSet
StringPrototypeSlice
Error
PbFOT
_flushFlag
__esModule
.desc.get
stop
destroyer
primordials
internal/errors
Pav
codes
internal/streams/utils
__esModule
fs/promises
PaF-M
path
PaZ
require
Pa"/
module
__filename
__dirname

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@anonrig I'm going to mark this as un-resolved again, since there was no response here as far as I can tell

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, I didn't see this up until now and didn't pressed the "resolve" button at all. Will address your concerns.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

didn't pressed the "resolve" button at all

Github begged to differ:

Screenshot 2026-06-14 005643

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AI "babysit" mode might be the root cause

Comment threadsrc/compile_cache.h
compiler_cache_store_;
// Lazily created zstd decompression context, reused across cache reads
// to avoid recreating its workspace for every file.
ZSTD_DCtx_s* zstd_dctx_ = nullptr;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can this be held in a unique_ptr instead of a raw pointer?

Comment threadsrc/compile_cache.cc
if (cctx != nullptr) {
ZSTD_freeCCtx(cctx);
}
});

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

unique_ptr for this as well

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

c++Issues and PRs that require attention from people who are familiar with C++.lib / srcIssues and PRs involving general changes in the lib/ or src/ directories.needs-ciPRs that need a full CI run.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@anonrig@nodejs-github-bot@joyeecheung@lemire@jasnell@addaleax
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

src: improve compile cache performance and size - #63861

Open
anonrig wants to merge 3 commits into
nodejs:mainfrom
anonrig:compile-cache-perf
Open

src: improve compile cache performance and size#63861
anonrig wants to merge 3 commits into
nodejs:mainfrom
anonrig:compile-cache-perf

Conversation

@anonrig

Copy link
Copy Markdown
Member

Improves the on-disk compile cache (NODE_COMPILE_CACHE / module.enableCompileCache()):

  • Read path: read cache files with a single exactly-sized read (using the file size from fstat) instead of an exponentially growing buffer, which previously cost O(log N) syscalls/allocations and ~2N bytes of copying per file.
  • Size: compress the cache content on disk with zstd (level 1, prioritizing speed since persistence happens at shutdown), falling back to raw storage when not compressible. Shrinks cache directories ~2-4x and makes the crc32 integrity check cheaper since it now runs over the compressed bytes. The magic number is bumped so files in the old format are discarded as cache misses and overwritten in place.
  • Consume path: hand the cache to V8 through a non-owning CachedData wrapper (BufferNotOwned) instead of copying the entire buffer on every cache hit. The underlying buffer is owned by the cache entry, which outlives the synchronous compilation (same pattern as the vm cached-data path in node_contextify.cc).

Corrupted cache files keep degrading to silent cache misses and are regenerated; a corrupted size header can no longer cause an oversized allocation since the zstd frame content size is cross-checked first. Added test/parallel/test-compile-cache-corrupted.js covering bad magic, truncation, content bit-flips, and header corruption.

No public API or documented behavior changes; the file format is private to src/compile_cache.cc. Benchmark numbers (cache size and warm-startup timings) to follow in a comment.

This change was developed with AI assistance (see Co-authored-by trailer).

@nodejs-github-botnodejs-github-bot added c++ Issues and PRs that require attention from people who are familiar with C++. lib / src Issues and PRs involving general changes in the lib/ or src/ directories. needs-ci PRs that need a full CI run. labels Jun 11, 2026
@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

Review requested:

  • @nodejs/loaders
  • @nodejs/vm

@anonrig

Copy link
Copy Markdown
MemberAuthor

Verification results (macOS arm64, release build, vs. baseline at the merge-base):

Tests: all 23 parallel/test-compile-cache* pass (22 existing unchanged + the new corruption test).

On-disk size

ScenarioBaselineThis PRReduction
10 MB snapshot/typescript.js fixture (single big CJS file)1,818,380 B752,924 B2.42x
300 small/medium CJS modules274,316 B179,915 B1.52x

Warm-start, self-controlled (same binary, cache on vs. off, interleaved 20 runs, trimmed means — avoids binary-layout skew between builds):

ScenarioCache benefit (baseline)Cache benefit (this PR)
big-file+26.4 ms+25.0 ms
many-modules+1.7 ms+1.5 ms

Warm-start with a hot page cache is neutral within noise (~1 ms on the pathological single-10 MB-file case, which is dominated by one large zstd decompress; with cold page cache or slower storage the 2.4x smaller read wins). Cold-start adds one-time compression at persist (~35 ms for the 1.8 MB blob at level 1, proportionally less for typical files).

The second commit reuses zstd contexts (one ZSTD_DCtx on the handler, one ZSTD_CCtx across Persist()), which removed most of the per-file decompression overhead observed with one-shot contexts on the many-modules scenario.

Improve the compile cache by:
- Reading cache files with a single exactly-sized read using the file
size from fstat instead of reading into an exponentially growing
buffer, which previously cost O(log N) syscalls and allocations and
about 2N bytes of copying per file.
- Compressing the cache content on disk with zstd at level 1, falling
back to raw storage when the data is not compressible. This shrinks
cache directories by about 2-4x. The magic number is bumped so that
files in the old format are discarded as cache misses and then
overwritten in place.
- Handing the cache to V8 through a non-owning CachedData wrapper
instead of copying the whole buffer on every cache hit.
Corrupted cache files keep degrading to silent cache misses and are
regenerated, now covered by a regression test.
Co-authored-by: Grok <grok@x.ai>
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
@anonrig
anonrigforce-pushed the compile-cache-perf branch from 816f4ef to 0cf8818CompareJune 11, 2026 22:38
Creating and freeing a zstd context for every cache file costs more
than the (de)compression itself for small caches. Lazily create one
decompression context on the handler and reuse it across reads, and
share one compression context across all entries in Persist().
Co-authored-by: Grok <grok@x.ai>
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
@anonrig
anonrigforce-pushed the compile-cache-perf branch from 0cf8818 to f427fb3CompareJune 11, 2026 22:40
@joyeecheung

joyeecheung commented Jun 12, 2026

Copy link
Copy Markdown
Member

It seems the performance gains are within noise and it's mostly only the compression changes size? Can you split them into different PRs and measure them individually?

  1. I am not sure if the read changes actually gives any wins for small files - in the happy path where the cache is small, it's better to just read once and resize rather than stat and read (which are two file system calls instead of just one). Also there is a TOCTOU risk in doing stat, and fstat is not realiable across platforms, so the loop condition should not be gated on the fstat result but we must always read until EOF is reached in case the size is not accurate.
  2. It doesn't appear that compression alone does much to performance or it might actually hurt but just got compensated by other changes. In that case it's better to make that configurable and let users choose instead.
  3. Avoiding the copy would be better and there were precedents in builtin caches that it actually helped. I suspect this was the only one that actually helps performance while the other two may not. Hence it's better to split and measure individually.

lemire added a commit to lemire/node that referenced this pull request Jun 12, 2026
Builds on the zstd compression in nodejs#63861 by embedding a small zstd
dictionary trained on a diverse corpus of real modules, so each
small/medium compile-cache entry compresses better. Per entry we keep
the smaller of the plain and dictionary-assisted frame, so the
dictionary only ever helps.
- Add src/compile_cache_zstd.dict (16 KiB). It is trained on V8 code
caches harvested (via vm.compileFunction, the same shape the CJS
loader produces) from a diverse corpus: bundled npm packages, lib/,
tools/ and a few deps.
- Add tools/generate_compile_cache_dict.py and a node.gyp action that
generates compile_cache_zstd_dict.h into SHARED_INTERMEDIATE_DIR at
build time; no generated header is checked in. libnode include_dirs
updated to pick it up.
- Prepare the CDict/DDict once per process (shared across all handlers
and Workers, matching the lazy-context approach from nodejs#63861) and use
them in Persist() and ReadCacheFile(). Persist() compresses the plain
and dict frames into separate buffers and selects the smaller, so the
written bytes and recorded size always agree. The dictionary is only
tried for entries up to 256 KiB; larger blobs never benefit, so the
second compression is skipped to avoid wasted work. Falls back to
plain zstd if dictionary preparation fails.
- The dictionary is embedded in the binary because the compile cache
must be usable early, portably, and without extra filesystem state.
- No on-disk format change: dict-assisted frames carry the dictID, plain
frames carry none, and a single DDict decompresses both.
- Size, measured on data held out from training (per-entry min policy):
diverse modules go from ~1.87x (plain zstd) to ~2.44x with the
dictionary (~24% smaller on disk); on test/parallel, which is not in
the training corpus at all, ~1.74x -> ~2.22x (~22% smaller). A real
end-to-end run (npm --version, ~70 modules) is ~15% smaller. Read
time is unchanged and the extra write-time work is negligible.
- Add a multi-module write/read roundtrip test and a startup benchmark
(standard createBenchmark harness).
Builds on the zstd compression in nodejs#63861 by embedding a small zstd
dictionary trained on a diverse corpus of real modules, so each
small/medium compile-cache entry compresses better. Per entry we keep
the smaller of the plain and dictionary-assisted frame, so the
dictionary only ever helps.
- Add src/compile_cache_zstd.dict (16 KiB). It is trained on V8 code
caches harvested (via vm.compileFunction, the same shape the CJS
loader produces) from a diverse corpus: bundled npm packages, lib/,
tools/ and a few deps.
- Add tools/generate_compile_cache_dict.py and a node.gyp action that
generates compile_cache_zstd_dict.h into SHARED_INTERMEDIATE_DIR at
build time; no generated header is checked in. libnode include_dirs
updated to pick it up.
- Prepare the CDict/DDict once per process (shared across all handlers
and Workers, matching the lazy-context approach from nodejs#63861) and use
them in Persist() and ReadCacheFile(). Persist() compresses the plain
and dict frames into separate buffers and selects the smaller, so the
written bytes and recorded size always agree. The dictionary is only
tried for entries up to 256 KiB; larger blobs never benefit, so the
second compression is skipped to avoid wasted work. Falls back to
plain zstd if dictionary preparation fails.
- The dictionary is embedded in the binary because the compile cache
must be usable early, portably, and without extra filesystem state.
- No on-disk format change: dict-assisted frames carry the dictID, plain
frames carry none, and a single DDict decompresses both.
- Size, measured on data held out from training (per-entry min policy):
diverse modules go from ~1.87x (plain zstd) to ~2.44x with the
dictionary (~24% smaller on disk); on test/parallel, which is not in
the training corpus at all, ~1.74x -> ~2.22x (~22% smaller). A real
end-to-end run (npm --version, ~70 modules) is ~15% smaller. Read
time is unchanged and the extra write-time work is negligible.
- Add a multi-module write/read roundtrip test and a startup benchmark
(standard createBenchmark harness).

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If we do include dictionary files in our source tree, I'd say we should either include instructions for how to (at least approximately) re-create this file, or even generate it at build time entirely. strings src/compile_cache_zstd.dict results in a surprising amount of Node.js-core specificity, which seems surprising given that the commit that adds it claims it was trained on public npm packages.

(strings src/compile_cache_zstd.dict output in the fold)
wUUu2
_HOFJW
[Z)S
nding
options
Paz
types
PaN
process
versions
openssl
internal/process/pre_execution
PdR`
prepareMainThreadExecution
Pc6Wl
markBootstrapComplete
internal/options
"<";
__createBinding
__setModuleDefault
PaZY4J
birname
note
signatures
ObjectSetPrototypeOf
PromisePrototypeThen
PromiseWithResolvers
RegExpPrototypeExec
RegExpPrototypeSymbolReplace
StringPrototypeSplit
StringPrototypeToLowerCase
_createBlob
_createBlobFromFilePath
getDataObject
kMaxLength
TextDecoder
markTransferMode
Pbvyo
isAnyArrayBuffer
PcBS
isArrayBufferView
PaZ
require
PaN"
util
InvalidArgumentError
ConnectTimeoutError
SessionCache
maybeNormalizeConnectError
PromiseWithResolvers
nonOpStart
Pf&
writableStreamMarkFirstWriteRequestInFlight
createPromiseCallback
Pb:6k
nonOpWrite
$Pg>
writableStreamDefaultWriterEnsureReadyPromiseRejected
writableStreamDefaultControllerClose
DOMException
Pc^v
extractHighWaterMark
extractSizeAlgorithm
Pb~:
kEmptyObject
wgIgAygCACIAIAQgAWtqIQUgASAAa0ECaiEGAkADQCABLQAAIABButUAai0AAEcNAyAAQQJGDQEgAEEBaiEAIAQgAUEBaiIBRw0ACyADIAU2AgAMwwILIAMoAgQhACADQgA3AwAgAyAAIAZBAWoiARAuIgBFBEBB4wEhAgyqAgsgA0H1ATYCHCADIAE2AhQgAyAANgIMQQAhAgzCAgtB9AEhAiABIARGDcECIAMoAgAiACAEIAFraiEFIAEgAGtBAWohBgJAA0AgAS0AACAAQbjVAGotAABHDQIgAEEBRg0BIABBAWohACAEIAFBAWoiAUcNAAsgAyAFNgIADMICCyADQYEEOwEoIAMoAgQhACADQgA3AwAgAyAAIAZBAWoiARAuIgANAwwCCyADQQA2AgALQQAhAiADQQA2AhwgAyABNgIUIANB5R82AhAgA0EINgIMDL8CC0HVASECDKUCCyADQfMBNgIcIAMgATYCFCADIAA2AgxBACECDL0CC0EAIQACQCADKAI4IgJFDQAgAigCQ
memoMethod
perf
calculatedSizcludes
kTypes
uvErrmapGet
ArrayPrototypeIndexOf
classRegExp
MathMax
ArrayIsArray
isWindows
ObjectIsExtensible
isPermissionModelError
getExpectedArgumentLength
ArrayPrototypeJoin
kIsNodeError
get source
get ports
CloseEvent
get listening
header.js.map
@[@b
node:path
./large-numbers.js
"<";
escapeHTML
PcB_C+
convertChangesToXML
#noProxy
#ProxyAgent
#getProxy
#timeoutConnection
Pc~}
#drainPendingRequests
connect
Pbv2&_
addRequest
createSocket
.get
.get
getMilestoneTimestamp
getTimeOriginTimestamp
PbjK
internalBinding
performance
constants
ERR_INVALID_ARG_TYPE
ERR_INVALID_ARG_VALUE
Pb:e
ERR_OUT_OF_RANGE
isErrorStackTraceLimitWritable
Buffer
Pa&
inspect
validateBoolean
validateFunction
validateNumber
validateString
validateOneOf
PbbH
validateObject
validateInteger
isReadableStream
isWritableStream
isNodeStream
utilColors
PbJ%
lazyUtilColors
PbNn6
__importStar
DQCABLQAAQSBHDUggBCABQQFqIgFHDQALQSUhAwxpC0ElIQMMaAsgAi0ALUEBcQRAQcMBIQMMTwsgAigCBCEAQQAhAyACQQA2AgQgAiAAIAEQKSIABEAgAkEmNgIcIAIgADYCDCACIAFBAWo2AhQMaAsgAUEBaiEBDFwLIAFBAWohASACLwEwIgBBgAFxBEBBACEAAkAgAigCOCIDRQ0AIAMoAlQiA0UNACACIAMRAAAhAAsgAEUNBiAAQRVHDR8gAkEFNgIcIAIgATYCFCACQfkXNgIQIAJBFTYCDEEAIQMMZwsCQCAAQaAEcUGgBEcNACACLQAtQQJxDQBBACEDIAJBADYCHCACIAE2AhQgAkGWEzYCECACQQQ2AgwMZwsgAgJ/IAIvATBBFHFBFEYEQEEBIAItAChBAUYNARogAi8BMkHlAEYMAQsgAi0AKUEFRgs6AC5BACEAAkAgAigCOCIDRQ0AIAMoAiQiA0UNACACIAMRAAAhAAsCQAJAAkACQAJAIAAOFgIBAAQEBAQEBAQE
get ok
get statusText
get headers
get body
get bodyUsed
clone
parsePullArgs
PcNs
validateBackpressure
primordials
PcJ=
internal/encoding
TextEncoder
internal/errors
Pav
codes
internal/util
internal/util/types
internal/validators
onComplete
E`w PaN
onError
node:assert
../core/symbols
../core/errors
../core/util
Pab
beep
ArrayFrom
ArrayPrototypeFilter
ArrayPrototypeIncludes
ArrayPrototypeMap
ArrayPrototypePush
PcB,t
ArrayPrototypePushApply
ArrayPrototypeSlice
ObjectDefineProperty
Pb>c
ObjectKeys
PdRc
ObjectPrototypeHasOwnProperty
Pbf)
ReflectGet
SafeMap
SafeSet
StringPrototypeSlice
Error
PbFOT
_flushFlag
__esModule
.desc.get
stop
destroyer
primordials
internal/errors
Pav
codes
internal/streams/utils
__esModule
fs/promises
PaF-M
path
PaZ
require
Pa"/
module
__filename
__dirname

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@anonrig I'm going to mark this as un-resolved again, since there was no response here as far as I can tell

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, I didn't see this up until now and didn't pressed the "resolve" button at all. Will address your concerns.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

didn't pressed the "resolve" button at all

Github begged to differ:

Screenshot 2026-06-14 005643

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AI "babysit" mode might be the root cause

Comment threadsrc/compile_cache.h
compiler_cache_store_;
// Lazily created zstd decompression context, reused across cache reads
// to avoid recreating its workspace for every file.
ZSTD_DCtx_s* zstd_dctx_ = nullptr;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can this be held in a unique_ptr instead of a raw pointer?

Comment threadsrc/compile_cache.cc
if (cctx != nullptr) {
ZSTD_freeCCtx(cctx);
}
});

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

unique_ptr for this as well

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

c++Issues and PRs that require attention from people who are familiar with C++.lib / srcIssues and PRs involving general changes in the lib/ or src/ directories.needs-ciPRs that need a full CI run.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@anonrig@nodejs-github-bot@joyeecheung@lemire@jasnell@addaleax
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

src: improve compile cache performance and size - #63861

Open
anonrig wants to merge 3 commits into
nodejs:mainfrom
anonrig:compile-cache-perf
Open

src: improve compile cache performance and size#63861
anonrig wants to merge 3 commits into
nodejs:mainfrom
anonrig:compile-cache-perf

Conversation

@anonrig

Copy link
Copy Markdown
Member

Improves the on-disk compile cache (NODE_COMPILE_CACHE / module.enableCompileCache()):

  • Read path: read cache files with a single exactly-sized read (using the file size from fstat) instead of an exponentially growing buffer, which previously cost O(log N) syscalls/allocations and ~2N bytes of copying per file.
  • Size: compress the cache content on disk with zstd (level 1, prioritizing speed since persistence happens at shutdown), falling back to raw storage when not compressible. Shrinks cache directories ~2-4x and makes the crc32 integrity check cheaper since it now runs over the compressed bytes. The magic number is bumped so files in the old format are discarded as cache misses and overwritten in place.
  • Consume path: hand the cache to V8 through a non-owning CachedData wrapper (BufferNotOwned) instead of copying the entire buffer on every cache hit. The underlying buffer is owned by the cache entry, which outlives the synchronous compilation (same pattern as the vm cached-data path in node_contextify.cc).

Corrupted cache files keep degrading to silent cache misses and are regenerated; a corrupted size header can no longer cause an oversized allocation since the zstd frame content size is cross-checked first. Added test/parallel/test-compile-cache-corrupted.js covering bad magic, truncation, content bit-flips, and header corruption.

No public API or documented behavior changes; the file format is private to src/compile_cache.cc. Benchmark numbers (cache size and warm-startup timings) to follow in a comment.

This change was developed with AI assistance (see Co-authored-by trailer).

@nodejs-github-botnodejs-github-bot added c++ Issues and PRs that require attention from people who are familiar with C++. lib / src Issues and PRs involving general changes in the lib/ or src/ directories. needs-ci PRs that need a full CI run. labels Jun 11, 2026
@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

Review requested:

  • @nodejs/loaders
  • @nodejs/vm

@anonrig

Copy link
Copy Markdown
MemberAuthor

Verification results (macOS arm64, release build, vs. baseline at the merge-base):

Tests: all 23 parallel/test-compile-cache* pass (22 existing unchanged + the new corruption test).

On-disk size

ScenarioBaselineThis PRReduction
10 MB snapshot/typescript.js fixture (single big CJS file)1,818,380 B752,924 B2.42x
300 small/medium CJS modules274,316 B179,915 B1.52x

Warm-start, self-controlled (same binary, cache on vs. off, interleaved 20 runs, trimmed means — avoids binary-layout skew between builds):

ScenarioCache benefit (baseline)Cache benefit (this PR)
big-file+26.4 ms+25.0 ms
many-modules+1.7 ms+1.5 ms

Warm-start with a hot page cache is neutral within noise (~1 ms on the pathological single-10 MB-file case, which is dominated by one large zstd decompress; with cold page cache or slower storage the 2.4x smaller read wins). Cold-start adds one-time compression at persist (~35 ms for the 1.8 MB blob at level 1, proportionally less for typical files).

The second commit reuses zstd contexts (one ZSTD_DCtx on the handler, one ZSTD_CCtx across Persist()), which removed most of the per-file decompression overhead observed with one-shot contexts on the many-modules scenario.

Improve the compile cache by:
- Reading cache files with a single exactly-sized read using the file
size from fstat instead of reading into an exponentially growing
buffer, which previously cost O(log N) syscalls and allocations and
about 2N bytes of copying per file.
- Compressing the cache content on disk with zstd at level 1, falling
back to raw storage when the data is not compressible. This shrinks
cache directories by about 2-4x. The magic number is bumped so that
files in the old format are discarded as cache misses and then
overwritten in place.
- Handing the cache to V8 through a non-owning CachedData wrapper
instead of copying the whole buffer on every cache hit.
Corrupted cache files keep degrading to silent cache misses and are
regenerated, now covered by a regression test.
Co-authored-by: Grok <grok@x.ai>
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
@anonrig
anonrigforce-pushed the compile-cache-perf branch from 816f4ef to 0cf8818CompareJune 11, 2026 22:38
Creating and freeing a zstd context for every cache file costs more
than the (de)compression itself for small caches. Lazily create one
decompression context on the handler and reuse it across reads, and
share one compression context across all entries in Persist().
Co-authored-by: Grok <grok@x.ai>
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
@anonrig
anonrigforce-pushed the compile-cache-perf branch from 0cf8818 to f427fb3CompareJune 11, 2026 22:40
@joyeecheung

joyeecheung commented Jun 12, 2026

Copy link
Copy Markdown
Member

It seems the performance gains are within noise and it's mostly only the compression changes size? Can you split them into different PRs and measure them individually?

  1. I am not sure if the read changes actually gives any wins for small files - in the happy path where the cache is small, it's better to just read once and resize rather than stat and read (which are two file system calls instead of just one). Also there is a TOCTOU risk in doing stat, and fstat is not realiable across platforms, so the loop condition should not be gated on the fstat result but we must always read until EOF is reached in case the size is not accurate.
  2. It doesn't appear that compression alone does much to performance or it might actually hurt but just got compensated by other changes. In that case it's better to make that configurable and let users choose instead.
  3. Avoiding the copy would be better and there were precedents in builtin caches that it actually helped. I suspect this was the only one that actually helps performance while the other two may not. Hence it's better to split and measure individually.

lemire added a commit to lemire/node that referenced this pull request Jun 12, 2026
Builds on the zstd compression in nodejs#63861 by embedding a small zstd
dictionary trained on a diverse corpus of real modules, so each
small/medium compile-cache entry compresses better. Per entry we keep
the smaller of the plain and dictionary-assisted frame, so the
dictionary only ever helps.
- Add src/compile_cache_zstd.dict (16 KiB). It is trained on V8 code
caches harvested (via vm.compileFunction, the same shape the CJS
loader produces) from a diverse corpus: bundled npm packages, lib/,
tools/ and a few deps.
- Add tools/generate_compile_cache_dict.py and a node.gyp action that
generates compile_cache_zstd_dict.h into SHARED_INTERMEDIATE_DIR at
build time; no generated header is checked in. libnode include_dirs
updated to pick it up.
- Prepare the CDict/DDict once per process (shared across all handlers
and Workers, matching the lazy-context approach from nodejs#63861) and use
them in Persist() and ReadCacheFile(). Persist() compresses the plain
and dict frames into separate buffers and selects the smaller, so the
written bytes and recorded size always agree. The dictionary is only
tried for entries up to 256 KiB; larger blobs never benefit, so the
second compression is skipped to avoid wasted work. Falls back to
plain zstd if dictionary preparation fails.
- The dictionary is embedded in the binary because the compile cache
must be usable early, portably, and without extra filesystem state.
- No on-disk format change: dict-assisted frames carry the dictID, plain
frames carry none, and a single DDict decompresses both.
- Size, measured on data held out from training (per-entry min policy):
diverse modules go from ~1.87x (plain zstd) to ~2.44x with the
dictionary (~24% smaller on disk); on test/parallel, which is not in
the training corpus at all, ~1.74x -> ~2.22x (~22% smaller). A real
end-to-end run (npm --version, ~70 modules) is ~15% smaller. Read
time is unchanged and the extra write-time work is negligible.
- Add a multi-module write/read roundtrip test and a startup benchmark
(standard createBenchmark harness).
Builds on the zstd compression in nodejs#63861 by embedding a small zstd
dictionary trained on a diverse corpus of real modules, so each
small/medium compile-cache entry compresses better. Per entry we keep
the smaller of the plain and dictionary-assisted frame, so the
dictionary only ever helps.
- Add src/compile_cache_zstd.dict (16 KiB). It is trained on V8 code
caches harvested (via vm.compileFunction, the same shape the CJS
loader produces) from a diverse corpus: bundled npm packages, lib/,
tools/ and a few deps.
- Add tools/generate_compile_cache_dict.py and a node.gyp action that
generates compile_cache_zstd_dict.h into SHARED_INTERMEDIATE_DIR at
build time; no generated header is checked in. libnode include_dirs
updated to pick it up.
- Prepare the CDict/DDict once per process (shared across all handlers
and Workers, matching the lazy-context approach from nodejs#63861) and use
them in Persist() and ReadCacheFile(). Persist() compresses the plain
and dict frames into separate buffers and selects the smaller, so the
written bytes and recorded size always agree. The dictionary is only
tried for entries up to 256 KiB; larger blobs never benefit, so the
second compression is skipped to avoid wasted work. Falls back to
plain zstd if dictionary preparation fails.
- The dictionary is embedded in the binary because the compile cache
must be usable early, portably, and without extra filesystem state.
- No on-disk format change: dict-assisted frames carry the dictID, plain
frames carry none, and a single DDict decompresses both.
- Size, measured on data held out from training (per-entry min policy):
diverse modules go from ~1.87x (plain zstd) to ~2.44x with the
dictionary (~24% smaller on disk); on test/parallel, which is not in
the training corpus at all, ~1.74x -> ~2.22x (~22% smaller). A real
end-to-end run (npm --version, ~70 modules) is ~15% smaller. Read
time is unchanged and the extra write-time work is negligible.
- Add a multi-module write/read roundtrip test and a startup benchmark
(standard createBenchmark harness).

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If we do include dictionary files in our source tree, I'd say we should either include instructions for how to (at least approximately) re-create this file, or even generate it at build time entirely. strings src/compile_cache_zstd.dict results in a surprising amount of Node.js-core specificity, which seems surprising given that the commit that adds it claims it was trained on public npm packages.

(strings src/compile_cache_zstd.dict output in the fold)
wUUu2
_HOFJW
[Z)S
nding
options
Paz
types
PaN
process
versions
openssl
internal/process/pre_execution
PdR`
prepareMainThreadExecution
Pc6Wl
markBootstrapComplete
internal/options
"<";
__createBinding
__setModuleDefault
PaZY4J
birname
note
signatures
ObjectSetPrototypeOf
PromisePrototypeThen
PromiseWithResolvers
RegExpPrototypeExec
RegExpPrototypeSymbolReplace
StringPrototypeSplit
StringPrototypeToLowerCase
_createBlob
_createBlobFromFilePath
getDataObject
kMaxLength
TextDecoder
markTransferMode
Pbvyo
isAnyArrayBuffer
PcBS
isArrayBufferView
PaZ
require
PaN"
util
InvalidArgumentError
ConnectTimeoutError
SessionCache
maybeNormalizeConnectError
PromiseWithResolvers
nonOpStart
Pf&
writableStreamMarkFirstWriteRequestInFlight
createPromiseCallback
Pb:6k
nonOpWrite
$Pg>
writableStreamDefaultWriterEnsureReadyPromiseRejected
writableStreamDefaultControllerClose
DOMException
Pc^v
extractHighWaterMark
extractSizeAlgorithm
Pb~:
kEmptyObject
wgIgAygCACIAIAQgAWtqIQUgASAAa0ECaiEGAkADQCABLQAAIABButUAai0AAEcNAyAAQQJGDQEgAEEBaiEAIAQgAUEBaiIBRw0ACyADIAU2AgAMwwILIAMoAgQhACADQgA3AwAgAyAAIAZBAWoiARAuIgBFBEBB4wEhAgyqAgsgA0H1ATYCHCADIAE2AhQgAyAANgIMQQAhAgzCAgtB9AEhAiABIARGDcECIAMoAgAiACAEIAFraiEFIAEgAGtBAWohBgJAA0AgAS0AACAAQbjVAGotAABHDQIgAEEBRg0BIABBAWohACAEIAFBAWoiAUcNAAsgAyAFNgIADMICCyADQYEEOwEoIAMoAgQhACADQgA3AwAgAyAAIAZBAWoiARAuIgANAwwCCyADQQA2AgALQQAhAiADQQA2AhwgAyABNgIUIANB5R82AhAgA0EINgIMDL8CC0HVASECDKUCCyADQfMBNgIcIAMgATYCFCADIAA2AgxBACECDL0CC0EAIQACQCADKAI4IgJFDQAgAigCQ
memoMethod
perf
calculatedSizcludes
kTypes
uvErrmapGet
ArrayPrototypeIndexOf
classRegExp
MathMax
ArrayIsArray
isWindows
ObjectIsExtensible
isPermissionModelError
getExpectedArgumentLength
ArrayPrototypeJoin
kIsNodeError
get source
get ports
CloseEvent
get listening
header.js.map
@[@b
node:path
./large-numbers.js
"<";
escapeHTML
PcB_C+
convertChangesToXML
#noProxy
#ProxyAgent
#getProxy
#timeoutConnection
Pc~}
#drainPendingRequests
connect
Pbv2&_
addRequest
createSocket
.get
.get
getMilestoneTimestamp
getTimeOriginTimestamp
PbjK
internalBinding
performance
constants
ERR_INVALID_ARG_TYPE
ERR_INVALID_ARG_VALUE
Pb:e
ERR_OUT_OF_RANGE
isErrorStackTraceLimitWritable
Buffer
Pa&
inspect
validateBoolean
validateFunction
validateNumber
validateString
validateOneOf
PbbH
validateObject
validateInteger
isReadableStream
isWritableStream
isNodeStream
utilColors
PbJ%
lazyUtilColors
PbNn6
__importStar
DQCABLQAAQSBHDUggBCABQQFqIgFHDQALQSUhAwxpC0ElIQMMaAsgAi0ALUEBcQRAQcMBIQMMTwsgAigCBCEAQQAhAyACQQA2AgQgAiAAIAEQKSIABEAgAkEmNgIcIAIgADYCDCACIAFBAWo2AhQMaAsgAUEBaiEBDFwLIAFBAWohASACLwEwIgBBgAFxBEBBACEAAkAgAigCOCIDRQ0AIAMoAlQiA0UNACACIAMRAAAhAAsgAEUNBiAAQRVHDR8gAkEFNgIcIAIgATYCFCACQfkXNgIQIAJBFTYCDEEAIQMMZwsCQCAAQaAEcUGgBEcNACACLQAtQQJxDQBBACEDIAJBADYCHCACIAE2AhQgAkGWEzYCECACQQQ2AgwMZwsgAgJ/IAIvATBBFHFBFEYEQEEBIAItAChBAUYNARogAi8BMkHlAEYMAQsgAi0AKUEFRgs6AC5BACEAAkAgAigCOCIDRQ0AIAMoAiQiA0UNACACIAMRAAAhAAsCQAJAAkACQAJAIAAOFgIBAAQEBAQEBAQE
get ok
get statusText
get headers
get body
get bodyUsed
clone
parsePullArgs
PcNs
validateBackpressure
primordials
PcJ=
internal/encoding
TextEncoder
internal/errors
Pav
codes
internal/util
internal/util/types
internal/validators
onComplete
E`w PaN
onError
node:assert
../core/symbols
../core/errors
../core/util
Pab
beep
ArrayFrom
ArrayPrototypeFilter
ArrayPrototypeIncludes
ArrayPrototypeMap
ArrayPrototypePush
PcB,t
ArrayPrototypePushApply
ArrayPrototypeSlice
ObjectDefineProperty
Pb>c
ObjectKeys
PdRc
ObjectPrototypeHasOwnProperty
Pbf)
ReflectGet
SafeMap
SafeSet
StringPrototypeSlice
Error
PbFOT
_flushFlag
__esModule
.desc.get
stop
destroyer
primordials
internal/errors
Pav
codes
internal/streams/utils
__esModule
fs/promises
PaF-M
path
PaZ
require
Pa"/
module
__filename
__dirname

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@anonrig I'm going to mark this as un-resolved again, since there was no response here as far as I can tell

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, I didn't see this up until now and didn't pressed the "resolve" button at all. Will address your concerns.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

didn't pressed the "resolve" button at all

Github begged to differ:

Screenshot 2026-06-14 005643

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AI "babysit" mode might be the root cause

Comment threadsrc/compile_cache.h
compiler_cache_store_;
// Lazily created zstd decompression context, reused across cache reads
// to avoid recreating its workspace for every file.
ZSTD_DCtx_s* zstd_dctx_ = nullptr;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can this be held in a unique_ptr instead of a raw pointer?

Comment threadsrc/compile_cache.cc
if (cctx != nullptr) {
ZSTD_freeCCtx(cctx);
}
});

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

unique_ptr for this as well

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

c++Issues and PRs that require attention from people who are familiar with C++.lib / srcIssues and PRs involving general changes in the lib/ or src/ directories.needs-ciPRs that need a full CI run.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@anonrig@nodejs-github-bot@joyeecheung@lemire@jasnell@addaleax
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

src: improve compile cache performance and size - #63861

Open
anonrig wants to merge 3 commits into
nodejs:mainfrom
anonrig:compile-cache-perf
Open

src: improve compile cache performance and size#63861
anonrig wants to merge 3 commits into
nodejs:mainfrom
anonrig:compile-cache-perf

Conversation

@anonrig

Copy link
Copy Markdown
Member

Improves the on-disk compile cache (NODE_COMPILE_CACHE / module.enableCompileCache()):

  • Read path: read cache files with a single exactly-sized read (using the file size from fstat) instead of an exponentially growing buffer, which previously cost O(log N) syscalls/allocations and ~2N bytes of copying per file.
  • Size: compress the cache content on disk with zstd (level 1, prioritizing speed since persistence happens at shutdown), falling back to raw storage when not compressible. Shrinks cache directories ~2-4x and makes the crc32 integrity check cheaper since it now runs over the compressed bytes. The magic number is bumped so files in the old format are discarded as cache misses and overwritten in place.
  • Consume path: hand the cache to V8 through a non-owning CachedData wrapper (BufferNotOwned) instead of copying the entire buffer on every cache hit. The underlying buffer is owned by the cache entry, which outlives the synchronous compilation (same pattern as the vm cached-data path in node_contextify.cc).

Corrupted cache files keep degrading to silent cache misses and are regenerated; a corrupted size header can no longer cause an oversized allocation since the zstd frame content size is cross-checked first. Added test/parallel/test-compile-cache-corrupted.js covering bad magic, truncation, content bit-flips, and header corruption.

No public API or documented behavior changes; the file format is private to src/compile_cache.cc. Benchmark numbers (cache size and warm-startup timings) to follow in a comment.

This change was developed with AI assistance (see Co-authored-by trailer).

@nodejs-github-botnodejs-github-bot added c++ Issues and PRs that require attention from people who are familiar with C++. lib / src Issues and PRs involving general changes in the lib/ or src/ directories. needs-ci PRs that need a full CI run. labels Jun 11, 2026
@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

Review requested:

  • @nodejs/loaders
  • @nodejs/vm

@anonrig

Copy link
Copy Markdown
MemberAuthor

Verification results (macOS arm64, release build, vs. baseline at the merge-base):

Tests: all 23 parallel/test-compile-cache* pass (22 existing unchanged + the new corruption test).

On-disk size

ScenarioBaselineThis PRReduction
10 MB snapshot/typescript.js fixture (single big CJS file)1,818,380 B752,924 B2.42x
300 small/medium CJS modules274,316 B179,915 B1.52x

Warm-start, self-controlled (same binary, cache on vs. off, interleaved 20 runs, trimmed means — avoids binary-layout skew between builds):

ScenarioCache benefit (baseline)Cache benefit (this PR)
big-file+26.4 ms+25.0 ms
many-modules+1.7 ms+1.5 ms

Warm-start with a hot page cache is neutral within noise (~1 ms on the pathological single-10 MB-file case, which is dominated by one large zstd decompress; with cold page cache or slower storage the 2.4x smaller read wins). Cold-start adds one-time compression at persist (~35 ms for the 1.8 MB blob at level 1, proportionally less for typical files).

The second commit reuses zstd contexts (one ZSTD_DCtx on the handler, one ZSTD_CCtx across Persist()), which removed most of the per-file decompression overhead observed with one-shot contexts on the many-modules scenario.

Improve the compile cache by:
- Reading cache files with a single exactly-sized read using the file
size from fstat instead of reading into an exponentially growing
buffer, which previously cost O(log N) syscalls and allocations and
about 2N bytes of copying per file.
- Compressing the cache content on disk with zstd at level 1, falling
back to raw storage when the data is not compressible. This shrinks
cache directories by about 2-4x. The magic number is bumped so that
files in the old format are discarded as cache misses and then
overwritten in place.
- Handing the cache to V8 through a non-owning CachedData wrapper
instead of copying the whole buffer on every cache hit.
Corrupted cache files keep degrading to silent cache misses and are
regenerated, now covered by a regression test.
Co-authored-by: Grok <grok@x.ai>
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
@anonrig
anonrigforce-pushed the compile-cache-perf branch from 816f4ef to 0cf8818CompareJune 11, 2026 22:38
Creating and freeing a zstd context for every cache file costs more
than the (de)compression itself for small caches. Lazily create one
decompression context on the handler and reuse it across reads, and
share one compression context across all entries in Persist().
Co-authored-by: Grok <grok@x.ai>
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
@anonrig
anonrigforce-pushed the compile-cache-perf branch from 0cf8818 to f427fb3CompareJune 11, 2026 22:40
@joyeecheung

joyeecheung commented Jun 12, 2026

Copy link
Copy Markdown
Member

It seems the performance gains are within noise and it's mostly only the compression changes size? Can you split them into different PRs and measure them individually?

  1. I am not sure if the read changes actually gives any wins for small files - in the happy path where the cache is small, it's better to just read once and resize rather than stat and read (which are two file system calls instead of just one). Also there is a TOCTOU risk in doing stat, and fstat is not realiable across platforms, so the loop condition should not be gated on the fstat result but we must always read until EOF is reached in case the size is not accurate.
  2. It doesn't appear that compression alone does much to performance or it might actually hurt but just got compensated by other changes. In that case it's better to make that configurable and let users choose instead.
  3. Avoiding the copy would be better and there were precedents in builtin caches that it actually helped. I suspect this was the only one that actually helps performance while the other two may not. Hence it's better to split and measure individually.

lemire added a commit to lemire/node that referenced this pull request Jun 12, 2026
Builds on the zstd compression in nodejs#63861 by embedding a small zstd
dictionary trained on a diverse corpus of real modules, so each
small/medium compile-cache entry compresses better. Per entry we keep
the smaller of the plain and dictionary-assisted frame, so the
dictionary only ever helps.
- Add src/compile_cache_zstd.dict (16 KiB). It is trained on V8 code
caches harvested (via vm.compileFunction, the same shape the CJS
loader produces) from a diverse corpus: bundled npm packages, lib/,
tools/ and a few deps.
- Add tools/generate_compile_cache_dict.py and a node.gyp action that
generates compile_cache_zstd_dict.h into SHARED_INTERMEDIATE_DIR at
build time; no generated header is checked in. libnode include_dirs
updated to pick it up.
- Prepare the CDict/DDict once per process (shared across all handlers
and Workers, matching the lazy-context approach from nodejs#63861) and use
them in Persist() and ReadCacheFile(). Persist() compresses the plain
and dict frames into separate buffers and selects the smaller, so the
written bytes and recorded size always agree. The dictionary is only
tried for entries up to 256 KiB; larger blobs never benefit, so the
second compression is skipped to avoid wasted work. Falls back to
plain zstd if dictionary preparation fails.
- The dictionary is embedded in the binary because the compile cache
must be usable early, portably, and without extra filesystem state.
- No on-disk format change: dict-assisted frames carry the dictID, plain
frames carry none, and a single DDict decompresses both.
- Size, measured on data held out from training (per-entry min policy):
diverse modules go from ~1.87x (plain zstd) to ~2.44x with the
dictionary (~24% smaller on disk); on test/parallel, which is not in
the training corpus at all, ~1.74x -> ~2.22x (~22% smaller). A real
end-to-end run (npm --version, ~70 modules) is ~15% smaller. Read
time is unchanged and the extra write-time work is negligible.
- Add a multi-module write/read roundtrip test and a startup benchmark
(standard createBenchmark harness).
Builds on the zstd compression in nodejs#63861 by embedding a small zstd
dictionary trained on a diverse corpus of real modules, so each
small/medium compile-cache entry compresses better. Per entry we keep
the smaller of the plain and dictionary-assisted frame, so the
dictionary only ever helps.
- Add src/compile_cache_zstd.dict (16 KiB). It is trained on V8 code
caches harvested (via vm.compileFunction, the same shape the CJS
loader produces) from a diverse corpus: bundled npm packages, lib/,
tools/ and a few deps.
- Add tools/generate_compile_cache_dict.py and a node.gyp action that
generates compile_cache_zstd_dict.h into SHARED_INTERMEDIATE_DIR at
build time; no generated header is checked in. libnode include_dirs
updated to pick it up.
- Prepare the CDict/DDict once per process (shared across all handlers
and Workers, matching the lazy-context approach from nodejs#63861) and use
them in Persist() and ReadCacheFile(). Persist() compresses the plain
and dict frames into separate buffers and selects the smaller, so the
written bytes and recorded size always agree. The dictionary is only
tried for entries up to 256 KiB; larger blobs never benefit, so the
second compression is skipped to avoid wasted work. Falls back to
plain zstd if dictionary preparation fails.
- The dictionary is embedded in the binary because the compile cache
must be usable early, portably, and without extra filesystem state.
- No on-disk format change: dict-assisted frames carry the dictID, plain
frames carry none, and a single DDict decompresses both.
- Size, measured on data held out from training (per-entry min policy):
diverse modules go from ~1.87x (plain zstd) to ~2.44x with the
dictionary (~24% smaller on disk); on test/parallel, which is not in
the training corpus at all, ~1.74x -> ~2.22x (~22% smaller). A real
end-to-end run (npm --version, ~70 modules) is ~15% smaller. Read
time is unchanged and the extra write-time work is negligible.
- Add a multi-module write/read roundtrip test and a startup benchmark
(standard createBenchmark harness).

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If we do include dictionary files in our source tree, I'd say we should either include instructions for how to (at least approximately) re-create this file, or even generate it at build time entirely. strings src/compile_cache_zstd.dict results in a surprising amount of Node.js-core specificity, which seems surprising given that the commit that adds it claims it was trained on public npm packages.

(strings src/compile_cache_zstd.dict output in the fold)
wUUu2
_HOFJW
[Z)S
nding
options
Paz
types
PaN
process
versions
openssl
internal/process/pre_execution
PdR`
prepareMainThreadExecution
Pc6Wl
markBootstrapComplete
internal/options
"<";
__createBinding
__setModuleDefault
PaZY4J
birname
note
signatures
ObjectSetPrototypeOf
PromisePrototypeThen
PromiseWithResolvers
RegExpPrototypeExec
RegExpPrototypeSymbolReplace
StringPrototypeSplit
StringPrototypeToLowerCase
_createBlob
_createBlobFromFilePath
getDataObject
kMaxLength
TextDecoder
markTransferMode
Pbvyo
isAnyArrayBuffer
PcBS
isArrayBufferView
PaZ
require
PaN"
util
InvalidArgumentError
ConnectTimeoutError
SessionCache
maybeNormalizeConnectError
PromiseWithResolvers
nonOpStart
Pf&
writableStreamMarkFirstWriteRequestInFlight
createPromiseCallback
Pb:6k
nonOpWrite
$Pg>
writableStreamDefaultWriterEnsureReadyPromiseRejected
writableStreamDefaultControllerClose
DOMException
Pc^v
extractHighWaterMark
extractSizeAlgorithm
Pb~:
kEmptyObject
wgIgAygCACIAIAQgAWtqIQUgASAAa0ECaiEGAkADQCABLQAAIABButUAai0AAEcNAyAAQQJGDQEgAEEBaiEAIAQgAUEBaiIBRw0ACyADIAU2AgAMwwILIAMoAgQhACADQgA3AwAgAyAAIAZBAWoiARAuIgBFBEBB4wEhAgyqAgsgA0H1ATYCHCADIAE2AhQgAyAANgIMQQAhAgzCAgtB9AEhAiABIARGDcECIAMoAgAiACAEIAFraiEFIAEgAGtBAWohBgJAA0AgAS0AACAAQbjVAGotAABHDQIgAEEBRg0BIABBAWohACAEIAFBAWoiAUcNAAsgAyAFNgIADMICCyADQYEEOwEoIAMoAgQhACADQgA3AwAgAyAAIAZBAWoiARAuIgANAwwCCyADQQA2AgALQQAhAiADQQA2AhwgAyABNgIUIANB5R82AhAgA0EINgIMDL8CC0HVASECDKUCCyADQfMBNgIcIAMgATYCFCADIAA2AgxBACECDL0CC0EAIQACQCADKAI4IgJFDQAgAigCQ
memoMethod
perf
calculatedSizcludes
kTypes
uvErrmapGet
ArrayPrototypeIndexOf
classRegExp
MathMax
ArrayIsArray
isWindows
ObjectIsExtensible
isPermissionModelError
getExpectedArgumentLength
ArrayPrototypeJoin
kIsNodeError
get source
get ports
CloseEvent
get listening
header.js.map
@[@b
node:path
./large-numbers.js
"<";
escapeHTML
PcB_C+
convertChangesToXML
#noProxy
#ProxyAgent
#getProxy
#timeoutConnection
Pc~}
#drainPendingRequests
connect
Pbv2&_
addRequest
createSocket
.get
.get
getMilestoneTimestamp
getTimeOriginTimestamp
PbjK
internalBinding
performance
constants
ERR_INVALID_ARG_TYPE
ERR_INVALID_ARG_VALUE
Pb:e
ERR_OUT_OF_RANGE
isErrorStackTraceLimitWritable
Buffer
Pa&
inspect
validateBoolean
validateFunction
validateNumber
validateString
validateOneOf
PbbH
validateObject
validateInteger
isReadableStream
isWritableStream
isNodeStream
utilColors
PbJ%
lazyUtilColors
PbNn6
__importStar
DQCABLQAAQSBHDUggBCABQQFqIgFHDQALQSUhAwxpC0ElIQMMaAsgAi0ALUEBcQRAQcMBIQMMTwsgAigCBCEAQQAhAyACQQA2AgQgAiAAIAEQKSIABEAgAkEmNgIcIAIgADYCDCACIAFBAWo2AhQMaAsgAUEBaiEBDFwLIAFBAWohASACLwEwIgBBgAFxBEBBACEAAkAgAigCOCIDRQ0AIAMoAlQiA0UNACACIAMRAAAhAAsgAEUNBiAAQRVHDR8gAkEFNgIcIAIgATYCFCACQfkXNgIQIAJBFTYCDEEAIQMMZwsCQCAAQaAEcUGgBEcNACACLQAtQQJxDQBBACEDIAJBADYCHCACIAE2AhQgAkGWEzYCECACQQQ2AgwMZwsgAgJ/IAIvATBBFHFBFEYEQEEBIAItAChBAUYNARogAi8BMkHlAEYMAQsgAi0AKUEFRgs6AC5BACEAAkAgAigCOCIDRQ0AIAMoAiQiA0UNACACIAMRAAAhAAsCQAJAAkACQAJAIAAOFgIBAAQEBAQEBAQE
get ok
get statusText
get headers
get body
get bodyUsed
clone
parsePullArgs
PcNs
validateBackpressure
primordials
PcJ=
internal/encoding
TextEncoder
internal/errors
Pav
codes
internal/util
internal/util/types
internal/validators
onComplete
E`w PaN
onError
node:assert
../core/symbols
../core/errors
../core/util
Pab
beep
ArrayFrom
ArrayPrototypeFilter
ArrayPrototypeIncludes
ArrayPrototypeMap
ArrayPrototypePush
PcB,t
ArrayPrototypePushApply
ArrayPrototypeSlice
ObjectDefineProperty
Pb>c
ObjectKeys
PdRc
ObjectPrototypeHasOwnProperty
Pbf)
ReflectGet
SafeMap
SafeSet
StringPrototypeSlice
Error
PbFOT
_flushFlag
__esModule
.desc.get
stop
destroyer
primordials
internal/errors
Pav
codes
internal/streams/utils
__esModule
fs/promises
PaF-M
path
PaZ
require
Pa"/
module
__filename
__dirname

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@anonrig I'm going to mark this as un-resolved again, since there was no response here as far as I can tell

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, I didn't see this up until now and didn't pressed the "resolve" button at all. Will address your concerns.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

didn't pressed the "resolve" button at all

Github begged to differ:

Screenshot 2026-06-14 005643

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AI "babysit" mode might be the root cause

Comment threadsrc/compile_cache.h
compiler_cache_store_;
// Lazily created zstd decompression context, reused across cache reads
// to avoid recreating its workspace for every file.
ZSTD_DCtx_s* zstd_dctx_ = nullptr;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can this be held in a unique_ptr instead of a raw pointer?

Comment threadsrc/compile_cache.cc
if (cctx != nullptr) {
ZSTD_freeCCtx(cctx);
}
});

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

unique_ptr for this as well

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

c++Issues and PRs that require attention from people who are familiar with C++.lib / srcIssues and PRs involving general changes in the lib/ or src/ directories.needs-ciPRs that need a full CI run.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@anonrig@nodejs-github-bot@joyeecheung@lemire@jasnell@addaleax
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

src: improve compile cache performance and size - #63861

Open
anonrig wants to merge 3 commits into
nodejs:mainfrom
anonrig:compile-cache-perf
Open

src: improve compile cache performance and size#63861
anonrig wants to merge 3 commits into
nodejs:mainfrom
anonrig:compile-cache-perf

Conversation

@anonrig

Copy link
Copy Markdown
Member

Improves the on-disk compile cache (NODE_COMPILE_CACHE / module.enableCompileCache()):

  • Read path: read cache files with a single exactly-sized read (using the file size from fstat) instead of an exponentially growing buffer, which previously cost O(log N) syscalls/allocations and ~2N bytes of copying per file.
  • Size: compress the cache content on disk with zstd (level 1, prioritizing speed since persistence happens at shutdown), falling back to raw storage when not compressible. Shrinks cache directories ~2-4x and makes the crc32 integrity check cheaper since it now runs over the compressed bytes. The magic number is bumped so files in the old format are discarded as cache misses and overwritten in place.
  • Consume path: hand the cache to V8 through a non-owning CachedData wrapper (BufferNotOwned) instead of copying the entire buffer on every cache hit. The underlying buffer is owned by the cache entry, which outlives the synchronous compilation (same pattern as the vm cached-data path in node_contextify.cc).

Corrupted cache files keep degrading to silent cache misses and are regenerated; a corrupted size header can no longer cause an oversized allocation since the zstd frame content size is cross-checked first. Added test/parallel/test-compile-cache-corrupted.js covering bad magic, truncation, content bit-flips, and header corruption.

No public API or documented behavior changes; the file format is private to src/compile_cache.cc. Benchmark numbers (cache size and warm-startup timings) to follow in a comment.

This change was developed with AI assistance (see Co-authored-by trailer).

@nodejs-github-botnodejs-github-bot added c++ Issues and PRs that require attention from people who are familiar with C++. lib / src Issues and PRs involving general changes in the lib/ or src/ directories. needs-ci PRs that need a full CI run. labels Jun 11, 2026
@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

Review requested:

  • @nodejs/loaders
  • @nodejs/vm

@anonrig

Copy link
Copy Markdown
MemberAuthor

Verification results (macOS arm64, release build, vs. baseline at the merge-base):

Tests: all 23 parallel/test-compile-cache* pass (22 existing unchanged + the new corruption test).

On-disk size

ScenarioBaselineThis PRReduction
10 MB snapshot/typescript.js fixture (single big CJS file)1,818,380 B752,924 B2.42x
300 small/medium CJS modules274,316 B179,915 B1.52x

Warm-start, self-controlled (same binary, cache on vs. off, interleaved 20 runs, trimmed means — avoids binary-layout skew between builds):

ScenarioCache benefit (baseline)Cache benefit (this PR)
big-file+26.4 ms+25.0 ms
many-modules+1.7 ms+1.5 ms

Warm-start with a hot page cache is neutral within noise (~1 ms on the pathological single-10 MB-file case, which is dominated by one large zstd decompress; with cold page cache or slower storage the 2.4x smaller read wins). Cold-start adds one-time compression at persist (~35 ms for the 1.8 MB blob at level 1, proportionally less for typical files).

The second commit reuses zstd contexts (one ZSTD_DCtx on the handler, one ZSTD_CCtx across Persist()), which removed most of the per-file decompression overhead observed with one-shot contexts on the many-modules scenario.

Improve the compile cache by:
- Reading cache files with a single exactly-sized read using the file
size from fstat instead of reading into an exponentially growing
buffer, which previously cost O(log N) syscalls and allocations and
about 2N bytes of copying per file.
- Compressing the cache content on disk with zstd at level 1, falling
back to raw storage when the data is not compressible. This shrinks
cache directories by about 2-4x. The magic number is bumped so that
files in the old format are discarded as cache misses and then
overwritten in place.
- Handing the cache to V8 through a non-owning CachedData wrapper
instead of copying the whole buffer on every cache hit.
Corrupted cache files keep degrading to silent cache misses and are
regenerated, now covered by a regression test.
Co-authored-by: Grok <grok@x.ai>
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
@anonrig
anonrigforce-pushed the compile-cache-perf branch from 816f4ef to 0cf8818CompareJune 11, 2026 22:38
Creating and freeing a zstd context for every cache file costs more
than the (de)compression itself for small caches. Lazily create one
decompression context on the handler and reuse it across reads, and
share one compression context across all entries in Persist().
Co-authored-by: Grok <grok@x.ai>
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
@anonrig
anonrigforce-pushed the compile-cache-perf branch from 0cf8818 to f427fb3CompareJune 11, 2026 22:40
@joyeecheung

joyeecheung commented Jun 12, 2026

Copy link
Copy Markdown
Member

It seems the performance gains are within noise and it's mostly only the compression changes size? Can you split them into different PRs and measure them individually?

  1. I am not sure if the read changes actually gives any wins for small files - in the happy path where the cache is small, it's better to just read once and resize rather than stat and read (which are two file system calls instead of just one). Also there is a TOCTOU risk in doing stat, and fstat is not realiable across platforms, so the loop condition should not be gated on the fstat result but we must always read until EOF is reached in case the size is not accurate.
  2. It doesn't appear that compression alone does much to performance or it might actually hurt but just got compensated by other changes. In that case it's better to make that configurable and let users choose instead.
  3. Avoiding the copy would be better and there were precedents in builtin caches that it actually helped. I suspect this was the only one that actually helps performance while the other two may not. Hence it's better to split and measure individually.

lemire added a commit to lemire/node that referenced this pull request Jun 12, 2026
Builds on the zstd compression in nodejs#63861 by embedding a small zstd
dictionary trained on a diverse corpus of real modules, so each
small/medium compile-cache entry compresses better. Per entry we keep
the smaller of the plain and dictionary-assisted frame, so the
dictionary only ever helps.
- Add src/compile_cache_zstd.dict (16 KiB). It is trained on V8 code
caches harvested (via vm.compileFunction, the same shape the CJS
loader produces) from a diverse corpus: bundled npm packages, lib/,
tools/ and a few deps.
- Add tools/generate_compile_cache_dict.py and a node.gyp action that
generates compile_cache_zstd_dict.h into SHARED_INTERMEDIATE_DIR at
build time; no generated header is checked in. libnode include_dirs
updated to pick it up.
- Prepare the CDict/DDict once per process (shared across all handlers
and Workers, matching the lazy-context approach from nodejs#63861) and use
them in Persist() and ReadCacheFile(). Persist() compresses the plain
and dict frames into separate buffers and selects the smaller, so the
written bytes and recorded size always agree. The dictionary is only
tried for entries up to 256 KiB; larger blobs never benefit, so the
second compression is skipped to avoid wasted work. Falls back to
plain zstd if dictionary preparation fails.
- The dictionary is embedded in the binary because the compile cache
must be usable early, portably, and without extra filesystem state.
- No on-disk format change: dict-assisted frames carry the dictID, plain
frames carry none, and a single DDict decompresses both.
- Size, measured on data held out from training (per-entry min policy):
diverse modules go from ~1.87x (plain zstd) to ~2.44x with the
dictionary (~24% smaller on disk); on test/parallel, which is not in
the training corpus at all, ~1.74x -> ~2.22x (~22% smaller). A real
end-to-end run (npm --version, ~70 modules) is ~15% smaller. Read
time is unchanged and the extra write-time work is negligible.
- Add a multi-module write/read roundtrip test and a startup benchmark
(standard createBenchmark harness).
Builds on the zstd compression in nodejs#63861 by embedding a small zstd
dictionary trained on a diverse corpus of real modules, so each
small/medium compile-cache entry compresses better. Per entry we keep
the smaller of the plain and dictionary-assisted frame, so the
dictionary only ever helps.
- Add src/compile_cache_zstd.dict (16 KiB). It is trained on V8 code
caches harvested (via vm.compileFunction, the same shape the CJS
loader produces) from a diverse corpus: bundled npm packages, lib/,
tools/ and a few deps.
- Add tools/generate_compile_cache_dict.py and a node.gyp action that
generates compile_cache_zstd_dict.h into SHARED_INTERMEDIATE_DIR at
build time; no generated header is checked in. libnode include_dirs
updated to pick it up.
- Prepare the CDict/DDict once per process (shared across all handlers
and Workers, matching the lazy-context approach from nodejs#63861) and use
them in Persist() and ReadCacheFile(). Persist() compresses the plain
and dict frames into separate buffers and selects the smaller, so the
written bytes and recorded size always agree. The dictionary is only
tried for entries up to 256 KiB; larger blobs never benefit, so the
second compression is skipped to avoid wasted work. Falls back to
plain zstd if dictionary preparation fails.
- The dictionary is embedded in the binary because the compile cache
must be usable early, portably, and without extra filesystem state.
- No on-disk format change: dict-assisted frames carry the dictID, plain
frames carry none, and a single DDict decompresses both.
- Size, measured on data held out from training (per-entry min policy):
diverse modules go from ~1.87x (plain zstd) to ~2.44x with the
dictionary (~24% smaller on disk); on test/parallel, which is not in
the training corpus at all, ~1.74x -> ~2.22x (~22% smaller). A real
end-to-end run (npm --version, ~70 modules) is ~15% smaller. Read
time is unchanged and the extra write-time work is negligible.
- Add a multi-module write/read roundtrip test and a startup benchmark
(standard createBenchmark harness).

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If we do include dictionary files in our source tree, I'd say we should either include instructions for how to (at least approximately) re-create this file, or even generate it at build time entirely. strings src/compile_cache_zstd.dict results in a surprising amount of Node.js-core specificity, which seems surprising given that the commit that adds it claims it was trained on public npm packages.

(strings src/compile_cache_zstd.dict output in the fold)
wUUu2
_HOFJW
[Z)S
nding
options
Paz
types
PaN
process
versions
openssl
internal/process/pre_execution
PdR`
prepareMainThreadExecution
Pc6Wl
markBootstrapComplete
internal/options
"<";
__createBinding
__setModuleDefault
PaZY4J
birname
note
signatures
ObjectSetPrototypeOf
PromisePrototypeThen
PromiseWithResolvers
RegExpPrototypeExec
RegExpPrototypeSymbolReplace
StringPrototypeSplit
StringPrototypeToLowerCase
_createBlob
_createBlobFromFilePath
getDataObject
kMaxLength
TextDecoder
markTransferMode
Pbvyo
isAnyArrayBuffer
PcBS
isArrayBufferView
PaZ
require
PaN"
util
InvalidArgumentError
ConnectTimeoutError
SessionCache
maybeNormalizeConnectError
PromiseWithResolvers
nonOpStart
Pf&
writableStreamMarkFirstWriteRequestInFlight
createPromiseCallback
Pb:6k
nonOpWrite
$Pg>
writableStreamDefaultWriterEnsureReadyPromiseRejected
writableStreamDefaultControllerClose
DOMException
Pc^v
extractHighWaterMark
extractSizeAlgorithm
Pb~:
kEmptyObject
wgIgAygCACIAIAQgAWtqIQUgASAAa0ECaiEGAkADQCABLQAAIABButUAai0AAEcNAyAAQQJGDQEgAEEBaiEAIAQgAUEBaiIBRw0ACyADIAU2AgAMwwILIAMoAgQhACADQgA3AwAgAyAAIAZBAWoiARAuIgBFBEBB4wEhAgyqAgsgA0H1ATYCHCADIAE2AhQgAyAANgIMQQAhAgzCAgtB9AEhAiABIARGDcECIAMoAgAiACAEIAFraiEFIAEgAGtBAWohBgJAA0AgAS0AACAAQbjVAGotAABHDQIgAEEBRg0BIABBAWohACAEIAFBAWoiAUcNAAsgAyAFNgIADMICCyADQYEEOwEoIAMoAgQhACADQgA3AwAgAyAAIAZBAWoiARAuIgANAwwCCyADQQA2AgALQQAhAiADQQA2AhwgAyABNgIUIANB5R82AhAgA0EINgIMDL8CC0HVASECDKUCCyADQfMBNgIcIAMgATYCFCADIAA2AgxBACECDL0CC0EAIQACQCADKAI4IgJFDQAgAigCQ
memoMethod
perf
calculatedSizcludes
kTypes
uvErrmapGet
ArrayPrototypeIndexOf
classRegExp
MathMax
ArrayIsArray
isWindows
ObjectIsExtensible
isPermissionModelError
getExpectedArgumentLength
ArrayPrototypeJoin
kIsNodeError
get source
get ports
CloseEvent
get listening
header.js.map
@[@b
node:path
./large-numbers.js
"<";
escapeHTML
PcB_C+
convertChangesToXML
#noProxy
#ProxyAgent
#getProxy
#timeoutConnection
Pc~}
#drainPendingRequests
connect
Pbv2&_
addRequest
createSocket
.get
.get
getMilestoneTimestamp
getTimeOriginTimestamp
PbjK
internalBinding
performance
constants
ERR_INVALID_ARG_TYPE
ERR_INVALID_ARG_VALUE
Pb:e
ERR_OUT_OF_RANGE
isErrorStackTraceLimitWritable
Buffer
Pa&
inspect
validateBoolean
validateFunction
validateNumber
validateString
validateOneOf
PbbH
validateObject
validateInteger
isReadableStream
isWritableStream
isNodeStream
utilColors
PbJ%
lazyUtilColors
PbNn6
__importStar
DQCABLQAAQSBHDUggBCABQQFqIgFHDQALQSUhAwxpC0ElIQMMaAsgAi0ALUEBcQRAQcMBIQMMTwsgAigCBCEAQQAhAyACQQA2AgQgAiAAIAEQKSIABEAgAkEmNgIcIAIgADYCDCACIAFBAWo2AhQMaAsgAUEBaiEBDFwLIAFBAWohASACLwEwIgBBgAFxBEBBACEAAkAgAigCOCIDRQ0AIAMoAlQiA0UNACACIAMRAAAhAAsgAEUNBiAAQRVHDR8gAkEFNgIcIAIgATYCFCACQfkXNgIQIAJBFTYCDEEAIQMMZwsCQCAAQaAEcUGgBEcNACACLQAtQQJxDQBBACEDIAJBADYCHCACIAE2AhQgAkGWEzYCECACQQQ2AgwMZwsgAgJ/IAIvATBBFHFBFEYEQEEBIAItAChBAUYNARogAi8BMkHlAEYMAQsgAi0AKUEFRgs6AC5BACEAAkAgAigCOCIDRQ0AIAMoAiQiA0UNACACIAMRAAAhAAsCQAJAAkACQAJAIAAOFgIBAAQEBAQEBAQE
get ok
get statusText
get headers
get body
get bodyUsed
clone
parsePullArgs
PcNs
validateBackpressure
primordials
PcJ=
internal/encoding
TextEncoder
internal/errors
Pav
codes
internal/util
internal/util/types
internal/validators
onComplete
E`w PaN
onError
node:assert
../core/symbols
../core/errors
../core/util
Pab
beep
ArrayFrom
ArrayPrototypeFilter
ArrayPrototypeIncludes
ArrayPrototypeMap
ArrayPrototypePush
PcB,t
ArrayPrototypePushApply
ArrayPrototypeSlice
ObjectDefineProperty
Pb>c
ObjectKeys
PdRc
ObjectPrototypeHasOwnProperty
Pbf)
ReflectGet
SafeMap
SafeSet
StringPrototypeSlice
Error
PbFOT
_flushFlag
__esModule
.desc.get
stop
destroyer
primordials
internal/errors
Pav
codes
internal/streams/utils
__esModule
fs/promises
PaF-M
path
PaZ
require
Pa"/
module
__filename
__dirname

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@anonrig I'm going to mark this as un-resolved again, since there was no response here as far as I can tell

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, I didn't see this up until now and didn't pressed the "resolve" button at all. Will address your concerns.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

didn't pressed the "resolve" button at all

Github begged to differ:

Screenshot 2026-06-14 005643

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AI "babysit" mode might be the root cause

Comment threadsrc/compile_cache.h
compiler_cache_store_;
// Lazily created zstd decompression context, reused across cache reads
// to avoid recreating its workspace for every file.
ZSTD_DCtx_s* zstd_dctx_ = nullptr;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can this be held in a unique_ptr instead of a raw pointer?

Comment threadsrc/compile_cache.cc
if (cctx != nullptr) {
ZSTD_freeCCtx(cctx);
}
});

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

unique_ptr for this as well

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

c++Issues and PRs that require attention from people who are familiar with C++.lib / srcIssues and PRs involving general changes in the lib/ or src/ directories.needs-ciPRs that need a full CI run.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@anonrig@nodejs-github-bot@joyeecheung@lemire@jasnell@addaleax
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

src: improve compile cache performance and size - #63861

Open
anonrig wants to merge 3 commits into
nodejs:mainfrom
anonrig:compile-cache-perf
Open

src: improve compile cache performance and size#63861
anonrig wants to merge 3 commits into
nodejs:mainfrom
anonrig:compile-cache-perf

Conversation

@anonrig

Copy link
Copy Markdown
Member

Improves the on-disk compile cache (NODE_COMPILE_CACHE / module.enableCompileCache()):

  • Read path: read cache files with a single exactly-sized read (using the file size from fstat) instead of an exponentially growing buffer, which previously cost O(log N) syscalls/allocations and ~2N bytes of copying per file.
  • Size: compress the cache content on disk with zstd (level 1, prioritizing speed since persistence happens at shutdown), falling back to raw storage when not compressible. Shrinks cache directories ~2-4x and makes the crc32 integrity check cheaper since it now runs over the compressed bytes. The magic number is bumped so files in the old format are discarded as cache misses and overwritten in place.
  • Consume path: hand the cache to V8 through a non-owning CachedData wrapper (BufferNotOwned) instead of copying the entire buffer on every cache hit. The underlying buffer is owned by the cache entry, which outlives the synchronous compilation (same pattern as the vm cached-data path in node_contextify.cc).

Corrupted cache files keep degrading to silent cache misses and are regenerated; a corrupted size header can no longer cause an oversized allocation since the zstd frame content size is cross-checked first. Added test/parallel/test-compile-cache-corrupted.js covering bad magic, truncation, content bit-flips, and header corruption.

No public API or documented behavior changes; the file format is private to src/compile_cache.cc. Benchmark numbers (cache size and warm-startup timings) to follow in a comment.

This change was developed with AI assistance (see Co-authored-by trailer).

@nodejs-github-botnodejs-github-bot added c++ Issues and PRs that require attention from people who are familiar with C++. lib / src Issues and PRs involving general changes in the lib/ or src/ directories. needs-ci PRs that need a full CI run. labels Jun 11, 2026
@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

Review requested:

  • @nodejs/loaders
  • @nodejs/vm

@anonrig

Copy link
Copy Markdown
MemberAuthor

Verification results (macOS arm64, release build, vs. baseline at the merge-base):

Tests: all 23 parallel/test-compile-cache* pass (22 existing unchanged + the new corruption test).

On-disk size

ScenarioBaselineThis PRReduction
10 MB snapshot/typescript.js fixture (single big CJS file)1,818,380 B752,924 B2.42x
300 small/medium CJS modules274,316 B179,915 B1.52x

Warm-start, self-controlled (same binary, cache on vs. off, interleaved 20 runs, trimmed means — avoids binary-layout skew between builds):

ScenarioCache benefit (baseline)Cache benefit (this PR)
big-file+26.4 ms+25.0 ms
many-modules+1.7 ms+1.5 ms

Warm-start with a hot page cache is neutral within noise (~1 ms on the pathological single-10 MB-file case, which is dominated by one large zstd decompress; with cold page cache or slower storage the 2.4x smaller read wins). Cold-start adds one-time compression at persist (~35 ms for the 1.8 MB blob at level 1, proportionally less for typical files).

The second commit reuses zstd contexts (one ZSTD_DCtx on the handler, one ZSTD_CCtx across Persist()), which removed most of the per-file decompression overhead observed with one-shot contexts on the many-modules scenario.

Improve the compile cache by:
- Reading cache files with a single exactly-sized read using the file
size from fstat instead of reading into an exponentially growing
buffer, which previously cost O(log N) syscalls and allocations and
about 2N bytes of copying per file.
- Compressing the cache content on disk with zstd at level 1, falling
back to raw storage when the data is not compressible. This shrinks
cache directories by about 2-4x. The magic number is bumped so that
files in the old format are discarded as cache misses and then
overwritten in place.
- Handing the cache to V8 through a non-owning CachedData wrapper
instead of copying the whole buffer on every cache hit.
Corrupted cache files keep degrading to silent cache misses and are
regenerated, now covered by a regression test.
Co-authored-by: Grok <grok@x.ai>
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
@anonrig
anonrigforce-pushed the compile-cache-perf branch from 816f4ef to 0cf8818CompareJune 11, 2026 22:38
Creating and freeing a zstd context for every cache file costs more
than the (de)compression itself for small caches. Lazily create one
decompression context on the handler and reuse it across reads, and
share one compression context across all entries in Persist().
Co-authored-by: Grok <grok@x.ai>
Signed-off-by: Yagiz Nizipli <yagiz@nizipli.com>
@anonrig
anonrigforce-pushed the compile-cache-perf branch from 0cf8818 to f427fb3CompareJune 11, 2026 22:40
@joyeecheung

joyeecheung commented Jun 12, 2026

Copy link
Copy Markdown
Member

It seems the performance gains are within noise and it's mostly only the compression changes size? Can you split them into different PRs and measure them individually?

  1. I am not sure if the read changes actually gives any wins for small files - in the happy path where the cache is small, it's better to just read once and resize rather than stat and read (which are two file system calls instead of just one). Also there is a TOCTOU risk in doing stat, and fstat is not realiable across platforms, so the loop condition should not be gated on the fstat result but we must always read until EOF is reached in case the size is not accurate.
  2. It doesn't appear that compression alone does much to performance or it might actually hurt but just got compensated by other changes. In that case it's better to make that configurable and let users choose instead.
  3. Avoiding the copy would be better and there were precedents in builtin caches that it actually helped. I suspect this was the only one that actually helps performance while the other two may not. Hence it's better to split and measure individually.

lemire added a commit to lemire/node that referenced this pull request Jun 12, 2026
Builds on the zstd compression in nodejs#63861 by embedding a small zstd
dictionary trained on a diverse corpus of real modules, so each
small/medium compile-cache entry compresses better. Per entry we keep
the smaller of the plain and dictionary-assisted frame, so the
dictionary only ever helps.
- Add src/compile_cache_zstd.dict (16 KiB). It is trained on V8 code
caches harvested (via vm.compileFunction, the same shape the CJS
loader produces) from a diverse corpus: bundled npm packages, lib/,
tools/ and a few deps.
- Add tools/generate_compile_cache_dict.py and a node.gyp action that
generates compile_cache_zstd_dict.h into SHARED_INTERMEDIATE_DIR at
build time; no generated header is checked in. libnode include_dirs
updated to pick it up.
- Prepare the CDict/DDict once per process (shared across all handlers
and Workers, matching the lazy-context approach from nodejs#63861) and use
them in Persist() and ReadCacheFile(). Persist() compresses the plain
and dict frames into separate buffers and selects the smaller, so the
written bytes and recorded size always agree. The dictionary is only
tried for entries up to 256 KiB; larger blobs never benefit, so the
second compression is skipped to avoid wasted work. Falls back to
plain zstd if dictionary preparation fails.
- The dictionary is embedded in the binary because the compile cache
must be usable early, portably, and without extra filesystem state.
- No on-disk format change: dict-assisted frames carry the dictID, plain
frames carry none, and a single DDict decompresses both.
- Size, measured on data held out from training (per-entry min policy):
diverse modules go from ~1.87x (plain zstd) to ~2.44x with the
dictionary (~24% smaller on disk); on test/parallel, which is not in
the training corpus at all, ~1.74x -> ~2.22x (~22% smaller). A real
end-to-end run (npm --version, ~70 modules) is ~15% smaller. Read
time is unchanged and the extra write-time work is negligible.
- Add a multi-module write/read roundtrip test and a startup benchmark
(standard createBenchmark harness).
Builds on the zstd compression in nodejs#63861 by embedding a small zstd
dictionary trained on a diverse corpus of real modules, so each
small/medium compile-cache entry compresses better. Per entry we keep
the smaller of the plain and dictionary-assisted frame, so the
dictionary only ever helps.
- Add src/compile_cache_zstd.dict (16 KiB). It is trained on V8 code
caches harvested (via vm.compileFunction, the same shape the CJS
loader produces) from a diverse corpus: bundled npm packages, lib/,
tools/ and a few deps.
- Add tools/generate_compile_cache_dict.py and a node.gyp action that
generates compile_cache_zstd_dict.h into SHARED_INTERMEDIATE_DIR at
build time; no generated header is checked in. libnode include_dirs
updated to pick it up.
- Prepare the CDict/DDict once per process (shared across all handlers
and Workers, matching the lazy-context approach from nodejs#63861) and use
them in Persist() and ReadCacheFile(). Persist() compresses the plain
and dict frames into separate buffers and selects the smaller, so the
written bytes and recorded size always agree. The dictionary is only
tried for entries up to 256 KiB; larger blobs never benefit, so the
second compression is skipped to avoid wasted work. Falls back to
plain zstd if dictionary preparation fails.
- The dictionary is embedded in the binary because the compile cache
must be usable early, portably, and without extra filesystem state.
- No on-disk format change: dict-assisted frames carry the dictID, plain
frames carry none, and a single DDict decompresses both.
- Size, measured on data held out from training (per-entry min policy):
diverse modules go from ~1.87x (plain zstd) to ~2.44x with the
dictionary (~24% smaller on disk); on test/parallel, which is not in
the training corpus at all, ~1.74x -> ~2.22x (~22% smaller). A real
end-to-end run (npm --version, ~70 modules) is ~15% smaller. Read
time is unchanged and the extra write-time work is negligible.
- Add a multi-module write/read roundtrip test and a startup benchmark
(standard createBenchmark harness).

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If we do include dictionary files in our source tree, I'd say we should either include instructions for how to (at least approximately) re-create this file, or even generate it at build time entirely. strings src/compile_cache_zstd.dict results in a surprising amount of Node.js-core specificity, which seems surprising given that the commit that adds it claims it was trained on public npm packages.

(strings src/compile_cache_zstd.dict output in the fold)
wUUu2
_HOFJW
[Z)S
nding
options
Paz
types
PaN
process
versions
openssl
internal/process/pre_execution
PdR`
prepareMainThreadExecution
Pc6Wl
markBootstrapComplete
internal/options
"<";
__createBinding
__setModuleDefault
PaZY4J
birname
note
signatures
ObjectSetPrototypeOf
PromisePrototypeThen
PromiseWithResolvers
RegExpPrototypeExec
RegExpPrototypeSymbolReplace
StringPrototypeSplit
StringPrototypeToLowerCase
_createBlob
_createBlobFromFilePath
getDataObject
kMaxLength
TextDecoder
markTransferMode
Pbvyo
isAnyArrayBuffer
PcBS
isArrayBufferView
PaZ
require
PaN"
util
InvalidArgumentError
ConnectTimeoutError
SessionCache
maybeNormalizeConnectError
PromiseWithResolvers
nonOpStart
Pf&
writableStreamMarkFirstWriteRequestInFlight
createPromiseCallback
Pb:6k
nonOpWrite
$Pg>
writableStreamDefaultWriterEnsureReadyPromiseRejected
writableStreamDefaultControllerClose
DOMException
Pc^v
extractHighWaterMark
extractSizeAlgorithm
Pb~:
kEmptyObject
wgIgAygCACIAIAQgAWtqIQUgASAAa0ECaiEGAkADQCABLQAAIABButUAai0AAEcNAyAAQQJGDQEgAEEBaiEAIAQgAUEBaiIBRw0ACyADIAU2AgAMwwILIAMoAgQhACADQgA3AwAgAyAAIAZBAWoiARAuIgBFBEBB4wEhAgyqAgsgA0H1ATYCHCADIAE2AhQgAyAANgIMQQAhAgzCAgtB9AEhAiABIARGDcECIAMoAgAiACAEIAFraiEFIAEgAGtBAWohBgJAA0AgAS0AACAAQbjVAGotAABHDQIgAEEBRg0BIABBAWohACAEIAFBAWoiAUcNAAsgAyAFNgIADMICCyADQYEEOwEoIAMoAgQhACADQgA3AwAgAyAAIAZBAWoiARAuIgANAwwCCyADQQA2AgALQQAhAiADQQA2AhwgAyABNgIUIANB5R82AhAgA0EINgIMDL8CC0HVASECDKUCCyADQfMBNgIcIAMgATYCFCADIAA2AgxBACECDL0CC0EAIQACQCADKAI4IgJFDQAgAigCQ
memoMethod
perf
calculatedSizcludes
kTypes
uvErrmapGet
ArrayPrototypeIndexOf
classRegExp
MathMax
ArrayIsArray
isWindows
ObjectIsExtensible
isPermissionModelError
getExpectedArgumentLength
ArrayPrototypeJoin
kIsNodeError
get source
get ports
CloseEvent
get listening
header.js.map
@[@b
node:path
./large-numbers.js
"<";
escapeHTML
PcB_C+
convertChangesToXML
#noProxy
#ProxyAgent
#getProxy
#timeoutConnection
Pc~}
#drainPendingRequests
connect
Pbv2&_
addRequest
createSocket
.get
.get
getMilestoneTimestamp
getTimeOriginTimestamp
PbjK
internalBinding
performance
constants
ERR_INVALID_ARG_TYPE
ERR_INVALID_ARG_VALUE
Pb:e
ERR_OUT_OF_RANGE
isErrorStackTraceLimitWritable
Buffer
Pa&
inspect
validateBoolean
validateFunction
validateNumber
validateString
validateOneOf
PbbH
validateObject
validateInteger
isReadableStream
isWritableStream
isNodeStream
utilColors
PbJ%
lazyUtilColors
PbNn6
__importStar
DQCABLQAAQSBHDUggBCABQQFqIgFHDQALQSUhAwxpC0ElIQMMaAsgAi0ALUEBcQRAQcMBIQMMTwsgAigCBCEAQQAhAyACQQA2AgQgAiAAIAEQKSIABEAgAkEmNgIcIAIgADYCDCACIAFBAWo2AhQMaAsgAUEBaiEBDFwLIAFBAWohASACLwEwIgBBgAFxBEBBACEAAkAgAigCOCIDRQ0AIAMoAlQiA0UNACACIAMRAAAhAAsgAEUNBiAAQRVHDR8gAkEFNgIcIAIgATYCFCACQfkXNgIQIAJBFTYCDEEAIQMMZwsCQCAAQaAEcUGgBEcNACACLQAtQQJxDQBBACEDIAJBADYCHCACIAE2AhQgAkGWEzYCECACQQQ2AgwMZwsgAgJ/IAIvATBBFHFBFEYEQEEBIAItAChBAUYNARogAi8BMkHlAEYMAQsgAi0AKUEFRgs6AC5BACEAAkAgAigCOCIDRQ0AIAMoAiQiA0UNACACIAMRAAAhAAsCQAJAAkACQAJAIAAOFgIBAAQEBAQEBAQE
get ok
get statusText
get headers
get body
get bodyUsed
clone
parsePullArgs
PcNs
validateBackpressure
primordials
PcJ=
internal/encoding
TextEncoder
internal/errors
Pav
codes
internal/util
internal/util/types
internal/validators
onComplete
E`w PaN
onError
node:assert
../core/symbols
../core/errors
../core/util
Pab
beep
ArrayFrom
ArrayPrototypeFilter
ArrayPrototypeIncludes
ArrayPrototypeMap
ArrayPrototypePush
PcB,t
ArrayPrototypePushApply
ArrayPrototypeSlice
ObjectDefineProperty
Pb>c
ObjectKeys
PdRc
ObjectPrototypeHasOwnProperty
Pbf)
ReflectGet
SafeMap
SafeSet
StringPrototypeSlice
Error
PbFOT
_flushFlag
__esModule
.desc.get
stop
destroyer
primordials
internal/errors
Pav
codes
internal/streams/utils
__esModule
fs/promises
PaF-M
path
PaZ
require
Pa"/
module
__filename
__dirname

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@anonrig I'm going to mark this as un-resolved again, since there was no response here as far as I can tell

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, I didn't see this up until now and didn't pressed the "resolve" button at all. Will address your concerns.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

didn't pressed the "resolve" button at all

Github begged to differ:

Screenshot 2026-06-14 005643

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AI "babysit" mode might be the root cause

Comment threadsrc/compile_cache.h
compiler_cache_store_;
// Lazily created zstd decompression context, reused across cache reads
// to avoid recreating its workspace for every file.
ZSTD_DCtx_s* zstd_dctx_ = nullptr;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can this be held in a unique_ptr instead of a raw pointer?

Comment threadsrc/compile_cache.cc
if (cctx != nullptr) {
ZSTD_freeCCtx(cctx);
}
});

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

unique_ptr for this as well

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

c++Issues and PRs that require attention from people who are familiar with C++.lib / srcIssues and PRs involving general changes in the lib/ or src/ directories.needs-ciPRs that need a full CI run.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@anonrig@nodejs-github-bot@joyeecheung@lemire@jasnell@addaleax