Adds zip support to the zlib module - #45651

Closed
arcanis wants to merge 2 commits into
nodejs:mainfrom
arcanis:mael/zip
Closed

Adds zip support to the zlib module#45651
arcanis wants to merge 2 commits into
nodejs:mainfrom
arcanis:mael/zip

Conversation

@arcanis

@arcanisarcanis commented Nov 28, 2022

Copy link
Copy Markdown
Contributor

Ref #45434

This PR adds a new ZipArchive class to the zlib module, which can be used to read and write content from zip archives. Its current API looks like this:

constfs=require(`fs`);const{ZipArchive}=require(`zlib`);// Creates a new in-memory archiveconstzip=newZipArchive();zip.addFile(`hello`,fs.readFileSync(__filename));constdata=zip.digest();fs.writeFileSync(`./archive.zip`,data);// The data obtained from `digest` can also be reopenedconstzip2=newZipArchive(data);console.log(zip2.getEntries());console.log(zip2.getEntries({withFileTypes: true}));constcontent=zip2.readEntry(0);console.log(content);

Maintenance cost

I kept the feature scope limited enough to cover most of the use cases but without increasing the maintenance cost or build cost. A few things have been cut from what the libzip would allow:

  • Opening files directly from the filesystem isn't supported, because it would bypass node:fs. The current API only works with memory buffers (according to my tests it doesn't have any negative impact even when compared to the wasm API which went through file descriptors).

  • Encryption isn't supported, because it's unclear how it should integrate with node:crypto. There's room for follow-up, but it didn't seem a required feature for the first iteration.

Performances

Keep in mind that raw performances aren't the main reason why zip support is important to have as a native feature. The speedup is nice, the simplified garbage collection is very nice, but the real benefit is having a stable cross-platform way to bundle files between platforms. It will be useful for cache mechanisms, transfer algorithms, user CLI generation, and more.

Still, I made some reasonable checks to make sure that no use case regressed. Size of the binary before / after:

before 89604801 85.45MB
after 89780257 85.62MB (+171KB)

Performance-wise, using Yarn as benchmark, the results show native being ~2x faster than wasm (keep in mind the wasm implementation isn't the most popular zip library; projects using jszip will see significantly larger differences):

YARN_EXPERIMENT_NATIVE_ZIPFS=0 PKG=gatsby
➤ YN0000: └ Completed in 15s 806ms
YARN_EXPERIMENT_NATIVE_ZIPFS=1 PKG=gatsby
➤ YN0000: └ Completed in 7s 404ms
YARN_EXPERIMENT_NATIVE_ZIPFS=0 PKG=typescript
➤ YN0000: └ Completed in 13s 676ms
YARN_EXPERIMENT_NATIVE_ZIPFS=1 PKG=typescript
➤ YN0000: └ Completed in 5s 351ms
YARN_EXPERIMENT_NATIVE_ZIPFS=0 PKG=next
➤ YN0000: └ Completed in 5s 923ms
YARN_EXPERIMENT_NATIVE_ZIPFS=1 PKG=next
➤ YN0000: └ Completed in 3s 512ms

To Do

  • Improve the documentation
  • Add more regression tests
  • Benchmark against the WASM libzip
  • API Bikeshedding

@nodejs-github-botnodejs-github-bot added build Issues and PRs related to Node.js builds or CI infrastructure. dependencies PRs that add, update, or configure Node.js dependencies. meta Issues and PRs related to the general management of the project. needs-ci PRs that need a full CI run. labels Nov 28, 2022
@arcanisarcanis mentioned this pull request Nov 28, 2022

@mscdexmscdex left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

-1 I know this is an unpopular opinion, but I'm not convinced this should live in node core. Sure it's a common file format, but then again so are tar, rar, 7z, zstd, xz, and others, which also don't belong in node core.

I feel like adding such modules to node core is further leading to feature creep. I get that other platforms like PHP and such may have zip modules, but they also include a ton of other modules that make them "kitchen sink" platforms, which I would hate to see node.js become.

@GeoffreyBooth

Copy link
Copy Markdown
Member

@mscdex Could the discussion of whether we should do this stay in #45434? And this PR discussion can focus on the implementation.

@addaleaxaddaleax left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'd appreciate some early review to let me know places where the code isn't compatible with Node.js' standards.

Might be jumping the gun here since yeah, there should be a discussion about whether this should happen at all, but, sure, gave it a first look. I do concur with @mscdex's concerns, fwiw.

Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc
}

void ZipArchive::MemoryInfo(MemoryTracker* tracker) const {
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

(might want to fill this out)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

How precise does it have to be? I updated the code to track the size of the input buffer + the buffer of any file that gets added later, but it doesn't include the small-ish libzip overhead, and gets confused if the same file is modified multiple times. If the value is indicative it might be fine?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It doesn’t have to be super precise, but it should be usable for debugging. You’ll probably want to track buf_ here via tracker->TrackField("buf", buf_);. If you can’t track or estimate memory owned by libzip (including memory for added entries?) then it’s probably fine to omit it, rather than to give numbers that are potentially very inaccurate (e.g. after repeated AddEntry() + DeleteEntry() calls).

Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc
}

zip_int64_t file_index = zip_file_add(zip->zip_, *path, file_source, ZIP_FL_OVERWRITE | ZIP_FL_ENC_UTF_8);
CHECK_GE(file_index, 0);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What if this call fails? Likely also applies elsewhere.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For now the code is overly strict and if any function fails, the program aborts on the CHECK_GE call. I have to replace most these calls by something that would just throw instead.

Comment threadsrc/node_zip.cc Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
@anonrig

Copy link
Copy Markdown
Member

Despite the comments from other reviewers, my main concern is about the usage of the buffer module. I strongly believe that the public API should consume necessary native buffers (TypedArray) instead of Node.js buffers.

@arcanis

Copy link
Copy Markdown
ContributorAuthor

I used Buffer since that's what the other main Node APIs tended to use (fs, crypto, zlib) - wouldn't it be surprising for users to return a regular typed array in just this API?

@anonrig

Copy link
Copy Markdown
Member

I used Buffer since that's what the other main Node APIs tended to use (fs, crypto, zlib) - wouldn't it be surprising for users to return a regular typed array in just this API?

There is an initiative to use native buffers in the new public APIs (referencing my personal talks with @addaleax and @jasnell), a @nodejs/tsc member can clarify if this is still a thing.

@addaleax

Copy link
Copy Markdown
Member

@anonrig I don’t know if that’s the best way forward here, but since this feels like a very broad conversation (“Should new Node.js APIs return Uint8Array or should they return Buffer?”), maybe it’s also best to handle that separately from this specific PR?

@jasnell

Copy link
Copy Markdown
Member

In this case, I think Buffer is fine given that it is consistent with the rest of the zlib module. I can see us eventually making a call on avoiding Buffer in the future (or standardizing on it) but this is not the place to decide that

@jasnell

Copy link
Copy Markdown
Member

Is a new top level module what we want here? As opposed to adding this to zlib? I know it's not based on zlib but neither is brotli.

@arcanis

Copy link
Copy Markdown
ContributorAuthor

I think @GeoffreyBooth had the same feedback. I don't have a strong opinion there, perhaps zlib would indeed better match user expectations.

@GeoffreyBooth

Copy link
Copy Markdown
Member

I would put it in zlib. In the future we could consider a friendlier name as an alias for zlib, like node:compression or something, but that's for later.

@tniessen

Copy link
Copy Markdown
Member

There is an initiative to use native buffers in the new public APIs

Just leaving this reference here: #41588

Comment threaddoc/api/zip.md Outdated
@tniessen

Copy link
Copy Markdown
Member

In the future we could consider a friendlier name as an alias for zlib, like node:compression or something, but that's for later.

That might make some sense for zip, which inherently supports compression, but if we do add other archive formats that don't, it won't fit. Also, except for zip, compression and archive formats are orthogonal even if related topics.

@arcanisarcanis changed the title Adds prototype zip moduleAdds prototype archive moduleDec 18, 2022
@arcanisarcanis changed the title Adds prototype archive moduleAdds zip support to the zlib moduleJan 6, 2023
@arcanis
arcanis marked this pull request as ready for review January 6, 2023 21:30
Comment threaddoc/api/zlib.md
Comment on lines +257 to +268
## Compressing multiple files together

<!-- YAML
added: REPLACEME
-->

The `zlib` library provides ways to compress individual objects, but not to
aggregate multiple ones into a single file suitable for redistribution (what
is often called archival).

To this end, `node:zip` provides the `ZipArchive` class which allows to create,
read, and modify zip archives:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I assume this section needs updating? Because of references to node:zip etc.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Indeed, it's a typo, the archive is now part of node:zlib (would it make sense to have a check in lint-md that all node:something identifiers must be valid?)

@bakkot

Copy link
Copy Markdown
Contributor

New effort at #64339

@panva

Copy link
Copy Markdown
Member

Superseded by #64339

@panvapanva closed this Aug 28, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

buildIssues and PRs related to Node.js builds or CI infrastructure.dependenciesPRs that add, update, or configure Node.js dependencies.metaIssues and PRs related to the general management of the project.needs-ciPRs that need a full CI run.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

12 participants

@arcanis@GeoffreyBooth@anonrig@addaleax@jasnell@tniessen@bakkot@panva@mcollina@mscdex@aduh95@nodejs-github-bot
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Adds zip support to the zlib module - #45651

Closed
arcanis wants to merge 2 commits into
nodejs:mainfrom
arcanis:mael/zip
Closed

Adds zip support to the zlib module#45651
arcanis wants to merge 2 commits into
nodejs:mainfrom
arcanis:mael/zip

Conversation

@arcanis

@arcanisarcanis commented Nov 28, 2022

Copy link
Copy Markdown
Contributor

Ref #45434

This PR adds a new ZipArchive class to the zlib module, which can be used to read and write content from zip archives. Its current API looks like this:

constfs=require(`fs`);const{ZipArchive}=require(`zlib`);// Creates a new in-memory archiveconstzip=newZipArchive();zip.addFile(`hello`,fs.readFileSync(__filename));constdata=zip.digest();fs.writeFileSync(`./archive.zip`,data);// The data obtained from `digest` can also be reopenedconstzip2=newZipArchive(data);console.log(zip2.getEntries());console.log(zip2.getEntries({withFileTypes: true}));constcontent=zip2.readEntry(0);console.log(content);

Maintenance cost

I kept the feature scope limited enough to cover most of the use cases but without increasing the maintenance cost or build cost. A few things have been cut from what the libzip would allow:

  • Opening files directly from the filesystem isn't supported, because it would bypass node:fs. The current API only works with memory buffers (according to my tests it doesn't have any negative impact even when compared to the wasm API which went through file descriptors).

  • Encryption isn't supported, because it's unclear how it should integrate with node:crypto. There's room for follow-up, but it didn't seem a required feature for the first iteration.

Performances

Keep in mind that raw performances aren't the main reason why zip support is important to have as a native feature. The speedup is nice, the simplified garbage collection is very nice, but the real benefit is having a stable cross-platform way to bundle files between platforms. It will be useful for cache mechanisms, transfer algorithms, user CLI generation, and more.

Still, I made some reasonable checks to make sure that no use case regressed. Size of the binary before / after:

before 89604801 85.45MB
after 89780257 85.62MB (+171KB)

Performance-wise, using Yarn as benchmark, the results show native being ~2x faster than wasm (keep in mind the wasm implementation isn't the most popular zip library; projects using jszip will see significantly larger differences):

YARN_EXPERIMENT_NATIVE_ZIPFS=0 PKG=gatsby
➤ YN0000: └ Completed in 15s 806ms
YARN_EXPERIMENT_NATIVE_ZIPFS=1 PKG=gatsby
➤ YN0000: └ Completed in 7s 404ms
YARN_EXPERIMENT_NATIVE_ZIPFS=0 PKG=typescript
➤ YN0000: └ Completed in 13s 676ms
YARN_EXPERIMENT_NATIVE_ZIPFS=1 PKG=typescript
➤ YN0000: └ Completed in 5s 351ms
YARN_EXPERIMENT_NATIVE_ZIPFS=0 PKG=next
➤ YN0000: └ Completed in 5s 923ms
YARN_EXPERIMENT_NATIVE_ZIPFS=1 PKG=next
➤ YN0000: └ Completed in 3s 512ms

To Do

  • Improve the documentation
  • Add more regression tests
  • Benchmark against the WASM libzip
  • API Bikeshedding

@nodejs-github-botnodejs-github-bot added build Issues and PRs related to Node.js builds or CI infrastructure. dependencies PRs that add, update, or configure Node.js dependencies. meta Issues and PRs related to the general management of the project. needs-ci PRs that need a full CI run. labels Nov 28, 2022
@arcanisarcanis mentioned this pull request Nov 28, 2022

@mscdexmscdex left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

-1 I know this is an unpopular opinion, but I'm not convinced this should live in node core. Sure it's a common file format, but then again so are tar, rar, 7z, zstd, xz, and others, which also don't belong in node core.

I feel like adding such modules to node core is further leading to feature creep. I get that other platforms like PHP and such may have zip modules, but they also include a ton of other modules that make them "kitchen sink" platforms, which I would hate to see node.js become.

@GeoffreyBooth

Copy link
Copy Markdown
Member

@mscdex Could the discussion of whether we should do this stay in #45434? And this PR discussion can focus on the implementation.

@addaleaxaddaleax left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'd appreciate some early review to let me know places where the code isn't compatible with Node.js' standards.

Might be jumping the gun here since yeah, there should be a discussion about whether this should happen at all, but, sure, gave it a first look. I do concur with @mscdex's concerns, fwiw.

Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc
}

void ZipArchive::MemoryInfo(MemoryTracker* tracker) const {
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

(might want to fill this out)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

How precise does it have to be? I updated the code to track the size of the input buffer + the buffer of any file that gets added later, but it doesn't include the small-ish libzip overhead, and gets confused if the same file is modified multiple times. If the value is indicative it might be fine?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It doesn’t have to be super precise, but it should be usable for debugging. You’ll probably want to track buf_ here via tracker->TrackField("buf", buf_);. If you can’t track or estimate memory owned by libzip (including memory for added entries?) then it’s probably fine to omit it, rather than to give numbers that are potentially very inaccurate (e.g. after repeated AddEntry() + DeleteEntry() calls).

Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc
}

zip_int64_t file_index = zip_file_add(zip->zip_, *path, file_source, ZIP_FL_OVERWRITE | ZIP_FL_ENC_UTF_8);
CHECK_GE(file_index, 0);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What if this call fails? Likely also applies elsewhere.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For now the code is overly strict and if any function fails, the program aborts on the CHECK_GE call. I have to replace most these calls by something that would just throw instead.

Comment threadsrc/node_zip.cc Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
@anonrig

Copy link
Copy Markdown
Member

Despite the comments from other reviewers, my main concern is about the usage of the buffer module. I strongly believe that the public API should consume necessary native buffers (TypedArray) instead of Node.js buffers.

@arcanis

Copy link
Copy Markdown
ContributorAuthor

I used Buffer since that's what the other main Node APIs tended to use (fs, crypto, zlib) - wouldn't it be surprising for users to return a regular typed array in just this API?

@anonrig

Copy link
Copy Markdown
Member

I used Buffer since that's what the other main Node APIs tended to use (fs, crypto, zlib) - wouldn't it be surprising for users to return a regular typed array in just this API?

There is an initiative to use native buffers in the new public APIs (referencing my personal talks with @addaleax and @jasnell), a @nodejs/tsc member can clarify if this is still a thing.

@addaleax

Copy link
Copy Markdown
Member

@anonrig I don’t know if that’s the best way forward here, but since this feels like a very broad conversation (“Should new Node.js APIs return Uint8Array or should they return Buffer?”), maybe it’s also best to handle that separately from this specific PR?

@jasnell

Copy link
Copy Markdown
Member

In this case, I think Buffer is fine given that it is consistent with the rest of the zlib module. I can see us eventually making a call on avoiding Buffer in the future (or standardizing on it) but this is not the place to decide that

@jasnell

Copy link
Copy Markdown
Member

Is a new top level module what we want here? As opposed to adding this to zlib? I know it's not based on zlib but neither is brotli.

@arcanis

Copy link
Copy Markdown
ContributorAuthor

I think @GeoffreyBooth had the same feedback. I don't have a strong opinion there, perhaps zlib would indeed better match user expectations.

@GeoffreyBooth

Copy link
Copy Markdown
Member

I would put it in zlib. In the future we could consider a friendlier name as an alias for zlib, like node:compression or something, but that's for later.

@tniessen

Copy link
Copy Markdown
Member

There is an initiative to use native buffers in the new public APIs

Just leaving this reference here: #41588

Comment threaddoc/api/zip.md Outdated
@tniessen

Copy link
Copy Markdown
Member

In the future we could consider a friendlier name as an alias for zlib, like node:compression or something, but that's for later.

That might make some sense for zip, which inherently supports compression, but if we do add other archive formats that don't, it won't fit. Also, except for zip, compression and archive formats are orthogonal even if related topics.

@arcanisarcanis changed the title Adds prototype zip moduleAdds prototype archive moduleDec 18, 2022
@arcanisarcanis changed the title Adds prototype archive moduleAdds zip support to the zlib moduleJan 6, 2023
@arcanis
arcanis marked this pull request as ready for review January 6, 2023 21:30
Comment threaddoc/api/zlib.md
Comment on lines +257 to +268
## Compressing multiple files together

<!-- YAML
added: REPLACEME
-->

The `zlib` library provides ways to compress individual objects, but not to
aggregate multiple ones into a single file suitable for redistribution (what
is often called archival).

To this end, `node:zip` provides the `ZipArchive` class which allows to create,
read, and modify zip archives:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I assume this section needs updating? Because of references to node:zip etc.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Indeed, it's a typo, the archive is now part of node:zlib (would it make sense to have a check in lint-md that all node:something identifiers must be valid?)

@bakkot

Copy link
Copy Markdown
Contributor

New effort at #64339

@panva

Copy link
Copy Markdown
Member

Superseded by #64339

@panvapanva closed this Aug 28, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

buildIssues and PRs related to Node.js builds or CI infrastructure.dependenciesPRs that add, update, or configure Node.js dependencies.metaIssues and PRs related to the general management of the project.needs-ciPRs that need a full CI run.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

12 participants

@arcanis@GeoffreyBooth@anonrig@addaleax@jasnell@tniessen@bakkot@panva@mcollina@mscdex@aduh95@nodejs-github-bot
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Adds zip support to the zlib module - #45651

Closed
arcanis wants to merge 2 commits into
nodejs:mainfrom
arcanis:mael/zip
Closed

Adds zip support to the zlib module#45651
arcanis wants to merge 2 commits into
nodejs:mainfrom
arcanis:mael/zip

Conversation

@arcanis

@arcanisarcanis commented Nov 28, 2022

Copy link
Copy Markdown
Contributor

Ref #45434

This PR adds a new ZipArchive class to the zlib module, which can be used to read and write content from zip archives. Its current API looks like this:

constfs=require(`fs`);const{ZipArchive}=require(`zlib`);// Creates a new in-memory archiveconstzip=newZipArchive();zip.addFile(`hello`,fs.readFileSync(__filename));constdata=zip.digest();fs.writeFileSync(`./archive.zip`,data);// The data obtained from `digest` can also be reopenedconstzip2=newZipArchive(data);console.log(zip2.getEntries());console.log(zip2.getEntries({withFileTypes: true}));constcontent=zip2.readEntry(0);console.log(content);

Maintenance cost

I kept the feature scope limited enough to cover most of the use cases but without increasing the maintenance cost or build cost. A few things have been cut from what the libzip would allow:

  • Opening files directly from the filesystem isn't supported, because it would bypass node:fs. The current API only works with memory buffers (according to my tests it doesn't have any negative impact even when compared to the wasm API which went through file descriptors).

  • Encryption isn't supported, because it's unclear how it should integrate with node:crypto. There's room for follow-up, but it didn't seem a required feature for the first iteration.

Performances

Keep in mind that raw performances aren't the main reason why zip support is important to have as a native feature. The speedup is nice, the simplified garbage collection is very nice, but the real benefit is having a stable cross-platform way to bundle files between platforms. It will be useful for cache mechanisms, transfer algorithms, user CLI generation, and more.

Still, I made some reasonable checks to make sure that no use case regressed. Size of the binary before / after:

before 89604801 85.45MB
after 89780257 85.62MB (+171KB)

Performance-wise, using Yarn as benchmark, the results show native being ~2x faster than wasm (keep in mind the wasm implementation isn't the most popular zip library; projects using jszip will see significantly larger differences):

YARN_EXPERIMENT_NATIVE_ZIPFS=0 PKG=gatsby
➤ YN0000: └ Completed in 15s 806ms
YARN_EXPERIMENT_NATIVE_ZIPFS=1 PKG=gatsby
➤ YN0000: └ Completed in 7s 404ms
YARN_EXPERIMENT_NATIVE_ZIPFS=0 PKG=typescript
➤ YN0000: └ Completed in 13s 676ms
YARN_EXPERIMENT_NATIVE_ZIPFS=1 PKG=typescript
➤ YN0000: └ Completed in 5s 351ms
YARN_EXPERIMENT_NATIVE_ZIPFS=0 PKG=next
➤ YN0000: └ Completed in 5s 923ms
YARN_EXPERIMENT_NATIVE_ZIPFS=1 PKG=next
➤ YN0000: └ Completed in 3s 512ms

To Do

  • Improve the documentation
  • Add more regression tests
  • Benchmark against the WASM libzip
  • API Bikeshedding

@nodejs-github-botnodejs-github-bot added build Issues and PRs related to Node.js builds or CI infrastructure. dependencies PRs that add, update, or configure Node.js dependencies. meta Issues and PRs related to the general management of the project. needs-ci PRs that need a full CI run. labels Nov 28, 2022
@arcanisarcanis mentioned this pull request Nov 28, 2022

@mscdexmscdex left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

-1 I know this is an unpopular opinion, but I'm not convinced this should live in node core. Sure it's a common file format, but then again so are tar, rar, 7z, zstd, xz, and others, which also don't belong in node core.

I feel like adding such modules to node core is further leading to feature creep. I get that other platforms like PHP and such may have zip modules, but they also include a ton of other modules that make them "kitchen sink" platforms, which I would hate to see node.js become.

@GeoffreyBooth

Copy link
Copy Markdown
Member

@mscdex Could the discussion of whether we should do this stay in #45434? And this PR discussion can focus on the implementation.

@addaleaxaddaleax left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'd appreciate some early review to let me know places where the code isn't compatible with Node.js' standards.

Might be jumping the gun here since yeah, there should be a discussion about whether this should happen at all, but, sure, gave it a first look. I do concur with @mscdex's concerns, fwiw.

Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc
}

void ZipArchive::MemoryInfo(MemoryTracker* tracker) const {
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

(might want to fill this out)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

How precise does it have to be? I updated the code to track the size of the input buffer + the buffer of any file that gets added later, but it doesn't include the small-ish libzip overhead, and gets confused if the same file is modified multiple times. If the value is indicative it might be fine?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It doesn’t have to be super precise, but it should be usable for debugging. You’ll probably want to track buf_ here via tracker->TrackField("buf", buf_);. If you can’t track or estimate memory owned by libzip (including memory for added entries?) then it’s probably fine to omit it, rather than to give numbers that are potentially very inaccurate (e.g. after repeated AddEntry() + DeleteEntry() calls).

Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc
}

zip_int64_t file_index = zip_file_add(zip->zip_, *path, file_source, ZIP_FL_OVERWRITE | ZIP_FL_ENC_UTF_8);
CHECK_GE(file_index, 0);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What if this call fails? Likely also applies elsewhere.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For now the code is overly strict and if any function fails, the program aborts on the CHECK_GE call. I have to replace most these calls by something that would just throw instead.

Comment threadsrc/node_zip.cc Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
@anonrig

Copy link
Copy Markdown
Member

Despite the comments from other reviewers, my main concern is about the usage of the buffer module. I strongly believe that the public API should consume necessary native buffers (TypedArray) instead of Node.js buffers.

@arcanis

Copy link
Copy Markdown
ContributorAuthor

I used Buffer since that's what the other main Node APIs tended to use (fs, crypto, zlib) - wouldn't it be surprising for users to return a regular typed array in just this API?

@anonrig

Copy link
Copy Markdown
Member

I used Buffer since that's what the other main Node APIs tended to use (fs, crypto, zlib) - wouldn't it be surprising for users to return a regular typed array in just this API?

There is an initiative to use native buffers in the new public APIs (referencing my personal talks with @addaleax and @jasnell), a @nodejs/tsc member can clarify if this is still a thing.

@addaleax

Copy link
Copy Markdown
Member

@anonrig I don’t know if that’s the best way forward here, but since this feels like a very broad conversation (“Should new Node.js APIs return Uint8Array or should they return Buffer?”), maybe it’s also best to handle that separately from this specific PR?

@jasnell

Copy link
Copy Markdown
Member

In this case, I think Buffer is fine given that it is consistent with the rest of the zlib module. I can see us eventually making a call on avoiding Buffer in the future (or standardizing on it) but this is not the place to decide that

@jasnell

Copy link
Copy Markdown
Member

Is a new top level module what we want here? As opposed to adding this to zlib? I know it's not based on zlib but neither is brotli.

@arcanis

Copy link
Copy Markdown
ContributorAuthor

I think @GeoffreyBooth had the same feedback. I don't have a strong opinion there, perhaps zlib would indeed better match user expectations.

@GeoffreyBooth

Copy link
Copy Markdown
Member

I would put it in zlib. In the future we could consider a friendlier name as an alias for zlib, like node:compression or something, but that's for later.

@tniessen

Copy link
Copy Markdown
Member

There is an initiative to use native buffers in the new public APIs

Just leaving this reference here: #41588

Comment threaddoc/api/zip.md Outdated
@tniessen

Copy link
Copy Markdown
Member

In the future we could consider a friendlier name as an alias for zlib, like node:compression or something, but that's for later.

That might make some sense for zip, which inherently supports compression, but if we do add other archive formats that don't, it won't fit. Also, except for zip, compression and archive formats are orthogonal even if related topics.

@arcanisarcanis changed the title Adds prototype zip moduleAdds prototype archive moduleDec 18, 2022
@arcanisarcanis changed the title Adds prototype archive moduleAdds zip support to the zlib moduleJan 6, 2023
@arcanis
arcanis marked this pull request as ready for review January 6, 2023 21:30
Comment threaddoc/api/zlib.md
Comment on lines +257 to +268
## Compressing multiple files together

<!-- YAML
added: REPLACEME
-->

The `zlib` library provides ways to compress individual objects, but not to
aggregate multiple ones into a single file suitable for redistribution (what
is often called archival).

To this end, `node:zip` provides the `ZipArchive` class which allows to create,
read, and modify zip archives:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I assume this section needs updating? Because of references to node:zip etc.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Indeed, it's a typo, the archive is now part of node:zlib (would it make sense to have a check in lint-md that all node:something identifiers must be valid?)

@bakkot

Copy link
Copy Markdown
Contributor

New effort at #64339

@panva

Copy link
Copy Markdown
Member

Superseded by #64339

@panvapanva closed this Aug 28, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

buildIssues and PRs related to Node.js builds or CI infrastructure.dependenciesPRs that add, update, or configure Node.js dependencies.metaIssues and PRs related to the general management of the project.needs-ciPRs that need a full CI run.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

12 participants

@arcanis@GeoffreyBooth@anonrig@addaleax@jasnell@tniessen@bakkot@panva@mcollina@mscdex@aduh95@nodejs-github-bot
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Adds zip support to the zlib module - #45651

Closed
arcanis wants to merge 2 commits into
nodejs:mainfrom
arcanis:mael/zip
Closed

Adds zip support to the zlib module#45651
arcanis wants to merge 2 commits into
nodejs:mainfrom
arcanis:mael/zip

Conversation

@arcanis

@arcanisarcanis commented Nov 28, 2022

Copy link
Copy Markdown
Contributor

Ref #45434

This PR adds a new ZipArchive class to the zlib module, which can be used to read and write content from zip archives. Its current API looks like this:

constfs=require(`fs`);const{ZipArchive}=require(`zlib`);// Creates a new in-memory archiveconstzip=newZipArchive();zip.addFile(`hello`,fs.readFileSync(__filename));constdata=zip.digest();fs.writeFileSync(`./archive.zip`,data);// The data obtained from `digest` can also be reopenedconstzip2=newZipArchive(data);console.log(zip2.getEntries());console.log(zip2.getEntries({withFileTypes: true}));constcontent=zip2.readEntry(0);console.log(content);

Maintenance cost

I kept the feature scope limited enough to cover most of the use cases but without increasing the maintenance cost or build cost. A few things have been cut from what the libzip would allow:

  • Opening files directly from the filesystem isn't supported, because it would bypass node:fs. The current API only works with memory buffers (according to my tests it doesn't have any negative impact even when compared to the wasm API which went through file descriptors).

  • Encryption isn't supported, because it's unclear how it should integrate with node:crypto. There's room for follow-up, but it didn't seem a required feature for the first iteration.

Performances

Keep in mind that raw performances aren't the main reason why zip support is important to have as a native feature. The speedup is nice, the simplified garbage collection is very nice, but the real benefit is having a stable cross-platform way to bundle files between platforms. It will be useful for cache mechanisms, transfer algorithms, user CLI generation, and more.

Still, I made some reasonable checks to make sure that no use case regressed. Size of the binary before / after:

before 89604801 85.45MB
after 89780257 85.62MB (+171KB)

Performance-wise, using Yarn as benchmark, the results show native being ~2x faster than wasm (keep in mind the wasm implementation isn't the most popular zip library; projects using jszip will see significantly larger differences):

YARN_EXPERIMENT_NATIVE_ZIPFS=0 PKG=gatsby
➤ YN0000: └ Completed in 15s 806ms
YARN_EXPERIMENT_NATIVE_ZIPFS=1 PKG=gatsby
➤ YN0000: └ Completed in 7s 404ms
YARN_EXPERIMENT_NATIVE_ZIPFS=0 PKG=typescript
➤ YN0000: └ Completed in 13s 676ms
YARN_EXPERIMENT_NATIVE_ZIPFS=1 PKG=typescript
➤ YN0000: └ Completed in 5s 351ms
YARN_EXPERIMENT_NATIVE_ZIPFS=0 PKG=next
➤ YN0000: └ Completed in 5s 923ms
YARN_EXPERIMENT_NATIVE_ZIPFS=1 PKG=next
➤ YN0000: └ Completed in 3s 512ms

To Do

  • Improve the documentation
  • Add more regression tests
  • Benchmark against the WASM libzip
  • API Bikeshedding

@nodejs-github-botnodejs-github-bot added build Issues and PRs related to Node.js builds or CI infrastructure. dependencies PRs that add, update, or configure Node.js dependencies. meta Issues and PRs related to the general management of the project. needs-ci PRs that need a full CI run. labels Nov 28, 2022
@arcanisarcanis mentioned this pull request Nov 28, 2022

@mscdexmscdex left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

-1 I know this is an unpopular opinion, but I'm not convinced this should live in node core. Sure it's a common file format, but then again so are tar, rar, 7z, zstd, xz, and others, which also don't belong in node core.

I feel like adding such modules to node core is further leading to feature creep. I get that other platforms like PHP and such may have zip modules, but they also include a ton of other modules that make them "kitchen sink" platforms, which I would hate to see node.js become.

@GeoffreyBooth

Copy link
Copy Markdown
Member

@mscdex Could the discussion of whether we should do this stay in #45434? And this PR discussion can focus on the implementation.

@addaleaxaddaleax left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'd appreciate some early review to let me know places where the code isn't compatible with Node.js' standards.

Might be jumping the gun here since yeah, there should be a discussion about whether this should happen at all, but, sure, gave it a first look. I do concur with @mscdex's concerns, fwiw.

Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc
}

void ZipArchive::MemoryInfo(MemoryTracker* tracker) const {
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

(might want to fill this out)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

How precise does it have to be? I updated the code to track the size of the input buffer + the buffer of any file that gets added later, but it doesn't include the small-ish libzip overhead, and gets confused if the same file is modified multiple times. If the value is indicative it might be fine?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It doesn’t have to be super precise, but it should be usable for debugging. You’ll probably want to track buf_ here via tracker->TrackField("buf", buf_);. If you can’t track or estimate memory owned by libzip (including memory for added entries?) then it’s probably fine to omit it, rather than to give numbers that are potentially very inaccurate (e.g. after repeated AddEntry() + DeleteEntry() calls).

Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc
}

zip_int64_t file_index = zip_file_add(zip->zip_, *path, file_source, ZIP_FL_OVERWRITE | ZIP_FL_ENC_UTF_8);
CHECK_GE(file_index, 0);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What if this call fails? Likely also applies elsewhere.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For now the code is overly strict and if any function fails, the program aborts on the CHECK_GE call. I have to replace most these calls by something that would just throw instead.

Comment threadsrc/node_zip.cc Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
@anonrig

Copy link
Copy Markdown
Member

Despite the comments from other reviewers, my main concern is about the usage of the buffer module. I strongly believe that the public API should consume necessary native buffers (TypedArray) instead of Node.js buffers.

@arcanis

Copy link
Copy Markdown
ContributorAuthor

I used Buffer since that's what the other main Node APIs tended to use (fs, crypto, zlib) - wouldn't it be surprising for users to return a regular typed array in just this API?

@anonrig

Copy link
Copy Markdown
Member

I used Buffer since that's what the other main Node APIs tended to use (fs, crypto, zlib) - wouldn't it be surprising for users to return a regular typed array in just this API?

There is an initiative to use native buffers in the new public APIs (referencing my personal talks with @addaleax and @jasnell), a @nodejs/tsc member can clarify if this is still a thing.

@addaleax

Copy link
Copy Markdown
Member

@anonrig I don’t know if that’s the best way forward here, but since this feels like a very broad conversation (“Should new Node.js APIs return Uint8Array or should they return Buffer?”), maybe it’s also best to handle that separately from this specific PR?

@jasnell

Copy link
Copy Markdown
Member

In this case, I think Buffer is fine given that it is consistent with the rest of the zlib module. I can see us eventually making a call on avoiding Buffer in the future (or standardizing on it) but this is not the place to decide that

@jasnell

Copy link
Copy Markdown
Member

Is a new top level module what we want here? As opposed to adding this to zlib? I know it's not based on zlib but neither is brotli.

@arcanis

Copy link
Copy Markdown
ContributorAuthor

I think @GeoffreyBooth had the same feedback. I don't have a strong opinion there, perhaps zlib would indeed better match user expectations.

@GeoffreyBooth

Copy link
Copy Markdown
Member

I would put it in zlib. In the future we could consider a friendlier name as an alias for zlib, like node:compression or something, but that's for later.

@tniessen

Copy link
Copy Markdown
Member

There is an initiative to use native buffers in the new public APIs

Just leaving this reference here: #41588

Comment threaddoc/api/zip.md Outdated
@tniessen

Copy link
Copy Markdown
Member

In the future we could consider a friendlier name as an alias for zlib, like node:compression or something, but that's for later.

That might make some sense for zip, which inherently supports compression, but if we do add other archive formats that don't, it won't fit. Also, except for zip, compression and archive formats are orthogonal even if related topics.

@arcanisarcanis changed the title Adds prototype zip moduleAdds prototype archive moduleDec 18, 2022
@arcanisarcanis changed the title Adds prototype archive moduleAdds zip support to the zlib moduleJan 6, 2023
@arcanis
arcanis marked this pull request as ready for review January 6, 2023 21:30
Comment threaddoc/api/zlib.md
Comment on lines +257 to +268
## Compressing multiple files together

<!-- YAML
added: REPLACEME
-->

The `zlib` library provides ways to compress individual objects, but not to
aggregate multiple ones into a single file suitable for redistribution (what
is often called archival).

To this end, `node:zip` provides the `ZipArchive` class which allows to create,
read, and modify zip archives:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I assume this section needs updating? Because of references to node:zip etc.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Indeed, it's a typo, the archive is now part of node:zlib (would it make sense to have a check in lint-md that all node:something identifiers must be valid?)

@bakkot

Copy link
Copy Markdown
Contributor

New effort at #64339

@panva

Copy link
Copy Markdown
Member

Superseded by #64339

@panvapanva closed this Aug 28, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

buildIssues and PRs related to Node.js builds or CI infrastructure.dependenciesPRs that add, update, or configure Node.js dependencies.metaIssues and PRs related to the general management of the project.needs-ciPRs that need a full CI run.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

12 participants

@arcanis@GeoffreyBooth@anonrig@addaleax@jasnell@tniessen@bakkot@panva@mcollina@mscdex@aduh95@nodejs-github-bot
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Adds zip support to the zlib module - #45651

Closed
arcanis wants to merge 2 commits into
nodejs:mainfrom
arcanis:mael/zip
Closed

Adds zip support to the zlib module#45651
arcanis wants to merge 2 commits into
nodejs:mainfrom
arcanis:mael/zip

Conversation

@arcanis

@arcanisarcanis commented Nov 28, 2022

Copy link
Copy Markdown
Contributor

Ref #45434

This PR adds a new ZipArchive class to the zlib module, which can be used to read and write content from zip archives. Its current API looks like this:

constfs=require(`fs`);const{ZipArchive}=require(`zlib`);// Creates a new in-memory archiveconstzip=newZipArchive();zip.addFile(`hello`,fs.readFileSync(__filename));constdata=zip.digest();fs.writeFileSync(`./archive.zip`,data);// The data obtained from `digest` can also be reopenedconstzip2=newZipArchive(data);console.log(zip2.getEntries());console.log(zip2.getEntries({withFileTypes: true}));constcontent=zip2.readEntry(0);console.log(content);

Maintenance cost

I kept the feature scope limited enough to cover most of the use cases but without increasing the maintenance cost or build cost. A few things have been cut from what the libzip would allow:

  • Opening files directly from the filesystem isn't supported, because it would bypass node:fs. The current API only works with memory buffers (according to my tests it doesn't have any negative impact even when compared to the wasm API which went through file descriptors).

  • Encryption isn't supported, because it's unclear how it should integrate with node:crypto. There's room for follow-up, but it didn't seem a required feature for the first iteration.

Performances

Keep in mind that raw performances aren't the main reason why zip support is important to have as a native feature. The speedup is nice, the simplified garbage collection is very nice, but the real benefit is having a stable cross-platform way to bundle files between platforms. It will be useful for cache mechanisms, transfer algorithms, user CLI generation, and more.

Still, I made some reasonable checks to make sure that no use case regressed. Size of the binary before / after:

before 89604801 85.45MB
after 89780257 85.62MB (+171KB)

Performance-wise, using Yarn as benchmark, the results show native being ~2x faster than wasm (keep in mind the wasm implementation isn't the most popular zip library; projects using jszip will see significantly larger differences):

YARN_EXPERIMENT_NATIVE_ZIPFS=0 PKG=gatsby
➤ YN0000: └ Completed in 15s 806ms
YARN_EXPERIMENT_NATIVE_ZIPFS=1 PKG=gatsby
➤ YN0000: └ Completed in 7s 404ms
YARN_EXPERIMENT_NATIVE_ZIPFS=0 PKG=typescript
➤ YN0000: └ Completed in 13s 676ms
YARN_EXPERIMENT_NATIVE_ZIPFS=1 PKG=typescript
➤ YN0000: └ Completed in 5s 351ms
YARN_EXPERIMENT_NATIVE_ZIPFS=0 PKG=next
➤ YN0000: └ Completed in 5s 923ms
YARN_EXPERIMENT_NATIVE_ZIPFS=1 PKG=next
➤ YN0000: └ Completed in 3s 512ms

To Do

  • Improve the documentation
  • Add more regression tests
  • Benchmark against the WASM libzip
  • API Bikeshedding

@nodejs-github-botnodejs-github-bot added build Issues and PRs related to Node.js builds or CI infrastructure. dependencies PRs that add, update, or configure Node.js dependencies. meta Issues and PRs related to the general management of the project. needs-ci PRs that need a full CI run. labels Nov 28, 2022
@arcanisarcanis mentioned this pull request Nov 28, 2022

@mscdexmscdex left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

-1 I know this is an unpopular opinion, but I'm not convinced this should live in node core. Sure it's a common file format, but then again so are tar, rar, 7z, zstd, xz, and others, which also don't belong in node core.

I feel like adding such modules to node core is further leading to feature creep. I get that other platforms like PHP and such may have zip modules, but they also include a ton of other modules that make them "kitchen sink" platforms, which I would hate to see node.js become.

@GeoffreyBooth

Copy link
Copy Markdown
Member

@mscdex Could the discussion of whether we should do this stay in #45434? And this PR discussion can focus on the implementation.

@addaleaxaddaleax left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'd appreciate some early review to let me know places where the code isn't compatible with Node.js' standards.

Might be jumping the gun here since yeah, there should be a discussion about whether this should happen at all, but, sure, gave it a first look. I do concur with @mscdex's concerns, fwiw.

Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc
}

void ZipArchive::MemoryInfo(MemoryTracker* tracker) const {
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

(might want to fill this out)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

How precise does it have to be? I updated the code to track the size of the input buffer + the buffer of any file that gets added later, but it doesn't include the small-ish libzip overhead, and gets confused if the same file is modified multiple times. If the value is indicative it might be fine?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It doesn’t have to be super precise, but it should be usable for debugging. You’ll probably want to track buf_ here via tracker->TrackField("buf", buf_);. If you can’t track or estimate memory owned by libzip (including memory for added entries?) then it’s probably fine to omit it, rather than to give numbers that are potentially very inaccurate (e.g. after repeated AddEntry() + DeleteEntry() calls).

Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc
}

zip_int64_t file_index = zip_file_add(zip->zip_, *path, file_source, ZIP_FL_OVERWRITE | ZIP_FL_ENC_UTF_8);
CHECK_GE(file_index, 0);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What if this call fails? Likely also applies elsewhere.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For now the code is overly strict and if any function fails, the program aborts on the CHECK_GE call. I have to replace most these calls by something that would just throw instead.

Comment threadsrc/node_zip.cc Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
@anonrig

Copy link
Copy Markdown
Member

Despite the comments from other reviewers, my main concern is about the usage of the buffer module. I strongly believe that the public API should consume necessary native buffers (TypedArray) instead of Node.js buffers.

@arcanis

Copy link
Copy Markdown
ContributorAuthor

I used Buffer since that's what the other main Node APIs tended to use (fs, crypto, zlib) - wouldn't it be surprising for users to return a regular typed array in just this API?

@anonrig

Copy link
Copy Markdown
Member

I used Buffer since that's what the other main Node APIs tended to use (fs, crypto, zlib) - wouldn't it be surprising for users to return a regular typed array in just this API?

There is an initiative to use native buffers in the new public APIs (referencing my personal talks with @addaleax and @jasnell), a @nodejs/tsc member can clarify if this is still a thing.

@addaleax

Copy link
Copy Markdown
Member

@anonrig I don’t know if that’s the best way forward here, but since this feels like a very broad conversation (“Should new Node.js APIs return Uint8Array or should they return Buffer?”), maybe it’s also best to handle that separately from this specific PR?

@jasnell

Copy link
Copy Markdown
Member

In this case, I think Buffer is fine given that it is consistent with the rest of the zlib module. I can see us eventually making a call on avoiding Buffer in the future (or standardizing on it) but this is not the place to decide that

@jasnell

Copy link
Copy Markdown
Member

Is a new top level module what we want here? As opposed to adding this to zlib? I know it's not based on zlib but neither is brotli.

@arcanis

Copy link
Copy Markdown
ContributorAuthor

I think @GeoffreyBooth had the same feedback. I don't have a strong opinion there, perhaps zlib would indeed better match user expectations.

@GeoffreyBooth

Copy link
Copy Markdown
Member

I would put it in zlib. In the future we could consider a friendlier name as an alias for zlib, like node:compression or something, but that's for later.

@tniessen

Copy link
Copy Markdown
Member

There is an initiative to use native buffers in the new public APIs

Just leaving this reference here: #41588

Comment threaddoc/api/zip.md Outdated
@tniessen

Copy link
Copy Markdown
Member

In the future we could consider a friendlier name as an alias for zlib, like node:compression or something, but that's for later.

That might make some sense for zip, which inherently supports compression, but if we do add other archive formats that don't, it won't fit. Also, except for zip, compression and archive formats are orthogonal even if related topics.

@arcanisarcanis changed the title Adds prototype zip moduleAdds prototype archive moduleDec 18, 2022
@arcanisarcanis changed the title Adds prototype archive moduleAdds zip support to the zlib moduleJan 6, 2023
@arcanis
arcanis marked this pull request as ready for review January 6, 2023 21:30
Comment threaddoc/api/zlib.md
Comment on lines +257 to +268
## Compressing multiple files together

<!-- YAML
added: REPLACEME
-->

The `zlib` library provides ways to compress individual objects, but not to
aggregate multiple ones into a single file suitable for redistribution (what
is often called archival).

To this end, `node:zip` provides the `ZipArchive` class which allows to create,
read, and modify zip archives:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I assume this section needs updating? Because of references to node:zip etc.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Indeed, it's a typo, the archive is now part of node:zlib (would it make sense to have a check in lint-md that all node:something identifiers must be valid?)

@bakkot

Copy link
Copy Markdown
Contributor

New effort at #64339

@panva

Copy link
Copy Markdown
Member

Superseded by #64339

@panvapanva closed this Aug 28, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

buildIssues and PRs related to Node.js builds or CI infrastructure.dependenciesPRs that add, update, or configure Node.js dependencies.metaIssues and PRs related to the general management of the project.needs-ciPRs that need a full CI run.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

12 participants

@arcanis@GeoffreyBooth@anonrig@addaleax@jasnell@tniessen@bakkot@panva@mcollina@mscdex@aduh95@nodejs-github-bot
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Adds zip support to the zlib module - #45651

Closed
arcanis wants to merge 2 commits into
nodejs:mainfrom
arcanis:mael/zip
Closed

Adds zip support to the zlib module#45651
arcanis wants to merge 2 commits into
nodejs:mainfrom
arcanis:mael/zip

Conversation

@arcanis

@arcanisarcanis commented Nov 28, 2022

Copy link
Copy Markdown
Contributor

Ref #45434

This PR adds a new ZipArchive class to the zlib module, which can be used to read and write content from zip archives. Its current API looks like this:

constfs=require(`fs`);const{ZipArchive}=require(`zlib`);// Creates a new in-memory archiveconstzip=newZipArchive();zip.addFile(`hello`,fs.readFileSync(__filename));constdata=zip.digest();fs.writeFileSync(`./archive.zip`,data);// The data obtained from `digest` can also be reopenedconstzip2=newZipArchive(data);console.log(zip2.getEntries());console.log(zip2.getEntries({withFileTypes: true}));constcontent=zip2.readEntry(0);console.log(content);

Maintenance cost

I kept the feature scope limited enough to cover most of the use cases but without increasing the maintenance cost or build cost. A few things have been cut from what the libzip would allow:

  • Opening files directly from the filesystem isn't supported, because it would bypass node:fs. The current API only works with memory buffers (according to my tests it doesn't have any negative impact even when compared to the wasm API which went through file descriptors).

  • Encryption isn't supported, because it's unclear how it should integrate with node:crypto. There's room for follow-up, but it didn't seem a required feature for the first iteration.

Performances

Keep in mind that raw performances aren't the main reason why zip support is important to have as a native feature. The speedup is nice, the simplified garbage collection is very nice, but the real benefit is having a stable cross-platform way to bundle files between platforms. It will be useful for cache mechanisms, transfer algorithms, user CLI generation, and more.

Still, I made some reasonable checks to make sure that no use case regressed. Size of the binary before / after:

before 89604801 85.45MB
after 89780257 85.62MB (+171KB)

Performance-wise, using Yarn as benchmark, the results show native being ~2x faster than wasm (keep in mind the wasm implementation isn't the most popular zip library; projects using jszip will see significantly larger differences):

YARN_EXPERIMENT_NATIVE_ZIPFS=0 PKG=gatsby
➤ YN0000: └ Completed in 15s 806ms
YARN_EXPERIMENT_NATIVE_ZIPFS=1 PKG=gatsby
➤ YN0000: └ Completed in 7s 404ms
YARN_EXPERIMENT_NATIVE_ZIPFS=0 PKG=typescript
➤ YN0000: └ Completed in 13s 676ms
YARN_EXPERIMENT_NATIVE_ZIPFS=1 PKG=typescript
➤ YN0000: └ Completed in 5s 351ms
YARN_EXPERIMENT_NATIVE_ZIPFS=0 PKG=next
➤ YN0000: └ Completed in 5s 923ms
YARN_EXPERIMENT_NATIVE_ZIPFS=1 PKG=next
➤ YN0000: └ Completed in 3s 512ms

To Do

  • Improve the documentation
  • Add more regression tests
  • Benchmark against the WASM libzip
  • API Bikeshedding

@nodejs-github-botnodejs-github-bot added build Issues and PRs related to Node.js builds or CI infrastructure. dependencies PRs that add, update, or configure Node.js dependencies. meta Issues and PRs related to the general management of the project. needs-ci PRs that need a full CI run. labels Nov 28, 2022
@arcanisarcanis mentioned this pull request Nov 28, 2022

@mscdexmscdex left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

-1 I know this is an unpopular opinion, but I'm not convinced this should live in node core. Sure it's a common file format, but then again so are tar, rar, 7z, zstd, xz, and others, which also don't belong in node core.

I feel like adding such modules to node core is further leading to feature creep. I get that other platforms like PHP and such may have zip modules, but they also include a ton of other modules that make them "kitchen sink" platforms, which I would hate to see node.js become.

@GeoffreyBooth

Copy link
Copy Markdown
Member

@mscdex Could the discussion of whether we should do this stay in #45434? And this PR discussion can focus on the implementation.

@addaleaxaddaleax left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'd appreciate some early review to let me know places where the code isn't compatible with Node.js' standards.

Might be jumping the gun here since yeah, there should be a discussion about whether this should happen at all, but, sure, gave it a first look. I do concur with @mscdex's concerns, fwiw.

Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc
}

void ZipArchive::MemoryInfo(MemoryTracker* tracker) const {
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

(might want to fill this out)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

How precise does it have to be? I updated the code to track the size of the input buffer + the buffer of any file that gets added later, but it doesn't include the small-ish libzip overhead, and gets confused if the same file is modified multiple times. If the value is indicative it might be fine?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It doesn’t have to be super precise, but it should be usable for debugging. You’ll probably want to track buf_ here via tracker->TrackField("buf", buf_);. If you can’t track or estimate memory owned by libzip (including memory for added entries?) then it’s probably fine to omit it, rather than to give numbers that are potentially very inaccurate (e.g. after repeated AddEntry() + DeleteEntry() calls).

Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc
}

zip_int64_t file_index = zip_file_add(zip->zip_, *path, file_source, ZIP_FL_OVERWRITE | ZIP_FL_ENC_UTF_8);
CHECK_GE(file_index, 0);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What if this call fails? Likely also applies elsewhere.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For now the code is overly strict and if any function fails, the program aborts on the CHECK_GE call. I have to replace most these calls by something that would just throw instead.

Comment threadsrc/node_zip.cc Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
@anonrig

Copy link
Copy Markdown
Member

Despite the comments from other reviewers, my main concern is about the usage of the buffer module. I strongly believe that the public API should consume necessary native buffers (TypedArray) instead of Node.js buffers.

@arcanis

Copy link
Copy Markdown
ContributorAuthor

I used Buffer since that's what the other main Node APIs tended to use (fs, crypto, zlib) - wouldn't it be surprising for users to return a regular typed array in just this API?

@anonrig

Copy link
Copy Markdown
Member

I used Buffer since that's what the other main Node APIs tended to use (fs, crypto, zlib) - wouldn't it be surprising for users to return a regular typed array in just this API?

There is an initiative to use native buffers in the new public APIs (referencing my personal talks with @addaleax and @jasnell), a @nodejs/tsc member can clarify if this is still a thing.

@addaleax

Copy link
Copy Markdown
Member

@anonrig I don’t know if that’s the best way forward here, but since this feels like a very broad conversation (“Should new Node.js APIs return Uint8Array or should they return Buffer?”), maybe it’s also best to handle that separately from this specific PR?

@jasnell

Copy link
Copy Markdown
Member

In this case, I think Buffer is fine given that it is consistent with the rest of the zlib module. I can see us eventually making a call on avoiding Buffer in the future (or standardizing on it) but this is not the place to decide that

@jasnell

Copy link
Copy Markdown
Member

Is a new top level module what we want here? As opposed to adding this to zlib? I know it's not based on zlib but neither is brotli.

@arcanis

Copy link
Copy Markdown
ContributorAuthor

I think @GeoffreyBooth had the same feedback. I don't have a strong opinion there, perhaps zlib would indeed better match user expectations.

@GeoffreyBooth

Copy link
Copy Markdown
Member

I would put it in zlib. In the future we could consider a friendlier name as an alias for zlib, like node:compression or something, but that's for later.

@tniessen

Copy link
Copy Markdown
Member

There is an initiative to use native buffers in the new public APIs

Just leaving this reference here: #41588

Comment threaddoc/api/zip.md Outdated
@tniessen

Copy link
Copy Markdown
Member

In the future we could consider a friendlier name as an alias for zlib, like node:compression or something, but that's for later.

That might make some sense for zip, which inherently supports compression, but if we do add other archive formats that don't, it won't fit. Also, except for zip, compression and archive formats are orthogonal even if related topics.

@arcanisarcanis changed the title Adds prototype zip moduleAdds prototype archive moduleDec 18, 2022
@arcanisarcanis changed the title Adds prototype archive moduleAdds zip support to the zlib moduleJan 6, 2023
@arcanis
arcanis marked this pull request as ready for review January 6, 2023 21:30
Comment threaddoc/api/zlib.md
Comment on lines +257 to +268
## Compressing multiple files together

<!-- YAML
added: REPLACEME
-->

The `zlib` library provides ways to compress individual objects, but not to
aggregate multiple ones into a single file suitable for redistribution (what
is often called archival).

To this end, `node:zip` provides the `ZipArchive` class which allows to create,
read, and modify zip archives:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I assume this section needs updating? Because of references to node:zip etc.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Indeed, it's a typo, the archive is now part of node:zlib (would it make sense to have a check in lint-md that all node:something identifiers must be valid?)

@bakkot

Copy link
Copy Markdown
Contributor

New effort at #64339

@panva

Copy link
Copy Markdown
Member

Superseded by #64339

@panvapanva closed this Aug 28, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

buildIssues and PRs related to Node.js builds or CI infrastructure.dependenciesPRs that add, update, or configure Node.js dependencies.metaIssues and PRs related to the general management of the project.needs-ciPRs that need a full CI run.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

12 participants

@arcanis@GeoffreyBooth@anonrig@addaleax@jasnell@tniessen@bakkot@panva@mcollina@mscdex@aduh95@nodejs-github-bot
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Adds zip support to the zlib module - #45651

Closed
arcanis wants to merge 2 commits into
nodejs:mainfrom
arcanis:mael/zip
Closed

Adds zip support to the zlib module#45651
arcanis wants to merge 2 commits into
nodejs:mainfrom
arcanis:mael/zip

Conversation

@arcanis

@arcanisarcanis commented Nov 28, 2022

Copy link
Copy Markdown
Contributor

Ref #45434

This PR adds a new ZipArchive class to the zlib module, which can be used to read and write content from zip archives. Its current API looks like this:

constfs=require(`fs`);const{ZipArchive}=require(`zlib`);// Creates a new in-memory archiveconstzip=newZipArchive();zip.addFile(`hello`,fs.readFileSync(__filename));constdata=zip.digest();fs.writeFileSync(`./archive.zip`,data);// The data obtained from `digest` can also be reopenedconstzip2=newZipArchive(data);console.log(zip2.getEntries());console.log(zip2.getEntries({withFileTypes: true}));constcontent=zip2.readEntry(0);console.log(content);

Maintenance cost

I kept the feature scope limited enough to cover most of the use cases but without increasing the maintenance cost or build cost. A few things have been cut from what the libzip would allow:

  • Opening files directly from the filesystem isn't supported, because it would bypass node:fs. The current API only works with memory buffers (according to my tests it doesn't have any negative impact even when compared to the wasm API which went through file descriptors).

  • Encryption isn't supported, because it's unclear how it should integrate with node:crypto. There's room for follow-up, but it didn't seem a required feature for the first iteration.

Performances

Keep in mind that raw performances aren't the main reason why zip support is important to have as a native feature. The speedup is nice, the simplified garbage collection is very nice, but the real benefit is having a stable cross-platform way to bundle files between platforms. It will be useful for cache mechanisms, transfer algorithms, user CLI generation, and more.

Still, I made some reasonable checks to make sure that no use case regressed. Size of the binary before / after:

before 89604801 85.45MB
after 89780257 85.62MB (+171KB)

Performance-wise, using Yarn as benchmark, the results show native being ~2x faster than wasm (keep in mind the wasm implementation isn't the most popular zip library; projects using jszip will see significantly larger differences):

YARN_EXPERIMENT_NATIVE_ZIPFS=0 PKG=gatsby
➤ YN0000: └ Completed in 15s 806ms
YARN_EXPERIMENT_NATIVE_ZIPFS=1 PKG=gatsby
➤ YN0000: └ Completed in 7s 404ms
YARN_EXPERIMENT_NATIVE_ZIPFS=0 PKG=typescript
➤ YN0000: └ Completed in 13s 676ms
YARN_EXPERIMENT_NATIVE_ZIPFS=1 PKG=typescript
➤ YN0000: └ Completed in 5s 351ms
YARN_EXPERIMENT_NATIVE_ZIPFS=0 PKG=next
➤ YN0000: └ Completed in 5s 923ms
YARN_EXPERIMENT_NATIVE_ZIPFS=1 PKG=next
➤ YN0000: └ Completed in 3s 512ms

To Do

  • Improve the documentation
  • Add more regression tests
  • Benchmark against the WASM libzip
  • API Bikeshedding

@nodejs-github-botnodejs-github-bot added build Issues and PRs related to Node.js builds or CI infrastructure. dependencies PRs that add, update, or configure Node.js dependencies. meta Issues and PRs related to the general management of the project. needs-ci PRs that need a full CI run. labels Nov 28, 2022
@arcanisarcanis mentioned this pull request Nov 28, 2022

@mscdexmscdex left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

-1 I know this is an unpopular opinion, but I'm not convinced this should live in node core. Sure it's a common file format, but then again so are tar, rar, 7z, zstd, xz, and others, which also don't belong in node core.

I feel like adding such modules to node core is further leading to feature creep. I get that other platforms like PHP and such may have zip modules, but they also include a ton of other modules that make them "kitchen sink" platforms, which I would hate to see node.js become.

@GeoffreyBooth

Copy link
Copy Markdown
Member

@mscdex Could the discussion of whether we should do this stay in #45434? And this PR discussion can focus on the implementation.

@addaleaxaddaleax left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'd appreciate some early review to let me know places where the code isn't compatible with Node.js' standards.

Might be jumping the gun here since yeah, there should be a discussion about whether this should happen at all, but, sure, gave it a first look. I do concur with @mscdex's concerns, fwiw.

Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc
}

void ZipArchive::MemoryInfo(MemoryTracker* tracker) const {
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

(might want to fill this out)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

How precise does it have to be? I updated the code to track the size of the input buffer + the buffer of any file that gets added later, but it doesn't include the small-ish libzip overhead, and gets confused if the same file is modified multiple times. If the value is indicative it might be fine?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It doesn’t have to be super precise, but it should be usable for debugging. You’ll probably want to track buf_ here via tracker->TrackField("buf", buf_);. If you can’t track or estimate memory owned by libzip (including memory for added entries?) then it’s probably fine to omit it, rather than to give numbers that are potentially very inaccurate (e.g. after repeated AddEntry() + DeleteEntry() calls).

Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc
}

zip_int64_t file_index = zip_file_add(zip->zip_, *path, file_source, ZIP_FL_OVERWRITE | ZIP_FL_ENC_UTF_8);
CHECK_GE(file_index, 0);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What if this call fails? Likely also applies elsewhere.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For now the code is overly strict and if any function fails, the program aborts on the CHECK_GE call. I have to replace most these calls by something that would just throw instead.

Comment threadsrc/node_zip.cc Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
@anonrig

Copy link
Copy Markdown
Member

Despite the comments from other reviewers, my main concern is about the usage of the buffer module. I strongly believe that the public API should consume necessary native buffers (TypedArray) instead of Node.js buffers.

@arcanis

Copy link
Copy Markdown
ContributorAuthor

I used Buffer since that's what the other main Node APIs tended to use (fs, crypto, zlib) - wouldn't it be surprising for users to return a regular typed array in just this API?

@anonrig

Copy link
Copy Markdown
Member

I used Buffer since that's what the other main Node APIs tended to use (fs, crypto, zlib) - wouldn't it be surprising for users to return a regular typed array in just this API?

There is an initiative to use native buffers in the new public APIs (referencing my personal talks with @addaleax and @jasnell), a @nodejs/tsc member can clarify if this is still a thing.

@addaleax

Copy link
Copy Markdown
Member

@anonrig I don’t know if that’s the best way forward here, but since this feels like a very broad conversation (“Should new Node.js APIs return Uint8Array or should they return Buffer?”), maybe it’s also best to handle that separately from this specific PR?

@jasnell

Copy link
Copy Markdown
Member

In this case, I think Buffer is fine given that it is consistent with the rest of the zlib module. I can see us eventually making a call on avoiding Buffer in the future (or standardizing on it) but this is not the place to decide that

@jasnell

Copy link
Copy Markdown
Member

Is a new top level module what we want here? As opposed to adding this to zlib? I know it's not based on zlib but neither is brotli.

@arcanis

Copy link
Copy Markdown
ContributorAuthor

I think @GeoffreyBooth had the same feedback. I don't have a strong opinion there, perhaps zlib would indeed better match user expectations.

@GeoffreyBooth

Copy link
Copy Markdown
Member

I would put it in zlib. In the future we could consider a friendlier name as an alias for zlib, like node:compression or something, but that's for later.

@tniessen

Copy link
Copy Markdown
Member

There is an initiative to use native buffers in the new public APIs

Just leaving this reference here: #41588

Comment threaddoc/api/zip.md Outdated
@tniessen

Copy link
Copy Markdown
Member

In the future we could consider a friendlier name as an alias for zlib, like node:compression or something, but that's for later.

That might make some sense for zip, which inherently supports compression, but if we do add other archive formats that don't, it won't fit. Also, except for zip, compression and archive formats are orthogonal even if related topics.

@arcanisarcanis changed the title Adds prototype zip moduleAdds prototype archive moduleDec 18, 2022
@arcanisarcanis changed the title Adds prototype archive moduleAdds zip support to the zlib moduleJan 6, 2023
@arcanis
arcanis marked this pull request as ready for review January 6, 2023 21:30
Comment threaddoc/api/zlib.md
Comment on lines +257 to +268
## Compressing multiple files together

<!-- YAML
added: REPLACEME
-->

The `zlib` library provides ways to compress individual objects, but not to
aggregate multiple ones into a single file suitable for redistribution (what
is often called archival).

To this end, `node:zip` provides the `ZipArchive` class which allows to create,
read, and modify zip archives:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I assume this section needs updating? Because of references to node:zip etc.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Indeed, it's a typo, the archive is now part of node:zlib (would it make sense to have a check in lint-md that all node:something identifiers must be valid?)

@bakkot

Copy link
Copy Markdown
Contributor

New effort at #64339

@panva

Copy link
Copy Markdown
Member

Superseded by #64339

@panvapanva closed this Aug 28, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

buildIssues and PRs related to Node.js builds or CI infrastructure.dependenciesPRs that add, update, or configure Node.js dependencies.metaIssues and PRs related to the general management of the project.needs-ciPRs that need a full CI run.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

12 participants

@arcanis@GeoffreyBooth@anonrig@addaleax@jasnell@tniessen@bakkot@panva@mcollina@mscdex@aduh95@nodejs-github-bot
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Adds zip support to the zlib module - #45651

Closed
arcanis wants to merge 2 commits into
nodejs:mainfrom
arcanis:mael/zip
Closed

Adds zip support to the zlib module#45651
arcanis wants to merge 2 commits into
nodejs:mainfrom
arcanis:mael/zip

Conversation

@arcanis

@arcanisarcanis commented Nov 28, 2022

Copy link
Copy Markdown
Contributor

Ref #45434

This PR adds a new ZipArchive class to the zlib module, which can be used to read and write content from zip archives. Its current API looks like this:

constfs=require(`fs`);const{ZipArchive}=require(`zlib`);// Creates a new in-memory archiveconstzip=newZipArchive();zip.addFile(`hello`,fs.readFileSync(__filename));constdata=zip.digest();fs.writeFileSync(`./archive.zip`,data);// The data obtained from `digest` can also be reopenedconstzip2=newZipArchive(data);console.log(zip2.getEntries());console.log(zip2.getEntries({withFileTypes: true}));constcontent=zip2.readEntry(0);console.log(content);

Maintenance cost

I kept the feature scope limited enough to cover most of the use cases but without increasing the maintenance cost or build cost. A few things have been cut from what the libzip would allow:

  • Opening files directly from the filesystem isn't supported, because it would bypass node:fs. The current API only works with memory buffers (according to my tests it doesn't have any negative impact even when compared to the wasm API which went through file descriptors).

  • Encryption isn't supported, because it's unclear how it should integrate with node:crypto. There's room for follow-up, but it didn't seem a required feature for the first iteration.

Performances

Keep in mind that raw performances aren't the main reason why zip support is important to have as a native feature. The speedup is nice, the simplified garbage collection is very nice, but the real benefit is having a stable cross-platform way to bundle files between platforms. It will be useful for cache mechanisms, transfer algorithms, user CLI generation, and more.

Still, I made some reasonable checks to make sure that no use case regressed. Size of the binary before / after:

before 89604801 85.45MB
after 89780257 85.62MB (+171KB)

Performance-wise, using Yarn as benchmark, the results show native being ~2x faster than wasm (keep in mind the wasm implementation isn't the most popular zip library; projects using jszip will see significantly larger differences):

YARN_EXPERIMENT_NATIVE_ZIPFS=0 PKG=gatsby
➤ YN0000: └ Completed in 15s 806ms
YARN_EXPERIMENT_NATIVE_ZIPFS=1 PKG=gatsby
➤ YN0000: └ Completed in 7s 404ms
YARN_EXPERIMENT_NATIVE_ZIPFS=0 PKG=typescript
➤ YN0000: └ Completed in 13s 676ms
YARN_EXPERIMENT_NATIVE_ZIPFS=1 PKG=typescript
➤ YN0000: └ Completed in 5s 351ms
YARN_EXPERIMENT_NATIVE_ZIPFS=0 PKG=next
➤ YN0000: └ Completed in 5s 923ms
YARN_EXPERIMENT_NATIVE_ZIPFS=1 PKG=next
➤ YN0000: └ Completed in 3s 512ms

To Do

  • Improve the documentation
  • Add more regression tests
  • Benchmark against the WASM libzip
  • API Bikeshedding

@nodejs-github-botnodejs-github-bot added build Issues and PRs related to Node.js builds or CI infrastructure. dependencies PRs that add, update, or configure Node.js dependencies. meta Issues and PRs related to the general management of the project. needs-ci PRs that need a full CI run. labels Nov 28, 2022
@arcanisarcanis mentioned this pull request Nov 28, 2022

@mscdexmscdex left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

-1 I know this is an unpopular opinion, but I'm not convinced this should live in node core. Sure it's a common file format, but then again so are tar, rar, 7z, zstd, xz, and others, which also don't belong in node core.

I feel like adding such modules to node core is further leading to feature creep. I get that other platforms like PHP and such may have zip modules, but they also include a ton of other modules that make them "kitchen sink" platforms, which I would hate to see node.js become.

@GeoffreyBooth

Copy link
Copy Markdown
Member

@mscdex Could the discussion of whether we should do this stay in #45434? And this PR discussion can focus on the implementation.

@addaleaxaddaleax left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'd appreciate some early review to let me know places where the code isn't compatible with Node.js' standards.

Might be jumping the gun here since yeah, there should be a discussion about whether this should happen at all, but, sure, gave it a first look. I do concur with @mscdex's concerns, fwiw.

Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc
}

void ZipArchive::MemoryInfo(MemoryTracker* tracker) const {
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

(might want to fill this out)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

How precise does it have to be? I updated the code to track the size of the input buffer + the buffer of any file that gets added later, but it doesn't include the small-ish libzip overhead, and gets confused if the same file is modified multiple times. If the value is indicative it might be fine?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It doesn’t have to be super precise, but it should be usable for debugging. You’ll probably want to track buf_ here via tracker->TrackField("buf", buf_);. If you can’t track or estimate memory owned by libzip (including memory for added entries?) then it’s probably fine to omit it, rather than to give numbers that are potentially very inaccurate (e.g. after repeated AddEntry() + DeleteEntry() calls).

Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc Outdated
Comment threadsrc/node_zip.cc
}

zip_int64_t file_index = zip_file_add(zip->zip_, *path, file_source, ZIP_FL_OVERWRITE | ZIP_FL_ENC_UTF_8);
CHECK_GE(file_index, 0);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What if this call fails? Likely also applies elsewhere.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For now the code is overly strict and if any function fails, the program aborts on the CHECK_GE call. I have to replace most these calls by something that would just throw instead.

Comment threadsrc/node_zip.cc Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
Comment threadlib/zip.js Outdated
@anonrig

Copy link
Copy Markdown
Member

Despite the comments from other reviewers, my main concern is about the usage of the buffer module. I strongly believe that the public API should consume necessary native buffers (TypedArray) instead of Node.js buffers.

@arcanis

Copy link
Copy Markdown
ContributorAuthor

I used Buffer since that's what the other main Node APIs tended to use (fs, crypto, zlib) - wouldn't it be surprising for users to return a regular typed array in just this API?

@anonrig

Copy link
Copy Markdown
Member

I used Buffer since that's what the other main Node APIs tended to use (fs, crypto, zlib) - wouldn't it be surprising for users to return a regular typed array in just this API?

There is an initiative to use native buffers in the new public APIs (referencing my personal talks with @addaleax and @jasnell), a @nodejs/tsc member can clarify if this is still a thing.

@addaleax

Copy link
Copy Markdown
Member

@anonrig I don’t know if that’s the best way forward here, but since this feels like a very broad conversation (“Should new Node.js APIs return Uint8Array or should they return Buffer?”), maybe it’s also best to handle that separately from this specific PR?

@jasnell

Copy link
Copy Markdown
Member

In this case, I think Buffer is fine given that it is consistent with the rest of the zlib module. I can see us eventually making a call on avoiding Buffer in the future (or standardizing on it) but this is not the place to decide that

@jasnell

Copy link
Copy Markdown
Member

Is a new top level module what we want here? As opposed to adding this to zlib? I know it's not based on zlib but neither is brotli.

@arcanis

Copy link
Copy Markdown
ContributorAuthor

I think @GeoffreyBooth had the same feedback. I don't have a strong opinion there, perhaps zlib would indeed better match user expectations.

@GeoffreyBooth

Copy link
Copy Markdown
Member

I would put it in zlib. In the future we could consider a friendlier name as an alias for zlib, like node:compression or something, but that's for later.

@tniessen

Copy link
Copy Markdown
Member

There is an initiative to use native buffers in the new public APIs

Just leaving this reference here: #41588

Comment threaddoc/api/zip.md Outdated
@tniessen

Copy link
Copy Markdown
Member

In the future we could consider a friendlier name as an alias for zlib, like node:compression or something, but that's for later.

That might make some sense for zip, which inherently supports compression, but if we do add other archive formats that don't, it won't fit. Also, except for zip, compression and archive formats are orthogonal even if related topics.

@arcanisarcanis changed the title Adds prototype zip moduleAdds prototype archive moduleDec 18, 2022
@arcanisarcanis changed the title Adds prototype archive moduleAdds zip support to the zlib moduleJan 6, 2023
@arcanis
arcanis marked this pull request as ready for review January 6, 2023 21:30
Comment threaddoc/api/zlib.md
Comment on lines +257 to +268
## Compressing multiple files together

<!-- YAML
added: REPLACEME
-->

The `zlib` library provides ways to compress individual objects, but not to
aggregate multiple ones into a single file suitable for redistribution (what
is often called archival).

To this end, `node:zip` provides the `ZipArchive` class which allows to create,
read, and modify zip archives:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I assume this section needs updating? Because of references to node:zip etc.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Indeed, it's a typo, the archive is now part of node:zlib (would it make sense to have a check in lint-md that all node:something identifiers must be valid?)

@bakkot

Copy link
Copy Markdown
Contributor

New effort at #64339

@panva

Copy link
Copy Markdown
Member

Superseded by #64339

@panvapanva closed this Aug 28, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

buildIssues and PRs related to Node.js builds or CI infrastructure.dependenciesPRs that add, update, or configure Node.js dependencies.metaIssues and PRs related to the general management of the project.needs-ciPRs that need a full CI run.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

12 participants

@arcanis@GeoffreyBooth@anonrig@addaleax@jasnell@tniessen@bakkot@panva@mcollina@mscdex@aduh95@nodejs-github-bot