Skip to content

Fix unzipping 4GB+ zip files - #77181

Merged
adamsitnik merged 8 commits into
dotnet:mainfrom
adamsitnik:zip4gb
Oct 28, 2022
Merged

Fix unzipping 4GB+ zip files#77181
adamsitnik merged 8 commits into
dotnet:mainfrom
adamsitnik:zip4gb

Conversation

@adamsitnik

@adamsitnikadamsitnik commented Oct 18, 2022

Copy link
Copy Markdown
Member

fixes#77159 without reverting #68106

I've tried to keep the changes as small as possible to make it easier to review and backport.

Explanation: for a file that is over 4 GB after zipping, some of the entries were reporting extraField.Size == 8, but readUncompressedSize, readCompressedSize and readStartDiskNumber equal false and readLocalHeaderOffset equal true. So what the code was doing, was reading first int64, not using it and exiting:

longvalue64=reader.ReadInt64();
if(readUncompressedSize)
zip64Block._uncompressedSize=value64;
if(ms.Position>extraField.Size-sizeof(long))
returntrue;

The fix is to allow advancing the stream (by reading from it) only when all 4 fields are provided (extraField.Size == 28) or when they are explicitly requested (the boolean flags set to true)

@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @dotnet/area-system-io-compression
See info in area-owners.md if you want to be subscribed.

Issue Details

fixes #77159 without reverting #68106

Author:adamsitnik
Assignees:-
Labels:

area-System.IO.Compression

Milestone:-

@ghostghost assigned adamsitnikOct 18, 2022
Comment threadsrc/libraries/System.IO.Compression/tests/ZipArchive/zip_LargeFiles.cs Outdated
@adamsitnik

Copy link
Copy Markdown
MemberAuthor

/azp list

@azure-pipelines

Copy link
Copy Markdown
CI/CD Pipelines for this repository:

@adamsitnik

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-libraries-coreclr outerloop

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

Comment threadsrc/libraries/System.IO.Compression/tests/ZipArchive/zip_LargeFiles.cs Outdated
Comment threadsrc/libraries/System.IO.Compression/tests/ZipArchive/zip_LargeFiles.cs Outdated
@danmoseley

Copy link
Copy Markdown
Contributor

Seems reasonable but I'd rather someone on the IO crew sign off on the change.

BTW if you touch this PR again maybe you could fix this line which I believe should check uncompressed length.
https://github.com/dotnet/runtime/pull/68106/files#diff-ea21d1af009443a658bc821a6eb47860f7887ff70155a68f11990ec64875a2dcR860

@adamsitnik

Copy link
Copy Markdown
MemberAuthor

Seems reasonable but I'd rather someone on the IO crew sign off on the change.

@jozkee PTAL

@jozkee

Copy link
Copy Markdown
Member

Taking a look...

Comment threadsrc/libraries/System.IO.Compression/tests/ZipArchive/zip_LargeFiles.cs Outdated
Comment threadsrc/libraries/System.IO.Compression/src/System/IO/Compression/ZipBlocks.cs Outdated
[OuterLoop("It requires almost 12 GB of free disk space")]
public static void UnzipOver4GBZipFile()
{
byte[] buffer = GC.AllocateUninitializedArray<byte>(1_000_000_000); // 1 GB

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Technically, this is not 1Gb, but since you are creating 6 files, the test is still valid.

I would rather create two tests, one that succeeds at the limit (say 3.999 Gb) without the fix and one that fails and then verify that both pass with your fix.

Since this is urgent for a backport, we can defer that.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Technically, this is not 1Gb, but since you are creating 6 files, the test is still valid.

You mean that $1000^3$ is not 1 GB but $1024^3$ is?

I would rather create two tests, one that succeeds at the limit (say 3.999 Gb) without the fix and one that fails and then verify that both pass with your fix.

My typical workflow is to add test that fails first, then make it pass. We have plenty of tests that verify < 4 GB archives so I don't see the value in adding it (also it would require us to use even more disk space and take longer to execute).

In the long term we could add more tests, especially such that verify that decompress(compress(x)) == x (which this test avoids, as it was taking 10 minutes to do so on my beefy PC).

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You mean that $1000^3$ is not 1 GB but $1024^3$ is?

Yes.

We have plenty of tests that verify < 4 GB archives so I don't see the value in adding it (also it would require us to use even more disk space and take longer to execute).

I meant to verify on the edge cases of the bug. In this case that would be 4Gb - 1, 4Gb, and 4Gb + 1; rather than a 6Gb case. But yes, that would multiply test execution time which is already quite large with this new test.

Comment on lines +202 to +207
// Advancing the stream (by reading from it) is possible only when:
// 1. There is an explicit ask to do that (valid files, corresponding boolean flag(s) set to true, #77159).
// 2. When the size indicates that all the information is available ("slightly invalid files", #49580).
bool readAllFields = extraField.Size >= sizeof(long) + sizeof(long) + sizeof(long) + sizeof(int);

long value64 = readUncompressedSize || readAllFields ? reader.ReadInt64() : -1;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could this fail if only 2 or 3 fields are available for reading but the bool params are false?

i.e: what if extraField.size = 16 and readUncompressedSize = false and readCompressedSize = false?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think that this is possible. @danmoseley has added following comment:

// The spec section 4.5.3:
// The order of the fields in the zip64 extended
// information record is fixed, but the fields MUST
// only appear if the corresponding Local or Central
// directory record field is set to 0xFFFF or 0xFFFFFFFF.
// However tools commonly write the fields anyway; the prevailing convention
// is to respect the size, but only actually use the values if their 32 bit
// values were all 0xFF.

My understanding is that we have two options:

  1. Correct archives: the fields are provided when the directory record fields are set (the right booleans are set to true, the size is their total size)
  2. "Slightly incorrect archives": the boolean fields are not set, but the provided size suggests that they are all present. Tests added in Fix reading slightly incorrect Zip files #68106 use such approach.

Comment threadsrc/libraries/System.IO.Compression/src/System/IO/Compression/ZipBlocks.cs Outdated
@adamsitnik

Copy link
Copy Markdown
MemberAuthor

}

// original values are unsigned, so implies value is too big to fit in signed integer
if (zip64Block._uncompressedSize < 0) throw new InvalidDataException(SR.FieldTooBigUncompressedSize);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These validations only occur if extraField.Size >= 28 since you are now returning if there's less than that.

Shouldn't we validate when reader.ReadInt* is called and/or if read of the field was requested?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is how it always was, I would prefer not to do it now as I intend to backport it to 7.0 and trying to make it as defensive/safe as possible.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The reason I was asking is because it wasn't always like that. It was changed to how it is now on #68106, which is the PR that introduced the regression.

[OuterLoop("It requires almost 12 GB of free disk space")]
public static void UnzipOver4GBZipFile()
{
byte[] buffer = GC.AllocateUninitializedArray<byte>(1_000_000_000); // 1 GB

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there a reasonable chance this will OOM? If so, consider catching that and skipping the test, rather than failing it.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there a reasonable chance this will OOM?

It should be executed only in Outerloop runs for the 64 bit "fast runtimes" and not in parallel with other tests so I don't think it's needed (at least for now).

[ConditionalFact(typeof(PlatformDetection),nameof(PlatformDetection.IsSpeedOptimized),nameof(PlatformDetection.Is64BitProcess))]// don't run it on slower runtimes

publicstaticboolIsSpeedOptimized=>!IsSizeOptimized;
publicstaticboolIsSizeOptimized=>IsBrowser||IsAndroid||IsAppleMobile;

@adamsitnik
adamsitnik merged commit 399c6dc into dotnet:mainOct 28, 2022
@adamsitnik

Copy link
Copy Markdown
MemberAuthor

/backport to release/7.0

@github-actions

Copy link
Copy Markdown
Contributor

Started backporting to release/7.0: https://github.com/dotnet/runtime/actions/runs/3347261332

@adamsitnik

Copy link
Copy Markdown
MemberAuthor

@danmoseley@jozkee@stephentoub thank you for the reviews!

@carlossanlopcarlossanlop left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Assuming the CI is green, this LGTM. I'd also wait for a sign-off from @jozkee.

Edit: Nevermind. This was merged while I was still adding a review. :)

}
else if (readAllFields)
{
_ = reader.ReadInt64();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For my own education: is there a perf improvement when using _ = ReadInt64(); to discard the value, compared to just not using _? Or is the idea here to explicitly signal that we are ignoring the value on purpose?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It shouldn't have a runtime effect, it's self documenting, both for humans and also for linters/analyzers that might otherwise flag an unused method result.

@ghostghost locked as resolved and limited conversation to collaborators Nov 28, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

DotNet 7 Regression when reading large Zip files

6 participants

@adamsitnik@danmoseley@jozkee@carlossanlop@stephentoub@iSazonov
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Fix unzipping 4GB+ zip files by adamsitnik · Pull Request #77181 · dotnet/runtime · GitHub
Skip to content

Fix unzipping 4GB+ zip files - #77181

Merged
adamsitnik merged 8 commits into
dotnet:mainfrom
adamsitnik:zip4gb
Oct 28, 2022
Merged

Fix unzipping 4GB+ zip files#77181
adamsitnik merged 8 commits into
dotnet:mainfrom
adamsitnik:zip4gb

Conversation

@adamsitnik

@adamsitnikadamsitnik commented Oct 18, 2022

Copy link
Copy Markdown
Member

fixes#77159 without reverting #68106

I've tried to keep the changes as small as possible to make it easier to review and backport.

Explanation: for a file that is over 4 GB after zipping, some of the entries were reporting extraField.Size == 8, but readUncompressedSize, readCompressedSize and readStartDiskNumber equal false and readLocalHeaderOffset equal true. So what the code was doing, was reading first int64, not using it and exiting:

longvalue64=reader.ReadInt64();
if(readUncompressedSize)
zip64Block._uncompressedSize=value64;
if(ms.Position>extraField.Size-sizeof(long))
returntrue;

The fix is to allow advancing the stream (by reading from it) only when all 4 fields are provided (extraField.Size == 28) or when they are explicitly requested (the boolean flags set to true)

@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @dotnet/area-system-io-compression
See info in area-owners.md if you want to be subscribed.

Issue Details

fixes #77159 without reverting #68106

Author:adamsitnik
Assignees:-
Labels:

area-System.IO.Compression

Milestone:-

@ghostghost assigned adamsitnikOct 18, 2022
Comment threadsrc/libraries/System.IO.Compression/tests/ZipArchive/zip_LargeFiles.cs Outdated
@adamsitnik

Copy link
Copy Markdown
MemberAuthor

/azp list

@azure-pipelines

Copy link
Copy Markdown
CI/CD Pipelines for this repository:

@adamsitnik

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-libraries-coreclr outerloop

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

Comment threadsrc/libraries/System.IO.Compression/tests/ZipArchive/zip_LargeFiles.cs Outdated
Comment threadsrc/libraries/System.IO.Compression/tests/ZipArchive/zip_LargeFiles.cs Outdated
@danmoseley

Copy link
Copy Markdown
Contributor

Seems reasonable but I'd rather someone on the IO crew sign off on the change.

BTW if you touch this PR again maybe you could fix this line which I believe should check uncompressed length.
https://github.com/dotnet/runtime/pull/68106/files#diff-ea21d1af009443a658bc821a6eb47860f7887ff70155a68f11990ec64875a2dcR860

@adamsitnik

Copy link
Copy Markdown
MemberAuthor

Seems reasonable but I'd rather someone on the IO crew sign off on the change.

@jozkee PTAL

@jozkee

Copy link
Copy Markdown
Member

Taking a look...

Comment threadsrc/libraries/System.IO.Compression/tests/ZipArchive/zip_LargeFiles.cs Outdated
Comment threadsrc/libraries/System.IO.Compression/src/System/IO/Compression/ZipBlocks.cs Outdated
[OuterLoop("It requires almost 12 GB of free disk space")]
public static void UnzipOver4GBZipFile()
{
byte[] buffer = GC.AllocateUninitializedArray<byte>(1_000_000_000); // 1 GB

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Technically, this is not 1Gb, but since you are creating 6 files, the test is still valid.

I would rather create two tests, one that succeeds at the limit (say 3.999 Gb) without the fix and one that fails and then verify that both pass with your fix.

Since this is urgent for a backport, we can defer that.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Technically, this is not 1Gb, but since you are creating 6 files, the test is still valid.

You mean that $1000^3$ is not 1 GB but $1024^3$ is?

I would rather create two tests, one that succeeds at the limit (say 3.999 Gb) without the fix and one that fails and then verify that both pass with your fix.

My typical workflow is to add test that fails first, then make it pass. We have plenty of tests that verify < 4 GB archives so I don't see the value in adding it (also it would require us to use even more disk space and take longer to execute).

In the long term we could add more tests, especially such that verify that decompress(compress(x)) == x (which this test avoids, as it was taking 10 minutes to do so on my beefy PC).

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You mean that $1000^3$ is not 1 GB but $1024^3$ is?

Yes.

We have plenty of tests that verify < 4 GB archives so I don't see the value in adding it (also it would require us to use even more disk space and take longer to execute).

I meant to verify on the edge cases of the bug. In this case that would be 4Gb - 1, 4Gb, and 4Gb + 1; rather than a 6Gb case. But yes, that would multiply test execution time which is already quite large with this new test.

Comment on lines +202 to +207
// Advancing the stream (by reading from it) is possible only when:
// 1. There is an explicit ask to do that (valid files, corresponding boolean flag(s) set to true, #77159).
// 2. When the size indicates that all the information is available ("slightly invalid files", #49580).
bool readAllFields = extraField.Size >= sizeof(long) + sizeof(long) + sizeof(long) + sizeof(int);

long value64 = readUncompressedSize || readAllFields ? reader.ReadInt64() : -1;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could this fail if only 2 or 3 fields are available for reading but the bool params are false?

i.e: what if extraField.size = 16 and readUncompressedSize = false and readCompressedSize = false?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think that this is possible. @danmoseley has added following comment:

// The spec section 4.5.3:
// The order of the fields in the zip64 extended
// information record is fixed, but the fields MUST
// only appear if the corresponding Local or Central
// directory record field is set to 0xFFFF or 0xFFFFFFFF.
// However tools commonly write the fields anyway; the prevailing convention
// is to respect the size, but only actually use the values if their 32 bit
// values were all 0xFF.

My understanding is that we have two options:

  1. Correct archives: the fields are provided when the directory record fields are set (the right booleans are set to true, the size is their total size)
  2. "Slightly incorrect archives": the boolean fields are not set, but the provided size suggests that they are all present. Tests added in Fix reading slightly incorrect Zip files #68106 use such approach.

Comment threadsrc/libraries/System.IO.Compression/src/System/IO/Compression/ZipBlocks.cs Outdated
@adamsitnik

Copy link
Copy Markdown
MemberAuthor

}

// original values are unsigned, so implies value is too big to fit in signed integer
if (zip64Block._uncompressedSize < 0) throw new InvalidDataException(SR.FieldTooBigUncompressedSize);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These validations only occur if extraField.Size >= 28 since you are now returning if there's less than that.

Shouldn't we validate when reader.ReadInt* is called and/or if read of the field was requested?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is how it always was, I would prefer not to do it now as I intend to backport it to 7.0 and trying to make it as defensive/safe as possible.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The reason I was asking is because it wasn't always like that. It was changed to how it is now on #68106, which is the PR that introduced the regression.

[OuterLoop("It requires almost 12 GB of free disk space")]
public static void UnzipOver4GBZipFile()
{
byte[] buffer = GC.AllocateUninitializedArray<byte>(1_000_000_000); // 1 GB

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there a reasonable chance this will OOM? If so, consider catching that and skipping the test, rather than failing it.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there a reasonable chance this will OOM?

It should be executed only in Outerloop runs for the 64 bit "fast runtimes" and not in parallel with other tests so I don't think it's needed (at least for now).

[ConditionalFact(typeof(PlatformDetection),nameof(PlatformDetection.IsSpeedOptimized),nameof(PlatformDetection.Is64BitProcess))]// don't run it on slower runtimes

publicstaticboolIsSpeedOptimized=>!IsSizeOptimized;
publicstaticboolIsSizeOptimized=>IsBrowser||IsAndroid||IsAppleMobile;

@adamsitnik
adamsitnik merged commit 399c6dc into dotnet:mainOct 28, 2022
@adamsitnik

Copy link
Copy Markdown
MemberAuthor

/backport to release/7.0

@github-actions

Copy link
Copy Markdown
Contributor

Started backporting to release/7.0: https://github.com/dotnet/runtime/actions/runs/3347261332

@adamsitnik

Copy link
Copy Markdown
MemberAuthor

@danmoseley@jozkee@stephentoub thank you for the reviews!

@carlossanlopcarlossanlop left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Assuming the CI is green, this LGTM. I'd also wait for a sign-off from @jozkee.

Edit: Nevermind. This was merged while I was still adding a review. :)

}
else if (readAllFields)
{
_ = reader.ReadInt64();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For my own education: is there a perf improvement when using _ = ReadInt64(); to discard the value, compared to just not using _? Or is the idea here to explicitly signal that we are ignoring the value on purpose?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It shouldn't have a runtime effect, it's self documenting, both for humans and also for linters/analyzers that might otherwise flag an unused method result.

@ghostghost locked as resolved and limited conversation to collaborators Nov 28, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

DotNet 7 Regression when reading large Zip files

6 participants

@adamsitnik@danmoseley@jozkee@carlossanlop@stephentoub@iSazonov
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Fix unzipping 4GB+ zip files by adamsitnik · Pull Request #77181 · dotnet/runtime · GitHub
Skip to content

Fix unzipping 4GB+ zip files - #77181

Merged
adamsitnik merged 8 commits into
dotnet:mainfrom
adamsitnik:zip4gb
Oct 28, 2022
Merged

Fix unzipping 4GB+ zip files#77181
adamsitnik merged 8 commits into
dotnet:mainfrom
adamsitnik:zip4gb

Conversation

@adamsitnik

@adamsitnikadamsitnik commented Oct 18, 2022

Copy link
Copy Markdown
Member

fixes#77159 without reverting #68106

I've tried to keep the changes as small as possible to make it easier to review and backport.

Explanation: for a file that is over 4 GB after zipping, some of the entries were reporting extraField.Size == 8, but readUncompressedSize, readCompressedSize and readStartDiskNumber equal false and readLocalHeaderOffset equal true. So what the code was doing, was reading first int64, not using it and exiting:

longvalue64=reader.ReadInt64();
if(readUncompressedSize)
zip64Block._uncompressedSize=value64;
if(ms.Position>extraField.Size-sizeof(long))
returntrue;

The fix is to allow advancing the stream (by reading from it) only when all 4 fields are provided (extraField.Size == 28) or when they are explicitly requested (the boolean flags set to true)

@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @dotnet/area-system-io-compression
See info in area-owners.md if you want to be subscribed.

Issue Details

fixes #77159 without reverting #68106

Author:adamsitnik
Assignees:-
Labels:

area-System.IO.Compression

Milestone:-

@ghostghost assigned adamsitnikOct 18, 2022
Comment threadsrc/libraries/System.IO.Compression/tests/ZipArchive/zip_LargeFiles.cs Outdated
@adamsitnik

Copy link
Copy Markdown
MemberAuthor

/azp list

@azure-pipelines

Copy link
Copy Markdown
CI/CD Pipelines for this repository:

@adamsitnik

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-libraries-coreclr outerloop

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

Comment threadsrc/libraries/System.IO.Compression/tests/ZipArchive/zip_LargeFiles.cs Outdated
Comment threadsrc/libraries/System.IO.Compression/tests/ZipArchive/zip_LargeFiles.cs Outdated
@danmoseley

Copy link
Copy Markdown
Contributor

Seems reasonable but I'd rather someone on the IO crew sign off on the change.

BTW if you touch this PR again maybe you could fix this line which I believe should check uncompressed length.
https://github.com/dotnet/runtime/pull/68106/files#diff-ea21d1af009443a658bc821a6eb47860f7887ff70155a68f11990ec64875a2dcR860

@adamsitnik

Copy link
Copy Markdown
MemberAuthor

Seems reasonable but I'd rather someone on the IO crew sign off on the change.

@jozkee PTAL

@jozkee

Copy link
Copy Markdown
Member

Taking a look...

Comment threadsrc/libraries/System.IO.Compression/tests/ZipArchive/zip_LargeFiles.cs Outdated
Comment threadsrc/libraries/System.IO.Compression/src/System/IO/Compression/ZipBlocks.cs Outdated
[OuterLoop("It requires almost 12 GB of free disk space")]
public static void UnzipOver4GBZipFile()
{
byte[] buffer = GC.AllocateUninitializedArray<byte>(1_000_000_000); // 1 GB

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Technically, this is not 1Gb, but since you are creating 6 files, the test is still valid.

I would rather create two tests, one that succeeds at the limit (say 3.999 Gb) without the fix and one that fails and then verify that both pass with your fix.

Since this is urgent for a backport, we can defer that.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Technically, this is not 1Gb, but since you are creating 6 files, the test is still valid.

You mean that $1000^3$ is not 1 GB but $1024^3$ is?

I would rather create two tests, one that succeeds at the limit (say 3.999 Gb) without the fix and one that fails and then verify that both pass with your fix.

My typical workflow is to add test that fails first, then make it pass. We have plenty of tests that verify < 4 GB archives so I don't see the value in adding it (also it would require us to use even more disk space and take longer to execute).

In the long term we could add more tests, especially such that verify that decompress(compress(x)) == x (which this test avoids, as it was taking 10 minutes to do so on my beefy PC).

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You mean that $1000^3$ is not 1 GB but $1024^3$ is?

Yes.

We have plenty of tests that verify < 4 GB archives so I don't see the value in adding it (also it would require us to use even more disk space and take longer to execute).

I meant to verify on the edge cases of the bug. In this case that would be 4Gb - 1, 4Gb, and 4Gb + 1; rather than a 6Gb case. But yes, that would multiply test execution time which is already quite large with this new test.

Comment on lines +202 to +207
// Advancing the stream (by reading from it) is possible only when:
// 1. There is an explicit ask to do that (valid files, corresponding boolean flag(s) set to true, #77159).
// 2. When the size indicates that all the information is available ("slightly invalid files", #49580).
bool readAllFields = extraField.Size >= sizeof(long) + sizeof(long) + sizeof(long) + sizeof(int);

long value64 = readUncompressedSize || readAllFields ? reader.ReadInt64() : -1;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could this fail if only 2 or 3 fields are available for reading but the bool params are false?

i.e: what if extraField.size = 16 and readUncompressedSize = false and readCompressedSize = false?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think that this is possible. @danmoseley has added following comment:

// The spec section 4.5.3:
// The order of the fields in the zip64 extended
// information record is fixed, but the fields MUST
// only appear if the corresponding Local or Central
// directory record field is set to 0xFFFF or 0xFFFFFFFF.
// However tools commonly write the fields anyway; the prevailing convention
// is to respect the size, but only actually use the values if their 32 bit
// values were all 0xFF.

My understanding is that we have two options:

  1. Correct archives: the fields are provided when the directory record fields are set (the right booleans are set to true, the size is their total size)
  2. "Slightly incorrect archives": the boolean fields are not set, but the provided size suggests that they are all present. Tests added in Fix reading slightly incorrect Zip files #68106 use such approach.

Comment threadsrc/libraries/System.IO.Compression/src/System/IO/Compression/ZipBlocks.cs Outdated
@adamsitnik

Copy link
Copy Markdown
MemberAuthor

}

// original values are unsigned, so implies value is too big to fit in signed integer
if (zip64Block._uncompressedSize < 0) throw new InvalidDataException(SR.FieldTooBigUncompressedSize);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These validations only occur if extraField.Size >= 28 since you are now returning if there's less than that.

Shouldn't we validate when reader.ReadInt* is called and/or if read of the field was requested?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is how it always was, I would prefer not to do it now as I intend to backport it to 7.0 and trying to make it as defensive/safe as possible.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The reason I was asking is because it wasn't always like that. It was changed to how it is now on #68106, which is the PR that introduced the regression.

[OuterLoop("It requires almost 12 GB of free disk space")]
public static void UnzipOver4GBZipFile()
{
byte[] buffer = GC.AllocateUninitializedArray<byte>(1_000_000_000); // 1 GB

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there a reasonable chance this will OOM? If so, consider catching that and skipping the test, rather than failing it.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there a reasonable chance this will OOM?

It should be executed only in Outerloop runs for the 64 bit "fast runtimes" and not in parallel with other tests so I don't think it's needed (at least for now).

[ConditionalFact(typeof(PlatformDetection),nameof(PlatformDetection.IsSpeedOptimized),nameof(PlatformDetection.Is64BitProcess))]// don't run it on slower runtimes

publicstaticboolIsSpeedOptimized=>!IsSizeOptimized;
publicstaticboolIsSizeOptimized=>IsBrowser||IsAndroid||IsAppleMobile;

@adamsitnik
adamsitnik merged commit 399c6dc into dotnet:mainOct 28, 2022
@adamsitnik

Copy link
Copy Markdown
MemberAuthor

/backport to release/7.0

@github-actions

Copy link
Copy Markdown
Contributor

Started backporting to release/7.0: https://github.com/dotnet/runtime/actions/runs/3347261332

@adamsitnik

Copy link
Copy Markdown
MemberAuthor

@danmoseley@jozkee@stephentoub thank you for the reviews!

@carlossanlopcarlossanlop left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Assuming the CI is green, this LGTM. I'd also wait for a sign-off from @jozkee.

Edit: Nevermind. This was merged while I was still adding a review. :)

}
else if (readAllFields)
{
_ = reader.ReadInt64();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For my own education: is there a perf improvement when using _ = ReadInt64(); to discard the value, compared to just not using _? Or is the idea here to explicitly signal that we are ignoring the value on purpose?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It shouldn't have a runtime effect, it's self documenting, both for humans and also for linters/analyzers that might otherwise flag an unused method result.

@ghostghost locked as resolved and limited conversation to collaborators Nov 28, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

DotNet 7 Regression when reading large Zip files

6 participants

@adamsitnik@danmoseley@jozkee@carlossanlop@stephentoub@iSazonov
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Fix unzipping 4GB+ zip files by adamsitnik · Pull Request #77181 · dotnet/runtime · GitHub
Skip to content

Fix unzipping 4GB+ zip files - #77181

Merged
adamsitnik merged 8 commits into
dotnet:mainfrom
adamsitnik:zip4gb
Oct 28, 2022
Merged

Fix unzipping 4GB+ zip files#77181
adamsitnik merged 8 commits into
dotnet:mainfrom
adamsitnik:zip4gb

Conversation

@adamsitnik

@adamsitnikadamsitnik commented Oct 18, 2022

Copy link
Copy Markdown
Member

fixes#77159 without reverting #68106

I've tried to keep the changes as small as possible to make it easier to review and backport.

Explanation: for a file that is over 4 GB after zipping, some of the entries were reporting extraField.Size == 8, but readUncompressedSize, readCompressedSize and readStartDiskNumber equal false and readLocalHeaderOffset equal true. So what the code was doing, was reading first int64, not using it and exiting:

longvalue64=reader.ReadInt64();
if(readUncompressedSize)
zip64Block._uncompressedSize=value64;
if(ms.Position>extraField.Size-sizeof(long))
returntrue;

The fix is to allow advancing the stream (by reading from it) only when all 4 fields are provided (extraField.Size == 28) or when they are explicitly requested (the boolean flags set to true)

@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @dotnet/area-system-io-compression
See info in area-owners.md if you want to be subscribed.

Issue Details

fixes #77159 without reverting #68106

Author:adamsitnik
Assignees:-
Labels:

area-System.IO.Compression

Milestone:-

@ghostghost assigned adamsitnikOct 18, 2022
Comment threadsrc/libraries/System.IO.Compression/tests/ZipArchive/zip_LargeFiles.cs Outdated
@adamsitnik

Copy link
Copy Markdown
MemberAuthor

/azp list

@azure-pipelines

Copy link
Copy Markdown
CI/CD Pipelines for this repository:

@adamsitnik

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-libraries-coreclr outerloop

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

Comment threadsrc/libraries/System.IO.Compression/tests/ZipArchive/zip_LargeFiles.cs Outdated
Comment threadsrc/libraries/System.IO.Compression/tests/ZipArchive/zip_LargeFiles.cs Outdated
@danmoseley

Copy link
Copy Markdown
Contributor

Seems reasonable but I'd rather someone on the IO crew sign off on the change.

BTW if you touch this PR again maybe you could fix this line which I believe should check uncompressed length.
https://github.com/dotnet/runtime/pull/68106/files#diff-ea21d1af009443a658bc821a6eb47860f7887ff70155a68f11990ec64875a2dcR860

@adamsitnik

Copy link
Copy Markdown
MemberAuthor

Seems reasonable but I'd rather someone on the IO crew sign off on the change.

@jozkee PTAL

@jozkee

Copy link
Copy Markdown
Member

Taking a look...

Comment threadsrc/libraries/System.IO.Compression/tests/ZipArchive/zip_LargeFiles.cs Outdated
Comment threadsrc/libraries/System.IO.Compression/src/System/IO/Compression/ZipBlocks.cs Outdated
[OuterLoop("It requires almost 12 GB of free disk space")]
public static void UnzipOver4GBZipFile()
{
byte[] buffer = GC.AllocateUninitializedArray<byte>(1_000_000_000); // 1 GB

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Technically, this is not 1Gb, but since you are creating 6 files, the test is still valid.

I would rather create two tests, one that succeeds at the limit (say 3.999 Gb) without the fix and one that fails and then verify that both pass with your fix.

Since this is urgent for a backport, we can defer that.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Technically, this is not 1Gb, but since you are creating 6 files, the test is still valid.

You mean that $1000^3$ is not 1 GB but $1024^3$ is?

I would rather create two tests, one that succeeds at the limit (say 3.999 Gb) without the fix and one that fails and then verify that both pass with your fix.

My typical workflow is to add test that fails first, then make it pass. We have plenty of tests that verify < 4 GB archives so I don't see the value in adding it (also it would require us to use even more disk space and take longer to execute).

In the long term we could add more tests, especially such that verify that decompress(compress(x)) == x (which this test avoids, as it was taking 10 minutes to do so on my beefy PC).

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You mean that $1000^3$ is not 1 GB but $1024^3$ is?

Yes.

We have plenty of tests that verify < 4 GB archives so I don't see the value in adding it (also it would require us to use even more disk space and take longer to execute).

I meant to verify on the edge cases of the bug. In this case that would be 4Gb - 1, 4Gb, and 4Gb + 1; rather than a 6Gb case. But yes, that would multiply test execution time which is already quite large with this new test.

Comment on lines +202 to +207
// Advancing the stream (by reading from it) is possible only when:
// 1. There is an explicit ask to do that (valid files, corresponding boolean flag(s) set to true, #77159).
// 2. When the size indicates that all the information is available ("slightly invalid files", #49580).
bool readAllFields = extraField.Size >= sizeof(long) + sizeof(long) + sizeof(long) + sizeof(int);

long value64 = readUncompressedSize || readAllFields ? reader.ReadInt64() : -1;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could this fail if only 2 or 3 fields are available for reading but the bool params are false?

i.e: what if extraField.size = 16 and readUncompressedSize = false and readCompressedSize = false?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think that this is possible. @danmoseley has added following comment:

// The spec section 4.5.3:
// The order of the fields in the zip64 extended
// information record is fixed, but the fields MUST
// only appear if the corresponding Local or Central
// directory record field is set to 0xFFFF or 0xFFFFFFFF.
// However tools commonly write the fields anyway; the prevailing convention
// is to respect the size, but only actually use the values if their 32 bit
// values were all 0xFF.

My understanding is that we have two options:

  1. Correct archives: the fields are provided when the directory record fields are set (the right booleans are set to true, the size is their total size)
  2. "Slightly incorrect archives": the boolean fields are not set, but the provided size suggests that they are all present. Tests added in Fix reading slightly incorrect Zip files #68106 use such approach.

Comment threadsrc/libraries/System.IO.Compression/src/System/IO/Compression/ZipBlocks.cs Outdated
@adamsitnik

Copy link
Copy Markdown
MemberAuthor

}

// original values are unsigned, so implies value is too big to fit in signed integer
if (zip64Block._uncompressedSize < 0) throw new InvalidDataException(SR.FieldTooBigUncompressedSize);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These validations only occur if extraField.Size >= 28 since you are now returning if there's less than that.

Shouldn't we validate when reader.ReadInt* is called and/or if read of the field was requested?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is how it always was, I would prefer not to do it now as I intend to backport it to 7.0 and trying to make it as defensive/safe as possible.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The reason I was asking is because it wasn't always like that. It was changed to how it is now on #68106, which is the PR that introduced the regression.

[OuterLoop("It requires almost 12 GB of free disk space")]
public static void UnzipOver4GBZipFile()
{
byte[] buffer = GC.AllocateUninitializedArray<byte>(1_000_000_000); // 1 GB

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there a reasonable chance this will OOM? If so, consider catching that and skipping the test, rather than failing it.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there a reasonable chance this will OOM?

It should be executed only in Outerloop runs for the 64 bit "fast runtimes" and not in parallel with other tests so I don't think it's needed (at least for now).

[ConditionalFact(typeof(PlatformDetection),nameof(PlatformDetection.IsSpeedOptimized),nameof(PlatformDetection.Is64BitProcess))]// don't run it on slower runtimes

publicstaticboolIsSpeedOptimized=>!IsSizeOptimized;
publicstaticboolIsSizeOptimized=>IsBrowser||IsAndroid||IsAppleMobile;

@adamsitnik
adamsitnik merged commit 399c6dc into dotnet:mainOct 28, 2022
@adamsitnik

Copy link
Copy Markdown
MemberAuthor

/backport to release/7.0

@github-actions

Copy link
Copy Markdown
Contributor

Started backporting to release/7.0: https://github.com/dotnet/runtime/actions/runs/3347261332

@adamsitnik

Copy link
Copy Markdown
MemberAuthor

@danmoseley@jozkee@stephentoub thank you for the reviews!

@carlossanlopcarlossanlop left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Assuming the CI is green, this LGTM. I'd also wait for a sign-off from @jozkee.

Edit: Nevermind. This was merged while I was still adding a review. :)

}
else if (readAllFields)
{
_ = reader.ReadInt64();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For my own education: is there a perf improvement when using _ = ReadInt64(); to discard the value, compared to just not using _? Or is the idea here to explicitly signal that we are ignoring the value on purpose?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It shouldn't have a runtime effect, it's self documenting, both for humans and also for linters/analyzers that might otherwise flag an unused method result.

@ghostghost locked as resolved and limited conversation to collaborators Nov 28, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

DotNet 7 Regression when reading large Zip files

6 participants

@adamsitnik@danmoseley@jozkee@carlossanlop@stephentoub@iSazonov
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' Fix unzipping 4GB+ zip files by adamsitnik · Pull Request #77181 · dotnet/runtime · GitHub
Skip to content

Fix unzipping 4GB+ zip files - #77181

Merged
adamsitnik merged 8 commits into
dotnet:mainfrom
adamsitnik:zip4gb
Oct 28, 2022
Merged

Fix unzipping 4GB+ zip files#77181
adamsitnik merged 8 commits into
dotnet:mainfrom
adamsitnik:zip4gb

Conversation

@adamsitnik

@adamsitnikadamsitnik commented Oct 18, 2022

Copy link
Copy Markdown
Member

fixes#77159 without reverting #68106

I've tried to keep the changes as small as possible to make it easier to review and backport.

Explanation: for a file that is over 4 GB after zipping, some of the entries were reporting extraField.Size == 8, but readUncompressedSize, readCompressedSize and readStartDiskNumber equal false and readLocalHeaderOffset equal true. So what the code was doing, was reading first int64, not using it and exiting:

longvalue64=reader.ReadInt64();
if(readUncompressedSize)
zip64Block._uncompressedSize=value64;
if(ms.Position>extraField.Size-sizeof(long))
returntrue;

The fix is to allow advancing the stream (by reading from it) only when all 4 fields are provided (extraField.Size == 28) or when they are explicitly requested (the boolean flags set to true)

@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @dotnet/area-system-io-compression
See info in area-owners.md if you want to be subscribed.

Issue Details

fixes #77159 without reverting #68106

Author:adamsitnik
Assignees:-
Labels:

area-System.IO.Compression

Milestone:-

@ghostghost assigned adamsitnikOct 18, 2022
Comment threadsrc/libraries/System.IO.Compression/tests/ZipArchive/zip_LargeFiles.cs Outdated
@adamsitnik

Copy link
Copy Markdown
MemberAuthor

/azp list

@azure-pipelines

Copy link
Copy Markdown
CI/CD Pipelines for this repository:

@adamsitnik

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-libraries-coreclr outerloop

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

Comment threadsrc/libraries/System.IO.Compression/tests/ZipArchive/zip_LargeFiles.cs Outdated
Comment threadsrc/libraries/System.IO.Compression/tests/ZipArchive/zip_LargeFiles.cs Outdated
@danmoseley

Copy link
Copy Markdown
Contributor

Seems reasonable but I'd rather someone on the IO crew sign off on the change.

BTW if you touch this PR again maybe you could fix this line which I believe should check uncompressed length.
https://github.com/dotnet/runtime/pull/68106/files#diff-ea21d1af009443a658bc821a6eb47860f7887ff70155a68f11990ec64875a2dcR860

@adamsitnik

Copy link
Copy Markdown
MemberAuthor

Seems reasonable but I'd rather someone on the IO crew sign off on the change.

@jozkee PTAL

@jozkee

Copy link
Copy Markdown
Member

Taking a look...

Comment threadsrc/libraries/System.IO.Compression/tests/ZipArchive/zip_LargeFiles.cs Outdated
Comment threadsrc/libraries/System.IO.Compression/src/System/IO/Compression/ZipBlocks.cs Outdated
[OuterLoop("It requires almost 12 GB of free disk space")]
public static void UnzipOver4GBZipFile()
{
byte[] buffer = GC.AllocateUninitializedArray<byte>(1_000_000_000); // 1 GB

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Technically, this is not 1Gb, but since you are creating 6 files, the test is still valid.

I would rather create two tests, one that succeeds at the limit (say 3.999 Gb) without the fix and one that fails and then verify that both pass with your fix.

Since this is urgent for a backport, we can defer that.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Technically, this is not 1Gb, but since you are creating 6 files, the test is still valid.

You mean that $1000^3$ is not 1 GB but $1024^3$ is?

I would rather create two tests, one that succeeds at the limit (say 3.999 Gb) without the fix and one that fails and then verify that both pass with your fix.

My typical workflow is to add test that fails first, then make it pass. We have plenty of tests that verify < 4 GB archives so I don't see the value in adding it (also it would require us to use even more disk space and take longer to execute).

In the long term we could add more tests, especially such that verify that decompress(compress(x)) == x (which this test avoids, as it was taking 10 minutes to do so on my beefy PC).

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You mean that $1000^3$ is not 1 GB but $1024^3$ is?

Yes.

We have plenty of tests that verify < 4 GB archives so I don't see the value in adding it (also it would require us to use even more disk space and take longer to execute).

I meant to verify on the edge cases of the bug. In this case that would be 4Gb - 1, 4Gb, and 4Gb + 1; rather than a 6Gb case. But yes, that would multiply test execution time which is already quite large with this new test.

Comment on lines +202 to +207
// Advancing the stream (by reading from it) is possible only when:
// 1. There is an explicit ask to do that (valid files, corresponding boolean flag(s) set to true, #77159).
// 2. When the size indicates that all the information is available ("slightly invalid files", #49580).
bool readAllFields = extraField.Size >= sizeof(long) + sizeof(long) + sizeof(long) + sizeof(int);

long value64 = readUncompressedSize || readAllFields ? reader.ReadInt64() : -1;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could this fail if only 2 or 3 fields are available for reading but the bool params are false?

i.e: what if extraField.size = 16 and readUncompressedSize = false and readCompressedSize = false?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think that this is possible. @danmoseley has added following comment:

// The spec section 4.5.3:
// The order of the fields in the zip64 extended
// information record is fixed, but the fields MUST
// only appear if the corresponding Local or Central
// directory record field is set to 0xFFFF or 0xFFFFFFFF.
// However tools commonly write the fields anyway; the prevailing convention
// is to respect the size, but only actually use the values if their 32 bit
// values were all 0xFF.

My understanding is that we have two options:

  1. Correct archives: the fields are provided when the directory record fields are set (the right booleans are set to true, the size is their total size)
  2. "Slightly incorrect archives": the boolean fields are not set, but the provided size suggests that they are all present. Tests added in Fix reading slightly incorrect Zip files #68106 use such approach.

Comment threadsrc/libraries/System.IO.Compression/src/System/IO/Compression/ZipBlocks.cs Outdated
@adamsitnik

Copy link
Copy Markdown
MemberAuthor

}

// original values are unsigned, so implies value is too big to fit in signed integer
if (zip64Block._uncompressedSize < 0) throw new InvalidDataException(SR.FieldTooBigUncompressedSize);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These validations only occur if extraField.Size >= 28 since you are now returning if there's less than that.

Shouldn't we validate when reader.ReadInt* is called and/or if read of the field was requested?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is how it always was, I would prefer not to do it now as I intend to backport it to 7.0 and trying to make it as defensive/safe as possible.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The reason I was asking is because it wasn't always like that. It was changed to how it is now on #68106, which is the PR that introduced the regression.

[OuterLoop("It requires almost 12 GB of free disk space")]
public static void UnzipOver4GBZipFile()
{
byte[] buffer = GC.AllocateUninitializedArray<byte>(1_000_000_000); // 1 GB

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there a reasonable chance this will OOM? If so, consider catching that and skipping the test, rather than failing it.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there a reasonable chance this will OOM?

It should be executed only in Outerloop runs for the 64 bit "fast runtimes" and not in parallel with other tests so I don't think it's needed (at least for now).

[ConditionalFact(typeof(PlatformDetection),nameof(PlatformDetection.IsSpeedOptimized),nameof(PlatformDetection.Is64BitProcess))]// don't run it on slower runtimes

publicstaticboolIsSpeedOptimized=>!IsSizeOptimized;
publicstaticboolIsSizeOptimized=>IsBrowser||IsAndroid||IsAppleMobile;

@adamsitnik
adamsitnik merged commit 399c6dc into dotnet:mainOct 28, 2022
@adamsitnik

Copy link
Copy Markdown
MemberAuthor

/backport to release/7.0

@github-actions

Copy link
Copy Markdown
Contributor

Started backporting to release/7.0: https://github.com/dotnet/runtime/actions/runs/3347261332

@adamsitnik

Copy link
Copy Markdown
MemberAuthor

@danmoseley@jozkee@stephentoub thank you for the reviews!

@carlossanlopcarlossanlop left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Assuming the CI is green, this LGTM. I'd also wait for a sign-off from @jozkee.

Edit: Nevermind. This was merged while I was still adding a review. :)

}
else if (readAllFields)
{
_ = reader.ReadInt64();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For my own education: is there a perf improvement when using _ = ReadInt64(); to discard the value, compared to just not using _? Or is the idea here to explicitly signal that we are ignoring the value on purpose?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It shouldn't have a runtime effect, it's self documenting, both for humans and also for linters/analyzers that might otherwise flag an unused method result.

@ghostghost locked as resolved and limited conversation to collaborators Nov 28, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

DotNet 7 Regression when reading large Zip files

6 participants

@adamsitnik@danmoseley@jozkee@carlossanlop@stephentoub@iSazonov
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Fix unzipping 4GB+ zip files by adamsitnik · Pull Request #77181 · dotnet/runtime · GitHub
Skip to content

Fix unzipping 4GB+ zip files - #77181

Merged
adamsitnik merged 8 commits into
dotnet:mainfrom
adamsitnik:zip4gb
Oct 28, 2022
Merged

Fix unzipping 4GB+ zip files#77181
adamsitnik merged 8 commits into
dotnet:mainfrom
adamsitnik:zip4gb

Conversation

@adamsitnik

@adamsitnikadamsitnik commented Oct 18, 2022

Copy link
Copy Markdown
Member

fixes#77159 without reverting #68106

I've tried to keep the changes as small as possible to make it easier to review and backport.

Explanation: for a file that is over 4 GB after zipping, some of the entries were reporting extraField.Size == 8, but readUncompressedSize, readCompressedSize and readStartDiskNumber equal false and readLocalHeaderOffset equal true. So what the code was doing, was reading first int64, not using it and exiting:

longvalue64=reader.ReadInt64();
if(readUncompressedSize)
zip64Block._uncompressedSize=value64;
if(ms.Position>extraField.Size-sizeof(long))
returntrue;

The fix is to allow advancing the stream (by reading from it) only when all 4 fields are provided (extraField.Size == 28) or when they are explicitly requested (the boolean flags set to true)

@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @dotnet/area-system-io-compression
See info in area-owners.md if you want to be subscribed.

Issue Details

fixes #77159 without reverting #68106

Author:adamsitnik
Assignees:-
Labels:

area-System.IO.Compression

Milestone:-

@ghostghost assigned adamsitnikOct 18, 2022
Comment threadsrc/libraries/System.IO.Compression/tests/ZipArchive/zip_LargeFiles.cs Outdated
@adamsitnik

Copy link
Copy Markdown
MemberAuthor

/azp list

@azure-pipelines

Copy link
Copy Markdown
CI/CD Pipelines for this repository:

@adamsitnik

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-libraries-coreclr outerloop

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

Comment threadsrc/libraries/System.IO.Compression/tests/ZipArchive/zip_LargeFiles.cs Outdated
Comment threadsrc/libraries/System.IO.Compression/tests/ZipArchive/zip_LargeFiles.cs Outdated
@danmoseley

Copy link
Copy Markdown
Contributor

Seems reasonable but I'd rather someone on the IO crew sign off on the change.

BTW if you touch this PR again maybe you could fix this line which I believe should check uncompressed length.
https://github.com/dotnet/runtime/pull/68106/files#diff-ea21d1af009443a658bc821a6eb47860f7887ff70155a68f11990ec64875a2dcR860

@adamsitnik

Copy link
Copy Markdown
MemberAuthor

Seems reasonable but I'd rather someone on the IO crew sign off on the change.

@jozkee PTAL

@jozkee

Copy link
Copy Markdown
Member

Taking a look...

Comment threadsrc/libraries/System.IO.Compression/tests/ZipArchive/zip_LargeFiles.cs Outdated
Comment threadsrc/libraries/System.IO.Compression/src/System/IO/Compression/ZipBlocks.cs Outdated
[OuterLoop("It requires almost 12 GB of free disk space")]
public static void UnzipOver4GBZipFile()
{
byte[] buffer = GC.AllocateUninitializedArray<byte>(1_000_000_000); // 1 GB

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Technically, this is not 1Gb, but since you are creating 6 files, the test is still valid.

I would rather create two tests, one that succeeds at the limit (say 3.999 Gb) without the fix and one that fails and then verify that both pass with your fix.

Since this is urgent for a backport, we can defer that.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Technically, this is not 1Gb, but since you are creating 6 files, the test is still valid.

You mean that $1000^3$ is not 1 GB but $1024^3$ is?

I would rather create two tests, one that succeeds at the limit (say 3.999 Gb) without the fix and one that fails and then verify that both pass with your fix.

My typical workflow is to add test that fails first, then make it pass. We have plenty of tests that verify < 4 GB archives so I don't see the value in adding it (also it would require us to use even more disk space and take longer to execute).

In the long term we could add more tests, especially such that verify that decompress(compress(x)) == x (which this test avoids, as it was taking 10 minutes to do so on my beefy PC).

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You mean that $1000^3$ is not 1 GB but $1024^3$ is?

Yes.

We have plenty of tests that verify < 4 GB archives so I don't see the value in adding it (also it would require us to use even more disk space and take longer to execute).

I meant to verify on the edge cases of the bug. In this case that would be 4Gb - 1, 4Gb, and 4Gb + 1; rather than a 6Gb case. But yes, that would multiply test execution time which is already quite large with this new test.

Comment on lines +202 to +207
// Advancing the stream (by reading from it) is possible only when:
// 1. There is an explicit ask to do that (valid files, corresponding boolean flag(s) set to true, #77159).
// 2. When the size indicates that all the information is available ("slightly invalid files", #49580).
bool readAllFields = extraField.Size >= sizeof(long) + sizeof(long) + sizeof(long) + sizeof(int);

long value64 = readUncompressedSize || readAllFields ? reader.ReadInt64() : -1;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could this fail if only 2 or 3 fields are available for reading but the bool params are false?

i.e: what if extraField.size = 16 and readUncompressedSize = false and readCompressedSize = false?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think that this is possible. @danmoseley has added following comment:

// The spec section 4.5.3:
// The order of the fields in the zip64 extended
// information record is fixed, but the fields MUST
// only appear if the corresponding Local or Central
// directory record field is set to 0xFFFF or 0xFFFFFFFF.
// However tools commonly write the fields anyway; the prevailing convention
// is to respect the size, but only actually use the values if their 32 bit
// values were all 0xFF.

My understanding is that we have two options:

  1. Correct archives: the fields are provided when the directory record fields are set (the right booleans are set to true, the size is their total size)
  2. "Slightly incorrect archives": the boolean fields are not set, but the provided size suggests that they are all present. Tests added in Fix reading slightly incorrect Zip files #68106 use such approach.

Comment threadsrc/libraries/System.IO.Compression/src/System/IO/Compression/ZipBlocks.cs Outdated
@adamsitnik

Copy link
Copy Markdown
MemberAuthor

}

// original values are unsigned, so implies value is too big to fit in signed integer
if (zip64Block._uncompressedSize < 0) throw new InvalidDataException(SR.FieldTooBigUncompressedSize);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These validations only occur if extraField.Size >= 28 since you are now returning if there's less than that.

Shouldn't we validate when reader.ReadInt* is called and/or if read of the field was requested?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is how it always was, I would prefer not to do it now as I intend to backport it to 7.0 and trying to make it as defensive/safe as possible.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The reason I was asking is because it wasn't always like that. It was changed to how it is now on #68106, which is the PR that introduced the regression.

[OuterLoop("It requires almost 12 GB of free disk space")]
public static void UnzipOver4GBZipFile()
{
byte[] buffer = GC.AllocateUninitializedArray<byte>(1_000_000_000); // 1 GB

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there a reasonable chance this will OOM? If so, consider catching that and skipping the test, rather than failing it.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there a reasonable chance this will OOM?

It should be executed only in Outerloop runs for the 64 bit "fast runtimes" and not in parallel with other tests so I don't think it's needed (at least for now).

[ConditionalFact(typeof(PlatformDetection),nameof(PlatformDetection.IsSpeedOptimized),nameof(PlatformDetection.Is64BitProcess))]// don't run it on slower runtimes

publicstaticboolIsSpeedOptimized=>!IsSizeOptimized;
publicstaticboolIsSizeOptimized=>IsBrowser||IsAndroid||IsAppleMobile;

@adamsitnik
adamsitnik merged commit 399c6dc into dotnet:mainOct 28, 2022
@adamsitnik

Copy link
Copy Markdown
MemberAuthor

/backport to release/7.0

@github-actions

Copy link
Copy Markdown
Contributor

Started backporting to release/7.0: https://github.com/dotnet/runtime/actions/runs/3347261332

@adamsitnik

Copy link
Copy Markdown
MemberAuthor

@danmoseley@jozkee@stephentoub thank you for the reviews!

@carlossanlopcarlossanlop left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Assuming the CI is green, this LGTM. I'd also wait for a sign-off from @jozkee.

Edit: Nevermind. This was merged while I was still adding a review. :)

}
else if (readAllFields)
{
_ = reader.ReadInt64();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For my own education: is there a perf improvement when using _ = ReadInt64(); to discard the value, compared to just not using _? Or is the idea here to explicitly signal that we are ignoring the value on purpose?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It shouldn't have a runtime effect, it's self documenting, both for humans and also for linters/analyzers that might otherwise flag an unused method result.

@ghostghost locked as resolved and limited conversation to collaborators Nov 28, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

DotNet 7 Regression when reading large Zip files

6 participants

@adamsitnik@danmoseley@jozkee@carlossanlop@stephentoub@iSazonov
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); })(); Fix unzipping 4GB+ zip files by adamsitnik · Pull Request #77181 · dotnet/runtime · GitHub
Skip to content

Fix unzipping 4GB+ zip files - #77181

Merged
adamsitnik merged 8 commits into
dotnet:mainfrom
adamsitnik:zip4gb
Oct 28, 2022
Merged

Fix unzipping 4GB+ zip files#77181
adamsitnik merged 8 commits into
dotnet:mainfrom
adamsitnik:zip4gb

Conversation

@adamsitnik

@adamsitnikadamsitnik commented Oct 18, 2022

Copy link
Copy Markdown
Member

fixes#77159 without reverting #68106

I've tried to keep the changes as small as possible to make it easier to review and backport.

Explanation: for a file that is over 4 GB after zipping, some of the entries were reporting extraField.Size == 8, but readUncompressedSize, readCompressedSize and readStartDiskNumber equal false and readLocalHeaderOffset equal true. So what the code was doing, was reading first int64, not using it and exiting:

longvalue64=reader.ReadInt64();
if(readUncompressedSize)
zip64Block._uncompressedSize=value64;
if(ms.Position>extraField.Size-sizeof(long))
returntrue;

The fix is to allow advancing the stream (by reading from it) only when all 4 fields are provided (extraField.Size == 28) or when they are explicitly requested (the boolean flags set to true)

@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @dotnet/area-system-io-compression
See info in area-owners.md if you want to be subscribed.

Issue Details

fixes #77159 without reverting #68106

Author:adamsitnik
Assignees:-
Labels:

area-System.IO.Compression

Milestone:-

@ghostghost assigned adamsitnikOct 18, 2022
Comment threadsrc/libraries/System.IO.Compression/tests/ZipArchive/zip_LargeFiles.cs Outdated
@adamsitnik

Copy link
Copy Markdown
MemberAuthor

/azp list

@azure-pipelines

Copy link
Copy Markdown
CI/CD Pipelines for this repository:

@adamsitnik

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-libraries-coreclr outerloop

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

Comment threadsrc/libraries/System.IO.Compression/tests/ZipArchive/zip_LargeFiles.cs Outdated
Comment threadsrc/libraries/System.IO.Compression/tests/ZipArchive/zip_LargeFiles.cs Outdated
@danmoseley

Copy link
Copy Markdown
Contributor

Seems reasonable but I'd rather someone on the IO crew sign off on the change.

BTW if you touch this PR again maybe you could fix this line which I believe should check uncompressed length.
https://github.com/dotnet/runtime/pull/68106/files#diff-ea21d1af009443a658bc821a6eb47860f7887ff70155a68f11990ec64875a2dcR860

@adamsitnik

Copy link
Copy Markdown
MemberAuthor

Seems reasonable but I'd rather someone on the IO crew sign off on the change.

@jozkee PTAL

@jozkee

Copy link
Copy Markdown
Member

Taking a look...

Comment threadsrc/libraries/System.IO.Compression/tests/ZipArchive/zip_LargeFiles.cs Outdated
Comment threadsrc/libraries/System.IO.Compression/src/System/IO/Compression/ZipBlocks.cs Outdated
[OuterLoop("It requires almost 12 GB of free disk space")]
public static void UnzipOver4GBZipFile()
{
byte[] buffer = GC.AllocateUninitializedArray<byte>(1_000_000_000); // 1 GB

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Technically, this is not 1Gb, but since you are creating 6 files, the test is still valid.

I would rather create two tests, one that succeeds at the limit (say 3.999 Gb) without the fix and one that fails and then verify that both pass with your fix.

Since this is urgent for a backport, we can defer that.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Technically, this is not 1Gb, but since you are creating 6 files, the test is still valid.

You mean that $1000^3$ is not 1 GB but $1024^3$ is?

I would rather create two tests, one that succeeds at the limit (say 3.999 Gb) without the fix and one that fails and then verify that both pass with your fix.

My typical workflow is to add test that fails first, then make it pass. We have plenty of tests that verify < 4 GB archives so I don't see the value in adding it (also it would require us to use even more disk space and take longer to execute).

In the long term we could add more tests, especially such that verify that decompress(compress(x)) == x (which this test avoids, as it was taking 10 minutes to do so on my beefy PC).

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You mean that $1000^3$ is not 1 GB but $1024^3$ is?

Yes.

We have plenty of tests that verify < 4 GB archives so I don't see the value in adding it (also it would require us to use even more disk space and take longer to execute).

I meant to verify on the edge cases of the bug. In this case that would be 4Gb - 1, 4Gb, and 4Gb + 1; rather than a 6Gb case. But yes, that would multiply test execution time which is already quite large with this new test.

Comment on lines +202 to +207
// Advancing the stream (by reading from it) is possible only when:
// 1. There is an explicit ask to do that (valid files, corresponding boolean flag(s) set to true, #77159).
// 2. When the size indicates that all the information is available ("slightly invalid files", #49580).
bool readAllFields = extraField.Size >= sizeof(long) + sizeof(long) + sizeof(long) + sizeof(int);

long value64 = readUncompressedSize || readAllFields ? reader.ReadInt64() : -1;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could this fail if only 2 or 3 fields are available for reading but the bool params are false?

i.e: what if extraField.size = 16 and readUncompressedSize = false and readCompressedSize = false?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think that this is possible. @danmoseley has added following comment:

// The spec section 4.5.3:
// The order of the fields in the zip64 extended
// information record is fixed, but the fields MUST
// only appear if the corresponding Local or Central
// directory record field is set to 0xFFFF or 0xFFFFFFFF.
// However tools commonly write the fields anyway; the prevailing convention
// is to respect the size, but only actually use the values if their 32 bit
// values were all 0xFF.

My understanding is that we have two options:

  1. Correct archives: the fields are provided when the directory record fields are set (the right booleans are set to true, the size is their total size)
  2. "Slightly incorrect archives": the boolean fields are not set, but the provided size suggests that they are all present. Tests added in Fix reading slightly incorrect Zip files #68106 use such approach.

Comment threadsrc/libraries/System.IO.Compression/src/System/IO/Compression/ZipBlocks.cs Outdated
@adamsitnik

Copy link
Copy Markdown
MemberAuthor

}

// original values are unsigned, so implies value is too big to fit in signed integer
if (zip64Block._uncompressedSize < 0) throw new InvalidDataException(SR.FieldTooBigUncompressedSize);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These validations only occur if extraField.Size >= 28 since you are now returning if there's less than that.

Shouldn't we validate when reader.ReadInt* is called and/or if read of the field was requested?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is how it always was, I would prefer not to do it now as I intend to backport it to 7.0 and trying to make it as defensive/safe as possible.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The reason I was asking is because it wasn't always like that. It was changed to how it is now on #68106, which is the PR that introduced the regression.

[OuterLoop("It requires almost 12 GB of free disk space")]
public static void UnzipOver4GBZipFile()
{
byte[] buffer = GC.AllocateUninitializedArray<byte>(1_000_000_000); // 1 GB

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there a reasonable chance this will OOM? If so, consider catching that and skipping the test, rather than failing it.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there a reasonable chance this will OOM?

It should be executed only in Outerloop runs for the 64 bit "fast runtimes" and not in parallel with other tests so I don't think it's needed (at least for now).

[ConditionalFact(typeof(PlatformDetection),nameof(PlatformDetection.IsSpeedOptimized),nameof(PlatformDetection.Is64BitProcess))]// don't run it on slower runtimes

publicstaticboolIsSpeedOptimized=>!IsSizeOptimized;
publicstaticboolIsSizeOptimized=>IsBrowser||IsAndroid||IsAppleMobile;

@adamsitnik
adamsitnik merged commit 399c6dc into dotnet:mainOct 28, 2022
@adamsitnik

Copy link
Copy Markdown
MemberAuthor

/backport to release/7.0

@github-actions

Copy link
Copy Markdown
Contributor

Started backporting to release/7.0: https://github.com/dotnet/runtime/actions/runs/3347261332

@adamsitnik

Copy link
Copy Markdown
MemberAuthor

@danmoseley@jozkee@stephentoub thank you for the reviews!

@carlossanlopcarlossanlop left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Assuming the CI is green, this LGTM. I'd also wait for a sign-off from @jozkee.

Edit: Nevermind. This was merged while I was still adding a review. :)

}
else if (readAllFields)
{
_ = reader.ReadInt64();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For my own education: is there a perf improvement when using _ = ReadInt64(); to discard the value, compared to just not using _? Or is the idea here to explicitly signal that we are ignoring the value on purpose?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It shouldn't have a runtime effect, it's self documenting, both for humans and also for linters/analyzers that might otherwise flag an unused method result.

@ghostghost locked as resolved and limited conversation to collaborators Nov 28, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

DotNet 7 Regression when reading large Zip files

6 participants

@adamsitnik@danmoseley@jozkee@carlossanlop@stephentoub@iSazonov