GH-40037: [C++][FS][Azure] Make attempted reads and writes against directories fail fast - #40119

Merged
felipecrv merged 20 commits into
apache:mainfrom
Tom-Newton:tomnewton/check_for_directory_marker_metadata/GH-40037
Feb 21, 2024
Merged

GH-40037: [C++][FS][Azure] Make attempted reads and writes against directories fail fast#40119
felipecrv merged 20 commits into
apache:mainfrom
Tom-Newton:tomnewton/check_for_directory_marker_metadata/GH-40037

Conversation

@Tom-Newton

@Tom-NewtonTom-Newton commented Feb 18, 2024

Copy link
Copy Markdown
Contributor

Rationale for this change

Prevent confusion if a user attempts to read or write a directory.

What changes are included in this PR?

  • Make ObjectAppendStream::Flush a noop if ObjectAppendStream::Init has not run successfully. This avoids an unhandled error when the destructor calls flush.
  • Check blob properties for directory marker metadata when initialising ObjectInputFile or ObjectAppendStream.
  • When initialising ObjectAppendStream call GetFileInfo if it is a flat namespace account.

Are these changes tested?

Add new tests DisallowReadingOrWritingDirectoryMarkers and DisallowCreatingFileAndDirectoryWithTheSameName to cover the new fail fast behaviour.
Also updated WriteMetadata to ensure that my changes to Flush didn't break setting metadata without calling Write on the stream.

Are there any user-facing changes?

Yes. Invalid read and write operations will now fail fast and gracefully. Previously could get into a confusing state where there were files and directories at the same path and there were some un-graceful failures.

@Tom-Newton

Copy link
Copy Markdown
ContributorAuthor

@felipecrv I expect you will be interested in reviewing this.

Comment threadcpp/src/arrow/filesystem/azurefs.cc Outdated
@github-actionsgithub-actionsBot added awaiting committer review Awaiting committer review and removed awaiting review Awaiting review labels Feb 19, 2024
DCHECK_GE(content_length_, 0);
pos_ = content_length_;
Status Init(const bool truncate,
std::function<Status()> ensure_not_flat_namespace_directory) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You can inject AzureFileSystem *azure_file_system here and not have to allocate a closure for this. You would call AzureFileSystem::Impl::EnsureNotFlatNamespaceDirectory(location) via azure_file_system->impl_ (accessible because the handles produced by the azure file system can be friends with the filesystem class).

@Tom-NewtonTom-NewtonFeb 20, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the extra info. I was planning to do this but I was struggling with the friends thing.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The trick is to forward-declare the class as well.

diff --git a/cpp/src/arrow/filesystem/azurefs.h b/cpp/src/arrow/filesystem/azurefs.h
index 2a131e40c..d48ef9dd7 100644
--- a/cpp/src/arrow/filesystem/azurefs.h
+++ b/cpp/src/arrow/filesystem/azurefs.h
@@ -44,6 +44,7 @@ classDataLakeServiceClient;
namespacearrow::fs {
+classObjectAppendStream;
classTestAzureFileSystem;
/// Options for the AzureFileSystem implementation.
@@ -180,6 +181,7 @@ classARROW_EXPORT AzureFileSystem : public FileSystem {
explicitAzureFileSystem(std::unique_ptr<Impl>&& impl);
+ friendclassObjectAppendStream;
friendclassTestAzureFileSystem;
voidForceCachedHierarchicalNamespaceSupport(int hns_support);

@Tom-NewtonTom-NewtonFeb 21, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think my main problem was that ObjectAppendStream is defined inside an anonymous namespace but I still haven't got it working as you describe.

Are you suggesting to use AzureFileSystem *azure_file_system or AzureFileSystem:Impl *azure_file_system as the argument to ObjectAppendStream::Impl. I don't know how I can get a AzureFileSystem pointer from inside AzureFileSystem::Impl and using AzureFileSystem::Impl as the argument leads to incomplete type errors which I don't think I can avoid.

Also if you wouldn't mind I would be interested to know what the disadvantage of a lambda function is compared to what you proposed.

Sorry about my lacking C++ knowledge here.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also if you wouldn't mind I would be interested to know what the disadvantage of a lambda function is compared to what you proposed.

To create the std::function, you heap allocate an object with copies of the values in the capture list and generate a lot more extra code in the binary:

class function {
T valuesfromthecpapturelist;
RetType operator()(ArgsType ...) {...};
}

When you think about an std::function this way (a pair of context data and a function), you realize the class you already serves that purpose.

But hey, this is becoming challenging, so I won't hold the PR anymore because of this. Moving to Init() was a big step in the right direction.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for explaining

Comment threadcpp/src/arrow/filesystem/azurefs.cc Outdated
DCHECK_GE(content_length_, 0);
pos_ = content_length_;
Status Init(const bool truncate,
std::function<Status()> ensure_not_flat_namespace_directory) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also if you wouldn't mind I would be interested to know what the disadvantage of a lambda function is compared to what you proposed.

To create the std::function, you heap allocate an object with copies of the values in the capture list and generate a lot more extra code in the binary:

class function {
T valuesfromthecpapturelist;
RetType operator()(ArgsType ...) {...};
}

When you think about an std::function this way (a pair of context data and a function), you realize the class you already serves that purpose.

But hey, this is becoming challenging, so I won't hold the PR anymore because of this. Moving to Init() was a big step in the right direction.

Comment on lines +1606 to +1607
TYPED_TEST(TestAzureFileSystemOnAllScenarios,
OpenOutputStreamWithMissingContainer) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is causing the linter to fail @Tom-Newton. Please fix and I will merge.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed

@felipecrv
felipecrv merged commit 8a62f30 into apache:mainFeb 21, 2024
@felipecrvfelipecrv removed the awaiting committer review Awaiting committer review label Feb 21, 2024
@github-actionsgithub-actionsBot added the awaiting committer review Awaiting committer review label Feb 21, 2024
@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@austin3dickey

Copy link
Copy Markdown
Contributor

My apologies for the noise. Looking into this now.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@austin3dickey

Copy link
Copy Markdown
Contributor

That should be the last one. Sorry again about the noise!

@felipecrv

Copy link
Copy Markdown
Contributor

@austin3dickey no worries :)

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[C++][FS][Azure] Prevent reading or writing directory marker blobs

3 participants

@Tom-Newton@austin3dickey@felipecrv
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

GH-40037: [C++][FS][Azure] Make attempted reads and writes against directories fail fast - #40119

Merged
felipecrv merged 20 commits into
apache:mainfrom
Tom-Newton:tomnewton/check_for_directory_marker_metadata/GH-40037
Feb 21, 2024
Merged

GH-40037: [C++][FS][Azure] Make attempted reads and writes against directories fail fast#40119
felipecrv merged 20 commits into
apache:mainfrom
Tom-Newton:tomnewton/check_for_directory_marker_metadata/GH-40037

Conversation

@Tom-Newton

@Tom-NewtonTom-Newton commented Feb 18, 2024

Copy link
Copy Markdown
Contributor

Rationale for this change

Prevent confusion if a user attempts to read or write a directory.

What changes are included in this PR?

  • Make ObjectAppendStream::Flush a noop if ObjectAppendStream::Init has not run successfully. This avoids an unhandled error when the destructor calls flush.
  • Check blob properties for directory marker metadata when initialising ObjectInputFile or ObjectAppendStream.
  • When initialising ObjectAppendStream call GetFileInfo if it is a flat namespace account.

Are these changes tested?

Add new tests DisallowReadingOrWritingDirectoryMarkers and DisallowCreatingFileAndDirectoryWithTheSameName to cover the new fail fast behaviour.
Also updated WriteMetadata to ensure that my changes to Flush didn't break setting metadata without calling Write on the stream.

Are there any user-facing changes?

Yes. Invalid read and write operations will now fail fast and gracefully. Previously could get into a confusing state where there were files and directories at the same path and there were some un-graceful failures.

@Tom-Newton

Copy link
Copy Markdown
ContributorAuthor

@felipecrv I expect you will be interested in reviewing this.

Comment threadcpp/src/arrow/filesystem/azurefs.cc Outdated
@github-actionsgithub-actionsBot added awaiting committer review Awaiting committer review and removed awaiting review Awaiting review labels Feb 19, 2024
DCHECK_GE(content_length_, 0);
pos_ = content_length_;
Status Init(const bool truncate,
std::function<Status()> ensure_not_flat_namespace_directory) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You can inject AzureFileSystem *azure_file_system here and not have to allocate a closure for this. You would call AzureFileSystem::Impl::EnsureNotFlatNamespaceDirectory(location) via azure_file_system->impl_ (accessible because the handles produced by the azure file system can be friends with the filesystem class).

@Tom-NewtonTom-NewtonFeb 20, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the extra info. I was planning to do this but I was struggling with the friends thing.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The trick is to forward-declare the class as well.

diff --git a/cpp/src/arrow/filesystem/azurefs.h b/cpp/src/arrow/filesystem/azurefs.h
index 2a131e40c..d48ef9dd7 100644
--- a/cpp/src/arrow/filesystem/azurefs.h
+++ b/cpp/src/arrow/filesystem/azurefs.h
@@ -44,6 +44,7 @@ classDataLakeServiceClient;
namespacearrow::fs {
+classObjectAppendStream;
classTestAzureFileSystem;
/// Options for the AzureFileSystem implementation.
@@ -180,6 +181,7 @@ classARROW_EXPORT AzureFileSystem : public FileSystem {
explicitAzureFileSystem(std::unique_ptr<Impl>&& impl);
+ friendclassObjectAppendStream;
friendclassTestAzureFileSystem;
voidForceCachedHierarchicalNamespaceSupport(int hns_support);

@Tom-NewtonTom-NewtonFeb 21, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think my main problem was that ObjectAppendStream is defined inside an anonymous namespace but I still haven't got it working as you describe.

Are you suggesting to use AzureFileSystem *azure_file_system or AzureFileSystem:Impl *azure_file_system as the argument to ObjectAppendStream::Impl. I don't know how I can get a AzureFileSystem pointer from inside AzureFileSystem::Impl and using AzureFileSystem::Impl as the argument leads to incomplete type errors which I don't think I can avoid.

Also if you wouldn't mind I would be interested to know what the disadvantage of a lambda function is compared to what you proposed.

Sorry about my lacking C++ knowledge here.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also if you wouldn't mind I would be interested to know what the disadvantage of a lambda function is compared to what you proposed.

To create the std::function, you heap allocate an object with copies of the values in the capture list and generate a lot more extra code in the binary:

class function {
T valuesfromthecpapturelist;
RetType operator()(ArgsType ...) {...};
}

When you think about an std::function this way (a pair of context data and a function), you realize the class you already serves that purpose.

But hey, this is becoming challenging, so I won't hold the PR anymore because of this. Moving to Init() was a big step in the right direction.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for explaining

Comment threadcpp/src/arrow/filesystem/azurefs.cc Outdated
DCHECK_GE(content_length_, 0);
pos_ = content_length_;
Status Init(const bool truncate,
std::function<Status()> ensure_not_flat_namespace_directory) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also if you wouldn't mind I would be interested to know what the disadvantage of a lambda function is compared to what you proposed.

To create the std::function, you heap allocate an object with copies of the values in the capture list and generate a lot more extra code in the binary:

class function {
T valuesfromthecpapturelist;
RetType operator()(ArgsType ...) {...};
}

When you think about an std::function this way (a pair of context data and a function), you realize the class you already serves that purpose.

But hey, this is becoming challenging, so I won't hold the PR anymore because of this. Moving to Init() was a big step in the right direction.

Comment on lines +1606 to +1607
TYPED_TEST(TestAzureFileSystemOnAllScenarios,
OpenOutputStreamWithMissingContainer) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is causing the linter to fail @Tom-Newton. Please fix and I will merge.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed

@felipecrv
felipecrv merged commit 8a62f30 into apache:mainFeb 21, 2024
@felipecrvfelipecrv removed the awaiting committer review Awaiting committer review label Feb 21, 2024
@github-actionsgithub-actionsBot added the awaiting committer review Awaiting committer review label Feb 21, 2024
@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@austin3dickey

Copy link
Copy Markdown
Contributor

My apologies for the noise. Looking into this now.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@austin3dickey

Copy link
Copy Markdown
Contributor

That should be the last one. Sorry again about the noise!

@felipecrv

Copy link
Copy Markdown
Contributor

@austin3dickey no worries :)

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[C++][FS][Azure] Prevent reading or writing directory marker blobs

3 participants

@Tom-Newton@austin3dickey@felipecrv
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

GH-40037: [C++][FS][Azure] Make attempted reads and writes against directories fail fast - #40119

Merged
felipecrv merged 20 commits into
apache:mainfrom
Tom-Newton:tomnewton/check_for_directory_marker_metadata/GH-40037
Feb 21, 2024
Merged

GH-40037: [C++][FS][Azure] Make attempted reads and writes against directories fail fast#40119
felipecrv merged 20 commits into
apache:mainfrom
Tom-Newton:tomnewton/check_for_directory_marker_metadata/GH-40037

Conversation

@Tom-Newton

@Tom-NewtonTom-Newton commented Feb 18, 2024

Copy link
Copy Markdown
Contributor

Rationale for this change

Prevent confusion if a user attempts to read or write a directory.

What changes are included in this PR?

  • Make ObjectAppendStream::Flush a noop if ObjectAppendStream::Init has not run successfully. This avoids an unhandled error when the destructor calls flush.
  • Check blob properties for directory marker metadata when initialising ObjectInputFile or ObjectAppendStream.
  • When initialising ObjectAppendStream call GetFileInfo if it is a flat namespace account.

Are these changes tested?

Add new tests DisallowReadingOrWritingDirectoryMarkers and DisallowCreatingFileAndDirectoryWithTheSameName to cover the new fail fast behaviour.
Also updated WriteMetadata to ensure that my changes to Flush didn't break setting metadata without calling Write on the stream.

Are there any user-facing changes?

Yes. Invalid read and write operations will now fail fast and gracefully. Previously could get into a confusing state where there were files and directories at the same path and there were some un-graceful failures.

@Tom-Newton

Copy link
Copy Markdown
ContributorAuthor

@felipecrv I expect you will be interested in reviewing this.

Comment threadcpp/src/arrow/filesystem/azurefs.cc Outdated
@github-actionsgithub-actionsBot added awaiting committer review Awaiting committer review and removed awaiting review Awaiting review labels Feb 19, 2024
DCHECK_GE(content_length_, 0);
pos_ = content_length_;
Status Init(const bool truncate,
std::function<Status()> ensure_not_flat_namespace_directory) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You can inject AzureFileSystem *azure_file_system here and not have to allocate a closure for this. You would call AzureFileSystem::Impl::EnsureNotFlatNamespaceDirectory(location) via azure_file_system->impl_ (accessible because the handles produced by the azure file system can be friends with the filesystem class).

@Tom-NewtonTom-NewtonFeb 20, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the extra info. I was planning to do this but I was struggling with the friends thing.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The trick is to forward-declare the class as well.

diff --git a/cpp/src/arrow/filesystem/azurefs.h b/cpp/src/arrow/filesystem/azurefs.h
index 2a131e40c..d48ef9dd7 100644
--- a/cpp/src/arrow/filesystem/azurefs.h
+++ b/cpp/src/arrow/filesystem/azurefs.h
@@ -44,6 +44,7 @@ classDataLakeServiceClient;
namespacearrow::fs {
+classObjectAppendStream;
classTestAzureFileSystem;
/// Options for the AzureFileSystem implementation.
@@ -180,6 +181,7 @@ classARROW_EXPORT AzureFileSystem : public FileSystem {
explicitAzureFileSystem(std::unique_ptr<Impl>&& impl);
+ friendclassObjectAppendStream;
friendclassTestAzureFileSystem;
voidForceCachedHierarchicalNamespaceSupport(int hns_support);

@Tom-NewtonTom-NewtonFeb 21, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think my main problem was that ObjectAppendStream is defined inside an anonymous namespace but I still haven't got it working as you describe.

Are you suggesting to use AzureFileSystem *azure_file_system or AzureFileSystem:Impl *azure_file_system as the argument to ObjectAppendStream::Impl. I don't know how I can get a AzureFileSystem pointer from inside AzureFileSystem::Impl and using AzureFileSystem::Impl as the argument leads to incomplete type errors which I don't think I can avoid.

Also if you wouldn't mind I would be interested to know what the disadvantage of a lambda function is compared to what you proposed.

Sorry about my lacking C++ knowledge here.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also if you wouldn't mind I would be interested to know what the disadvantage of a lambda function is compared to what you proposed.

To create the std::function, you heap allocate an object with copies of the values in the capture list and generate a lot more extra code in the binary:

class function {
T valuesfromthecpapturelist;
RetType operator()(ArgsType ...) {...};
}

When you think about an std::function this way (a pair of context data and a function), you realize the class you already serves that purpose.

But hey, this is becoming challenging, so I won't hold the PR anymore because of this. Moving to Init() was a big step in the right direction.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for explaining

Comment threadcpp/src/arrow/filesystem/azurefs.cc Outdated
DCHECK_GE(content_length_, 0);
pos_ = content_length_;
Status Init(const bool truncate,
std::function<Status()> ensure_not_flat_namespace_directory) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also if you wouldn't mind I would be interested to know what the disadvantage of a lambda function is compared to what you proposed.

To create the std::function, you heap allocate an object with copies of the values in the capture list and generate a lot more extra code in the binary:

class function {
T valuesfromthecpapturelist;
RetType operator()(ArgsType ...) {...};
}

When you think about an std::function this way (a pair of context data and a function), you realize the class you already serves that purpose.

But hey, this is becoming challenging, so I won't hold the PR anymore because of this. Moving to Init() was a big step in the right direction.

Comment on lines +1606 to +1607
TYPED_TEST(TestAzureFileSystemOnAllScenarios,
OpenOutputStreamWithMissingContainer) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is causing the linter to fail @Tom-Newton. Please fix and I will merge.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed

@felipecrv
felipecrv merged commit 8a62f30 into apache:mainFeb 21, 2024
@felipecrvfelipecrv removed the awaiting committer review Awaiting committer review label Feb 21, 2024
@github-actionsgithub-actionsBot added the awaiting committer review Awaiting committer review label Feb 21, 2024
@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@austin3dickey

Copy link
Copy Markdown
Contributor

My apologies for the noise. Looking into this now.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@austin3dickey

Copy link
Copy Markdown
Contributor

That should be the last one. Sorry again about the noise!

@felipecrv

Copy link
Copy Markdown
Contributor

@austin3dickey no worries :)

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[C++][FS][Azure] Prevent reading or writing directory marker blobs

3 participants

@Tom-Newton@austin3dickey@felipecrv
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

GH-40037: [C++][FS][Azure] Make attempted reads and writes against directories fail fast - #40119

Merged
felipecrv merged 20 commits into
apache:mainfrom
Tom-Newton:tomnewton/check_for_directory_marker_metadata/GH-40037
Feb 21, 2024
Merged

GH-40037: [C++][FS][Azure] Make attempted reads and writes against directories fail fast#40119
felipecrv merged 20 commits into
apache:mainfrom
Tom-Newton:tomnewton/check_for_directory_marker_metadata/GH-40037

Conversation

@Tom-Newton

@Tom-NewtonTom-Newton commented Feb 18, 2024

Copy link
Copy Markdown
Contributor

Rationale for this change

Prevent confusion if a user attempts to read or write a directory.

What changes are included in this PR?

  • Make ObjectAppendStream::Flush a noop if ObjectAppendStream::Init has not run successfully. This avoids an unhandled error when the destructor calls flush.
  • Check blob properties for directory marker metadata when initialising ObjectInputFile or ObjectAppendStream.
  • When initialising ObjectAppendStream call GetFileInfo if it is a flat namespace account.

Are these changes tested?

Add new tests DisallowReadingOrWritingDirectoryMarkers and DisallowCreatingFileAndDirectoryWithTheSameName to cover the new fail fast behaviour.
Also updated WriteMetadata to ensure that my changes to Flush didn't break setting metadata without calling Write on the stream.

Are there any user-facing changes?

Yes. Invalid read and write operations will now fail fast and gracefully. Previously could get into a confusing state where there were files and directories at the same path and there were some un-graceful failures.

@Tom-Newton

Copy link
Copy Markdown
ContributorAuthor

@felipecrv I expect you will be interested in reviewing this.

Comment threadcpp/src/arrow/filesystem/azurefs.cc Outdated
@github-actionsgithub-actionsBot added awaiting committer review Awaiting committer review and removed awaiting review Awaiting review labels Feb 19, 2024
DCHECK_GE(content_length_, 0);
pos_ = content_length_;
Status Init(const bool truncate,
std::function<Status()> ensure_not_flat_namespace_directory) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You can inject AzureFileSystem *azure_file_system here and not have to allocate a closure for this. You would call AzureFileSystem::Impl::EnsureNotFlatNamespaceDirectory(location) via azure_file_system->impl_ (accessible because the handles produced by the azure file system can be friends with the filesystem class).

@Tom-NewtonTom-NewtonFeb 20, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the extra info. I was planning to do this but I was struggling with the friends thing.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The trick is to forward-declare the class as well.

diff --git a/cpp/src/arrow/filesystem/azurefs.h b/cpp/src/arrow/filesystem/azurefs.h
index 2a131e40c..d48ef9dd7 100644
--- a/cpp/src/arrow/filesystem/azurefs.h
+++ b/cpp/src/arrow/filesystem/azurefs.h
@@ -44,6 +44,7 @@ classDataLakeServiceClient;
namespacearrow::fs {
+classObjectAppendStream;
classTestAzureFileSystem;
/// Options for the AzureFileSystem implementation.
@@ -180,6 +181,7 @@ classARROW_EXPORT AzureFileSystem : public FileSystem {
explicitAzureFileSystem(std::unique_ptr<Impl>&& impl);
+ friendclassObjectAppendStream;
friendclassTestAzureFileSystem;
voidForceCachedHierarchicalNamespaceSupport(int hns_support);

@Tom-NewtonTom-NewtonFeb 21, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think my main problem was that ObjectAppendStream is defined inside an anonymous namespace but I still haven't got it working as you describe.

Are you suggesting to use AzureFileSystem *azure_file_system or AzureFileSystem:Impl *azure_file_system as the argument to ObjectAppendStream::Impl. I don't know how I can get a AzureFileSystem pointer from inside AzureFileSystem::Impl and using AzureFileSystem::Impl as the argument leads to incomplete type errors which I don't think I can avoid.

Also if you wouldn't mind I would be interested to know what the disadvantage of a lambda function is compared to what you proposed.

Sorry about my lacking C++ knowledge here.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also if you wouldn't mind I would be interested to know what the disadvantage of a lambda function is compared to what you proposed.

To create the std::function, you heap allocate an object with copies of the values in the capture list and generate a lot more extra code in the binary:

class function {
T valuesfromthecpapturelist;
RetType operator()(ArgsType ...) {...};
}

When you think about an std::function this way (a pair of context data and a function), you realize the class you already serves that purpose.

But hey, this is becoming challenging, so I won't hold the PR anymore because of this. Moving to Init() was a big step in the right direction.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for explaining

Comment threadcpp/src/arrow/filesystem/azurefs.cc Outdated
DCHECK_GE(content_length_, 0);
pos_ = content_length_;
Status Init(const bool truncate,
std::function<Status()> ensure_not_flat_namespace_directory) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also if you wouldn't mind I would be interested to know what the disadvantage of a lambda function is compared to what you proposed.

To create the std::function, you heap allocate an object with copies of the values in the capture list and generate a lot more extra code in the binary:

class function {
T valuesfromthecpapturelist;
RetType operator()(ArgsType ...) {...};
}

When you think about an std::function this way (a pair of context data and a function), you realize the class you already serves that purpose.

But hey, this is becoming challenging, so I won't hold the PR anymore because of this. Moving to Init() was a big step in the right direction.

Comment on lines +1606 to +1607
TYPED_TEST(TestAzureFileSystemOnAllScenarios,
OpenOutputStreamWithMissingContainer) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is causing the linter to fail @Tom-Newton. Please fix and I will merge.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed

@felipecrv
felipecrv merged commit 8a62f30 into apache:mainFeb 21, 2024
@felipecrvfelipecrv removed the awaiting committer review Awaiting committer review label Feb 21, 2024
@github-actionsgithub-actionsBot added the awaiting committer review Awaiting committer review label Feb 21, 2024
@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@austin3dickey

Copy link
Copy Markdown
Contributor

My apologies for the noise. Looking into this now.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@austin3dickey

Copy link
Copy Markdown
Contributor

That should be the last one. Sorry again about the noise!

@felipecrv

Copy link
Copy Markdown
Contributor

@austin3dickey no worries :)

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[C++][FS][Azure] Prevent reading or writing directory marker blobs

3 participants

@Tom-Newton@austin3dickey@felipecrv
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

GH-40037: [C++][FS][Azure] Make attempted reads and writes against directories fail fast - #40119

Merged
felipecrv merged 20 commits into
apache:mainfrom
Tom-Newton:tomnewton/check_for_directory_marker_metadata/GH-40037
Feb 21, 2024
Merged

GH-40037: [C++][FS][Azure] Make attempted reads and writes against directories fail fast#40119
felipecrv merged 20 commits into
apache:mainfrom
Tom-Newton:tomnewton/check_for_directory_marker_metadata/GH-40037

Conversation

@Tom-Newton

@Tom-NewtonTom-Newton commented Feb 18, 2024

Copy link
Copy Markdown
Contributor

Rationale for this change

Prevent confusion if a user attempts to read or write a directory.

What changes are included in this PR?

  • Make ObjectAppendStream::Flush a noop if ObjectAppendStream::Init has not run successfully. This avoids an unhandled error when the destructor calls flush.
  • Check blob properties for directory marker metadata when initialising ObjectInputFile or ObjectAppendStream.
  • When initialising ObjectAppendStream call GetFileInfo if it is a flat namespace account.

Are these changes tested?

Add new tests DisallowReadingOrWritingDirectoryMarkers and DisallowCreatingFileAndDirectoryWithTheSameName to cover the new fail fast behaviour.
Also updated WriteMetadata to ensure that my changes to Flush didn't break setting metadata without calling Write on the stream.

Are there any user-facing changes?

Yes. Invalid read and write operations will now fail fast and gracefully. Previously could get into a confusing state where there were files and directories at the same path and there were some un-graceful failures.

@Tom-Newton

Copy link
Copy Markdown
ContributorAuthor

@felipecrv I expect you will be interested in reviewing this.

Comment threadcpp/src/arrow/filesystem/azurefs.cc Outdated
@github-actionsgithub-actionsBot added awaiting committer review Awaiting committer review and removed awaiting review Awaiting review labels Feb 19, 2024
DCHECK_GE(content_length_, 0);
pos_ = content_length_;
Status Init(const bool truncate,
std::function<Status()> ensure_not_flat_namespace_directory) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You can inject AzureFileSystem *azure_file_system here and not have to allocate a closure for this. You would call AzureFileSystem::Impl::EnsureNotFlatNamespaceDirectory(location) via azure_file_system->impl_ (accessible because the handles produced by the azure file system can be friends with the filesystem class).

@Tom-NewtonTom-NewtonFeb 20, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the extra info. I was planning to do this but I was struggling with the friends thing.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The trick is to forward-declare the class as well.

diff --git a/cpp/src/arrow/filesystem/azurefs.h b/cpp/src/arrow/filesystem/azurefs.h
index 2a131e40c..d48ef9dd7 100644
--- a/cpp/src/arrow/filesystem/azurefs.h
+++ b/cpp/src/arrow/filesystem/azurefs.h
@@ -44,6 +44,7 @@ classDataLakeServiceClient;
namespacearrow::fs {
+classObjectAppendStream;
classTestAzureFileSystem;
/// Options for the AzureFileSystem implementation.
@@ -180,6 +181,7 @@ classARROW_EXPORT AzureFileSystem : public FileSystem {
explicitAzureFileSystem(std::unique_ptr<Impl>&& impl);
+ friendclassObjectAppendStream;
friendclassTestAzureFileSystem;
voidForceCachedHierarchicalNamespaceSupport(int hns_support);

@Tom-NewtonTom-NewtonFeb 21, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think my main problem was that ObjectAppendStream is defined inside an anonymous namespace but I still haven't got it working as you describe.

Are you suggesting to use AzureFileSystem *azure_file_system or AzureFileSystem:Impl *azure_file_system as the argument to ObjectAppendStream::Impl. I don't know how I can get a AzureFileSystem pointer from inside AzureFileSystem::Impl and using AzureFileSystem::Impl as the argument leads to incomplete type errors which I don't think I can avoid.

Also if you wouldn't mind I would be interested to know what the disadvantage of a lambda function is compared to what you proposed.

Sorry about my lacking C++ knowledge here.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also if you wouldn't mind I would be interested to know what the disadvantage of a lambda function is compared to what you proposed.

To create the std::function, you heap allocate an object with copies of the values in the capture list and generate a lot more extra code in the binary:

class function {
T valuesfromthecpapturelist;
RetType operator()(ArgsType ...) {...};
}

When you think about an std::function this way (a pair of context data and a function), you realize the class you already serves that purpose.

But hey, this is becoming challenging, so I won't hold the PR anymore because of this. Moving to Init() was a big step in the right direction.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for explaining

Comment threadcpp/src/arrow/filesystem/azurefs.cc Outdated
DCHECK_GE(content_length_, 0);
pos_ = content_length_;
Status Init(const bool truncate,
std::function<Status()> ensure_not_flat_namespace_directory) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also if you wouldn't mind I would be interested to know what the disadvantage of a lambda function is compared to what you proposed.

To create the std::function, you heap allocate an object with copies of the values in the capture list and generate a lot more extra code in the binary:

class function {
T valuesfromthecpapturelist;
RetType operator()(ArgsType ...) {...};
}

When you think about an std::function this way (a pair of context data and a function), you realize the class you already serves that purpose.

But hey, this is becoming challenging, so I won't hold the PR anymore because of this. Moving to Init() was a big step in the right direction.

Comment on lines +1606 to +1607
TYPED_TEST(TestAzureFileSystemOnAllScenarios,
OpenOutputStreamWithMissingContainer) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is causing the linter to fail @Tom-Newton. Please fix and I will merge.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed

@felipecrv
felipecrv merged commit 8a62f30 into apache:mainFeb 21, 2024
@felipecrvfelipecrv removed the awaiting committer review Awaiting committer review label Feb 21, 2024
@github-actionsgithub-actionsBot added the awaiting committer review Awaiting committer review label Feb 21, 2024
@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@austin3dickey

Copy link
Copy Markdown
Contributor

My apologies for the noise. Looking into this now.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@austin3dickey

Copy link
Copy Markdown
Contributor

That should be the last one. Sorry again about the noise!

@felipecrv

Copy link
Copy Markdown
Contributor

@austin3dickey no worries :)

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[C++][FS][Azure] Prevent reading or writing directory marker blobs

3 participants

@Tom-Newton@austin3dickey@felipecrv
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

GH-40037: [C++][FS][Azure] Make attempted reads and writes against directories fail fast - #40119

Merged
felipecrv merged 20 commits into
apache:mainfrom
Tom-Newton:tomnewton/check_for_directory_marker_metadata/GH-40037
Feb 21, 2024
Merged

GH-40037: [C++][FS][Azure] Make attempted reads and writes against directories fail fast#40119
felipecrv merged 20 commits into
apache:mainfrom
Tom-Newton:tomnewton/check_for_directory_marker_metadata/GH-40037

Conversation

@Tom-Newton

@Tom-NewtonTom-Newton commented Feb 18, 2024

Copy link
Copy Markdown
Contributor

Rationale for this change

Prevent confusion if a user attempts to read or write a directory.

What changes are included in this PR?

  • Make ObjectAppendStream::Flush a noop if ObjectAppendStream::Init has not run successfully. This avoids an unhandled error when the destructor calls flush.
  • Check blob properties for directory marker metadata when initialising ObjectInputFile or ObjectAppendStream.
  • When initialising ObjectAppendStream call GetFileInfo if it is a flat namespace account.

Are these changes tested?

Add new tests DisallowReadingOrWritingDirectoryMarkers and DisallowCreatingFileAndDirectoryWithTheSameName to cover the new fail fast behaviour.
Also updated WriteMetadata to ensure that my changes to Flush didn't break setting metadata without calling Write on the stream.

Are there any user-facing changes?

Yes. Invalid read and write operations will now fail fast and gracefully. Previously could get into a confusing state where there were files and directories at the same path and there were some un-graceful failures.

@Tom-Newton

Copy link
Copy Markdown
ContributorAuthor

@felipecrv I expect you will be interested in reviewing this.

Comment threadcpp/src/arrow/filesystem/azurefs.cc Outdated
@github-actionsgithub-actionsBot added awaiting committer review Awaiting committer review and removed awaiting review Awaiting review labels Feb 19, 2024
DCHECK_GE(content_length_, 0);
pos_ = content_length_;
Status Init(const bool truncate,
std::function<Status()> ensure_not_flat_namespace_directory) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You can inject AzureFileSystem *azure_file_system here and not have to allocate a closure for this. You would call AzureFileSystem::Impl::EnsureNotFlatNamespaceDirectory(location) via azure_file_system->impl_ (accessible because the handles produced by the azure file system can be friends with the filesystem class).

@Tom-NewtonTom-NewtonFeb 20, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the extra info. I was planning to do this but I was struggling with the friends thing.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The trick is to forward-declare the class as well.

diff --git a/cpp/src/arrow/filesystem/azurefs.h b/cpp/src/arrow/filesystem/azurefs.h
index 2a131e40c..d48ef9dd7 100644
--- a/cpp/src/arrow/filesystem/azurefs.h
+++ b/cpp/src/arrow/filesystem/azurefs.h
@@ -44,6 +44,7 @@ classDataLakeServiceClient;
namespacearrow::fs {
+classObjectAppendStream;
classTestAzureFileSystem;
/// Options for the AzureFileSystem implementation.
@@ -180,6 +181,7 @@ classARROW_EXPORT AzureFileSystem : public FileSystem {
explicitAzureFileSystem(std::unique_ptr<Impl>&& impl);
+ friendclassObjectAppendStream;
friendclassTestAzureFileSystem;
voidForceCachedHierarchicalNamespaceSupport(int hns_support);

@Tom-NewtonTom-NewtonFeb 21, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think my main problem was that ObjectAppendStream is defined inside an anonymous namespace but I still haven't got it working as you describe.

Are you suggesting to use AzureFileSystem *azure_file_system or AzureFileSystem:Impl *azure_file_system as the argument to ObjectAppendStream::Impl. I don't know how I can get a AzureFileSystem pointer from inside AzureFileSystem::Impl and using AzureFileSystem::Impl as the argument leads to incomplete type errors which I don't think I can avoid.

Also if you wouldn't mind I would be interested to know what the disadvantage of a lambda function is compared to what you proposed.

Sorry about my lacking C++ knowledge here.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also if you wouldn't mind I would be interested to know what the disadvantage of a lambda function is compared to what you proposed.

To create the std::function, you heap allocate an object with copies of the values in the capture list and generate a lot more extra code in the binary:

class function {
T valuesfromthecpapturelist;
RetType operator()(ArgsType ...) {...};
}

When you think about an std::function this way (a pair of context data and a function), you realize the class you already serves that purpose.

But hey, this is becoming challenging, so I won't hold the PR anymore because of this. Moving to Init() was a big step in the right direction.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for explaining

Comment threadcpp/src/arrow/filesystem/azurefs.cc Outdated
DCHECK_GE(content_length_, 0);
pos_ = content_length_;
Status Init(const bool truncate,
std::function<Status()> ensure_not_flat_namespace_directory) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also if you wouldn't mind I would be interested to know what the disadvantage of a lambda function is compared to what you proposed.

To create the std::function, you heap allocate an object with copies of the values in the capture list and generate a lot more extra code in the binary:

class function {
T valuesfromthecpapturelist;
RetType operator()(ArgsType ...) {...};
}

When you think about an std::function this way (a pair of context data and a function), you realize the class you already serves that purpose.

But hey, this is becoming challenging, so I won't hold the PR anymore because of this. Moving to Init() was a big step in the right direction.

Comment on lines +1606 to +1607
TYPED_TEST(TestAzureFileSystemOnAllScenarios,
OpenOutputStreamWithMissingContainer) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is causing the linter to fail @Tom-Newton. Please fix and I will merge.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed

@felipecrv
felipecrv merged commit 8a62f30 into apache:mainFeb 21, 2024
@felipecrvfelipecrv removed the awaiting committer review Awaiting committer review label Feb 21, 2024
@github-actionsgithub-actionsBot added the awaiting committer review Awaiting committer review label Feb 21, 2024
@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@austin3dickey

Copy link
Copy Markdown
Contributor

My apologies for the noise. Looking into this now.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@austin3dickey

Copy link
Copy Markdown
Contributor

That should be the last one. Sorry again about the noise!

@felipecrv

Copy link
Copy Markdown
Contributor

@austin3dickey no worries :)

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[C++][FS][Azure] Prevent reading or writing directory marker blobs

3 participants

@Tom-Newton@austin3dickey@felipecrv
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

GH-40037: [C++][FS][Azure] Make attempted reads and writes against directories fail fast - #40119

Merged
felipecrv merged 20 commits into
apache:mainfrom
Tom-Newton:tomnewton/check_for_directory_marker_metadata/GH-40037
Feb 21, 2024
Merged

GH-40037: [C++][FS][Azure] Make attempted reads and writes against directories fail fast#40119
felipecrv merged 20 commits into
apache:mainfrom
Tom-Newton:tomnewton/check_for_directory_marker_metadata/GH-40037

Conversation

@Tom-Newton

@Tom-NewtonTom-Newton commented Feb 18, 2024

Copy link
Copy Markdown
Contributor

Rationale for this change

Prevent confusion if a user attempts to read or write a directory.

What changes are included in this PR?

  • Make ObjectAppendStream::Flush a noop if ObjectAppendStream::Init has not run successfully. This avoids an unhandled error when the destructor calls flush.
  • Check blob properties for directory marker metadata when initialising ObjectInputFile or ObjectAppendStream.
  • When initialising ObjectAppendStream call GetFileInfo if it is a flat namespace account.

Are these changes tested?

Add new tests DisallowReadingOrWritingDirectoryMarkers and DisallowCreatingFileAndDirectoryWithTheSameName to cover the new fail fast behaviour.
Also updated WriteMetadata to ensure that my changes to Flush didn't break setting metadata without calling Write on the stream.

Are there any user-facing changes?

Yes. Invalid read and write operations will now fail fast and gracefully. Previously could get into a confusing state where there were files and directories at the same path and there were some un-graceful failures.

@Tom-Newton

Copy link
Copy Markdown
ContributorAuthor

@felipecrv I expect you will be interested in reviewing this.

Comment threadcpp/src/arrow/filesystem/azurefs.cc Outdated
@github-actionsgithub-actionsBot added awaiting committer review Awaiting committer review and removed awaiting review Awaiting review labels Feb 19, 2024
DCHECK_GE(content_length_, 0);
pos_ = content_length_;
Status Init(const bool truncate,
std::function<Status()> ensure_not_flat_namespace_directory) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You can inject AzureFileSystem *azure_file_system here and not have to allocate a closure for this. You would call AzureFileSystem::Impl::EnsureNotFlatNamespaceDirectory(location) via azure_file_system->impl_ (accessible because the handles produced by the azure file system can be friends with the filesystem class).

@Tom-NewtonTom-NewtonFeb 20, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the extra info. I was planning to do this but I was struggling with the friends thing.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The trick is to forward-declare the class as well.

diff --git a/cpp/src/arrow/filesystem/azurefs.h b/cpp/src/arrow/filesystem/azurefs.h
index 2a131e40c..d48ef9dd7 100644
--- a/cpp/src/arrow/filesystem/azurefs.h
+++ b/cpp/src/arrow/filesystem/azurefs.h
@@ -44,6 +44,7 @@ classDataLakeServiceClient;
namespacearrow::fs {
+classObjectAppendStream;
classTestAzureFileSystem;
/// Options for the AzureFileSystem implementation.
@@ -180,6 +181,7 @@ classARROW_EXPORT AzureFileSystem : public FileSystem {
explicitAzureFileSystem(std::unique_ptr<Impl>&& impl);
+ friendclassObjectAppendStream;
friendclassTestAzureFileSystem;
voidForceCachedHierarchicalNamespaceSupport(int hns_support);

@Tom-NewtonTom-NewtonFeb 21, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think my main problem was that ObjectAppendStream is defined inside an anonymous namespace but I still haven't got it working as you describe.

Are you suggesting to use AzureFileSystem *azure_file_system or AzureFileSystem:Impl *azure_file_system as the argument to ObjectAppendStream::Impl. I don't know how I can get a AzureFileSystem pointer from inside AzureFileSystem::Impl and using AzureFileSystem::Impl as the argument leads to incomplete type errors which I don't think I can avoid.

Also if you wouldn't mind I would be interested to know what the disadvantage of a lambda function is compared to what you proposed.

Sorry about my lacking C++ knowledge here.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also if you wouldn't mind I would be interested to know what the disadvantage of a lambda function is compared to what you proposed.

To create the std::function, you heap allocate an object with copies of the values in the capture list and generate a lot more extra code in the binary:

class function {
T valuesfromthecpapturelist;
RetType operator()(ArgsType ...) {...};
}

When you think about an std::function this way (a pair of context data and a function), you realize the class you already serves that purpose.

But hey, this is becoming challenging, so I won't hold the PR anymore because of this. Moving to Init() was a big step in the right direction.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for explaining

Comment threadcpp/src/arrow/filesystem/azurefs.cc Outdated
DCHECK_GE(content_length_, 0);
pos_ = content_length_;
Status Init(const bool truncate,
std::function<Status()> ensure_not_flat_namespace_directory) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also if you wouldn't mind I would be interested to know what the disadvantage of a lambda function is compared to what you proposed.

To create the std::function, you heap allocate an object with copies of the values in the capture list and generate a lot more extra code in the binary:

class function {
T valuesfromthecpapturelist;
RetType operator()(ArgsType ...) {...};
}

When you think about an std::function this way (a pair of context data and a function), you realize the class you already serves that purpose.

But hey, this is becoming challenging, so I won't hold the PR anymore because of this. Moving to Init() was a big step in the right direction.

Comment on lines +1606 to +1607
TYPED_TEST(TestAzureFileSystemOnAllScenarios,
OpenOutputStreamWithMissingContainer) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is causing the linter to fail @Tom-Newton. Please fix and I will merge.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed

@felipecrv
felipecrv merged commit 8a62f30 into apache:mainFeb 21, 2024
@felipecrvfelipecrv removed the awaiting committer review Awaiting committer review label Feb 21, 2024
@github-actionsgithub-actionsBot added the awaiting committer review Awaiting committer review label Feb 21, 2024
@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@austin3dickey

Copy link
Copy Markdown
Contributor

My apologies for the noise. Looking into this now.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@austin3dickey

Copy link
Copy Markdown
Contributor

That should be the last one. Sorry again about the noise!

@felipecrv

Copy link
Copy Markdown
Contributor

@austin3dickey no worries :)

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[C++][FS][Azure] Prevent reading or writing directory marker blobs

3 participants

@Tom-Newton@austin3dickey@felipecrv
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

GH-40037: [C++][FS][Azure] Make attempted reads and writes against directories fail fast - #40119

Merged
felipecrv merged 20 commits into
apache:mainfrom
Tom-Newton:tomnewton/check_for_directory_marker_metadata/GH-40037
Feb 21, 2024
Merged

GH-40037: [C++][FS][Azure] Make attempted reads and writes against directories fail fast#40119
felipecrv merged 20 commits into
apache:mainfrom
Tom-Newton:tomnewton/check_for_directory_marker_metadata/GH-40037

Conversation

@Tom-Newton

@Tom-NewtonTom-Newton commented Feb 18, 2024

Copy link
Copy Markdown
Contributor

Rationale for this change

Prevent confusion if a user attempts to read or write a directory.

What changes are included in this PR?

  • Make ObjectAppendStream::Flush a noop if ObjectAppendStream::Init has not run successfully. This avoids an unhandled error when the destructor calls flush.
  • Check blob properties for directory marker metadata when initialising ObjectInputFile or ObjectAppendStream.
  • When initialising ObjectAppendStream call GetFileInfo if it is a flat namespace account.

Are these changes tested?

Add new tests DisallowReadingOrWritingDirectoryMarkers and DisallowCreatingFileAndDirectoryWithTheSameName to cover the new fail fast behaviour.
Also updated WriteMetadata to ensure that my changes to Flush didn't break setting metadata without calling Write on the stream.

Are there any user-facing changes?

Yes. Invalid read and write operations will now fail fast and gracefully. Previously could get into a confusing state where there were files and directories at the same path and there were some un-graceful failures.

@Tom-Newton

Copy link
Copy Markdown
ContributorAuthor

@felipecrv I expect you will be interested in reviewing this.

Comment threadcpp/src/arrow/filesystem/azurefs.cc Outdated
@github-actionsgithub-actionsBot added awaiting committer review Awaiting committer review and removed awaiting review Awaiting review labels Feb 19, 2024
DCHECK_GE(content_length_, 0);
pos_ = content_length_;
Status Init(const bool truncate,
std::function<Status()> ensure_not_flat_namespace_directory) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You can inject AzureFileSystem *azure_file_system here and not have to allocate a closure for this. You would call AzureFileSystem::Impl::EnsureNotFlatNamespaceDirectory(location) via azure_file_system->impl_ (accessible because the handles produced by the azure file system can be friends with the filesystem class).

@Tom-NewtonTom-NewtonFeb 20, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the extra info. I was planning to do this but I was struggling with the friends thing.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The trick is to forward-declare the class as well.

diff --git a/cpp/src/arrow/filesystem/azurefs.h b/cpp/src/arrow/filesystem/azurefs.h
index 2a131e40c..d48ef9dd7 100644
--- a/cpp/src/arrow/filesystem/azurefs.h
+++ b/cpp/src/arrow/filesystem/azurefs.h
@@ -44,6 +44,7 @@ classDataLakeServiceClient;
namespacearrow::fs {
+classObjectAppendStream;
classTestAzureFileSystem;
/// Options for the AzureFileSystem implementation.
@@ -180,6 +181,7 @@ classARROW_EXPORT AzureFileSystem : public FileSystem {
explicitAzureFileSystem(std::unique_ptr<Impl>&& impl);
+ friendclassObjectAppendStream;
friendclassTestAzureFileSystem;
voidForceCachedHierarchicalNamespaceSupport(int hns_support);

@Tom-NewtonTom-NewtonFeb 21, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think my main problem was that ObjectAppendStream is defined inside an anonymous namespace but I still haven't got it working as you describe.

Are you suggesting to use AzureFileSystem *azure_file_system or AzureFileSystem:Impl *azure_file_system as the argument to ObjectAppendStream::Impl. I don't know how I can get a AzureFileSystem pointer from inside AzureFileSystem::Impl and using AzureFileSystem::Impl as the argument leads to incomplete type errors which I don't think I can avoid.

Also if you wouldn't mind I would be interested to know what the disadvantage of a lambda function is compared to what you proposed.

Sorry about my lacking C++ knowledge here.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also if you wouldn't mind I would be interested to know what the disadvantage of a lambda function is compared to what you proposed.

To create the std::function, you heap allocate an object with copies of the values in the capture list and generate a lot more extra code in the binary:

class function {
T valuesfromthecpapturelist;
RetType operator()(ArgsType ...) {...};
}

When you think about an std::function this way (a pair of context data and a function), you realize the class you already serves that purpose.

But hey, this is becoming challenging, so I won't hold the PR anymore because of this. Moving to Init() was a big step in the right direction.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for explaining

Comment threadcpp/src/arrow/filesystem/azurefs.cc Outdated
DCHECK_GE(content_length_, 0);
pos_ = content_length_;
Status Init(const bool truncate,
std::function<Status()> ensure_not_flat_namespace_directory) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also if you wouldn't mind I would be interested to know what the disadvantage of a lambda function is compared to what you proposed.

To create the std::function, you heap allocate an object with copies of the values in the capture list and generate a lot more extra code in the binary:

class function {
T valuesfromthecpapturelist;
RetType operator()(ArgsType ...) {...};
}

When you think about an std::function this way (a pair of context data and a function), you realize the class you already serves that purpose.

But hey, this is becoming challenging, so I won't hold the PR anymore because of this. Moving to Init() was a big step in the right direction.

Comment on lines +1606 to +1607
TYPED_TEST(TestAzureFileSystemOnAllScenarios,
OpenOutputStreamWithMissingContainer) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is causing the linter to fail @Tom-Newton. Please fix and I will merge.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed

@felipecrv
felipecrv merged commit 8a62f30 into apache:mainFeb 21, 2024
@felipecrvfelipecrv removed the awaiting committer review Awaiting committer review label Feb 21, 2024
@github-actionsgithub-actionsBot added the awaiting committer review Awaiting committer review label Feb 21, 2024
@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@austin3dickey

Copy link
Copy Markdown
Contributor

My apologies for the noise. Looking into this now.

@conbench-apache-arrow

Copy link
Copy Markdown

After merging your PR, Conbench analyzed the 7 benchmarking runs that have been run so far on merge-commit 8a62f30.

There was 1 benchmark result with an error:

There were 2 benchmark results indicating a performance regression:

The full Conbench report has more details. It also includes information about 8 possible false positives for unstable benchmarks that are known to sometimes produce them.

@austin3dickey

Copy link
Copy Markdown
Contributor

That should be the last one. Sorry again about the noise!

@felipecrv

Copy link
Copy Markdown
Contributor

@austin3dickey no worries :)

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[C++][FS][Azure] Prevent reading or writing directory marker blobs

3 participants

@Tom-Newton@austin3dickey@felipecrv