ARROW-12288: [C++] Create Scanner interface - #9947

Closed
westonpace wants to merge 13 commits into
apache:masterfrom
westonpace:feature/arrow-12288
Closed

ARROW-12288: [C++] Create Scanner interface#9947
westonpace wants to merge 13 commits into
apache:masterfrom
westonpace:feature/arrow-12288

Conversation

@westonpace

@westonpacewestonpace commented Apr 8, 2021

Copy link
Copy Markdown
Member

To prepare for the AsyncScanner this PR creates a Scanner interface and, along the way, simplifies the current Scanner API so that the new scanner won't need to match.

What is removed:

  • Scanner::GetFragments was only used in FileSystemDataset::Write. The correct source of truth for fragments is the Dataset. Note: The python implementation exposed this method but it was not documented or used in any unit test. I think it can be safely removed and we need not worry about deprecation.
  • Scanner::schema is redundant and ambiguous. There are two schemas at the scan level. The dataset schema (the unified master schema that we expect all fragment schemas to be a subset of) and the projection schema (a combination of the dataset schema and the projection expression). Both of these are available on the scan options object and there is an accessor for these options so the caller might as well get them from there. This schema function was exposed via R and used internally there but I think any uses can be easily changed to using the options.
  • FileFormat::splittable and Fragment::splittable. These were intended to advertise that batch readahead was available on the given fragment/format. However, there is no need to advertise this. They are not used by the SyncScanner and the AsyncScanner will just assume that the format/fragment's will utilize readahead if they can (respecting the readahead options in ScanOptions)
  • Direct instantiation of Scanner. All Scanner creation should go through ScannerBuilder now. This allows the ScannerBuilder to determine what implementation to use. This was mostly the way things were implemented already. Only a few tests instantiated a Scanner directly.

What is deprecated

  • Scanner::Scan is going to be deprecated (ARROW-11797). It will not be implemented by AsyncScanner. I do not actually deprecate it in this PR as I reserve that for ARROW-11797. Unfortunately, this method was exposed via python & R and likely was used so deprecation is recommended over outright removal.

What is new

  • Scanner::ScanBatches and Scanner::ScanBatchesUnordered have been added. These functions will be the new preferred "scan" method going forward. This allows the parallelization (batch readahead, file readahead, etc.) to be handled by C++ and simplifies the user's life.
  • ScanOptions::batch_readahead and ScanOptions::fragment_readahead options allow more fine grained control over how to perform readahead. One technicality is that these options will not be respected well by the SyncScanner (although I think the current ARROW-11797 utilizes batch readahead) so they are more placeholders for when we implement AsyncScanner.
  • ScanOptions::cpu_executor and ScanOptions::io_context are added and should be fairly self explanatory.
  • ScanOptions::use_async will toggle which scanner to use.

@github-actions

Copy link
Copy Markdown

@westonpace

Copy link
Copy Markdown
MemberAuthor

@ursabot please benchmark

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@fsaintjacquesfsaintjacques left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Very nice cleanup.

Comment threadcpp/src/arrow/dataset/scanner.h Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Indeed.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'll backport this into #9589 since they need to be synced up at some point.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry about that. If you want to land #9589 first I can rebase this on top of that. I'm pretty sure the two are fairly compatible although there is some work to adjust to the new structs.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've already backported it so whichever one gets reviewed first should get merged first :) (though I'm about to go back and remove operator== as per your point below.)

Comment threadcpp/src/arrow/dataset/scanner.h Outdated
Comment threadcpp/src/arrow/dataset/scanner.cc Outdated
Comment threadcpp/src/arrow/dataset/scanner.cc Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall this state pattern is common enough for iterators/async generators that I feel like we should encapsulate it instead of defining it ad-hoc every time. (Not for this PR, but as a future task.)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Already done for generators (#9945). It's a bit trickier here with the iterator here since it's enumerating both fragments and record batches at the same time. In the async version we enumerate them separately so it made more sense to have a general utility. Although this comment made me realize this should be named EnumeratingIterator and not TaggingIterator (I'm using "tagging" to refer to attaching the fragment to the record batch which is not what I'm doing here).

@lidavidm

Copy link
Copy Markdown
Member

Should we actually merge this before ARROW-11797? Since that PR adds some more methods to the interface. Plus then we can kick things off in parallel.

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = c0ce2b1 and contender = f88a9172bb7ace47efc83a2eb53c895d7fa5d958. Results will be available as each benchmark for each run completes:
[Finished] ursa-i9-9960x: https://conbench.ursa.dev/compare/runs/ae62d394-72d3-4748-97e1-803b871167ec...40e7e6bc-9177-422c-baa4-9e45a96509ea/
[Finished] ursa-thinkcentre-m75q: https://conbench.ursa.dev/compare/runs/dd805af5-4077-47ef-acc1-4b49532ae41e...dc7601de-9ac2-44a3-8d68-766fa45f9d98/
[Finished] ec2-t3-large-us-east-2: https://conbench.ursa.dev/compare/runs/0b51231e-fc1d-4888-834f-c3bccad1e026...83d518fe-63c5-437a-bf47-4141b9e6ae8e/
[Finished] ec2-t3-xlarge-us-east-2: https://conbench.ursa.dev/compare/runs/c4009461-4c7d-4f82-b535-9b5dd71e101c...c9451681-a295-4306-bd59-83eb135c8567/

@lidavidm

Copy link
Copy Markdown
Member

This needs rebasing & I'll take a final look but otherwise I think we can get this at least into 4.0? And then we can hopefully also get the deprecation of Scan() in, which should set us up for 5.0.0.

@jorisvandenbossche

Copy link
Copy Markdown
Member

[about Scanner::GetFragments being removed] Note: The python implementation exposed this method but it was not documented or used in any unit test. I think it can be safely removed and we need not worry about deprecation.

Agreed, I don't think we need to worry about deprecating this one.

@bkietzbkietz left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some minor comments, thanks for doing this:

Comment threadcpp/src/arrow/dataset/scanner.h Outdated
Comment threadcpp/src/arrow/dataset/scanner.h Outdated
Comment threadcpp/src/arrow/dataset/test_util.h Outdated
Comment threadcpp/src/arrow/dataset/test_util.h Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit:

Suggested change
voidAssertBatchEquals(RecordBatchReader*expected, RecordBatch*batch) {
voidAssertBatchEquals(RecordBatchReader*expected, constRecordBatch&batch) {

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done. Although why isn't RecordBatchReader passed by reference? Is it simply the fact that RecordBatch is const?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We generally favor pointers over non-const reference parameters, IIRC to make it clear what is used mutably, but I can't find the reference for this (I had thought it was in the Google C++ Style Guide).

Comment threadcpp/src/arrow/dataset/test_util.h Outdated
Comment threadcpp/src/arrow/dataset/test_util.h Outdated
Comment threadcpp/src/arrow/dataset/file_base.cc Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This seems to remove parallel execution of batch partitioning. If that's intended then please call it out and add a comment with the follow up JIRA noted

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I had taken out the equivalent from #9589/ARROW-11797 but I can restore it.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

When I rebased I restored the old version which relies on Scan (@lidavidm maybe the easiest thing to do is just add the pragma avoiding warning for deprecation. I'll add some kind of switch for the async version to use ScanBatches).

However, the old version created one scan task per fragment which relied on Scanner::GetFragments (which is actually the only reason I modified this function in this PR at all).

So when restoring the older version it just calls Scan and collects the iterator into a vector. My guess is the older version was a hangover from the older version of dataset write which didn't use the scanner? @bkietz do you mind thinking this through?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also, this wasn't purely to put the parallel partitioning back in. The ScanBatches implementation introduced the ever-annoying nested parallelism back in. I had avoided it in #9892 by converting synchronous scan task execution into a future (using TaskGroup::FinishAsync) and then combining that future with any asynchronous scan task executions. I bring this up because @lidavidm is probably going to run into the same problem with ToTable when rebasing ARROW-12208.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In ARROW-11797 I renamed Scan to ScanInternal and made it private for this purpose so I'll fix that once merged. (I think this can/should merge first since it's gone through review.)

Comment threadcpp/src/arrow/dataset/scanner.h Outdated
@ursabot

ursabot commented Apr 12, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 12, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 12, 2021

Copy link
Copy Markdown

@lidavidmlidavidm left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The two test failures look unrelated (MinIO failing to initialize on Windows and Maven encountering a network hiccup).

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@westonpace@ursabot@lidavidm@jorisvandenbossche@fsaintjacques@bkietz
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all \u003cpre\u003e\u003ccode\u003e blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks"); } } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); } })(); (function(){ try { var __m = "github.com"; var __re = new RegExp('^' + "github\\.com" + '
Skip to content

ARROW-12288: [C++] Create Scanner interface - #9947

Closed
westonpace wants to merge 13 commits into
apache:masterfrom
westonpace:feature/arrow-12288
Closed

ARROW-12288: [C++] Create Scanner interface#9947
westonpace wants to merge 13 commits into
apache:masterfrom
westonpace:feature/arrow-12288

Conversation

@westonpace

@westonpacewestonpace commented Apr 8, 2021

Copy link
Copy Markdown
Member

To prepare for the AsyncScanner this PR creates a Scanner interface and, along the way, simplifies the current Scanner API so that the new scanner won't need to match.

What is removed:

  • Scanner::GetFragments was only used in FileSystemDataset::Write. The correct source of truth for fragments is the Dataset. Note: The python implementation exposed this method but it was not documented or used in any unit test. I think it can be safely removed and we need not worry about deprecation.
  • Scanner::schema is redundant and ambiguous. There are two schemas at the scan level. The dataset schema (the unified master schema that we expect all fragment schemas to be a subset of) and the projection schema (a combination of the dataset schema and the projection expression). Both of these are available on the scan options object and there is an accessor for these options so the caller might as well get them from there. This schema function was exposed via R and used internally there but I think any uses can be easily changed to using the options.
  • FileFormat::splittable and Fragment::splittable. These were intended to advertise that batch readahead was available on the given fragment/format. However, there is no need to advertise this. They are not used by the SyncScanner and the AsyncScanner will just assume that the format/fragment's will utilize readahead if they can (respecting the readahead options in ScanOptions)
  • Direct instantiation of Scanner. All Scanner creation should go through ScannerBuilder now. This allows the ScannerBuilder to determine what implementation to use. This was mostly the way things were implemented already. Only a few tests instantiated a Scanner directly.

What is deprecated

  • Scanner::Scan is going to be deprecated (ARROW-11797). It will not be implemented by AsyncScanner. I do not actually deprecate it in this PR as I reserve that for ARROW-11797. Unfortunately, this method was exposed via python & R and likely was used so deprecation is recommended over outright removal.

What is new

  • Scanner::ScanBatches and Scanner::ScanBatchesUnordered have been added. These functions will be the new preferred "scan" method going forward. This allows the parallelization (batch readahead, file readahead, etc.) to be handled by C++ and simplifies the user's life.
  • ScanOptions::batch_readahead and ScanOptions::fragment_readahead options allow more fine grained control over how to perform readahead. One technicality is that these options will not be respected well by the SyncScanner (although I think the current ARROW-11797 utilizes batch readahead) so they are more placeholders for when we implement AsyncScanner.
  • ScanOptions::cpu_executor and ScanOptions::io_context are added and should be fairly self explanatory.
  • ScanOptions::use_async will toggle which scanner to use.

@github-actions

Copy link
Copy Markdown

@westonpace

Copy link
Copy Markdown
MemberAuthor

@ursabot please benchmark

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@fsaintjacquesfsaintjacques left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Very nice cleanup.

Comment threadcpp/src/arrow/dataset/scanner.h Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Indeed.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'll backport this into #9589 since they need to be synced up at some point.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry about that. If you want to land #9589 first I can rebase this on top of that. I'm pretty sure the two are fairly compatible although there is some work to adjust to the new structs.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've already backported it so whichever one gets reviewed first should get merged first :) (though I'm about to go back and remove operator== as per your point below.)

Comment threadcpp/src/arrow/dataset/scanner.h Outdated
Comment threadcpp/src/arrow/dataset/scanner.cc Outdated
Comment threadcpp/src/arrow/dataset/scanner.cc Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall this state pattern is common enough for iterators/async generators that I feel like we should encapsulate it instead of defining it ad-hoc every time. (Not for this PR, but as a future task.)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Already done for generators (#9945). It's a bit trickier here with the iterator here since it's enumerating both fragments and record batches at the same time. In the async version we enumerate them separately so it made more sense to have a general utility. Although this comment made me realize this should be named EnumeratingIterator and not TaggingIterator (I'm using "tagging" to refer to attaching the fragment to the record batch which is not what I'm doing here).

@lidavidm

Copy link
Copy Markdown
Member

Should we actually merge this before ARROW-11797? Since that PR adds some more methods to the interface. Plus then we can kick things off in parallel.

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = c0ce2b1 and contender = f88a9172bb7ace47efc83a2eb53c895d7fa5d958. Results will be available as each benchmark for each run completes:
[Finished] ursa-i9-9960x: https://conbench.ursa.dev/compare/runs/ae62d394-72d3-4748-97e1-803b871167ec...40e7e6bc-9177-422c-baa4-9e45a96509ea/
[Finished] ursa-thinkcentre-m75q: https://conbench.ursa.dev/compare/runs/dd805af5-4077-47ef-acc1-4b49532ae41e...dc7601de-9ac2-44a3-8d68-766fa45f9d98/
[Finished] ec2-t3-large-us-east-2: https://conbench.ursa.dev/compare/runs/0b51231e-fc1d-4888-834f-c3bccad1e026...83d518fe-63c5-437a-bf47-4141b9e6ae8e/
[Finished] ec2-t3-xlarge-us-east-2: https://conbench.ursa.dev/compare/runs/c4009461-4c7d-4f82-b535-9b5dd71e101c...c9451681-a295-4306-bd59-83eb135c8567/

@lidavidm

Copy link
Copy Markdown
Member

This needs rebasing & I'll take a final look but otherwise I think we can get this at least into 4.0? And then we can hopefully also get the deprecation of Scan() in, which should set us up for 5.0.0.

@jorisvandenbossche

Copy link
Copy Markdown
Member

[about Scanner::GetFragments being removed] Note: The python implementation exposed this method but it was not documented or used in any unit test. I think it can be safely removed and we need not worry about deprecation.

Agreed, I don't think we need to worry about deprecating this one.

@bkietzbkietz left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some minor comments, thanks for doing this:

Comment threadcpp/src/arrow/dataset/scanner.h Outdated
Comment threadcpp/src/arrow/dataset/scanner.h Outdated
Comment threadcpp/src/arrow/dataset/test_util.h Outdated
Comment threadcpp/src/arrow/dataset/test_util.h Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit:

Suggested change
voidAssertBatchEquals(RecordBatchReader*expected, RecordBatch*batch) {
voidAssertBatchEquals(RecordBatchReader*expected, constRecordBatch&batch) {

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done. Although why isn't RecordBatchReader passed by reference? Is it simply the fact that RecordBatch is const?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We generally favor pointers over non-const reference parameters, IIRC to make it clear what is used mutably, but I can't find the reference for this (I had thought it was in the Google C++ Style Guide).

Comment threadcpp/src/arrow/dataset/test_util.h Outdated
Comment threadcpp/src/arrow/dataset/test_util.h Outdated
Comment threadcpp/src/arrow/dataset/file_base.cc Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This seems to remove parallel execution of batch partitioning. If that's intended then please call it out and add a comment with the follow up JIRA noted

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I had taken out the equivalent from #9589/ARROW-11797 but I can restore it.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

When I rebased I restored the old version which relies on Scan (@lidavidm maybe the easiest thing to do is just add the pragma avoiding warning for deprecation. I'll add some kind of switch for the async version to use ScanBatches).

However, the old version created one scan task per fragment which relied on Scanner::GetFragments (which is actually the only reason I modified this function in this PR at all).

So when restoring the older version it just calls Scan and collects the iterator into a vector. My guess is the older version was a hangover from the older version of dataset write which didn't use the scanner? @bkietz do you mind thinking this through?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also, this wasn't purely to put the parallel partitioning back in. The ScanBatches implementation introduced the ever-annoying nested parallelism back in. I had avoided it in #9892 by converting synchronous scan task execution into a future (using TaskGroup::FinishAsync) and then combining that future with any asynchronous scan task executions. I bring this up because @lidavidm is probably going to run into the same problem with ToTable when rebasing ARROW-12208.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In ARROW-11797 I renamed Scan to ScanInternal and made it private for this purpose so I'll fix that once merged. (I think this can/should merge first since it's gone through review.)

Comment threadcpp/src/arrow/dataset/scanner.h Outdated
@ursabot

ursabot commented Apr 12, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 12, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 12, 2021

Copy link
Copy Markdown

@lidavidmlidavidm left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The two test failures look unrelated (MinIO failing to initialize on Windows and Maven encountering a network hiccup).

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@westonpace@ursabot@lidavidm@jorisvandenbossche@fsaintjacques@bkietz
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ARROW-12288: [C++] Create Scanner interface - #9947

Closed
westonpace wants to merge 13 commits into
apache:masterfrom
westonpace:feature/arrow-12288
Closed

ARROW-12288: [C++] Create Scanner interface#9947
westonpace wants to merge 13 commits into
apache:masterfrom
westonpace:feature/arrow-12288

Conversation

@westonpace

@westonpacewestonpace commented Apr 8, 2021

Copy link
Copy Markdown
Member

To prepare for the AsyncScanner this PR creates a Scanner interface and, along the way, simplifies the current Scanner API so that the new scanner won't need to match.

What is removed:

  • Scanner::GetFragments was only used in FileSystemDataset::Write. The correct source of truth for fragments is the Dataset. Note: The python implementation exposed this method but it was not documented or used in any unit test. I think it can be safely removed and we need not worry about deprecation.
  • Scanner::schema is redundant and ambiguous. There are two schemas at the scan level. The dataset schema (the unified master schema that we expect all fragment schemas to be a subset of) and the projection schema (a combination of the dataset schema and the projection expression). Both of these are available on the scan options object and there is an accessor for these options so the caller might as well get them from there. This schema function was exposed via R and used internally there but I think any uses can be easily changed to using the options.
  • FileFormat::splittable and Fragment::splittable. These were intended to advertise that batch readahead was available on the given fragment/format. However, there is no need to advertise this. They are not used by the SyncScanner and the AsyncScanner will just assume that the format/fragment's will utilize readahead if they can (respecting the readahead options in ScanOptions)
  • Direct instantiation of Scanner. All Scanner creation should go through ScannerBuilder now. This allows the ScannerBuilder to determine what implementation to use. This was mostly the way things were implemented already. Only a few tests instantiated a Scanner directly.

What is deprecated

  • Scanner::Scan is going to be deprecated (ARROW-11797). It will not be implemented by AsyncScanner. I do not actually deprecate it in this PR as I reserve that for ARROW-11797. Unfortunately, this method was exposed via python & R and likely was used so deprecation is recommended over outright removal.

What is new

  • Scanner::ScanBatches and Scanner::ScanBatchesUnordered have been added. These functions will be the new preferred "scan" method going forward. This allows the parallelization (batch readahead, file readahead, etc.) to be handled by C++ and simplifies the user's life.
  • ScanOptions::batch_readahead and ScanOptions::fragment_readahead options allow more fine grained control over how to perform readahead. One technicality is that these options will not be respected well by the SyncScanner (although I think the current ARROW-11797 utilizes batch readahead) so they are more placeholders for when we implement AsyncScanner.
  • ScanOptions::cpu_executor and ScanOptions::io_context are added and should be fairly self explanatory.
  • ScanOptions::use_async will toggle which scanner to use.

@github-actions

Copy link
Copy Markdown

@westonpace

Copy link
Copy Markdown
MemberAuthor

@ursabot please benchmark

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@fsaintjacquesfsaintjacques left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Very nice cleanup.

Comment threadcpp/src/arrow/dataset/scanner.h Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Indeed.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'll backport this into #9589 since they need to be synced up at some point.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry about that. If you want to land #9589 first I can rebase this on top of that. I'm pretty sure the two are fairly compatible although there is some work to adjust to the new structs.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've already backported it so whichever one gets reviewed first should get merged first :) (though I'm about to go back and remove operator== as per your point below.)

Comment threadcpp/src/arrow/dataset/scanner.h Outdated
Comment threadcpp/src/arrow/dataset/scanner.cc Outdated
Comment threadcpp/src/arrow/dataset/scanner.cc Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall this state pattern is common enough for iterators/async generators that I feel like we should encapsulate it instead of defining it ad-hoc every time. (Not for this PR, but as a future task.)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Already done for generators (#9945). It's a bit trickier here with the iterator here since it's enumerating both fragments and record batches at the same time. In the async version we enumerate them separately so it made more sense to have a general utility. Although this comment made me realize this should be named EnumeratingIterator and not TaggingIterator (I'm using "tagging" to refer to attaching the fragment to the record batch which is not what I'm doing here).

@lidavidm

Copy link
Copy Markdown
Member

Should we actually merge this before ARROW-11797? Since that PR adds some more methods to the interface. Plus then we can kick things off in parallel.

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = c0ce2b1 and contender = f88a9172bb7ace47efc83a2eb53c895d7fa5d958. Results will be available as each benchmark for each run completes:
[Finished] ursa-i9-9960x: https://conbench.ursa.dev/compare/runs/ae62d394-72d3-4748-97e1-803b871167ec...40e7e6bc-9177-422c-baa4-9e45a96509ea/
[Finished] ursa-thinkcentre-m75q: https://conbench.ursa.dev/compare/runs/dd805af5-4077-47ef-acc1-4b49532ae41e...dc7601de-9ac2-44a3-8d68-766fa45f9d98/
[Finished] ec2-t3-large-us-east-2: https://conbench.ursa.dev/compare/runs/0b51231e-fc1d-4888-834f-c3bccad1e026...83d518fe-63c5-437a-bf47-4141b9e6ae8e/
[Finished] ec2-t3-xlarge-us-east-2: https://conbench.ursa.dev/compare/runs/c4009461-4c7d-4f82-b535-9b5dd71e101c...c9451681-a295-4306-bd59-83eb135c8567/

@lidavidm

Copy link
Copy Markdown
Member

This needs rebasing & I'll take a final look but otherwise I think we can get this at least into 4.0? And then we can hopefully also get the deprecation of Scan() in, which should set us up for 5.0.0.

@jorisvandenbossche

Copy link
Copy Markdown
Member

[about Scanner::GetFragments being removed] Note: The python implementation exposed this method but it was not documented or used in any unit test. I think it can be safely removed and we need not worry about deprecation.

Agreed, I don't think we need to worry about deprecating this one.

@bkietzbkietz left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some minor comments, thanks for doing this:

Comment threadcpp/src/arrow/dataset/scanner.h Outdated
Comment threadcpp/src/arrow/dataset/scanner.h Outdated
Comment threadcpp/src/arrow/dataset/test_util.h Outdated
Comment threadcpp/src/arrow/dataset/test_util.h Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit:

Suggested change
voidAssertBatchEquals(RecordBatchReader*expected, RecordBatch*batch) {
voidAssertBatchEquals(RecordBatchReader*expected, constRecordBatch&batch) {

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done. Although why isn't RecordBatchReader passed by reference? Is it simply the fact that RecordBatch is const?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We generally favor pointers over non-const reference parameters, IIRC to make it clear what is used mutably, but I can't find the reference for this (I had thought it was in the Google C++ Style Guide).

Comment threadcpp/src/arrow/dataset/test_util.h Outdated
Comment threadcpp/src/arrow/dataset/test_util.h Outdated
Comment threadcpp/src/arrow/dataset/file_base.cc Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This seems to remove parallel execution of batch partitioning. If that's intended then please call it out and add a comment with the follow up JIRA noted

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I had taken out the equivalent from #9589/ARROW-11797 but I can restore it.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

When I rebased I restored the old version which relies on Scan (@lidavidm maybe the easiest thing to do is just add the pragma avoiding warning for deprecation. I'll add some kind of switch for the async version to use ScanBatches).

However, the old version created one scan task per fragment which relied on Scanner::GetFragments (which is actually the only reason I modified this function in this PR at all).

So when restoring the older version it just calls Scan and collects the iterator into a vector. My guess is the older version was a hangover from the older version of dataset write which didn't use the scanner? @bkietz do you mind thinking this through?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also, this wasn't purely to put the parallel partitioning back in. The ScanBatches implementation introduced the ever-annoying nested parallelism back in. I had avoided it in #9892 by converting synchronous scan task execution into a future (using TaskGroup::FinishAsync) and then combining that future with any asynchronous scan task executions. I bring this up because @lidavidm is probably going to run into the same problem with ToTable when rebasing ARROW-12208.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In ARROW-11797 I renamed Scan to ScanInternal and made it private for this purpose so I'll fix that once merged. (I think this can/should merge first since it's gone through review.)

Comment threadcpp/src/arrow/dataset/scanner.h Outdated
@ursabot

ursabot commented Apr 12, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 12, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 12, 2021

Copy link
Copy Markdown

@lidavidmlidavidm left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The two test failures look unrelated (MinIO failing to initialize on Windows and Maven encountering a network hiccup).

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@westonpace@ursabot@lidavidm@jorisvandenbossche@fsaintjacques@bkietz
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length \u003e 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ARROW-12288: [C++] Create Scanner interface - #9947

Closed
westonpace wants to merge 13 commits into
apache:masterfrom
westonpace:feature/arrow-12288
Closed

ARROW-12288: [C++] Create Scanner interface#9947
westonpace wants to merge 13 commits into
apache:masterfrom
westonpace:feature/arrow-12288

Conversation

@westonpace

@westonpacewestonpace commented Apr 8, 2021

Copy link
Copy Markdown
Member

To prepare for the AsyncScanner this PR creates a Scanner interface and, along the way, simplifies the current Scanner API so that the new scanner won't need to match.

What is removed:

  • Scanner::GetFragments was only used in FileSystemDataset::Write. The correct source of truth for fragments is the Dataset. Note: The python implementation exposed this method but it was not documented or used in any unit test. I think it can be safely removed and we need not worry about deprecation.
  • Scanner::schema is redundant and ambiguous. There are two schemas at the scan level. The dataset schema (the unified master schema that we expect all fragment schemas to be a subset of) and the projection schema (a combination of the dataset schema and the projection expression). Both of these are available on the scan options object and there is an accessor for these options so the caller might as well get them from there. This schema function was exposed via R and used internally there but I think any uses can be easily changed to using the options.
  • FileFormat::splittable and Fragment::splittable. These were intended to advertise that batch readahead was available on the given fragment/format. However, there is no need to advertise this. They are not used by the SyncScanner and the AsyncScanner will just assume that the format/fragment's will utilize readahead if they can (respecting the readahead options in ScanOptions)
  • Direct instantiation of Scanner. All Scanner creation should go through ScannerBuilder now. This allows the ScannerBuilder to determine what implementation to use. This was mostly the way things were implemented already. Only a few tests instantiated a Scanner directly.

What is deprecated

  • Scanner::Scan is going to be deprecated (ARROW-11797). It will not be implemented by AsyncScanner. I do not actually deprecate it in this PR as I reserve that for ARROW-11797. Unfortunately, this method was exposed via python & R and likely was used so deprecation is recommended over outright removal.

What is new

  • Scanner::ScanBatches and Scanner::ScanBatchesUnordered have been added. These functions will be the new preferred "scan" method going forward. This allows the parallelization (batch readahead, file readahead, etc.) to be handled by C++ and simplifies the user's life.
  • ScanOptions::batch_readahead and ScanOptions::fragment_readahead options allow more fine grained control over how to perform readahead. One technicality is that these options will not be respected well by the SyncScanner (although I think the current ARROW-11797 utilizes batch readahead) so they are more placeholders for when we implement AsyncScanner.
  • ScanOptions::cpu_executor and ScanOptions::io_context are added and should be fairly self explanatory.
  • ScanOptions::use_async will toggle which scanner to use.

@github-actions

Copy link
Copy Markdown

@westonpace

Copy link
Copy Markdown
MemberAuthor

@ursabot please benchmark

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@fsaintjacquesfsaintjacques left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Very nice cleanup.

Comment threadcpp/src/arrow/dataset/scanner.h Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Indeed.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'll backport this into #9589 since they need to be synced up at some point.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry about that. If you want to land #9589 first I can rebase this on top of that. I'm pretty sure the two are fairly compatible although there is some work to adjust to the new structs.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've already backported it so whichever one gets reviewed first should get merged first :) (though I'm about to go back and remove operator== as per your point below.)

Comment threadcpp/src/arrow/dataset/scanner.h Outdated
Comment threadcpp/src/arrow/dataset/scanner.cc Outdated
Comment threadcpp/src/arrow/dataset/scanner.cc Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall this state pattern is common enough for iterators/async generators that I feel like we should encapsulate it instead of defining it ad-hoc every time. (Not for this PR, but as a future task.)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Already done for generators (#9945). It's a bit trickier here with the iterator here since it's enumerating both fragments and record batches at the same time. In the async version we enumerate them separately so it made more sense to have a general utility. Although this comment made me realize this should be named EnumeratingIterator and not TaggingIterator (I'm using "tagging" to refer to attaching the fragment to the record batch which is not what I'm doing here).

@lidavidm

Copy link
Copy Markdown
Member

Should we actually merge this before ARROW-11797? Since that PR adds some more methods to the interface. Plus then we can kick things off in parallel.

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = c0ce2b1 and contender = f88a9172bb7ace47efc83a2eb53c895d7fa5d958. Results will be available as each benchmark for each run completes:
[Finished] ursa-i9-9960x: https://conbench.ursa.dev/compare/runs/ae62d394-72d3-4748-97e1-803b871167ec...40e7e6bc-9177-422c-baa4-9e45a96509ea/
[Finished] ursa-thinkcentre-m75q: https://conbench.ursa.dev/compare/runs/dd805af5-4077-47ef-acc1-4b49532ae41e...dc7601de-9ac2-44a3-8d68-766fa45f9d98/
[Finished] ec2-t3-large-us-east-2: https://conbench.ursa.dev/compare/runs/0b51231e-fc1d-4888-834f-c3bccad1e026...83d518fe-63c5-437a-bf47-4141b9e6ae8e/
[Finished] ec2-t3-xlarge-us-east-2: https://conbench.ursa.dev/compare/runs/c4009461-4c7d-4f82-b535-9b5dd71e101c...c9451681-a295-4306-bd59-83eb135c8567/

@lidavidm

Copy link
Copy Markdown
Member

This needs rebasing & I'll take a final look but otherwise I think we can get this at least into 4.0? And then we can hopefully also get the deprecation of Scan() in, which should set us up for 5.0.0.

@jorisvandenbossche

Copy link
Copy Markdown
Member

[about Scanner::GetFragments being removed] Note: The python implementation exposed this method but it was not documented or used in any unit test. I think it can be safely removed and we need not worry about deprecation.

Agreed, I don't think we need to worry about deprecating this one.

@bkietzbkietz left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some minor comments, thanks for doing this:

Comment threadcpp/src/arrow/dataset/scanner.h Outdated
Comment threadcpp/src/arrow/dataset/scanner.h Outdated
Comment threadcpp/src/arrow/dataset/test_util.h Outdated
Comment threadcpp/src/arrow/dataset/test_util.h Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit:

Suggested change
voidAssertBatchEquals(RecordBatchReader*expected, RecordBatch*batch) {
voidAssertBatchEquals(RecordBatchReader*expected, constRecordBatch&batch) {

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done. Although why isn't RecordBatchReader passed by reference? Is it simply the fact that RecordBatch is const?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We generally favor pointers over non-const reference parameters, IIRC to make it clear what is used mutably, but I can't find the reference for this (I had thought it was in the Google C++ Style Guide).

Comment threadcpp/src/arrow/dataset/test_util.h Outdated
Comment threadcpp/src/arrow/dataset/test_util.h Outdated
Comment threadcpp/src/arrow/dataset/file_base.cc Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This seems to remove parallel execution of batch partitioning. If that's intended then please call it out and add a comment with the follow up JIRA noted

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I had taken out the equivalent from #9589/ARROW-11797 but I can restore it.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

When I rebased I restored the old version which relies on Scan (@lidavidm maybe the easiest thing to do is just add the pragma avoiding warning for deprecation. I'll add some kind of switch for the async version to use ScanBatches).

However, the old version created one scan task per fragment which relied on Scanner::GetFragments (which is actually the only reason I modified this function in this PR at all).

So when restoring the older version it just calls Scan and collects the iterator into a vector. My guess is the older version was a hangover from the older version of dataset write which didn't use the scanner? @bkietz do you mind thinking this through?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also, this wasn't purely to put the parallel partitioning back in. The ScanBatches implementation introduced the ever-annoying nested parallelism back in. I had avoided it in #9892 by converting synchronous scan task execution into a future (using TaskGroup::FinishAsync) and then combining that future with any asynchronous scan task executions. I bring this up because @lidavidm is probably going to run into the same problem with ToTable when rebasing ARROW-12208.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In ARROW-11797 I renamed Scan to ScanInternal and made it private for this purpose so I'll fix that once merged. (I think this can/should merge first since it's gone through review.)

Comment threadcpp/src/arrow/dataset/scanner.h Outdated
@ursabot

ursabot commented Apr 12, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 12, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 12, 2021

Copy link
Copy Markdown

@lidavidmlidavidm left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The two test failures look unrelated (MinIO failing to initialize on Windows and Maven encountering a network hiccup).

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@westonpace@ursabot@lidavidm@jorisvandenbossche@fsaintjacques@bkietz
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

ARROW-12288: [C++] Create Scanner interface - #9947

Closed
westonpace wants to merge 13 commits into
apache:masterfrom
westonpace:feature/arrow-12288
Closed

ARROW-12288: [C++] Create Scanner interface#9947
westonpace wants to merge 13 commits into
apache:masterfrom
westonpace:feature/arrow-12288

Conversation

@westonpace

@westonpacewestonpace commented Apr 8, 2021

Copy link
Copy Markdown
Member

To prepare for the AsyncScanner this PR creates a Scanner interface and, along the way, simplifies the current Scanner API so that the new scanner won't need to match.

What is removed:

  • Scanner::GetFragments was only used in FileSystemDataset::Write. The correct source of truth for fragments is the Dataset. Note: The python implementation exposed this method but it was not documented or used in any unit test. I think it can be safely removed and we need not worry about deprecation.
  • Scanner::schema is redundant and ambiguous. There are two schemas at the scan level. The dataset schema (the unified master schema that we expect all fragment schemas to be a subset of) and the projection schema (a combination of the dataset schema and the projection expression). Both of these are available on the scan options object and there is an accessor for these options so the caller might as well get them from there. This schema function was exposed via R and used internally there but I think any uses can be easily changed to using the options.
  • FileFormat::splittable and Fragment::splittable. These were intended to advertise that batch readahead was available on the given fragment/format. However, there is no need to advertise this. They are not used by the SyncScanner and the AsyncScanner will just assume that the format/fragment's will utilize readahead if they can (respecting the readahead options in ScanOptions)
  • Direct instantiation of Scanner. All Scanner creation should go through ScannerBuilder now. This allows the ScannerBuilder to determine what implementation to use. This was mostly the way things were implemented already. Only a few tests instantiated a Scanner directly.

What is deprecated

  • Scanner::Scan is going to be deprecated (ARROW-11797). It will not be implemented by AsyncScanner. I do not actually deprecate it in this PR as I reserve that for ARROW-11797. Unfortunately, this method was exposed via python & R and likely was used so deprecation is recommended over outright removal.

What is new

  • Scanner::ScanBatches and Scanner::ScanBatchesUnordered have been added. These functions will be the new preferred "scan" method going forward. This allows the parallelization (batch readahead, file readahead, etc.) to be handled by C++ and simplifies the user's life.
  • ScanOptions::batch_readahead and ScanOptions::fragment_readahead options allow more fine grained control over how to perform readahead. One technicality is that these options will not be respected well by the SyncScanner (although I think the current ARROW-11797 utilizes batch readahead) so they are more placeholders for when we implement AsyncScanner.
  • ScanOptions::cpu_executor and ScanOptions::io_context are added and should be fairly self explanatory.
  • ScanOptions::use_async will toggle which scanner to use.

@github-actions

Copy link
Copy Markdown

@westonpace

Copy link
Copy Markdown
MemberAuthor

@ursabot please benchmark

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@fsaintjacquesfsaintjacques left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Very nice cleanup.

Comment threadcpp/src/arrow/dataset/scanner.h Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Indeed.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'll backport this into #9589 since they need to be synced up at some point.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry about that. If you want to land #9589 first I can rebase this on top of that. I'm pretty sure the two are fairly compatible although there is some work to adjust to the new structs.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've already backported it so whichever one gets reviewed first should get merged first :) (though I'm about to go back and remove operator== as per your point below.)

Comment threadcpp/src/arrow/dataset/scanner.h Outdated
Comment threadcpp/src/arrow/dataset/scanner.cc Outdated
Comment threadcpp/src/arrow/dataset/scanner.cc Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall this state pattern is common enough for iterators/async generators that I feel like we should encapsulate it instead of defining it ad-hoc every time. (Not for this PR, but as a future task.)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Already done for generators (#9945). It's a bit trickier here with the iterator here since it's enumerating both fragments and record batches at the same time. In the async version we enumerate them separately so it made more sense to have a general utility. Although this comment made me realize this should be named EnumeratingIterator and not TaggingIterator (I'm using "tagging" to refer to attaching the fragment to the record batch which is not what I'm doing here).

@lidavidm

Copy link
Copy Markdown
Member

Should we actually merge this before ARROW-11797? Since that PR adds some more methods to the interface. Plus then we can kick things off in parallel.

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = c0ce2b1 and contender = f88a9172bb7ace47efc83a2eb53c895d7fa5d958. Results will be available as each benchmark for each run completes:
[Finished] ursa-i9-9960x: https://conbench.ursa.dev/compare/runs/ae62d394-72d3-4748-97e1-803b871167ec...40e7e6bc-9177-422c-baa4-9e45a96509ea/
[Finished] ursa-thinkcentre-m75q: https://conbench.ursa.dev/compare/runs/dd805af5-4077-47ef-acc1-4b49532ae41e...dc7601de-9ac2-44a3-8d68-766fa45f9d98/
[Finished] ec2-t3-large-us-east-2: https://conbench.ursa.dev/compare/runs/0b51231e-fc1d-4888-834f-c3bccad1e026...83d518fe-63c5-437a-bf47-4141b9e6ae8e/
[Finished] ec2-t3-xlarge-us-east-2: https://conbench.ursa.dev/compare/runs/c4009461-4c7d-4f82-b535-9b5dd71e101c...c9451681-a295-4306-bd59-83eb135c8567/

@lidavidm

Copy link
Copy Markdown
Member

This needs rebasing & I'll take a final look but otherwise I think we can get this at least into 4.0? And then we can hopefully also get the deprecation of Scan() in, which should set us up for 5.0.0.

@jorisvandenbossche

Copy link
Copy Markdown
Member

[about Scanner::GetFragments being removed] Note: The python implementation exposed this method but it was not documented or used in any unit test. I think it can be safely removed and we need not worry about deprecation.

Agreed, I don't think we need to worry about deprecating this one.

@bkietzbkietz left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some minor comments, thanks for doing this:

Comment threadcpp/src/arrow/dataset/scanner.h Outdated
Comment threadcpp/src/arrow/dataset/scanner.h Outdated
Comment threadcpp/src/arrow/dataset/test_util.h Outdated
Comment threadcpp/src/arrow/dataset/test_util.h Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit:

Suggested change
voidAssertBatchEquals(RecordBatchReader*expected, RecordBatch*batch) {
voidAssertBatchEquals(RecordBatchReader*expected, constRecordBatch&batch) {

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done. Although why isn't RecordBatchReader passed by reference? Is it simply the fact that RecordBatch is const?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We generally favor pointers over non-const reference parameters, IIRC to make it clear what is used mutably, but I can't find the reference for this (I had thought it was in the Google C++ Style Guide).

Comment threadcpp/src/arrow/dataset/test_util.h Outdated
Comment threadcpp/src/arrow/dataset/test_util.h Outdated
Comment threadcpp/src/arrow/dataset/file_base.cc Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This seems to remove parallel execution of batch partitioning. If that's intended then please call it out and add a comment with the follow up JIRA noted

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I had taken out the equivalent from #9589/ARROW-11797 but I can restore it.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

When I rebased I restored the old version which relies on Scan (@lidavidm maybe the easiest thing to do is just add the pragma avoiding warning for deprecation. I'll add some kind of switch for the async version to use ScanBatches).

However, the old version created one scan task per fragment which relied on Scanner::GetFragments (which is actually the only reason I modified this function in this PR at all).

So when restoring the older version it just calls Scan and collects the iterator into a vector. My guess is the older version was a hangover from the older version of dataset write which didn't use the scanner? @bkietz do you mind thinking this through?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also, this wasn't purely to put the parallel partitioning back in. The ScanBatches implementation introduced the ever-annoying nested parallelism back in. I had avoided it in #9892 by converting synchronous scan task execution into a future (using TaskGroup::FinishAsync) and then combining that future with any asynchronous scan task executions. I bring this up because @lidavidm is probably going to run into the same problem with ToTable when rebasing ARROW-12208.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In ARROW-11797 I renamed Scan to ScanInternal and made it private for this purpose so I'll fix that once merged. (I think this can/should merge first since it's gone through review.)

Comment threadcpp/src/arrow/dataset/scanner.h Outdated
@ursabot

ursabot commented Apr 12, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 12, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 12, 2021

Copy link
Copy Markdown

@lidavidmlidavidm left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The two test failures look unrelated (MinIO failing to initialize on Windows and Maven encountering a network hiccup).

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@westonpace@ursabot@lidavidm@jorisvandenbossche@fsaintjacques@bkietz
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ARROW-12288: [C++] Create Scanner interface - #9947

Closed
westonpace wants to merge 13 commits into
apache:masterfrom
westonpace:feature/arrow-12288
Closed

ARROW-12288: [C++] Create Scanner interface#9947
westonpace wants to merge 13 commits into
apache:masterfrom
westonpace:feature/arrow-12288

Conversation

@westonpace

@westonpacewestonpace commented Apr 8, 2021

Copy link
Copy Markdown
Member

To prepare for the AsyncScanner this PR creates a Scanner interface and, along the way, simplifies the current Scanner API so that the new scanner won't need to match.

What is removed:

  • Scanner::GetFragments was only used in FileSystemDataset::Write. The correct source of truth for fragments is the Dataset. Note: The python implementation exposed this method but it was not documented or used in any unit test. I think it can be safely removed and we need not worry about deprecation.
  • Scanner::schema is redundant and ambiguous. There are two schemas at the scan level. The dataset schema (the unified master schema that we expect all fragment schemas to be a subset of) and the projection schema (a combination of the dataset schema and the projection expression). Both of these are available on the scan options object and there is an accessor for these options so the caller might as well get them from there. This schema function was exposed via R and used internally there but I think any uses can be easily changed to using the options.
  • FileFormat::splittable and Fragment::splittable. These were intended to advertise that batch readahead was available on the given fragment/format. However, there is no need to advertise this. They are not used by the SyncScanner and the AsyncScanner will just assume that the format/fragment's will utilize readahead if they can (respecting the readahead options in ScanOptions)
  • Direct instantiation of Scanner. All Scanner creation should go through ScannerBuilder now. This allows the ScannerBuilder to determine what implementation to use. This was mostly the way things were implemented already. Only a few tests instantiated a Scanner directly.

What is deprecated

  • Scanner::Scan is going to be deprecated (ARROW-11797). It will not be implemented by AsyncScanner. I do not actually deprecate it in this PR as I reserve that for ARROW-11797. Unfortunately, this method was exposed via python & R and likely was used so deprecation is recommended over outright removal.

What is new

  • Scanner::ScanBatches and Scanner::ScanBatchesUnordered have been added. These functions will be the new preferred "scan" method going forward. This allows the parallelization (batch readahead, file readahead, etc.) to be handled by C++ and simplifies the user's life.
  • ScanOptions::batch_readahead and ScanOptions::fragment_readahead options allow more fine grained control over how to perform readahead. One technicality is that these options will not be respected well by the SyncScanner (although I think the current ARROW-11797 utilizes batch readahead) so they are more placeholders for when we implement AsyncScanner.
  • ScanOptions::cpu_executor and ScanOptions::io_context are added and should be fairly self explanatory.
  • ScanOptions::use_async will toggle which scanner to use.

@github-actions

Copy link
Copy Markdown

@westonpace

Copy link
Copy Markdown
MemberAuthor

@ursabot please benchmark

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@fsaintjacquesfsaintjacques left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Very nice cleanup.

Comment threadcpp/src/arrow/dataset/scanner.h Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Indeed.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'll backport this into #9589 since they need to be synced up at some point.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry about that. If you want to land #9589 first I can rebase this on top of that. I'm pretty sure the two are fairly compatible although there is some work to adjust to the new structs.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've already backported it so whichever one gets reviewed first should get merged first :) (though I'm about to go back and remove operator== as per your point below.)

Comment threadcpp/src/arrow/dataset/scanner.h Outdated
Comment threadcpp/src/arrow/dataset/scanner.cc Outdated
Comment threadcpp/src/arrow/dataset/scanner.cc Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall this state pattern is common enough for iterators/async generators that I feel like we should encapsulate it instead of defining it ad-hoc every time. (Not for this PR, but as a future task.)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Already done for generators (#9945). It's a bit trickier here with the iterator here since it's enumerating both fragments and record batches at the same time. In the async version we enumerate them separately so it made more sense to have a general utility. Although this comment made me realize this should be named EnumeratingIterator and not TaggingIterator (I'm using "tagging" to refer to attaching the fragment to the record batch which is not what I'm doing here).

@lidavidm

Copy link
Copy Markdown
Member

Should we actually merge this before ARROW-11797? Since that PR adds some more methods to the interface. Plus then we can kick things off in parallel.

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = c0ce2b1 and contender = f88a9172bb7ace47efc83a2eb53c895d7fa5d958. Results will be available as each benchmark for each run completes:
[Finished] ursa-i9-9960x: https://conbench.ursa.dev/compare/runs/ae62d394-72d3-4748-97e1-803b871167ec...40e7e6bc-9177-422c-baa4-9e45a96509ea/
[Finished] ursa-thinkcentre-m75q: https://conbench.ursa.dev/compare/runs/dd805af5-4077-47ef-acc1-4b49532ae41e...dc7601de-9ac2-44a3-8d68-766fa45f9d98/
[Finished] ec2-t3-large-us-east-2: https://conbench.ursa.dev/compare/runs/0b51231e-fc1d-4888-834f-c3bccad1e026...83d518fe-63c5-437a-bf47-4141b9e6ae8e/
[Finished] ec2-t3-xlarge-us-east-2: https://conbench.ursa.dev/compare/runs/c4009461-4c7d-4f82-b535-9b5dd71e101c...c9451681-a295-4306-bd59-83eb135c8567/

@lidavidm

Copy link
Copy Markdown
Member

This needs rebasing & I'll take a final look but otherwise I think we can get this at least into 4.0? And then we can hopefully also get the deprecation of Scan() in, which should set us up for 5.0.0.

@jorisvandenbossche

Copy link
Copy Markdown
Member

[about Scanner::GetFragments being removed] Note: The python implementation exposed this method but it was not documented or used in any unit test. I think it can be safely removed and we need not worry about deprecation.

Agreed, I don't think we need to worry about deprecating this one.

@bkietzbkietz left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some minor comments, thanks for doing this:

Comment threadcpp/src/arrow/dataset/scanner.h Outdated
Comment threadcpp/src/arrow/dataset/scanner.h Outdated
Comment threadcpp/src/arrow/dataset/test_util.h Outdated
Comment threadcpp/src/arrow/dataset/test_util.h Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit:

Suggested change
voidAssertBatchEquals(RecordBatchReader*expected, RecordBatch*batch) {
voidAssertBatchEquals(RecordBatchReader*expected, constRecordBatch&batch) {

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done. Although why isn't RecordBatchReader passed by reference? Is it simply the fact that RecordBatch is const?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We generally favor pointers over non-const reference parameters, IIRC to make it clear what is used mutably, but I can't find the reference for this (I had thought it was in the Google C++ Style Guide).

Comment threadcpp/src/arrow/dataset/test_util.h Outdated
Comment threadcpp/src/arrow/dataset/test_util.h Outdated
Comment threadcpp/src/arrow/dataset/file_base.cc Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This seems to remove parallel execution of batch partitioning. If that's intended then please call it out and add a comment with the follow up JIRA noted

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I had taken out the equivalent from #9589/ARROW-11797 but I can restore it.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

When I rebased I restored the old version which relies on Scan (@lidavidm maybe the easiest thing to do is just add the pragma avoiding warning for deprecation. I'll add some kind of switch for the async version to use ScanBatches).

However, the old version created one scan task per fragment which relied on Scanner::GetFragments (which is actually the only reason I modified this function in this PR at all).

So when restoring the older version it just calls Scan and collects the iterator into a vector. My guess is the older version was a hangover from the older version of dataset write which didn't use the scanner? @bkietz do you mind thinking this through?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also, this wasn't purely to put the parallel partitioning back in. The ScanBatches implementation introduced the ever-annoying nested parallelism back in. I had avoided it in #9892 by converting synchronous scan task execution into a future (using TaskGroup::FinishAsync) and then combining that future with any asynchronous scan task executions. I bring this up because @lidavidm is probably going to run into the same problem with ToTable when rebasing ARROW-12208.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In ARROW-11797 I renamed Scan to ScanInternal and made it private for this purpose so I'll fix that once merged. (I think this can/should merge first since it's gone through review.)

Comment threadcpp/src/arrow/dataset/scanner.h Outdated
@ursabot

ursabot commented Apr 12, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 12, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 12, 2021

Copy link
Copy Markdown

@lidavidmlidavidm left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The two test failures look unrelated (MinIO failing to initialize on Windows and Maven encountering a network hiccup).

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@westonpace@ursabot@lidavidm@jorisvandenbossche@fsaintjacques@bkietz
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ARROW-12288: [C++] Create Scanner interface - #9947

Closed
westonpace wants to merge 13 commits into
apache:masterfrom
westonpace:feature/arrow-12288
Closed

ARROW-12288: [C++] Create Scanner interface#9947
westonpace wants to merge 13 commits into
apache:masterfrom
westonpace:feature/arrow-12288

Conversation

@westonpace

@westonpacewestonpace commented Apr 8, 2021

Copy link
Copy Markdown
Member

To prepare for the AsyncScanner this PR creates a Scanner interface and, along the way, simplifies the current Scanner API so that the new scanner won't need to match.

What is removed:

  • Scanner::GetFragments was only used in FileSystemDataset::Write. The correct source of truth for fragments is the Dataset. Note: The python implementation exposed this method but it was not documented or used in any unit test. I think it can be safely removed and we need not worry about deprecation.
  • Scanner::schema is redundant and ambiguous. There are two schemas at the scan level. The dataset schema (the unified master schema that we expect all fragment schemas to be a subset of) and the projection schema (a combination of the dataset schema and the projection expression). Both of these are available on the scan options object and there is an accessor for these options so the caller might as well get them from there. This schema function was exposed via R and used internally there but I think any uses can be easily changed to using the options.
  • FileFormat::splittable and Fragment::splittable. These were intended to advertise that batch readahead was available on the given fragment/format. However, there is no need to advertise this. They are not used by the SyncScanner and the AsyncScanner will just assume that the format/fragment's will utilize readahead if they can (respecting the readahead options in ScanOptions)
  • Direct instantiation of Scanner. All Scanner creation should go through ScannerBuilder now. This allows the ScannerBuilder to determine what implementation to use. This was mostly the way things were implemented already. Only a few tests instantiated a Scanner directly.

What is deprecated

  • Scanner::Scan is going to be deprecated (ARROW-11797). It will not be implemented by AsyncScanner. I do not actually deprecate it in this PR as I reserve that for ARROW-11797. Unfortunately, this method was exposed via python & R and likely was used so deprecation is recommended over outright removal.

What is new

  • Scanner::ScanBatches and Scanner::ScanBatchesUnordered have been added. These functions will be the new preferred "scan" method going forward. This allows the parallelization (batch readahead, file readahead, etc.) to be handled by C++ and simplifies the user's life.
  • ScanOptions::batch_readahead and ScanOptions::fragment_readahead options allow more fine grained control over how to perform readahead. One technicality is that these options will not be respected well by the SyncScanner (although I think the current ARROW-11797 utilizes batch readahead) so they are more placeholders for when we implement AsyncScanner.
  • ScanOptions::cpu_executor and ScanOptions::io_context are added and should be fairly self explanatory.
  • ScanOptions::use_async will toggle which scanner to use.

@github-actions

Copy link
Copy Markdown

@westonpace

Copy link
Copy Markdown
MemberAuthor

@ursabot please benchmark

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@fsaintjacquesfsaintjacques left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Very nice cleanup.

Comment threadcpp/src/arrow/dataset/scanner.h Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Indeed.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'll backport this into #9589 since they need to be synced up at some point.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry about that. If you want to land #9589 first I can rebase this on top of that. I'm pretty sure the two are fairly compatible although there is some work to adjust to the new structs.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've already backported it so whichever one gets reviewed first should get merged first :) (though I'm about to go back and remove operator== as per your point below.)

Comment threadcpp/src/arrow/dataset/scanner.h Outdated
Comment threadcpp/src/arrow/dataset/scanner.cc Outdated
Comment threadcpp/src/arrow/dataset/scanner.cc Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall this state pattern is common enough for iterators/async generators that I feel like we should encapsulate it instead of defining it ad-hoc every time. (Not for this PR, but as a future task.)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Already done for generators (#9945). It's a bit trickier here with the iterator here since it's enumerating both fragments and record batches at the same time. In the async version we enumerate them separately so it made more sense to have a general utility. Although this comment made me realize this should be named EnumeratingIterator and not TaggingIterator (I'm using "tagging" to refer to attaching the fragment to the record batch which is not what I'm doing here).

@lidavidm

Copy link
Copy Markdown
Member

Should we actually merge this before ARROW-11797? Since that PR adds some more methods to the interface. Plus then we can kick things off in parallel.

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = c0ce2b1 and contender = f88a9172bb7ace47efc83a2eb53c895d7fa5d958. Results will be available as each benchmark for each run completes:
[Finished] ursa-i9-9960x: https://conbench.ursa.dev/compare/runs/ae62d394-72d3-4748-97e1-803b871167ec...40e7e6bc-9177-422c-baa4-9e45a96509ea/
[Finished] ursa-thinkcentre-m75q: https://conbench.ursa.dev/compare/runs/dd805af5-4077-47ef-acc1-4b49532ae41e...dc7601de-9ac2-44a3-8d68-766fa45f9d98/
[Finished] ec2-t3-large-us-east-2: https://conbench.ursa.dev/compare/runs/0b51231e-fc1d-4888-834f-c3bccad1e026...83d518fe-63c5-437a-bf47-4141b9e6ae8e/
[Finished] ec2-t3-xlarge-us-east-2: https://conbench.ursa.dev/compare/runs/c4009461-4c7d-4f82-b535-9b5dd71e101c...c9451681-a295-4306-bd59-83eb135c8567/

@lidavidm

Copy link
Copy Markdown
Member

This needs rebasing & I'll take a final look but otherwise I think we can get this at least into 4.0? And then we can hopefully also get the deprecation of Scan() in, which should set us up for 5.0.0.

@jorisvandenbossche

Copy link
Copy Markdown
Member

[about Scanner::GetFragments being removed] Note: The python implementation exposed this method but it was not documented or used in any unit test. I think it can be safely removed and we need not worry about deprecation.

Agreed, I don't think we need to worry about deprecating this one.

@bkietzbkietz left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some minor comments, thanks for doing this:

Comment threadcpp/src/arrow/dataset/scanner.h Outdated
Comment threadcpp/src/arrow/dataset/scanner.h Outdated
Comment threadcpp/src/arrow/dataset/test_util.h Outdated
Comment threadcpp/src/arrow/dataset/test_util.h Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit:

Suggested change
voidAssertBatchEquals(RecordBatchReader*expected, RecordBatch*batch) {
voidAssertBatchEquals(RecordBatchReader*expected, constRecordBatch&batch) {

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done. Although why isn't RecordBatchReader passed by reference? Is it simply the fact that RecordBatch is const?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We generally favor pointers over non-const reference parameters, IIRC to make it clear what is used mutably, but I can't find the reference for this (I had thought it was in the Google C++ Style Guide).

Comment threadcpp/src/arrow/dataset/test_util.h Outdated
Comment threadcpp/src/arrow/dataset/test_util.h Outdated
Comment threadcpp/src/arrow/dataset/file_base.cc Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This seems to remove parallel execution of batch partitioning. If that's intended then please call it out and add a comment with the follow up JIRA noted

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I had taken out the equivalent from #9589/ARROW-11797 but I can restore it.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

When I rebased I restored the old version which relies on Scan (@lidavidm maybe the easiest thing to do is just add the pragma avoiding warning for deprecation. I'll add some kind of switch for the async version to use ScanBatches).

However, the old version created one scan task per fragment which relied on Scanner::GetFragments (which is actually the only reason I modified this function in this PR at all).

So when restoring the older version it just calls Scan and collects the iterator into a vector. My guess is the older version was a hangover from the older version of dataset write which didn't use the scanner? @bkietz do you mind thinking this through?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also, this wasn't purely to put the parallel partitioning back in. The ScanBatches implementation introduced the ever-annoying nested parallelism back in. I had avoided it in #9892 by converting synchronous scan task execution into a future (using TaskGroup::FinishAsync) and then combining that future with any asynchronous scan task executions. I bring this up because @lidavidm is probably going to run into the same problem with ToTable when rebasing ARROW-12208.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In ARROW-11797 I renamed Scan to ScanInternal and made it private for this purpose so I'll fix that once merged. (I think this can/should merge first since it's gone through review.)

Comment threadcpp/src/arrow/dataset/scanner.h Outdated
@ursabot

ursabot commented Apr 12, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 12, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 12, 2021

Copy link
Copy Markdown

@lidavidmlidavidm left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The two test failures look unrelated (MinIO failing to initialize on Windows and Maven encountering a network hiccup).

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@westonpace@ursabot@lidavidm@jorisvandenbossche@fsaintjacques@bkietz
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

ARROW-12288: [C++] Create Scanner interface - #9947

Closed
westonpace wants to merge 13 commits into
apache:masterfrom
westonpace:feature/arrow-12288
Closed

ARROW-12288: [C++] Create Scanner interface#9947
westonpace wants to merge 13 commits into
apache:masterfrom
westonpace:feature/arrow-12288

Conversation

@westonpace

@westonpacewestonpace commented Apr 8, 2021

Copy link
Copy Markdown
Member

To prepare for the AsyncScanner this PR creates a Scanner interface and, along the way, simplifies the current Scanner API so that the new scanner won't need to match.

What is removed:

  • Scanner::GetFragments was only used in FileSystemDataset::Write. The correct source of truth for fragments is the Dataset. Note: The python implementation exposed this method but it was not documented or used in any unit test. I think it can be safely removed and we need not worry about deprecation.
  • Scanner::schema is redundant and ambiguous. There are two schemas at the scan level. The dataset schema (the unified master schema that we expect all fragment schemas to be a subset of) and the projection schema (a combination of the dataset schema and the projection expression). Both of these are available on the scan options object and there is an accessor for these options so the caller might as well get them from there. This schema function was exposed via R and used internally there but I think any uses can be easily changed to using the options.
  • FileFormat::splittable and Fragment::splittable. These were intended to advertise that batch readahead was available on the given fragment/format. However, there is no need to advertise this. They are not used by the SyncScanner and the AsyncScanner will just assume that the format/fragment's will utilize readahead if they can (respecting the readahead options in ScanOptions)
  • Direct instantiation of Scanner. All Scanner creation should go through ScannerBuilder now. This allows the ScannerBuilder to determine what implementation to use. This was mostly the way things were implemented already. Only a few tests instantiated a Scanner directly.

What is deprecated

  • Scanner::Scan is going to be deprecated (ARROW-11797). It will not be implemented by AsyncScanner. I do not actually deprecate it in this PR as I reserve that for ARROW-11797. Unfortunately, this method was exposed via python & R and likely was used so deprecation is recommended over outright removal.

What is new

  • Scanner::ScanBatches and Scanner::ScanBatchesUnordered have been added. These functions will be the new preferred "scan" method going forward. This allows the parallelization (batch readahead, file readahead, etc.) to be handled by C++ and simplifies the user's life.
  • ScanOptions::batch_readahead and ScanOptions::fragment_readahead options allow more fine grained control over how to perform readahead. One technicality is that these options will not be respected well by the SyncScanner (although I think the current ARROW-11797 utilizes batch readahead) so they are more placeholders for when we implement AsyncScanner.
  • ScanOptions::cpu_executor and ScanOptions::io_context are added and should be fairly self explanatory.
  • ScanOptions::use_async will toggle which scanner to use.

@github-actions

Copy link
Copy Markdown

@westonpace

Copy link
Copy Markdown
MemberAuthor

@ursabot please benchmark

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@fsaintjacquesfsaintjacques left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Very nice cleanup.

Comment threadcpp/src/arrow/dataset/scanner.h Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Indeed.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'll backport this into #9589 since they need to be synced up at some point.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry about that. If you want to land #9589 first I can rebase this on top of that. I'm pretty sure the two are fairly compatible although there is some work to adjust to the new structs.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've already backported it so whichever one gets reviewed first should get merged first :) (though I'm about to go back and remove operator== as per your point below.)

Comment threadcpp/src/arrow/dataset/scanner.h Outdated
Comment threadcpp/src/arrow/dataset/scanner.cc Outdated
Comment threadcpp/src/arrow/dataset/scanner.cc Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall this state pattern is common enough for iterators/async generators that I feel like we should encapsulate it instead of defining it ad-hoc every time. (Not for this PR, but as a future task.)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Already done for generators (#9945). It's a bit trickier here with the iterator here since it's enumerating both fragments and record batches at the same time. In the async version we enumerate them separately so it made more sense to have a general utility. Although this comment made me realize this should be named EnumeratingIterator and not TaggingIterator (I'm using "tagging" to refer to attaching the fragment to the record batch which is not what I'm doing here).

@lidavidm

Copy link
Copy Markdown
Member

Should we actually merge this before ARROW-11797? Since that PR adds some more methods to the interface. Plus then we can kick things off in parallel.

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 9, 2021

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = c0ce2b1 and contender = f88a9172bb7ace47efc83a2eb53c895d7fa5d958. Results will be available as each benchmark for each run completes:
[Finished] ursa-i9-9960x: https://conbench.ursa.dev/compare/runs/ae62d394-72d3-4748-97e1-803b871167ec...40e7e6bc-9177-422c-baa4-9e45a96509ea/
[Finished] ursa-thinkcentre-m75q: https://conbench.ursa.dev/compare/runs/dd805af5-4077-47ef-acc1-4b49532ae41e...dc7601de-9ac2-44a3-8d68-766fa45f9d98/
[Finished] ec2-t3-large-us-east-2: https://conbench.ursa.dev/compare/runs/0b51231e-fc1d-4888-834f-c3bccad1e026...83d518fe-63c5-437a-bf47-4141b9e6ae8e/
[Finished] ec2-t3-xlarge-us-east-2: https://conbench.ursa.dev/compare/runs/c4009461-4c7d-4f82-b535-9b5dd71e101c...c9451681-a295-4306-bd59-83eb135c8567/

@lidavidm

Copy link
Copy Markdown
Member

This needs rebasing & I'll take a final look but otherwise I think we can get this at least into 4.0? And then we can hopefully also get the deprecation of Scan() in, which should set us up for 5.0.0.

@jorisvandenbossche

Copy link
Copy Markdown
Member

[about Scanner::GetFragments being removed] Note: The python implementation exposed this method but it was not documented or used in any unit test. I think it can be safely removed and we need not worry about deprecation.

Agreed, I don't think we need to worry about deprecating this one.

@bkietzbkietz left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some minor comments, thanks for doing this:

Comment threadcpp/src/arrow/dataset/scanner.h Outdated
Comment threadcpp/src/arrow/dataset/scanner.h Outdated
Comment threadcpp/src/arrow/dataset/test_util.h Outdated
Comment threadcpp/src/arrow/dataset/test_util.h Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit:

Suggested change
voidAssertBatchEquals(RecordBatchReader*expected, RecordBatch*batch) {
voidAssertBatchEquals(RecordBatchReader*expected, constRecordBatch&batch) {

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done. Although why isn't RecordBatchReader passed by reference? Is it simply the fact that RecordBatch is const?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We generally favor pointers over non-const reference parameters, IIRC to make it clear what is used mutably, but I can't find the reference for this (I had thought it was in the Google C++ Style Guide).

Comment threadcpp/src/arrow/dataset/test_util.h Outdated
Comment threadcpp/src/arrow/dataset/test_util.h Outdated
Comment threadcpp/src/arrow/dataset/file_base.cc Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This seems to remove parallel execution of batch partitioning. If that's intended then please call it out and add a comment with the follow up JIRA noted

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I had taken out the equivalent from #9589/ARROW-11797 but I can restore it.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

When I rebased I restored the old version which relies on Scan (@lidavidm maybe the easiest thing to do is just add the pragma avoiding warning for deprecation. I'll add some kind of switch for the async version to use ScanBatches).

However, the old version created one scan task per fragment which relied on Scanner::GetFragments (which is actually the only reason I modified this function in this PR at all).

So when restoring the older version it just calls Scan and collects the iterator into a vector. My guess is the older version was a hangover from the older version of dataset write which didn't use the scanner? @bkietz do you mind thinking this through?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also, this wasn't purely to put the parallel partitioning back in. The ScanBatches implementation introduced the ever-annoying nested parallelism back in. I had avoided it in #9892 by converting synchronous scan task execution into a future (using TaskGroup::FinishAsync) and then combining that future with any asynchronous scan task executions. I bring this up because @lidavidm is probably going to run into the same problem with ToTable when rebasing ARROW-12208.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In ARROW-11797 I renamed Scan to ScanInternal and made it private for this purpose so I'll fix that once merged. (I think this can/should merge first since it's gone through review.)

Comment threadcpp/src/arrow/dataset/scanner.h Outdated
@ursabot

ursabot commented Apr 12, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 12, 2021

Copy link
Copy Markdown

@ursabot

ursabot commented Apr 12, 2021

Copy link
Copy Markdown

@lidavidmlidavidm left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The two test failures look unrelated (MinIO failing to initialize on Windows and Maven encountering a network hiccup).

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@westonpace@ursabot@lidavidm@jorisvandenbossche@fsaintjacques@bkietz