ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet reader - #5641

Closed
andygrove wants to merge 11 commits into
apache:masterfrom
andygrove:ARROW-6700
Closed

ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet reader#5641
andygrove wants to merge 11 commits into
apache:masterfrom
andygrove:ARROW-6700

Conversation

@andygrove

@andygroveandygrove commented Oct 13, 2019

Copy link
Copy Markdown
Member

Replaces the DataFusion Parquet reader with the new Arrow reader in the parquet crate.

@github-actions

Copy link
Copy Markdown

@liurenjie1024liurenjie1024 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some part should be replaced by new ArrowReader.

Comment threadrust/parquet/src/arrow/mod.rs Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

array_reader is not designed for public use. This PR #5523 contains public api and doc example. Essentially, you should use ArrowReader

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please use new ArrowReader here.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Schema should also consider projection.

@andygrove

Copy link
Copy Markdown
MemberAuthor

@liurenjie1024 I updated to use ArrowReader. When I run cargo test I actually get a SIGSEGV:

error: process didn't exit successfully: `/home/andy/git/andygrove/arrow/rust/target/debug/deps/datafusion-39547aa10aa86781` (signal: 11, SIGSEGV: invalid memory reference)

@liurenjie1024

Copy link
Copy Markdown
Contributor

@andygrove I pull your request and run the tests. The root cause is that currently arrow reader doesn't support some data types(e.g., UTF8) and it caused program to crash.


fn next(&mut self) -> Result<Option<RecordBatch>> {
match self.request_tx.send(()) {
Ok(_) => match self.response_rx.recv() {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why we need another thread here? This send request, wait response model is also blocked here waiting for IO.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We need the threading because the Parquet structs/traits do not implement Sync + Send and cannot be sent between threads.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I would much prefer it if we could make Parquet safe to use in multi threaded environments.

@andygrove

Copy link
Copy Markdown
MemberAuthor

@andygrove I pull your request and run the tests. The root cause is that currently arrow reader doesn't support some data types(e.g., UTF8) and it caused program to crash.

We should make the code fail gracefully by returning an Err for unsupported types.

@liurenjie1024

Copy link
Copy Markdown
Contributor

@andygrove After debugging, I found that this is not caused by unwrap call, but by error in supporting UTF8. I'll add support for utf8 and this will be fixed.

@andygroveandygrove changed the title ARROW-6700: [Rust] [DataFusion] Use new Arrow reader [WIP]ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet reader [WIP]Nov 18, 2019
@andygrove

Copy link
Copy Markdown
MemberAuthor

Hi @liurenjie1024 is there any update on adding support for UTF8?

@liurenjie1024

Copy link
Copy Markdown
Contributor

@andygrove Sorry, almost forgot about this. I'll address it this week.

@andygrove

Copy link
Copy Markdown
MemberAuthor

Hi @liurenjie1024 the next issue is that I have a regression due to Reading Timestamp(Nanosecond) type from parquet is not supported yet. Do you think this is an easy one to add? It requires a new Int96Type and Int96Converter.

@andygrove

Copy link
Copy Markdown
MemberAuthor

@liurenjie1024 I attempted to add support to the array reader for TimestampNanoseconds. It runs but returns the wrong results currently. Could you take a look?

@nevi-me

Copy link
Copy Markdown
Contributor

@andygrove isn't it the same issue that timestamps are stored in 96bit values? How are you converting from the 96bit values?

@andygrove

Copy link
Copy Markdown
MemberAuthor

@nevi-me Yes, I guess I need to write a custom converter here rather than try and use the CastConverter. I'll have a go at that.

@liurenjie1024

Copy link
Copy Markdown
Contributor

@andygrove I'll take a look this week.

@nevi-me

Copy link
Copy Markdown
Contributor

I've created a complex converter for int96, and I've fixed the binary reads. The problem with the binary reads was that when I introduced StringType, I missed a part in the parquet code where we were supposed to convert a binary physical type with no logical type, to BinaryType.

All tests are passing locally for me.

@nevi-me

Copy link
Copy Markdown
Contributor

Oops, I forgot examples. @andygrove, the alltypes_plain.parquet file has saved the string columns as binary. Instead of a quick fix of converting the binary data to utf8, this might be an opportunity for us to create a cast kernel, or some string kernel that converts BinaryArray to StringArray.

Then in DataFusion we could do an implicit cast to string in the SQL code. What do you think about this?

@andygrove

Copy link
Copy Markdown
MemberAuthor

Wow @nevi-me that's awesome! If the columns are stored as binary then DataFusion should also treat them as binary unless the user adds an explicit CAST in the SQL. I pushed a quick change for that and all tests are passing for me now.

I know there are some things I need to clean up in this PR so I'll start working on those tomorrow.

@andygroveandygrove changed the title ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet reader [WIP]ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet readerDec 10, 2019
@andygroveandygrove removed the WIP PR is work in progress label Dec 10, 2019
if array.is_null(i) {
b.append(false)?;
} else {
b.append_value(str::from_utf8(from.value(i)).unwrap())?;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

from_utf8 can panic if the data is not valid utf8 bytes. Perhaps it's better to convert failures to null values instead? This would behave similarly to other casts that return nulls on overflowing data.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why we can't return error? I don't it's correct to return null rather than error.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We've had the broader discussion around casts, on whether to return an error on overflows/invalid data, or whether to expose an option to the user to define their desired behaviour. I haven't opened a JIRA for this, perhaps I should, so we can address the overall cast behaviour there?

For now, an expedient would actually not be to downcast to BinaryArray, but to create a StringArray from the binary data, but maybe that's a premature optimisation.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Another discussion about configuring behaviors of error handling would be fine. But I can't understand for now why we can't just return error? Inconsistency with existing behavior?

@nevi-menevi-meDec 10, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, inconsistency. For example, if you cast a i64 to u8, negative values and values that don't fit into u8 will be cast to nulls. It doesn't return an error.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For this PR, I agree that we can return null to match existing behavior. But I don't think it's good idea to make returning null as default behavior, but returning error should be. Please open an jira ticket to track this.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Comment threadrust/datafusion/src/sql/planner.rs Outdated
SQLType::Float(_) | SQLType::Real => Ok(DataType::Float64),
SQLType::Double => Ok(DataType::Float64),
SQLType::Char(_) | SQLType::Varchar(_) => Ok(DataType::Utf8),
SQLType::Custom(t) if t.to_lowercase() == "string" => Ok(DataType::Utf8),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We probably don't need this just yet, as cast(binary as varchar) works without any changes.

@paddyhoranpaddyhoran left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm a little out of the loop (day job, etc.), but LGTM at a high level. Thanks @andygrove@nevi-me and @liurenjie1024, great progress.

);
}

//TODO assertions

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are you going to add assertions in this PR?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes. I'm away on a business trip today and tomorrow but intend on adding assertions this weekend.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@andygrove@liurenjie1024@nevi-me@alippai@paddyhoran
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet reader - #5641

Closed
andygrove wants to merge 11 commits into
apache:masterfrom
andygrove:ARROW-6700
Closed

ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet reader#5641
andygrove wants to merge 11 commits into
apache:masterfrom
andygrove:ARROW-6700

Conversation

@andygrove

@andygroveandygrove commented Oct 13, 2019

Copy link
Copy Markdown
Member

Replaces the DataFusion Parquet reader with the new Arrow reader in the parquet crate.

@github-actions

Copy link
Copy Markdown

@liurenjie1024liurenjie1024 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some part should be replaced by new ArrowReader.

Comment threadrust/parquet/src/arrow/mod.rs Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

array_reader is not designed for public use. This PR #5523 contains public api and doc example. Essentially, you should use ArrowReader

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please use new ArrowReader here.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Schema should also consider projection.

@andygrove

Copy link
Copy Markdown
MemberAuthor

@liurenjie1024 I updated to use ArrowReader. When I run cargo test I actually get a SIGSEGV:

error: process didn't exit successfully: `/home/andy/git/andygrove/arrow/rust/target/debug/deps/datafusion-39547aa10aa86781` (signal: 11, SIGSEGV: invalid memory reference)

@liurenjie1024

Copy link
Copy Markdown
Contributor

@andygrove I pull your request and run the tests. The root cause is that currently arrow reader doesn't support some data types(e.g., UTF8) and it caused program to crash.


fn next(&mut self) -> Result<Option<RecordBatch>> {
match self.request_tx.send(()) {
Ok(_) => match self.response_rx.recv() {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why we need another thread here? This send request, wait response model is also blocked here waiting for IO.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We need the threading because the Parquet structs/traits do not implement Sync + Send and cannot be sent between threads.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I would much prefer it if we could make Parquet safe to use in multi threaded environments.

@andygrove

Copy link
Copy Markdown
MemberAuthor

@andygrove I pull your request and run the tests. The root cause is that currently arrow reader doesn't support some data types(e.g., UTF8) and it caused program to crash.

We should make the code fail gracefully by returning an Err for unsupported types.

@liurenjie1024

Copy link
Copy Markdown
Contributor

@andygrove After debugging, I found that this is not caused by unwrap call, but by error in supporting UTF8. I'll add support for utf8 and this will be fixed.

@andygroveandygrove changed the title ARROW-6700: [Rust] [DataFusion] Use new Arrow reader [WIP]ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet reader [WIP]Nov 18, 2019
@andygrove

Copy link
Copy Markdown
MemberAuthor

Hi @liurenjie1024 is there any update on adding support for UTF8?

@liurenjie1024

Copy link
Copy Markdown
Contributor

@andygrove Sorry, almost forgot about this. I'll address it this week.

@andygrove

Copy link
Copy Markdown
MemberAuthor

Hi @liurenjie1024 the next issue is that I have a regression due to Reading Timestamp(Nanosecond) type from parquet is not supported yet. Do you think this is an easy one to add? It requires a new Int96Type and Int96Converter.

@andygrove

Copy link
Copy Markdown
MemberAuthor

@liurenjie1024 I attempted to add support to the array reader for TimestampNanoseconds. It runs but returns the wrong results currently. Could you take a look?

@nevi-me

Copy link
Copy Markdown
Contributor

@andygrove isn't it the same issue that timestamps are stored in 96bit values? How are you converting from the 96bit values?

@andygrove

Copy link
Copy Markdown
MemberAuthor

@nevi-me Yes, I guess I need to write a custom converter here rather than try and use the CastConverter. I'll have a go at that.

@liurenjie1024

Copy link
Copy Markdown
Contributor

@andygrove I'll take a look this week.

@nevi-me

Copy link
Copy Markdown
Contributor

I've created a complex converter for int96, and I've fixed the binary reads. The problem with the binary reads was that when I introduced StringType, I missed a part in the parquet code where we were supposed to convert a binary physical type with no logical type, to BinaryType.

All tests are passing locally for me.

@nevi-me

Copy link
Copy Markdown
Contributor

Oops, I forgot examples. @andygrove, the alltypes_plain.parquet file has saved the string columns as binary. Instead of a quick fix of converting the binary data to utf8, this might be an opportunity for us to create a cast kernel, or some string kernel that converts BinaryArray to StringArray.

Then in DataFusion we could do an implicit cast to string in the SQL code. What do you think about this?

@andygrove

Copy link
Copy Markdown
MemberAuthor

Wow @nevi-me that's awesome! If the columns are stored as binary then DataFusion should also treat them as binary unless the user adds an explicit CAST in the SQL. I pushed a quick change for that and all tests are passing for me now.

I know there are some things I need to clean up in this PR so I'll start working on those tomorrow.

@andygroveandygrove changed the title ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet reader [WIP]ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet readerDec 10, 2019
@andygroveandygrove removed the WIP PR is work in progress label Dec 10, 2019
if array.is_null(i) {
b.append(false)?;
} else {
b.append_value(str::from_utf8(from.value(i)).unwrap())?;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

from_utf8 can panic if the data is not valid utf8 bytes. Perhaps it's better to convert failures to null values instead? This would behave similarly to other casts that return nulls on overflowing data.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why we can't return error? I don't it's correct to return null rather than error.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We've had the broader discussion around casts, on whether to return an error on overflows/invalid data, or whether to expose an option to the user to define their desired behaviour. I haven't opened a JIRA for this, perhaps I should, so we can address the overall cast behaviour there?

For now, an expedient would actually not be to downcast to BinaryArray, but to create a StringArray from the binary data, but maybe that's a premature optimisation.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Another discussion about configuring behaviors of error handling would be fine. But I can't understand for now why we can't just return error? Inconsistency with existing behavior?

@nevi-menevi-meDec 10, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, inconsistency. For example, if you cast a i64 to u8, negative values and values that don't fit into u8 will be cast to nulls. It doesn't return an error.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For this PR, I agree that we can return null to match existing behavior. But I don't think it's good idea to make returning null as default behavior, but returning error should be. Please open an jira ticket to track this.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Comment threadrust/datafusion/src/sql/planner.rs Outdated
SQLType::Float(_) | SQLType::Real => Ok(DataType::Float64),
SQLType::Double => Ok(DataType::Float64),
SQLType::Char(_) | SQLType::Varchar(_) => Ok(DataType::Utf8),
SQLType::Custom(t) if t.to_lowercase() == "string" => Ok(DataType::Utf8),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We probably don't need this just yet, as cast(binary as varchar) works without any changes.

@paddyhoranpaddyhoran left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm a little out of the loop (day job, etc.), but LGTM at a high level. Thanks @andygrove@nevi-me and @liurenjie1024, great progress.

);
}

//TODO assertions

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are you going to add assertions in this PR?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes. I'm away on a business trip today and tomorrow but intend on adding assertions this weekend.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@andygrove@liurenjie1024@nevi-me@alippai@paddyhoran
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet reader - #5641

Closed
andygrove wants to merge 11 commits into
apache:masterfrom
andygrove:ARROW-6700
Closed

ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet reader#5641
andygrove wants to merge 11 commits into
apache:masterfrom
andygrove:ARROW-6700

Conversation

@andygrove

@andygroveandygrove commented Oct 13, 2019

Copy link
Copy Markdown
Member

Replaces the DataFusion Parquet reader with the new Arrow reader in the parquet crate.

@github-actions

Copy link
Copy Markdown

@liurenjie1024liurenjie1024 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some part should be replaced by new ArrowReader.

Comment threadrust/parquet/src/arrow/mod.rs Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

array_reader is not designed for public use. This PR #5523 contains public api and doc example. Essentially, you should use ArrowReader

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please use new ArrowReader here.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Schema should also consider projection.

@andygrove

Copy link
Copy Markdown
MemberAuthor

@liurenjie1024 I updated to use ArrowReader. When I run cargo test I actually get a SIGSEGV:

error: process didn't exit successfully: `/home/andy/git/andygrove/arrow/rust/target/debug/deps/datafusion-39547aa10aa86781` (signal: 11, SIGSEGV: invalid memory reference)

@liurenjie1024

Copy link
Copy Markdown
Contributor

@andygrove I pull your request and run the tests. The root cause is that currently arrow reader doesn't support some data types(e.g., UTF8) and it caused program to crash.


fn next(&mut self) -> Result<Option<RecordBatch>> {
match self.request_tx.send(()) {
Ok(_) => match self.response_rx.recv() {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why we need another thread here? This send request, wait response model is also blocked here waiting for IO.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We need the threading because the Parquet structs/traits do not implement Sync + Send and cannot be sent between threads.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I would much prefer it if we could make Parquet safe to use in multi threaded environments.

@andygrove

Copy link
Copy Markdown
MemberAuthor

@andygrove I pull your request and run the tests. The root cause is that currently arrow reader doesn't support some data types(e.g., UTF8) and it caused program to crash.

We should make the code fail gracefully by returning an Err for unsupported types.

@liurenjie1024

Copy link
Copy Markdown
Contributor

@andygrove After debugging, I found that this is not caused by unwrap call, but by error in supporting UTF8. I'll add support for utf8 and this will be fixed.

@andygroveandygrove changed the title ARROW-6700: [Rust] [DataFusion] Use new Arrow reader [WIP]ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet reader [WIP]Nov 18, 2019
@andygrove

Copy link
Copy Markdown
MemberAuthor

Hi @liurenjie1024 is there any update on adding support for UTF8?

@liurenjie1024

Copy link
Copy Markdown
Contributor

@andygrove Sorry, almost forgot about this. I'll address it this week.

@andygrove

Copy link
Copy Markdown
MemberAuthor

Hi @liurenjie1024 the next issue is that I have a regression due to Reading Timestamp(Nanosecond) type from parquet is not supported yet. Do you think this is an easy one to add? It requires a new Int96Type and Int96Converter.

@andygrove

Copy link
Copy Markdown
MemberAuthor

@liurenjie1024 I attempted to add support to the array reader for TimestampNanoseconds. It runs but returns the wrong results currently. Could you take a look?

@nevi-me

Copy link
Copy Markdown
Contributor

@andygrove isn't it the same issue that timestamps are stored in 96bit values? How are you converting from the 96bit values?

@andygrove

Copy link
Copy Markdown
MemberAuthor

@nevi-me Yes, I guess I need to write a custom converter here rather than try and use the CastConverter. I'll have a go at that.

@liurenjie1024

Copy link
Copy Markdown
Contributor

@andygrove I'll take a look this week.

@nevi-me

Copy link
Copy Markdown
Contributor

I've created a complex converter for int96, and I've fixed the binary reads. The problem with the binary reads was that when I introduced StringType, I missed a part in the parquet code where we were supposed to convert a binary physical type with no logical type, to BinaryType.

All tests are passing locally for me.

@nevi-me

Copy link
Copy Markdown
Contributor

Oops, I forgot examples. @andygrove, the alltypes_plain.parquet file has saved the string columns as binary. Instead of a quick fix of converting the binary data to utf8, this might be an opportunity for us to create a cast kernel, or some string kernel that converts BinaryArray to StringArray.

Then in DataFusion we could do an implicit cast to string in the SQL code. What do you think about this?

@andygrove

Copy link
Copy Markdown
MemberAuthor

Wow @nevi-me that's awesome! If the columns are stored as binary then DataFusion should also treat them as binary unless the user adds an explicit CAST in the SQL. I pushed a quick change for that and all tests are passing for me now.

I know there are some things I need to clean up in this PR so I'll start working on those tomorrow.

@andygroveandygrove changed the title ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet reader [WIP]ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet readerDec 10, 2019
@andygroveandygrove removed the WIP PR is work in progress label Dec 10, 2019
if array.is_null(i) {
b.append(false)?;
} else {
b.append_value(str::from_utf8(from.value(i)).unwrap())?;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

from_utf8 can panic if the data is not valid utf8 bytes. Perhaps it's better to convert failures to null values instead? This would behave similarly to other casts that return nulls on overflowing data.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why we can't return error? I don't it's correct to return null rather than error.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We've had the broader discussion around casts, on whether to return an error on overflows/invalid data, or whether to expose an option to the user to define their desired behaviour. I haven't opened a JIRA for this, perhaps I should, so we can address the overall cast behaviour there?

For now, an expedient would actually not be to downcast to BinaryArray, but to create a StringArray from the binary data, but maybe that's a premature optimisation.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Another discussion about configuring behaviors of error handling would be fine. But I can't understand for now why we can't just return error? Inconsistency with existing behavior?

@nevi-menevi-meDec 10, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, inconsistency. For example, if you cast a i64 to u8, negative values and values that don't fit into u8 will be cast to nulls. It doesn't return an error.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For this PR, I agree that we can return null to match existing behavior. But I don't think it's good idea to make returning null as default behavior, but returning error should be. Please open an jira ticket to track this.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Comment threadrust/datafusion/src/sql/planner.rs Outdated
SQLType::Float(_) | SQLType::Real => Ok(DataType::Float64),
SQLType::Double => Ok(DataType::Float64),
SQLType::Char(_) | SQLType::Varchar(_) => Ok(DataType::Utf8),
SQLType::Custom(t) if t.to_lowercase() == "string" => Ok(DataType::Utf8),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We probably don't need this just yet, as cast(binary as varchar) works without any changes.

@paddyhoranpaddyhoran left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm a little out of the loop (day job, etc.), but LGTM at a high level. Thanks @andygrove@nevi-me and @liurenjie1024, great progress.

);
}

//TODO assertions

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are you going to add assertions in this PR?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes. I'm away on a business trip today and tomorrow but intend on adding assertions this weekend.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@andygrove@liurenjie1024@nevi-me@alippai@paddyhoran
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet reader - #5641

Closed
andygrove wants to merge 11 commits into
apache:masterfrom
andygrove:ARROW-6700
Closed

ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet reader#5641
andygrove wants to merge 11 commits into
apache:masterfrom
andygrove:ARROW-6700

Conversation

@andygrove

@andygroveandygrove commented Oct 13, 2019

Copy link
Copy Markdown
Member

Replaces the DataFusion Parquet reader with the new Arrow reader in the parquet crate.

@github-actions

Copy link
Copy Markdown

@liurenjie1024liurenjie1024 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some part should be replaced by new ArrowReader.

Comment threadrust/parquet/src/arrow/mod.rs Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

array_reader is not designed for public use. This PR #5523 contains public api and doc example. Essentially, you should use ArrowReader

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please use new ArrowReader here.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Schema should also consider projection.

@andygrove

Copy link
Copy Markdown
MemberAuthor

@liurenjie1024 I updated to use ArrowReader. When I run cargo test I actually get a SIGSEGV:

error: process didn't exit successfully: `/home/andy/git/andygrove/arrow/rust/target/debug/deps/datafusion-39547aa10aa86781` (signal: 11, SIGSEGV: invalid memory reference)

@liurenjie1024

Copy link
Copy Markdown
Contributor

@andygrove I pull your request and run the tests. The root cause is that currently arrow reader doesn't support some data types(e.g., UTF8) and it caused program to crash.


fn next(&mut self) -> Result<Option<RecordBatch>> {
match self.request_tx.send(()) {
Ok(_) => match self.response_rx.recv() {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why we need another thread here? This send request, wait response model is also blocked here waiting for IO.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We need the threading because the Parquet structs/traits do not implement Sync + Send and cannot be sent between threads.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I would much prefer it if we could make Parquet safe to use in multi threaded environments.

@andygrove

Copy link
Copy Markdown
MemberAuthor

@andygrove I pull your request and run the tests. The root cause is that currently arrow reader doesn't support some data types(e.g., UTF8) and it caused program to crash.

We should make the code fail gracefully by returning an Err for unsupported types.

@liurenjie1024

Copy link
Copy Markdown
Contributor

@andygrove After debugging, I found that this is not caused by unwrap call, but by error in supporting UTF8. I'll add support for utf8 and this will be fixed.

@andygroveandygrove changed the title ARROW-6700: [Rust] [DataFusion] Use new Arrow reader [WIP]ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet reader [WIP]Nov 18, 2019
@andygrove

Copy link
Copy Markdown
MemberAuthor

Hi @liurenjie1024 is there any update on adding support for UTF8?

@liurenjie1024

Copy link
Copy Markdown
Contributor

@andygrove Sorry, almost forgot about this. I'll address it this week.

@andygrove

Copy link
Copy Markdown
MemberAuthor

Hi @liurenjie1024 the next issue is that I have a regression due to Reading Timestamp(Nanosecond) type from parquet is not supported yet. Do you think this is an easy one to add? It requires a new Int96Type and Int96Converter.

@andygrove

Copy link
Copy Markdown
MemberAuthor

@liurenjie1024 I attempted to add support to the array reader for TimestampNanoseconds. It runs but returns the wrong results currently. Could you take a look?

@nevi-me

Copy link
Copy Markdown
Contributor

@andygrove isn't it the same issue that timestamps are stored in 96bit values? How are you converting from the 96bit values?

@andygrove

Copy link
Copy Markdown
MemberAuthor

@nevi-me Yes, I guess I need to write a custom converter here rather than try and use the CastConverter. I'll have a go at that.

@liurenjie1024

Copy link
Copy Markdown
Contributor

@andygrove I'll take a look this week.

@nevi-me

Copy link
Copy Markdown
Contributor

I've created a complex converter for int96, and I've fixed the binary reads. The problem with the binary reads was that when I introduced StringType, I missed a part in the parquet code where we were supposed to convert a binary physical type with no logical type, to BinaryType.

All tests are passing locally for me.

@nevi-me

Copy link
Copy Markdown
Contributor

Oops, I forgot examples. @andygrove, the alltypes_plain.parquet file has saved the string columns as binary. Instead of a quick fix of converting the binary data to utf8, this might be an opportunity for us to create a cast kernel, or some string kernel that converts BinaryArray to StringArray.

Then in DataFusion we could do an implicit cast to string in the SQL code. What do you think about this?

@andygrove

Copy link
Copy Markdown
MemberAuthor

Wow @nevi-me that's awesome! If the columns are stored as binary then DataFusion should also treat them as binary unless the user adds an explicit CAST in the SQL. I pushed a quick change for that and all tests are passing for me now.

I know there are some things I need to clean up in this PR so I'll start working on those tomorrow.

@andygroveandygrove changed the title ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet reader [WIP]ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet readerDec 10, 2019
@andygroveandygrove removed the WIP PR is work in progress label Dec 10, 2019
if array.is_null(i) {
b.append(false)?;
} else {
b.append_value(str::from_utf8(from.value(i)).unwrap())?;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

from_utf8 can panic if the data is not valid utf8 bytes. Perhaps it's better to convert failures to null values instead? This would behave similarly to other casts that return nulls on overflowing data.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why we can't return error? I don't it's correct to return null rather than error.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We've had the broader discussion around casts, on whether to return an error on overflows/invalid data, or whether to expose an option to the user to define their desired behaviour. I haven't opened a JIRA for this, perhaps I should, so we can address the overall cast behaviour there?

For now, an expedient would actually not be to downcast to BinaryArray, but to create a StringArray from the binary data, but maybe that's a premature optimisation.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Another discussion about configuring behaviors of error handling would be fine. But I can't understand for now why we can't just return error? Inconsistency with existing behavior?

@nevi-menevi-meDec 10, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, inconsistency. For example, if you cast a i64 to u8, negative values and values that don't fit into u8 will be cast to nulls. It doesn't return an error.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For this PR, I agree that we can return null to match existing behavior. But I don't think it's good idea to make returning null as default behavior, but returning error should be. Please open an jira ticket to track this.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Comment threadrust/datafusion/src/sql/planner.rs Outdated
SQLType::Float(_) | SQLType::Real => Ok(DataType::Float64),
SQLType::Double => Ok(DataType::Float64),
SQLType::Char(_) | SQLType::Varchar(_) => Ok(DataType::Utf8),
SQLType::Custom(t) if t.to_lowercase() == "string" => Ok(DataType::Utf8),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We probably don't need this just yet, as cast(binary as varchar) works without any changes.

@paddyhoranpaddyhoran left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm a little out of the loop (day job, etc.), but LGTM at a high level. Thanks @andygrove@nevi-me and @liurenjie1024, great progress.

);
}

//TODO assertions

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are you going to add assertions in this PR?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes. I'm away on a business trip today and tomorrow but intend on adding assertions this weekend.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@andygrove@liurenjie1024@nevi-me@alippai@paddyhoran
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet reader - #5641

Closed
andygrove wants to merge 11 commits into
apache:masterfrom
andygrove:ARROW-6700
Closed

ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet reader#5641
andygrove wants to merge 11 commits into
apache:masterfrom
andygrove:ARROW-6700

Conversation

@andygrove

@andygroveandygrove commented Oct 13, 2019

Copy link
Copy Markdown
Member

Replaces the DataFusion Parquet reader with the new Arrow reader in the parquet crate.

@github-actions

Copy link
Copy Markdown

@liurenjie1024liurenjie1024 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some part should be replaced by new ArrowReader.

Comment threadrust/parquet/src/arrow/mod.rs Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

array_reader is not designed for public use. This PR #5523 contains public api and doc example. Essentially, you should use ArrowReader

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please use new ArrowReader here.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Schema should also consider projection.

@andygrove

Copy link
Copy Markdown
MemberAuthor

@liurenjie1024 I updated to use ArrowReader. When I run cargo test I actually get a SIGSEGV:

error: process didn't exit successfully: `/home/andy/git/andygrove/arrow/rust/target/debug/deps/datafusion-39547aa10aa86781` (signal: 11, SIGSEGV: invalid memory reference)

@liurenjie1024

Copy link
Copy Markdown
Contributor

@andygrove I pull your request and run the tests. The root cause is that currently arrow reader doesn't support some data types(e.g., UTF8) and it caused program to crash.


fn next(&mut self) -> Result<Option<RecordBatch>> {
match self.request_tx.send(()) {
Ok(_) => match self.response_rx.recv() {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why we need another thread here? This send request, wait response model is also blocked here waiting for IO.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We need the threading because the Parquet structs/traits do not implement Sync + Send and cannot be sent between threads.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I would much prefer it if we could make Parquet safe to use in multi threaded environments.

@andygrove

Copy link
Copy Markdown
MemberAuthor

@andygrove I pull your request and run the tests. The root cause is that currently arrow reader doesn't support some data types(e.g., UTF8) and it caused program to crash.

We should make the code fail gracefully by returning an Err for unsupported types.

@liurenjie1024

Copy link
Copy Markdown
Contributor

@andygrove After debugging, I found that this is not caused by unwrap call, but by error in supporting UTF8. I'll add support for utf8 and this will be fixed.

@andygroveandygrove changed the title ARROW-6700: [Rust] [DataFusion] Use new Arrow reader [WIP]ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet reader [WIP]Nov 18, 2019
@andygrove

Copy link
Copy Markdown
MemberAuthor

Hi @liurenjie1024 is there any update on adding support for UTF8?

@liurenjie1024

Copy link
Copy Markdown
Contributor

@andygrove Sorry, almost forgot about this. I'll address it this week.

@andygrove

Copy link
Copy Markdown
MemberAuthor

Hi @liurenjie1024 the next issue is that I have a regression due to Reading Timestamp(Nanosecond) type from parquet is not supported yet. Do you think this is an easy one to add? It requires a new Int96Type and Int96Converter.

@andygrove

Copy link
Copy Markdown
MemberAuthor

@liurenjie1024 I attempted to add support to the array reader for TimestampNanoseconds. It runs but returns the wrong results currently. Could you take a look?

@nevi-me

Copy link
Copy Markdown
Contributor

@andygrove isn't it the same issue that timestamps are stored in 96bit values? How are you converting from the 96bit values?

@andygrove

Copy link
Copy Markdown
MemberAuthor

@nevi-me Yes, I guess I need to write a custom converter here rather than try and use the CastConverter. I'll have a go at that.

@liurenjie1024

Copy link
Copy Markdown
Contributor

@andygrove I'll take a look this week.

@nevi-me

Copy link
Copy Markdown
Contributor

I've created a complex converter for int96, and I've fixed the binary reads. The problem with the binary reads was that when I introduced StringType, I missed a part in the parquet code where we were supposed to convert a binary physical type with no logical type, to BinaryType.

All tests are passing locally for me.

@nevi-me

Copy link
Copy Markdown
Contributor

Oops, I forgot examples. @andygrove, the alltypes_plain.parquet file has saved the string columns as binary. Instead of a quick fix of converting the binary data to utf8, this might be an opportunity for us to create a cast kernel, or some string kernel that converts BinaryArray to StringArray.

Then in DataFusion we could do an implicit cast to string in the SQL code. What do you think about this?

@andygrove

Copy link
Copy Markdown
MemberAuthor

Wow @nevi-me that's awesome! If the columns are stored as binary then DataFusion should also treat them as binary unless the user adds an explicit CAST in the SQL. I pushed a quick change for that and all tests are passing for me now.

I know there are some things I need to clean up in this PR so I'll start working on those tomorrow.

@andygroveandygrove changed the title ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet reader [WIP]ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet readerDec 10, 2019
@andygroveandygrove removed the WIP PR is work in progress label Dec 10, 2019
if array.is_null(i) {
b.append(false)?;
} else {
b.append_value(str::from_utf8(from.value(i)).unwrap())?;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

from_utf8 can panic if the data is not valid utf8 bytes. Perhaps it's better to convert failures to null values instead? This would behave similarly to other casts that return nulls on overflowing data.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why we can't return error? I don't it's correct to return null rather than error.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We've had the broader discussion around casts, on whether to return an error on overflows/invalid data, or whether to expose an option to the user to define their desired behaviour. I haven't opened a JIRA for this, perhaps I should, so we can address the overall cast behaviour there?

For now, an expedient would actually not be to downcast to BinaryArray, but to create a StringArray from the binary data, but maybe that's a premature optimisation.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Another discussion about configuring behaviors of error handling would be fine. But I can't understand for now why we can't just return error? Inconsistency with existing behavior?

@nevi-menevi-meDec 10, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, inconsistency. For example, if you cast a i64 to u8, negative values and values that don't fit into u8 will be cast to nulls. It doesn't return an error.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For this PR, I agree that we can return null to match existing behavior. But I don't think it's good idea to make returning null as default behavior, but returning error should be. Please open an jira ticket to track this.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Comment threadrust/datafusion/src/sql/planner.rs Outdated
SQLType::Float(_) | SQLType::Real => Ok(DataType::Float64),
SQLType::Double => Ok(DataType::Float64),
SQLType::Char(_) | SQLType::Varchar(_) => Ok(DataType::Utf8),
SQLType::Custom(t) if t.to_lowercase() == "string" => Ok(DataType::Utf8),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We probably don't need this just yet, as cast(binary as varchar) works without any changes.

@paddyhoranpaddyhoran left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm a little out of the loop (day job, etc.), but LGTM at a high level. Thanks @andygrove@nevi-me and @liurenjie1024, great progress.

);
}

//TODO assertions

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are you going to add assertions in this PR?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes. I'm away on a business trip today and tomorrow but intend on adding assertions this weekend.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@andygrove@liurenjie1024@nevi-me@alippai@paddyhoran
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet reader - #5641

Closed
andygrove wants to merge 11 commits into
apache:masterfrom
andygrove:ARROW-6700
Closed

ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet reader#5641
andygrove wants to merge 11 commits into
apache:masterfrom
andygrove:ARROW-6700

Conversation

@andygrove

@andygroveandygrove commented Oct 13, 2019

Copy link
Copy Markdown
Member

Replaces the DataFusion Parquet reader with the new Arrow reader in the parquet crate.

@github-actions

Copy link
Copy Markdown

@liurenjie1024liurenjie1024 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some part should be replaced by new ArrowReader.

Comment threadrust/parquet/src/arrow/mod.rs Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

array_reader is not designed for public use. This PR #5523 contains public api and doc example. Essentially, you should use ArrowReader

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please use new ArrowReader here.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Schema should also consider projection.

@andygrove

Copy link
Copy Markdown
MemberAuthor

@liurenjie1024 I updated to use ArrowReader. When I run cargo test I actually get a SIGSEGV:

error: process didn't exit successfully: `/home/andy/git/andygrove/arrow/rust/target/debug/deps/datafusion-39547aa10aa86781` (signal: 11, SIGSEGV: invalid memory reference)

@liurenjie1024

Copy link
Copy Markdown
Contributor

@andygrove I pull your request and run the tests. The root cause is that currently arrow reader doesn't support some data types(e.g., UTF8) and it caused program to crash.


fn next(&mut self) -> Result<Option<RecordBatch>> {
match self.request_tx.send(()) {
Ok(_) => match self.response_rx.recv() {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why we need another thread here? This send request, wait response model is also blocked here waiting for IO.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We need the threading because the Parquet structs/traits do not implement Sync + Send and cannot be sent between threads.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I would much prefer it if we could make Parquet safe to use in multi threaded environments.

@andygrove

Copy link
Copy Markdown
MemberAuthor

@andygrove I pull your request and run the tests. The root cause is that currently arrow reader doesn't support some data types(e.g., UTF8) and it caused program to crash.

We should make the code fail gracefully by returning an Err for unsupported types.

@liurenjie1024

Copy link
Copy Markdown
Contributor

@andygrove After debugging, I found that this is not caused by unwrap call, but by error in supporting UTF8. I'll add support for utf8 and this will be fixed.

@andygroveandygrove changed the title ARROW-6700: [Rust] [DataFusion] Use new Arrow reader [WIP]ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet reader [WIP]Nov 18, 2019
@andygrove

Copy link
Copy Markdown
MemberAuthor

Hi @liurenjie1024 is there any update on adding support for UTF8?

@liurenjie1024

Copy link
Copy Markdown
Contributor

@andygrove Sorry, almost forgot about this. I'll address it this week.

@andygrove

Copy link
Copy Markdown
MemberAuthor

Hi @liurenjie1024 the next issue is that I have a regression due to Reading Timestamp(Nanosecond) type from parquet is not supported yet. Do you think this is an easy one to add? It requires a new Int96Type and Int96Converter.

@andygrove

Copy link
Copy Markdown
MemberAuthor

@liurenjie1024 I attempted to add support to the array reader for TimestampNanoseconds. It runs but returns the wrong results currently. Could you take a look?

@nevi-me

Copy link
Copy Markdown
Contributor

@andygrove isn't it the same issue that timestamps are stored in 96bit values? How are you converting from the 96bit values?

@andygrove

Copy link
Copy Markdown
MemberAuthor

@nevi-me Yes, I guess I need to write a custom converter here rather than try and use the CastConverter. I'll have a go at that.

@liurenjie1024

Copy link
Copy Markdown
Contributor

@andygrove I'll take a look this week.

@nevi-me

Copy link
Copy Markdown
Contributor

I've created a complex converter for int96, and I've fixed the binary reads. The problem with the binary reads was that when I introduced StringType, I missed a part in the parquet code where we were supposed to convert a binary physical type with no logical type, to BinaryType.

All tests are passing locally for me.

@nevi-me

Copy link
Copy Markdown
Contributor

Oops, I forgot examples. @andygrove, the alltypes_plain.parquet file has saved the string columns as binary. Instead of a quick fix of converting the binary data to utf8, this might be an opportunity for us to create a cast kernel, or some string kernel that converts BinaryArray to StringArray.

Then in DataFusion we could do an implicit cast to string in the SQL code. What do you think about this?

@andygrove

Copy link
Copy Markdown
MemberAuthor

Wow @nevi-me that's awesome! If the columns are stored as binary then DataFusion should also treat them as binary unless the user adds an explicit CAST in the SQL. I pushed a quick change for that and all tests are passing for me now.

I know there are some things I need to clean up in this PR so I'll start working on those tomorrow.

@andygroveandygrove changed the title ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet reader [WIP]ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet readerDec 10, 2019
@andygroveandygrove removed the WIP PR is work in progress label Dec 10, 2019
if array.is_null(i) {
b.append(false)?;
} else {
b.append_value(str::from_utf8(from.value(i)).unwrap())?;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

from_utf8 can panic if the data is not valid utf8 bytes. Perhaps it's better to convert failures to null values instead? This would behave similarly to other casts that return nulls on overflowing data.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why we can't return error? I don't it's correct to return null rather than error.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We've had the broader discussion around casts, on whether to return an error on overflows/invalid data, or whether to expose an option to the user to define their desired behaviour. I haven't opened a JIRA for this, perhaps I should, so we can address the overall cast behaviour there?

For now, an expedient would actually not be to downcast to BinaryArray, but to create a StringArray from the binary data, but maybe that's a premature optimisation.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Another discussion about configuring behaviors of error handling would be fine. But I can't understand for now why we can't just return error? Inconsistency with existing behavior?

@nevi-menevi-meDec 10, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, inconsistency. For example, if you cast a i64 to u8, negative values and values that don't fit into u8 will be cast to nulls. It doesn't return an error.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For this PR, I agree that we can return null to match existing behavior. But I don't think it's good idea to make returning null as default behavior, but returning error should be. Please open an jira ticket to track this.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Comment threadrust/datafusion/src/sql/planner.rs Outdated
SQLType::Float(_) | SQLType::Real => Ok(DataType::Float64),
SQLType::Double => Ok(DataType::Float64),
SQLType::Char(_) | SQLType::Varchar(_) => Ok(DataType::Utf8),
SQLType::Custom(t) if t.to_lowercase() == "string" => Ok(DataType::Utf8),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We probably don't need this just yet, as cast(binary as varchar) works without any changes.

@paddyhoranpaddyhoran left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm a little out of the loop (day job, etc.), but LGTM at a high level. Thanks @andygrove@nevi-me and @liurenjie1024, great progress.

);
}

//TODO assertions

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are you going to add assertions in this PR?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes. I'm away on a business trip today and tomorrow but intend on adding assertions this weekend.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@andygrove@liurenjie1024@nevi-me@alippai@paddyhoran
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet reader - #5641

Closed
andygrove wants to merge 11 commits into
apache:masterfrom
andygrove:ARROW-6700
Closed

ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet reader#5641
andygrove wants to merge 11 commits into
apache:masterfrom
andygrove:ARROW-6700

Conversation

@andygrove

@andygroveandygrove commented Oct 13, 2019

Copy link
Copy Markdown
Member

Replaces the DataFusion Parquet reader with the new Arrow reader in the parquet crate.

@github-actions

Copy link
Copy Markdown

@liurenjie1024liurenjie1024 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some part should be replaced by new ArrowReader.

Comment threadrust/parquet/src/arrow/mod.rs Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

array_reader is not designed for public use. This PR #5523 contains public api and doc example. Essentially, you should use ArrowReader

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please use new ArrowReader here.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Schema should also consider projection.

@andygrove

Copy link
Copy Markdown
MemberAuthor

@liurenjie1024 I updated to use ArrowReader. When I run cargo test I actually get a SIGSEGV:

error: process didn't exit successfully: `/home/andy/git/andygrove/arrow/rust/target/debug/deps/datafusion-39547aa10aa86781` (signal: 11, SIGSEGV: invalid memory reference)

@liurenjie1024

Copy link
Copy Markdown
Contributor

@andygrove I pull your request and run the tests. The root cause is that currently arrow reader doesn't support some data types(e.g., UTF8) and it caused program to crash.


fn next(&mut self) -> Result<Option<RecordBatch>> {
match self.request_tx.send(()) {
Ok(_) => match self.response_rx.recv() {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why we need another thread here? This send request, wait response model is also blocked here waiting for IO.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We need the threading because the Parquet structs/traits do not implement Sync + Send and cannot be sent between threads.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I would much prefer it if we could make Parquet safe to use in multi threaded environments.

@andygrove

Copy link
Copy Markdown
MemberAuthor

@andygrove I pull your request and run the tests. The root cause is that currently arrow reader doesn't support some data types(e.g., UTF8) and it caused program to crash.

We should make the code fail gracefully by returning an Err for unsupported types.

@liurenjie1024

Copy link
Copy Markdown
Contributor

@andygrove After debugging, I found that this is not caused by unwrap call, but by error in supporting UTF8. I'll add support for utf8 and this will be fixed.

@andygroveandygrove changed the title ARROW-6700: [Rust] [DataFusion] Use new Arrow reader [WIP]ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet reader [WIP]Nov 18, 2019
@andygrove

Copy link
Copy Markdown
MemberAuthor

Hi @liurenjie1024 is there any update on adding support for UTF8?

@liurenjie1024

Copy link
Copy Markdown
Contributor

@andygrove Sorry, almost forgot about this. I'll address it this week.

@andygrove

Copy link
Copy Markdown
MemberAuthor

Hi @liurenjie1024 the next issue is that I have a regression due to Reading Timestamp(Nanosecond) type from parquet is not supported yet. Do you think this is an easy one to add? It requires a new Int96Type and Int96Converter.

@andygrove

Copy link
Copy Markdown
MemberAuthor

@liurenjie1024 I attempted to add support to the array reader for TimestampNanoseconds. It runs but returns the wrong results currently. Could you take a look?

@nevi-me

Copy link
Copy Markdown
Contributor

@andygrove isn't it the same issue that timestamps are stored in 96bit values? How are you converting from the 96bit values?

@andygrove

Copy link
Copy Markdown
MemberAuthor

@nevi-me Yes, I guess I need to write a custom converter here rather than try and use the CastConverter. I'll have a go at that.

@liurenjie1024

Copy link
Copy Markdown
Contributor

@andygrove I'll take a look this week.

@nevi-me

Copy link
Copy Markdown
Contributor

I've created a complex converter for int96, and I've fixed the binary reads. The problem with the binary reads was that when I introduced StringType, I missed a part in the parquet code where we were supposed to convert a binary physical type with no logical type, to BinaryType.

All tests are passing locally for me.

@nevi-me

Copy link
Copy Markdown
Contributor

Oops, I forgot examples. @andygrove, the alltypes_plain.parquet file has saved the string columns as binary. Instead of a quick fix of converting the binary data to utf8, this might be an opportunity for us to create a cast kernel, or some string kernel that converts BinaryArray to StringArray.

Then in DataFusion we could do an implicit cast to string in the SQL code. What do you think about this?

@andygrove

Copy link
Copy Markdown
MemberAuthor

Wow @nevi-me that's awesome! If the columns are stored as binary then DataFusion should also treat them as binary unless the user adds an explicit CAST in the SQL. I pushed a quick change for that and all tests are passing for me now.

I know there are some things I need to clean up in this PR so I'll start working on those tomorrow.

@andygroveandygrove changed the title ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet reader [WIP]ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet readerDec 10, 2019
@andygroveandygrove removed the WIP PR is work in progress label Dec 10, 2019
if array.is_null(i) {
b.append(false)?;
} else {
b.append_value(str::from_utf8(from.value(i)).unwrap())?;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

from_utf8 can panic if the data is not valid utf8 bytes. Perhaps it's better to convert failures to null values instead? This would behave similarly to other casts that return nulls on overflowing data.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why we can't return error? I don't it's correct to return null rather than error.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We've had the broader discussion around casts, on whether to return an error on overflows/invalid data, or whether to expose an option to the user to define their desired behaviour. I haven't opened a JIRA for this, perhaps I should, so we can address the overall cast behaviour there?

For now, an expedient would actually not be to downcast to BinaryArray, but to create a StringArray from the binary data, but maybe that's a premature optimisation.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Another discussion about configuring behaviors of error handling would be fine. But I can't understand for now why we can't just return error? Inconsistency with existing behavior?

@nevi-menevi-meDec 10, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, inconsistency. For example, if you cast a i64 to u8, negative values and values that don't fit into u8 will be cast to nulls. It doesn't return an error.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For this PR, I agree that we can return null to match existing behavior. But I don't think it's good idea to make returning null as default behavior, but returning error should be. Please open an jira ticket to track this.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Comment threadrust/datafusion/src/sql/planner.rs Outdated
SQLType::Float(_) | SQLType::Real => Ok(DataType::Float64),
SQLType::Double => Ok(DataType::Float64),
SQLType::Char(_) | SQLType::Varchar(_) => Ok(DataType::Utf8),
SQLType::Custom(t) if t.to_lowercase() == "string" => Ok(DataType::Utf8),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We probably don't need this just yet, as cast(binary as varchar) works without any changes.

@paddyhoranpaddyhoran left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm a little out of the loop (day job, etc.), but LGTM at a high level. Thanks @andygrove@nevi-me and @liurenjie1024, great progress.

);
}

//TODO assertions

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are you going to add assertions in this PR?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes. I'm away on a business trip today and tomorrow but intend on adding assertions this weekend.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@andygrove@liurenjie1024@nevi-me@alippai@paddyhoran
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet reader - #5641

Closed
andygrove wants to merge 11 commits into
apache:masterfrom
andygrove:ARROW-6700
Closed

ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet reader#5641
andygrove wants to merge 11 commits into
apache:masterfrom
andygrove:ARROW-6700

Conversation

@andygrove

@andygroveandygrove commented Oct 13, 2019

Copy link
Copy Markdown
Member

Replaces the DataFusion Parquet reader with the new Arrow reader in the parquet crate.

@github-actions

Copy link
Copy Markdown

@liurenjie1024liurenjie1024 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some part should be replaced by new ArrowReader.

Comment threadrust/parquet/src/arrow/mod.rs Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

array_reader is not designed for public use. This PR #5523 contains public api and doc example. Essentially, you should use ArrowReader

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please use new ArrowReader here.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Schema should also consider projection.

@andygrove

Copy link
Copy Markdown
MemberAuthor

@liurenjie1024 I updated to use ArrowReader. When I run cargo test I actually get a SIGSEGV:

error: process didn't exit successfully: `/home/andy/git/andygrove/arrow/rust/target/debug/deps/datafusion-39547aa10aa86781` (signal: 11, SIGSEGV: invalid memory reference)

@liurenjie1024

Copy link
Copy Markdown
Contributor

@andygrove I pull your request and run the tests. The root cause is that currently arrow reader doesn't support some data types(e.g., UTF8) and it caused program to crash.


fn next(&mut self) -> Result<Option<RecordBatch>> {
match self.request_tx.send(()) {
Ok(_) => match self.response_rx.recv() {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why we need another thread here? This send request, wait response model is also blocked here waiting for IO.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We need the threading because the Parquet structs/traits do not implement Sync + Send and cannot be sent between threads.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I would much prefer it if we could make Parquet safe to use in multi threaded environments.

@andygrove

Copy link
Copy Markdown
MemberAuthor

@andygrove I pull your request and run the tests. The root cause is that currently arrow reader doesn't support some data types(e.g., UTF8) and it caused program to crash.

We should make the code fail gracefully by returning an Err for unsupported types.

@liurenjie1024

Copy link
Copy Markdown
Contributor

@andygrove After debugging, I found that this is not caused by unwrap call, but by error in supporting UTF8. I'll add support for utf8 and this will be fixed.

@andygroveandygrove changed the title ARROW-6700: [Rust] [DataFusion] Use new Arrow reader [WIP]ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet reader [WIP]Nov 18, 2019
@andygrove

Copy link
Copy Markdown
MemberAuthor

Hi @liurenjie1024 is there any update on adding support for UTF8?

@liurenjie1024

Copy link
Copy Markdown
Contributor

@andygrove Sorry, almost forgot about this. I'll address it this week.

@andygrove

Copy link
Copy Markdown
MemberAuthor

Hi @liurenjie1024 the next issue is that I have a regression due to Reading Timestamp(Nanosecond) type from parquet is not supported yet. Do you think this is an easy one to add? It requires a new Int96Type and Int96Converter.

@andygrove

Copy link
Copy Markdown
MemberAuthor

@liurenjie1024 I attempted to add support to the array reader for TimestampNanoseconds. It runs but returns the wrong results currently. Could you take a look?

@nevi-me

Copy link
Copy Markdown
Contributor

@andygrove isn't it the same issue that timestamps are stored in 96bit values? How are you converting from the 96bit values?

@andygrove

Copy link
Copy Markdown
MemberAuthor

@nevi-me Yes, I guess I need to write a custom converter here rather than try and use the CastConverter. I'll have a go at that.

@liurenjie1024

Copy link
Copy Markdown
Contributor

@andygrove I'll take a look this week.

@nevi-me

Copy link
Copy Markdown
Contributor

I've created a complex converter for int96, and I've fixed the binary reads. The problem with the binary reads was that when I introduced StringType, I missed a part in the parquet code where we were supposed to convert a binary physical type with no logical type, to BinaryType.

All tests are passing locally for me.

@nevi-me

Copy link
Copy Markdown
Contributor

Oops, I forgot examples. @andygrove, the alltypes_plain.parquet file has saved the string columns as binary. Instead of a quick fix of converting the binary data to utf8, this might be an opportunity for us to create a cast kernel, or some string kernel that converts BinaryArray to StringArray.

Then in DataFusion we could do an implicit cast to string in the SQL code. What do you think about this?

@andygrove

Copy link
Copy Markdown
MemberAuthor

Wow @nevi-me that's awesome! If the columns are stored as binary then DataFusion should also treat them as binary unless the user adds an explicit CAST in the SQL. I pushed a quick change for that and all tests are passing for me now.

I know there are some things I need to clean up in this PR so I'll start working on those tomorrow.

@andygroveandygrove changed the title ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet reader [WIP]ARROW-6700: [Rust] [DataFusion] Use new Arrow Parquet readerDec 10, 2019
@andygroveandygrove removed the WIP PR is work in progress label Dec 10, 2019
if array.is_null(i) {
b.append(false)?;
} else {
b.append_value(str::from_utf8(from.value(i)).unwrap())?;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

from_utf8 can panic if the data is not valid utf8 bytes. Perhaps it's better to convert failures to null values instead? This would behave similarly to other casts that return nulls on overflowing data.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why we can't return error? I don't it's correct to return null rather than error.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We've had the broader discussion around casts, on whether to return an error on overflows/invalid data, or whether to expose an option to the user to define their desired behaviour. I haven't opened a JIRA for this, perhaps I should, so we can address the overall cast behaviour there?

For now, an expedient would actually not be to downcast to BinaryArray, but to create a StringArray from the binary data, but maybe that's a premature optimisation.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Another discussion about configuring behaviors of error handling would be fine. But I can't understand for now why we can't just return error? Inconsistency with existing behavior?

@nevi-menevi-meDec 10, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, inconsistency. For example, if you cast a i64 to u8, negative values and values that don't fit into u8 will be cast to nulls. It doesn't return an error.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For this PR, I agree that we can return null to match existing behavior. But I don't think it's good idea to make returning null as default behavior, but returning error should be. Please open an jira ticket to track this.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Comment threadrust/datafusion/src/sql/planner.rs Outdated
SQLType::Float(_) | SQLType::Real => Ok(DataType::Float64),
SQLType::Double => Ok(DataType::Float64),
SQLType::Char(_) | SQLType::Varchar(_) => Ok(DataType::Utf8),
SQLType::Custom(t) if t.to_lowercase() == "string" => Ok(DataType::Utf8),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We probably don't need this just yet, as cast(binary as varchar) works without any changes.

@paddyhoranpaddyhoran left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm a little out of the loop (day job, etc.), but LGTM at a high level. Thanks @andygrove@nevi-me and @liurenjie1024, great progress.

);
}

//TODO assertions

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are you going to add assertions in this PR?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes. I'm away on a business trip today and tomorrow but intend on adding assertions this weekend.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@andygrove@liurenjie1024@nevi-me@alippai@paddyhoran