ARROW-12105: [R] Replace vars_select, vars_rename with eval_select, eval_rename - #14371

Merged
thisisnic merged 22 commits into
apache:masterfrom
thisisnic:ARROW-12105_eval_select
Oct 14, 2022
Merged

ARROW-12105: [R] Replace vars_select, vars_rename with eval_select, eval_rename#14371
thisisnic merged 22 commits into
apache:masterfrom
thisisnic:ARROW-12105_eval_select

Conversation

@thisisnic

Copy link
Copy Markdown
Member

No description provided.

)
})

test_that("multiple select/rename and group_by", {

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I added these in because they are needed to make sure the implementation of the utility function column_select() is working properly.

@nealrichardsonnealrichardson left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for taking this on!

Comment threadr/R/arrow-package.R Outdated
#' @importFrom rlang quo_set_env quo_get_env is_formula quo_is_call f_rhs parse_expr f_env new_quosure
#' @importFrom rlang new_quosures expr_text
#' @importFrom tidyselect vars_pull vars_rename vars_select eval_select
#' @importFrom tidyselect vars_pull vars_select eval_select eval_rename

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we need to update the other usages of vars_select in the package too?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep - this is still on my to-do list, but I think all other feedback has been addressed.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done now

Comment threadr/R/util.R Outdated
Comment threadr/data-raw/docgen.R
# TODO(ARROW-17384): implement where
"Use of `where()` selection helper not yet supported"
)
docs[["dplyr::across"]] <- character(0)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎉

Comment threadr/R/dplyr-select.R Outdated
Comment threadr/R/util.R Outdated
}

simulate_data_frame <- function(schema) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Add a comment explaining why we need this function.

And do we need to export this? I think I saw some discussion between @paleolimbot and @krlmlr about needing this function in DBI or adbc.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In DBI, we need to infer the SQL data types from a RecordBatchReader, even if the DBI backend is not aware of Arrow (yet).

I wonder if there should be a way to convert a schema to a zero-row Arrow table. One way to do this is via schema -> zero-row table -> data frame. For DBI, schema -> data frame is sufficient, but perhaps schema -> zero-row table is easier to implemented in C++, and the table -> data frame operation is already efficient enough for zero-row tables.

Comment threadr/R/util.R Outdated
Comment threadr/R/util.R Outdated
abort(msg, call = call)
}

simulate_data_frame <- function(schema) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Since this is going to be called every time someone does select/rename/relocate, I'd like for this function to be cheaper. In other PRs I've been noticing the overhead of creating R6 objects, which generally is not terrible (~150 microseconds on my machine) but it adds up. And here, we're creating lots of objects we're throwing away: for each column, we create a Field, then a DataType from that, then in concat_arrays, we create a null DataType, an Array with that, and a new Array that is cast to the correct DataType. That adds up to around 1ms per column, every time this function is called. That's enough to get noticed.

Can we move this to C++? Should be a simple enough switch statement to map Arrow type ids to the corresponding R length-0 vector.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just to confirm, I benchmarked this function on the schema of the taxi dataset (20 columns), and the median time was 15ms, so a little under 1ms per column.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for confirming! Currently working on a C++ simplification though may ask for help to get it over the line ahead of the release if I can't figure it out by the end of tomorrow.

Comment threadr/src/table.cpp
Co-authored-by: Neal Richardson <neal.p.richardson@gmail.com>
@thisisnic

thisisnic commented Oct 13, 2022

Copy link
Copy Markdown
MemberAuthor

Currently getting this error message in the unit tests, due to weirdness with trying to create 0-length arrays using extension types 😬

Will try to create a nice reprex/solution/workaround, but if anyone has any suggestions in the meantime, let me know.

Error (test-extension.R:340:3): Dataset/arrow_dplyr_query can roundtrip extension types
Error: NotImplemented: MakeBuilder: cannot construct builder for type <arrow_custom_vctr[0]>
/home/nic2/arrow/cpp/src/arrow/builder.cc:289 VisitTypeInline(*type, &impl)
Backtrace:
1. ... %>% dplyr::collect()
at test-extension.R:340:2
4. arrow::select.Dataset(., number, letter, extension)
5. arrow::column_select(.data, enquos(...), op = "select")
at r/R/dplyr-select.R:24:2
6. arrow::simulate_data_frame(implicit_schema(.data))
at r/R/dplyr-select.R:97:2
8. arrow::Table__from_schema(schema)

It can be reproduced here via:

extension_vec <- vctrs::new_vctr(letters[1:10], class = "arrow_custom_vctr")
df <- tibble::tibble(x = extension_vec)
arrow_df <- arrow_table(df)
simulate_data_frame(arrow_df$schema)
Error: NotImplemented: MakeBuilder: cannot construct builder for type <arrow_custom_vctr[0]>
/home/nic2/arrow/cpp/src/arrow/builder.cc:289 VisitTypeInline(*type, &impl)

@nealrichardson

Copy link
Copy Markdown
Member

I'm sure it's solvable since we can roundtrip extension type data with >0 elements. Dewey may be best positioned to advise. Can you defer that to a followup, and find some workaround here, perhaps swap in null() type for extension types?

@thisisnic

Copy link
Copy Markdown
MemberAuthor

defer that to a followup

My favourite 5 words. Opened ARROW-18043

@thisisnic

Copy link
Copy Markdown
MemberAuthor

@nealrichardson Mind giving this another look over? I'm going to get to the "swap vars_select for eval_select" remaining bits tomorrow morning, but it'd be good to have you look at the rest of it in case there's any changes I need to make there too.

@nealrichardson

Copy link
Copy Markdown
Member

Looks great to me!

@thisisnic

Copy link
Copy Markdown
MemberAuthor

@nealrichardson Waiting on the CI, but otherwise I think this is ready for a final round of reviewing :D

@thisisnic
thisisnic merged commit 2cbf489 into apache:masterOct 14, 2022
@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = bd785c9 and contender = 2cbf489. 2cbf489 is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed ⬇️53.33% ⬆️0.0%] test-mac-arm
[Failed ⬇️28.22% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.39% ⬆️0.04%] ursa-thinkcentre-m75q
Buildkite builds:
[Finished] 2cbf4891 ec2-t3-xlarge-us-east-2
[Failed] 2cbf4891 test-mac-arm
[Failed] 2cbf4891 ursa-i9-9960x
[Finished] 2cbf4891 ursa-thinkcentre-m75q
[Finished] bd785c99 ec2-t3-xlarge-us-east-2
[Failed] bd785c99 test-mac-arm
[Failed] bd785c99 ursa-i9-9960x
[Finished] bd785c99 ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

@ursabot

Copy link
Copy Markdown

['Python', 'R'] benchmarks have high level of regressions.
test-mac-arm
ursa-i9-9960x

@thisisnic

Copy link
Copy Markdown
MemberAuthor

@nealrichardson I think I may need to refactor this in a follow-up PR - check out the regressions :\ Any suggestion for what I can do instead?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@thisisnic@nealrichardson@ursabot@krlmlr
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

ARROW-12105: [R] Replace vars_select, vars_rename with eval_select, eval_rename - #14371

Merged
thisisnic merged 22 commits into
apache:masterfrom
thisisnic:ARROW-12105_eval_select
Oct 14, 2022
Merged

ARROW-12105: [R] Replace vars_select, vars_rename with eval_select, eval_rename#14371
thisisnic merged 22 commits into
apache:masterfrom
thisisnic:ARROW-12105_eval_select

Conversation

@thisisnic

Copy link
Copy Markdown
Member

No description provided.

)
})

test_that("multiple select/rename and group_by", {

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I added these in because they are needed to make sure the implementation of the utility function column_select() is working properly.

@nealrichardsonnealrichardson left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for taking this on!

Comment threadr/R/arrow-package.R Outdated
#' @importFrom rlang quo_set_env quo_get_env is_formula quo_is_call f_rhs parse_expr f_env new_quosure
#' @importFrom rlang new_quosures expr_text
#' @importFrom tidyselect vars_pull vars_rename vars_select eval_select
#' @importFrom tidyselect vars_pull vars_select eval_select eval_rename

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we need to update the other usages of vars_select in the package too?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep - this is still on my to-do list, but I think all other feedback has been addressed.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done now

Comment threadr/R/util.R Outdated
Comment threadr/data-raw/docgen.R
# TODO(ARROW-17384): implement where
"Use of `where()` selection helper not yet supported"
)
docs[["dplyr::across"]] <- character(0)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎉

Comment threadr/R/dplyr-select.R Outdated
Comment threadr/R/util.R Outdated
}

simulate_data_frame <- function(schema) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Add a comment explaining why we need this function.

And do we need to export this? I think I saw some discussion between @paleolimbot and @krlmlr about needing this function in DBI or adbc.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In DBI, we need to infer the SQL data types from a RecordBatchReader, even if the DBI backend is not aware of Arrow (yet).

I wonder if there should be a way to convert a schema to a zero-row Arrow table. One way to do this is via schema -> zero-row table -> data frame. For DBI, schema -> data frame is sufficient, but perhaps schema -> zero-row table is easier to implemented in C++, and the table -> data frame operation is already efficient enough for zero-row tables.

Comment threadr/R/util.R Outdated
Comment threadr/R/util.R Outdated
abort(msg, call = call)
}

simulate_data_frame <- function(schema) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Since this is going to be called every time someone does select/rename/relocate, I'd like for this function to be cheaper. In other PRs I've been noticing the overhead of creating R6 objects, which generally is not terrible (~150 microseconds on my machine) but it adds up. And here, we're creating lots of objects we're throwing away: for each column, we create a Field, then a DataType from that, then in concat_arrays, we create a null DataType, an Array with that, and a new Array that is cast to the correct DataType. That adds up to around 1ms per column, every time this function is called. That's enough to get noticed.

Can we move this to C++? Should be a simple enough switch statement to map Arrow type ids to the corresponding R length-0 vector.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just to confirm, I benchmarked this function on the schema of the taxi dataset (20 columns), and the median time was 15ms, so a little under 1ms per column.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for confirming! Currently working on a C++ simplification though may ask for help to get it over the line ahead of the release if I can't figure it out by the end of tomorrow.

Comment threadr/src/table.cpp
Co-authored-by: Neal Richardson <neal.p.richardson@gmail.com>
@thisisnic

thisisnic commented Oct 13, 2022

Copy link
Copy Markdown
MemberAuthor

Currently getting this error message in the unit tests, due to weirdness with trying to create 0-length arrays using extension types 😬

Will try to create a nice reprex/solution/workaround, but if anyone has any suggestions in the meantime, let me know.

Error (test-extension.R:340:3): Dataset/arrow_dplyr_query can roundtrip extension types
Error: NotImplemented: MakeBuilder: cannot construct builder for type <arrow_custom_vctr[0]>
/home/nic2/arrow/cpp/src/arrow/builder.cc:289 VisitTypeInline(*type, &impl)
Backtrace:
1. ... %>% dplyr::collect()
at test-extension.R:340:2
4. arrow::select.Dataset(., number, letter, extension)
5. arrow::column_select(.data, enquos(...), op = "select")
at r/R/dplyr-select.R:24:2
6. arrow::simulate_data_frame(implicit_schema(.data))
at r/R/dplyr-select.R:97:2
8. arrow::Table__from_schema(schema)

It can be reproduced here via:

extension_vec <- vctrs::new_vctr(letters[1:10], class = "arrow_custom_vctr")
df <- tibble::tibble(x = extension_vec)
arrow_df <- arrow_table(df)
simulate_data_frame(arrow_df$schema)
Error: NotImplemented: MakeBuilder: cannot construct builder for type <arrow_custom_vctr[0]>
/home/nic2/arrow/cpp/src/arrow/builder.cc:289 VisitTypeInline(*type, &impl)

@nealrichardson

Copy link
Copy Markdown
Member

I'm sure it's solvable since we can roundtrip extension type data with >0 elements. Dewey may be best positioned to advise. Can you defer that to a followup, and find some workaround here, perhaps swap in null() type for extension types?

@thisisnic

Copy link
Copy Markdown
MemberAuthor

defer that to a followup

My favourite 5 words. Opened ARROW-18043

@thisisnic

Copy link
Copy Markdown
MemberAuthor

@nealrichardson Mind giving this another look over? I'm going to get to the "swap vars_select for eval_select" remaining bits tomorrow morning, but it'd be good to have you look at the rest of it in case there's any changes I need to make there too.

@nealrichardson

Copy link
Copy Markdown
Member

Looks great to me!

@thisisnic

Copy link
Copy Markdown
MemberAuthor

@nealrichardson Waiting on the CI, but otherwise I think this is ready for a final round of reviewing :D

@thisisnic
thisisnic merged commit 2cbf489 into apache:masterOct 14, 2022
@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = bd785c9 and contender = 2cbf489. 2cbf489 is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed ⬇️53.33% ⬆️0.0%] test-mac-arm
[Failed ⬇️28.22% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.39% ⬆️0.04%] ursa-thinkcentre-m75q
Buildkite builds:
[Finished] 2cbf4891 ec2-t3-xlarge-us-east-2
[Failed] 2cbf4891 test-mac-arm
[Failed] 2cbf4891 ursa-i9-9960x
[Finished] 2cbf4891 ursa-thinkcentre-m75q
[Finished] bd785c99 ec2-t3-xlarge-us-east-2
[Failed] bd785c99 test-mac-arm
[Failed] bd785c99 ursa-i9-9960x
[Finished] bd785c99 ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

@ursabot

Copy link
Copy Markdown

['Python', 'R'] benchmarks have high level of regressions.
test-mac-arm
ursa-i9-9960x

@thisisnic

Copy link
Copy Markdown
MemberAuthor

@nealrichardson I think I may need to refactor this in a follow-up PR - check out the regressions :\ Any suggestion for what I can do instead?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@thisisnic@nealrichardson@ursabot@krlmlr
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ARROW-12105: [R] Replace vars_select, vars_rename with eval_select, eval_rename - #14371

Merged
thisisnic merged 22 commits into
apache:masterfrom
thisisnic:ARROW-12105_eval_select
Oct 14, 2022
Merged

ARROW-12105: [R] Replace vars_select, vars_rename with eval_select, eval_rename#14371
thisisnic merged 22 commits into
apache:masterfrom
thisisnic:ARROW-12105_eval_select

Conversation

@thisisnic

Copy link
Copy Markdown
Member

No description provided.

)
})

test_that("multiple select/rename and group_by", {

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I added these in because they are needed to make sure the implementation of the utility function column_select() is working properly.

@nealrichardsonnealrichardson left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for taking this on!

Comment threadr/R/arrow-package.R Outdated
#' @importFrom rlang quo_set_env quo_get_env is_formula quo_is_call f_rhs parse_expr f_env new_quosure
#' @importFrom rlang new_quosures expr_text
#' @importFrom tidyselect vars_pull vars_rename vars_select eval_select
#' @importFrom tidyselect vars_pull vars_select eval_select eval_rename

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we need to update the other usages of vars_select in the package too?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep - this is still on my to-do list, but I think all other feedback has been addressed.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done now

Comment threadr/R/util.R Outdated
Comment threadr/data-raw/docgen.R
# TODO(ARROW-17384): implement where
"Use of `where()` selection helper not yet supported"
)
docs[["dplyr::across"]] <- character(0)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎉

Comment threadr/R/dplyr-select.R Outdated
Comment threadr/R/util.R Outdated
}

simulate_data_frame <- function(schema) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Add a comment explaining why we need this function.

And do we need to export this? I think I saw some discussion between @paleolimbot and @krlmlr about needing this function in DBI or adbc.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In DBI, we need to infer the SQL data types from a RecordBatchReader, even if the DBI backend is not aware of Arrow (yet).

I wonder if there should be a way to convert a schema to a zero-row Arrow table. One way to do this is via schema -> zero-row table -> data frame. For DBI, schema -> data frame is sufficient, but perhaps schema -> zero-row table is easier to implemented in C++, and the table -> data frame operation is already efficient enough for zero-row tables.

Comment threadr/R/util.R Outdated
Comment threadr/R/util.R Outdated
abort(msg, call = call)
}

simulate_data_frame <- function(schema) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Since this is going to be called every time someone does select/rename/relocate, I'd like for this function to be cheaper. In other PRs I've been noticing the overhead of creating R6 objects, which generally is not terrible (~150 microseconds on my machine) but it adds up. And here, we're creating lots of objects we're throwing away: for each column, we create a Field, then a DataType from that, then in concat_arrays, we create a null DataType, an Array with that, and a new Array that is cast to the correct DataType. That adds up to around 1ms per column, every time this function is called. That's enough to get noticed.

Can we move this to C++? Should be a simple enough switch statement to map Arrow type ids to the corresponding R length-0 vector.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just to confirm, I benchmarked this function on the schema of the taxi dataset (20 columns), and the median time was 15ms, so a little under 1ms per column.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for confirming! Currently working on a C++ simplification though may ask for help to get it over the line ahead of the release if I can't figure it out by the end of tomorrow.

Comment threadr/src/table.cpp
Co-authored-by: Neal Richardson <neal.p.richardson@gmail.com>
@thisisnic

thisisnic commented Oct 13, 2022

Copy link
Copy Markdown
MemberAuthor

Currently getting this error message in the unit tests, due to weirdness with trying to create 0-length arrays using extension types 😬

Will try to create a nice reprex/solution/workaround, but if anyone has any suggestions in the meantime, let me know.

Error (test-extension.R:340:3): Dataset/arrow_dplyr_query can roundtrip extension types
Error: NotImplemented: MakeBuilder: cannot construct builder for type <arrow_custom_vctr[0]>
/home/nic2/arrow/cpp/src/arrow/builder.cc:289 VisitTypeInline(*type, &impl)
Backtrace:
1. ... %>% dplyr::collect()
at test-extension.R:340:2
4. arrow::select.Dataset(., number, letter, extension)
5. arrow::column_select(.data, enquos(...), op = "select")
at r/R/dplyr-select.R:24:2
6. arrow::simulate_data_frame(implicit_schema(.data))
at r/R/dplyr-select.R:97:2
8. arrow::Table__from_schema(schema)

It can be reproduced here via:

extension_vec <- vctrs::new_vctr(letters[1:10], class = "arrow_custom_vctr")
df <- tibble::tibble(x = extension_vec)
arrow_df <- arrow_table(df)
simulate_data_frame(arrow_df$schema)
Error: NotImplemented: MakeBuilder: cannot construct builder for type <arrow_custom_vctr[0]>
/home/nic2/arrow/cpp/src/arrow/builder.cc:289 VisitTypeInline(*type, &impl)

@nealrichardson

Copy link
Copy Markdown
Member

I'm sure it's solvable since we can roundtrip extension type data with >0 elements. Dewey may be best positioned to advise. Can you defer that to a followup, and find some workaround here, perhaps swap in null() type for extension types?

@thisisnic

Copy link
Copy Markdown
MemberAuthor

defer that to a followup

My favourite 5 words. Opened ARROW-18043

@thisisnic

Copy link
Copy Markdown
MemberAuthor

@nealrichardson Mind giving this another look over? I'm going to get to the "swap vars_select for eval_select" remaining bits tomorrow morning, but it'd be good to have you look at the rest of it in case there's any changes I need to make there too.

@nealrichardson

Copy link
Copy Markdown
Member

Looks great to me!

@thisisnic

Copy link
Copy Markdown
MemberAuthor

@nealrichardson Waiting on the CI, but otherwise I think this is ready for a final round of reviewing :D

@thisisnic
thisisnic merged commit 2cbf489 into apache:masterOct 14, 2022
@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = bd785c9 and contender = 2cbf489. 2cbf489 is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed ⬇️53.33% ⬆️0.0%] test-mac-arm
[Failed ⬇️28.22% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.39% ⬆️0.04%] ursa-thinkcentre-m75q
Buildkite builds:
[Finished] 2cbf4891 ec2-t3-xlarge-us-east-2
[Failed] 2cbf4891 test-mac-arm
[Failed] 2cbf4891 ursa-i9-9960x
[Finished] 2cbf4891 ursa-thinkcentre-m75q
[Finished] bd785c99 ec2-t3-xlarge-us-east-2
[Failed] bd785c99 test-mac-arm
[Failed] bd785c99 ursa-i9-9960x
[Finished] bd785c99 ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

@ursabot

Copy link
Copy Markdown

['Python', 'R'] benchmarks have high level of regressions.
test-mac-arm
ursa-i9-9960x

@thisisnic

Copy link
Copy Markdown
MemberAuthor

@nealrichardson I think I may need to refactor this in a follow-up PR - check out the regressions :\ Any suggestion for what I can do instead?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@thisisnic@nealrichardson@ursabot@krlmlr
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ARROW-12105: [R] Replace vars_select, vars_rename with eval_select, eval_rename - #14371

Merged
thisisnic merged 22 commits into
apache:masterfrom
thisisnic:ARROW-12105_eval_select
Oct 14, 2022
Merged

ARROW-12105: [R] Replace vars_select, vars_rename with eval_select, eval_rename#14371
thisisnic merged 22 commits into
apache:masterfrom
thisisnic:ARROW-12105_eval_select

Conversation

@thisisnic

Copy link
Copy Markdown
Member

No description provided.

)
})

test_that("multiple select/rename and group_by", {

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I added these in because they are needed to make sure the implementation of the utility function column_select() is working properly.

@nealrichardsonnealrichardson left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for taking this on!

Comment threadr/R/arrow-package.R Outdated
#' @importFrom rlang quo_set_env quo_get_env is_formula quo_is_call f_rhs parse_expr f_env new_quosure
#' @importFrom rlang new_quosures expr_text
#' @importFrom tidyselect vars_pull vars_rename vars_select eval_select
#' @importFrom tidyselect vars_pull vars_select eval_select eval_rename

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we need to update the other usages of vars_select in the package too?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep - this is still on my to-do list, but I think all other feedback has been addressed.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done now

Comment threadr/R/util.R Outdated
Comment threadr/data-raw/docgen.R
# TODO(ARROW-17384): implement where
"Use of `where()` selection helper not yet supported"
)
docs[["dplyr::across"]] <- character(0)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎉

Comment threadr/R/dplyr-select.R Outdated
Comment threadr/R/util.R Outdated
}

simulate_data_frame <- function(schema) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Add a comment explaining why we need this function.

And do we need to export this? I think I saw some discussion between @paleolimbot and @krlmlr about needing this function in DBI or adbc.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In DBI, we need to infer the SQL data types from a RecordBatchReader, even if the DBI backend is not aware of Arrow (yet).

I wonder if there should be a way to convert a schema to a zero-row Arrow table. One way to do this is via schema -> zero-row table -> data frame. For DBI, schema -> data frame is sufficient, but perhaps schema -> zero-row table is easier to implemented in C++, and the table -> data frame operation is already efficient enough for zero-row tables.

Comment threadr/R/util.R Outdated
Comment threadr/R/util.R Outdated
abort(msg, call = call)
}

simulate_data_frame <- function(schema) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Since this is going to be called every time someone does select/rename/relocate, I'd like for this function to be cheaper. In other PRs I've been noticing the overhead of creating R6 objects, which generally is not terrible (~150 microseconds on my machine) but it adds up. And here, we're creating lots of objects we're throwing away: for each column, we create a Field, then a DataType from that, then in concat_arrays, we create a null DataType, an Array with that, and a new Array that is cast to the correct DataType. That adds up to around 1ms per column, every time this function is called. That's enough to get noticed.

Can we move this to C++? Should be a simple enough switch statement to map Arrow type ids to the corresponding R length-0 vector.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just to confirm, I benchmarked this function on the schema of the taxi dataset (20 columns), and the median time was 15ms, so a little under 1ms per column.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for confirming! Currently working on a C++ simplification though may ask for help to get it over the line ahead of the release if I can't figure it out by the end of tomorrow.

Comment threadr/src/table.cpp
Co-authored-by: Neal Richardson <neal.p.richardson@gmail.com>
@thisisnic

thisisnic commented Oct 13, 2022

Copy link
Copy Markdown
MemberAuthor

Currently getting this error message in the unit tests, due to weirdness with trying to create 0-length arrays using extension types 😬

Will try to create a nice reprex/solution/workaround, but if anyone has any suggestions in the meantime, let me know.

Error (test-extension.R:340:3): Dataset/arrow_dplyr_query can roundtrip extension types
Error: NotImplemented: MakeBuilder: cannot construct builder for type <arrow_custom_vctr[0]>
/home/nic2/arrow/cpp/src/arrow/builder.cc:289 VisitTypeInline(*type, &impl)
Backtrace:
1. ... %>% dplyr::collect()
at test-extension.R:340:2
4. arrow::select.Dataset(., number, letter, extension)
5. arrow::column_select(.data, enquos(...), op = "select")
at r/R/dplyr-select.R:24:2
6. arrow::simulate_data_frame(implicit_schema(.data))
at r/R/dplyr-select.R:97:2
8. arrow::Table__from_schema(schema)

It can be reproduced here via:

extension_vec <- vctrs::new_vctr(letters[1:10], class = "arrow_custom_vctr")
df <- tibble::tibble(x = extension_vec)
arrow_df <- arrow_table(df)
simulate_data_frame(arrow_df$schema)
Error: NotImplemented: MakeBuilder: cannot construct builder for type <arrow_custom_vctr[0]>
/home/nic2/arrow/cpp/src/arrow/builder.cc:289 VisitTypeInline(*type, &impl)

@nealrichardson

Copy link
Copy Markdown
Member

I'm sure it's solvable since we can roundtrip extension type data with >0 elements. Dewey may be best positioned to advise. Can you defer that to a followup, and find some workaround here, perhaps swap in null() type for extension types?

@thisisnic

Copy link
Copy Markdown
MemberAuthor

defer that to a followup

My favourite 5 words. Opened ARROW-18043

@thisisnic

Copy link
Copy Markdown
MemberAuthor

@nealrichardson Mind giving this another look over? I'm going to get to the "swap vars_select for eval_select" remaining bits tomorrow morning, but it'd be good to have you look at the rest of it in case there's any changes I need to make there too.

@nealrichardson

Copy link
Copy Markdown
Member

Looks great to me!

@thisisnic

Copy link
Copy Markdown
MemberAuthor

@nealrichardson Waiting on the CI, but otherwise I think this is ready for a final round of reviewing :D

@thisisnic
thisisnic merged commit 2cbf489 into apache:masterOct 14, 2022
@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = bd785c9 and contender = 2cbf489. 2cbf489 is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed ⬇️53.33% ⬆️0.0%] test-mac-arm
[Failed ⬇️28.22% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.39% ⬆️0.04%] ursa-thinkcentre-m75q
Buildkite builds:
[Finished] 2cbf4891 ec2-t3-xlarge-us-east-2
[Failed] 2cbf4891 test-mac-arm
[Failed] 2cbf4891 ursa-i9-9960x
[Finished] 2cbf4891 ursa-thinkcentre-m75q
[Finished] bd785c99 ec2-t3-xlarge-us-east-2
[Failed] bd785c99 test-mac-arm
[Failed] bd785c99 ursa-i9-9960x
[Finished] bd785c99 ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

@ursabot

Copy link
Copy Markdown

['Python', 'R'] benchmarks have high level of regressions.
test-mac-arm
ursa-i9-9960x

@thisisnic

Copy link
Copy Markdown
MemberAuthor

@nealrichardson I think I may need to refactor this in a follow-up PR - check out the regressions :\ Any suggestion for what I can do instead?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@thisisnic@nealrichardson@ursabot@krlmlr
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

ARROW-12105: [R] Replace vars_select, vars_rename with eval_select, eval_rename - #14371

Merged
thisisnic merged 22 commits into
apache:masterfrom
thisisnic:ARROW-12105_eval_select
Oct 14, 2022
Merged

ARROW-12105: [R] Replace vars_select, vars_rename with eval_select, eval_rename#14371
thisisnic merged 22 commits into
apache:masterfrom
thisisnic:ARROW-12105_eval_select

Conversation

@thisisnic

Copy link
Copy Markdown
Member

No description provided.

)
})

test_that("multiple select/rename and group_by", {

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I added these in because they are needed to make sure the implementation of the utility function column_select() is working properly.

@nealrichardsonnealrichardson left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for taking this on!

Comment threadr/R/arrow-package.R Outdated
#' @importFrom rlang quo_set_env quo_get_env is_formula quo_is_call f_rhs parse_expr f_env new_quosure
#' @importFrom rlang new_quosures expr_text
#' @importFrom tidyselect vars_pull vars_rename vars_select eval_select
#' @importFrom tidyselect vars_pull vars_select eval_select eval_rename

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we need to update the other usages of vars_select in the package too?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep - this is still on my to-do list, but I think all other feedback has been addressed.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done now

Comment threadr/R/util.R Outdated
Comment threadr/data-raw/docgen.R
# TODO(ARROW-17384): implement where
"Use of `where()` selection helper not yet supported"
)
docs[["dplyr::across"]] <- character(0)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎉

Comment threadr/R/dplyr-select.R Outdated
Comment threadr/R/util.R Outdated
}

simulate_data_frame <- function(schema) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Add a comment explaining why we need this function.

And do we need to export this? I think I saw some discussion between @paleolimbot and @krlmlr about needing this function in DBI or adbc.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In DBI, we need to infer the SQL data types from a RecordBatchReader, even if the DBI backend is not aware of Arrow (yet).

I wonder if there should be a way to convert a schema to a zero-row Arrow table. One way to do this is via schema -> zero-row table -> data frame. For DBI, schema -> data frame is sufficient, but perhaps schema -> zero-row table is easier to implemented in C++, and the table -> data frame operation is already efficient enough for zero-row tables.

Comment threadr/R/util.R Outdated
Comment threadr/R/util.R Outdated
abort(msg, call = call)
}

simulate_data_frame <- function(schema) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Since this is going to be called every time someone does select/rename/relocate, I'd like for this function to be cheaper. In other PRs I've been noticing the overhead of creating R6 objects, which generally is not terrible (~150 microseconds on my machine) but it adds up. And here, we're creating lots of objects we're throwing away: for each column, we create a Field, then a DataType from that, then in concat_arrays, we create a null DataType, an Array with that, and a new Array that is cast to the correct DataType. That adds up to around 1ms per column, every time this function is called. That's enough to get noticed.

Can we move this to C++? Should be a simple enough switch statement to map Arrow type ids to the corresponding R length-0 vector.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just to confirm, I benchmarked this function on the schema of the taxi dataset (20 columns), and the median time was 15ms, so a little under 1ms per column.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for confirming! Currently working on a C++ simplification though may ask for help to get it over the line ahead of the release if I can't figure it out by the end of tomorrow.

Comment threadr/src/table.cpp
Co-authored-by: Neal Richardson <neal.p.richardson@gmail.com>
@thisisnic

thisisnic commented Oct 13, 2022

Copy link
Copy Markdown
MemberAuthor

Currently getting this error message in the unit tests, due to weirdness with trying to create 0-length arrays using extension types 😬

Will try to create a nice reprex/solution/workaround, but if anyone has any suggestions in the meantime, let me know.

Error (test-extension.R:340:3): Dataset/arrow_dplyr_query can roundtrip extension types
Error: NotImplemented: MakeBuilder: cannot construct builder for type <arrow_custom_vctr[0]>
/home/nic2/arrow/cpp/src/arrow/builder.cc:289 VisitTypeInline(*type, &impl)
Backtrace:
1. ... %>% dplyr::collect()
at test-extension.R:340:2
4. arrow::select.Dataset(., number, letter, extension)
5. arrow::column_select(.data, enquos(...), op = "select")
at r/R/dplyr-select.R:24:2
6. arrow::simulate_data_frame(implicit_schema(.data))
at r/R/dplyr-select.R:97:2
8. arrow::Table__from_schema(schema)

It can be reproduced here via:

extension_vec <- vctrs::new_vctr(letters[1:10], class = "arrow_custom_vctr")
df <- tibble::tibble(x = extension_vec)
arrow_df <- arrow_table(df)
simulate_data_frame(arrow_df$schema)
Error: NotImplemented: MakeBuilder: cannot construct builder for type <arrow_custom_vctr[0]>
/home/nic2/arrow/cpp/src/arrow/builder.cc:289 VisitTypeInline(*type, &impl)

@nealrichardson

Copy link
Copy Markdown
Member

I'm sure it's solvable since we can roundtrip extension type data with >0 elements. Dewey may be best positioned to advise. Can you defer that to a followup, and find some workaround here, perhaps swap in null() type for extension types?

@thisisnic

Copy link
Copy Markdown
MemberAuthor

defer that to a followup

My favourite 5 words. Opened ARROW-18043

@thisisnic

Copy link
Copy Markdown
MemberAuthor

@nealrichardson Mind giving this another look over? I'm going to get to the "swap vars_select for eval_select" remaining bits tomorrow morning, but it'd be good to have you look at the rest of it in case there's any changes I need to make there too.

@nealrichardson

Copy link
Copy Markdown
Member

Looks great to me!

@thisisnic

Copy link
Copy Markdown
MemberAuthor

@nealrichardson Waiting on the CI, but otherwise I think this is ready for a final round of reviewing :D

@thisisnic
thisisnic merged commit 2cbf489 into apache:masterOct 14, 2022
@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = bd785c9 and contender = 2cbf489. 2cbf489 is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed ⬇️53.33% ⬆️0.0%] test-mac-arm
[Failed ⬇️28.22% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.39% ⬆️0.04%] ursa-thinkcentre-m75q
Buildkite builds:
[Finished] 2cbf4891 ec2-t3-xlarge-us-east-2
[Failed] 2cbf4891 test-mac-arm
[Failed] 2cbf4891 ursa-i9-9960x
[Finished] 2cbf4891 ursa-thinkcentre-m75q
[Finished] bd785c99 ec2-t3-xlarge-us-east-2
[Failed] bd785c99 test-mac-arm
[Failed] bd785c99 ursa-i9-9960x
[Finished] bd785c99 ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

@ursabot

Copy link
Copy Markdown

['Python', 'R'] benchmarks have high level of regressions.
test-mac-arm
ursa-i9-9960x

@thisisnic

Copy link
Copy Markdown
MemberAuthor

@nealrichardson I think I may need to refactor this in a follow-up PR - check out the regressions :\ Any suggestion for what I can do instead?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@thisisnic@nealrichardson@ursabot@krlmlr
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ARROW-12105: [R] Replace vars_select, vars_rename with eval_select, eval_rename - #14371

Merged
thisisnic merged 22 commits into
apache:masterfrom
thisisnic:ARROW-12105_eval_select
Oct 14, 2022
Merged

ARROW-12105: [R] Replace vars_select, vars_rename with eval_select, eval_rename#14371
thisisnic merged 22 commits into
apache:masterfrom
thisisnic:ARROW-12105_eval_select

Conversation

@thisisnic

Copy link
Copy Markdown
Member

No description provided.

)
})

test_that("multiple select/rename and group_by", {

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I added these in because they are needed to make sure the implementation of the utility function column_select() is working properly.

@nealrichardsonnealrichardson left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for taking this on!

Comment threadr/R/arrow-package.R Outdated
#' @importFrom rlang quo_set_env quo_get_env is_formula quo_is_call f_rhs parse_expr f_env new_quosure
#' @importFrom rlang new_quosures expr_text
#' @importFrom tidyselect vars_pull vars_rename vars_select eval_select
#' @importFrom tidyselect vars_pull vars_select eval_select eval_rename

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we need to update the other usages of vars_select in the package too?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep - this is still on my to-do list, but I think all other feedback has been addressed.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done now

Comment threadr/R/util.R Outdated
Comment threadr/data-raw/docgen.R
# TODO(ARROW-17384): implement where
"Use of `where()` selection helper not yet supported"
)
docs[["dplyr::across"]] <- character(0)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎉

Comment threadr/R/dplyr-select.R Outdated
Comment threadr/R/util.R Outdated
}

simulate_data_frame <- function(schema) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Add a comment explaining why we need this function.

And do we need to export this? I think I saw some discussion between @paleolimbot and @krlmlr about needing this function in DBI or adbc.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In DBI, we need to infer the SQL data types from a RecordBatchReader, even if the DBI backend is not aware of Arrow (yet).

I wonder if there should be a way to convert a schema to a zero-row Arrow table. One way to do this is via schema -> zero-row table -> data frame. For DBI, schema -> data frame is sufficient, but perhaps schema -> zero-row table is easier to implemented in C++, and the table -> data frame operation is already efficient enough for zero-row tables.

Comment threadr/R/util.R Outdated
Comment threadr/R/util.R Outdated
abort(msg, call = call)
}

simulate_data_frame <- function(schema) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Since this is going to be called every time someone does select/rename/relocate, I'd like for this function to be cheaper. In other PRs I've been noticing the overhead of creating R6 objects, which generally is not terrible (~150 microseconds on my machine) but it adds up. And here, we're creating lots of objects we're throwing away: for each column, we create a Field, then a DataType from that, then in concat_arrays, we create a null DataType, an Array with that, and a new Array that is cast to the correct DataType. That adds up to around 1ms per column, every time this function is called. That's enough to get noticed.

Can we move this to C++? Should be a simple enough switch statement to map Arrow type ids to the corresponding R length-0 vector.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just to confirm, I benchmarked this function on the schema of the taxi dataset (20 columns), and the median time was 15ms, so a little under 1ms per column.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for confirming! Currently working on a C++ simplification though may ask for help to get it over the line ahead of the release if I can't figure it out by the end of tomorrow.

Comment threadr/src/table.cpp
Co-authored-by: Neal Richardson <neal.p.richardson@gmail.com>
@thisisnic

thisisnic commented Oct 13, 2022

Copy link
Copy Markdown
MemberAuthor

Currently getting this error message in the unit tests, due to weirdness with trying to create 0-length arrays using extension types 😬

Will try to create a nice reprex/solution/workaround, but if anyone has any suggestions in the meantime, let me know.

Error (test-extension.R:340:3): Dataset/arrow_dplyr_query can roundtrip extension types
Error: NotImplemented: MakeBuilder: cannot construct builder for type <arrow_custom_vctr[0]>
/home/nic2/arrow/cpp/src/arrow/builder.cc:289 VisitTypeInline(*type, &impl)
Backtrace:
1. ... %>% dplyr::collect()
at test-extension.R:340:2
4. arrow::select.Dataset(., number, letter, extension)
5. arrow::column_select(.data, enquos(...), op = "select")
at r/R/dplyr-select.R:24:2
6. arrow::simulate_data_frame(implicit_schema(.data))
at r/R/dplyr-select.R:97:2
8. arrow::Table__from_schema(schema)

It can be reproduced here via:

extension_vec <- vctrs::new_vctr(letters[1:10], class = "arrow_custom_vctr")
df <- tibble::tibble(x = extension_vec)
arrow_df <- arrow_table(df)
simulate_data_frame(arrow_df$schema)
Error: NotImplemented: MakeBuilder: cannot construct builder for type <arrow_custom_vctr[0]>
/home/nic2/arrow/cpp/src/arrow/builder.cc:289 VisitTypeInline(*type, &impl)

@nealrichardson

Copy link
Copy Markdown
Member

I'm sure it's solvable since we can roundtrip extension type data with >0 elements. Dewey may be best positioned to advise. Can you defer that to a followup, and find some workaround here, perhaps swap in null() type for extension types?

@thisisnic

Copy link
Copy Markdown
MemberAuthor

defer that to a followup

My favourite 5 words. Opened ARROW-18043

@thisisnic

Copy link
Copy Markdown
MemberAuthor

@nealrichardson Mind giving this another look over? I'm going to get to the "swap vars_select for eval_select" remaining bits tomorrow morning, but it'd be good to have you look at the rest of it in case there's any changes I need to make there too.

@nealrichardson

Copy link
Copy Markdown
Member

Looks great to me!

@thisisnic

Copy link
Copy Markdown
MemberAuthor

@nealrichardson Waiting on the CI, but otherwise I think this is ready for a final round of reviewing :D

@thisisnic
thisisnic merged commit 2cbf489 into apache:masterOct 14, 2022
@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = bd785c9 and contender = 2cbf489. 2cbf489 is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed ⬇️53.33% ⬆️0.0%] test-mac-arm
[Failed ⬇️28.22% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.39% ⬆️0.04%] ursa-thinkcentre-m75q
Buildkite builds:
[Finished] 2cbf4891 ec2-t3-xlarge-us-east-2
[Failed] 2cbf4891 test-mac-arm
[Failed] 2cbf4891 ursa-i9-9960x
[Finished] 2cbf4891 ursa-thinkcentre-m75q
[Finished] bd785c99 ec2-t3-xlarge-us-east-2
[Failed] bd785c99 test-mac-arm
[Failed] bd785c99 ursa-i9-9960x
[Finished] bd785c99 ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

@ursabot

Copy link
Copy Markdown

['Python', 'R'] benchmarks have high level of regressions.
test-mac-arm
ursa-i9-9960x

@thisisnic

Copy link
Copy Markdown
MemberAuthor

@nealrichardson I think I may need to refactor this in a follow-up PR - check out the regressions :\ Any suggestion for what I can do instead?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@thisisnic@nealrichardson@ursabot@krlmlr
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ARROW-12105: [R] Replace vars_select, vars_rename with eval_select, eval_rename - #14371

Merged
thisisnic merged 22 commits into
apache:masterfrom
thisisnic:ARROW-12105_eval_select
Oct 14, 2022
Merged

ARROW-12105: [R] Replace vars_select, vars_rename with eval_select, eval_rename#14371
thisisnic merged 22 commits into
apache:masterfrom
thisisnic:ARROW-12105_eval_select

Conversation

@thisisnic

Copy link
Copy Markdown
Member

No description provided.

)
})

test_that("multiple select/rename and group_by", {

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I added these in because they are needed to make sure the implementation of the utility function column_select() is working properly.

@nealrichardsonnealrichardson left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for taking this on!

Comment threadr/R/arrow-package.R Outdated
#' @importFrom rlang quo_set_env quo_get_env is_formula quo_is_call f_rhs parse_expr f_env new_quosure
#' @importFrom rlang new_quosures expr_text
#' @importFrom tidyselect vars_pull vars_rename vars_select eval_select
#' @importFrom tidyselect vars_pull vars_select eval_select eval_rename

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we need to update the other usages of vars_select in the package too?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep - this is still on my to-do list, but I think all other feedback has been addressed.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done now

Comment threadr/R/util.R Outdated
Comment threadr/data-raw/docgen.R
# TODO(ARROW-17384): implement where
"Use of `where()` selection helper not yet supported"
)
docs[["dplyr::across"]] <- character(0)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎉

Comment threadr/R/dplyr-select.R Outdated
Comment threadr/R/util.R Outdated
}

simulate_data_frame <- function(schema) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Add a comment explaining why we need this function.

And do we need to export this? I think I saw some discussion between @paleolimbot and @krlmlr about needing this function in DBI or adbc.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In DBI, we need to infer the SQL data types from a RecordBatchReader, even if the DBI backend is not aware of Arrow (yet).

I wonder if there should be a way to convert a schema to a zero-row Arrow table. One way to do this is via schema -> zero-row table -> data frame. For DBI, schema -> data frame is sufficient, but perhaps schema -> zero-row table is easier to implemented in C++, and the table -> data frame operation is already efficient enough for zero-row tables.

Comment threadr/R/util.R Outdated
Comment threadr/R/util.R Outdated
abort(msg, call = call)
}

simulate_data_frame <- function(schema) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Since this is going to be called every time someone does select/rename/relocate, I'd like for this function to be cheaper. In other PRs I've been noticing the overhead of creating R6 objects, which generally is not terrible (~150 microseconds on my machine) but it adds up. And here, we're creating lots of objects we're throwing away: for each column, we create a Field, then a DataType from that, then in concat_arrays, we create a null DataType, an Array with that, and a new Array that is cast to the correct DataType. That adds up to around 1ms per column, every time this function is called. That's enough to get noticed.

Can we move this to C++? Should be a simple enough switch statement to map Arrow type ids to the corresponding R length-0 vector.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just to confirm, I benchmarked this function on the schema of the taxi dataset (20 columns), and the median time was 15ms, so a little under 1ms per column.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for confirming! Currently working on a C++ simplification though may ask for help to get it over the line ahead of the release if I can't figure it out by the end of tomorrow.

Comment threadr/src/table.cpp
Co-authored-by: Neal Richardson <neal.p.richardson@gmail.com>
@thisisnic

thisisnic commented Oct 13, 2022

Copy link
Copy Markdown
MemberAuthor

Currently getting this error message in the unit tests, due to weirdness with trying to create 0-length arrays using extension types 😬

Will try to create a nice reprex/solution/workaround, but if anyone has any suggestions in the meantime, let me know.

Error (test-extension.R:340:3): Dataset/arrow_dplyr_query can roundtrip extension types
Error: NotImplemented: MakeBuilder: cannot construct builder for type <arrow_custom_vctr[0]>
/home/nic2/arrow/cpp/src/arrow/builder.cc:289 VisitTypeInline(*type, &impl)
Backtrace:
1. ... %>% dplyr::collect()
at test-extension.R:340:2
4. arrow::select.Dataset(., number, letter, extension)
5. arrow::column_select(.data, enquos(...), op = "select")
at r/R/dplyr-select.R:24:2
6. arrow::simulate_data_frame(implicit_schema(.data))
at r/R/dplyr-select.R:97:2
8. arrow::Table__from_schema(schema)

It can be reproduced here via:

extension_vec <- vctrs::new_vctr(letters[1:10], class = "arrow_custom_vctr")
df <- tibble::tibble(x = extension_vec)
arrow_df <- arrow_table(df)
simulate_data_frame(arrow_df$schema)
Error: NotImplemented: MakeBuilder: cannot construct builder for type <arrow_custom_vctr[0]>
/home/nic2/arrow/cpp/src/arrow/builder.cc:289 VisitTypeInline(*type, &impl)

@nealrichardson

Copy link
Copy Markdown
Member

I'm sure it's solvable since we can roundtrip extension type data with >0 elements. Dewey may be best positioned to advise. Can you defer that to a followup, and find some workaround here, perhaps swap in null() type for extension types?

@thisisnic

Copy link
Copy Markdown
MemberAuthor

defer that to a followup

My favourite 5 words. Opened ARROW-18043

@thisisnic

Copy link
Copy Markdown
MemberAuthor

@nealrichardson Mind giving this another look over? I'm going to get to the "swap vars_select for eval_select" remaining bits tomorrow morning, but it'd be good to have you look at the rest of it in case there's any changes I need to make there too.

@nealrichardson

Copy link
Copy Markdown
Member

Looks great to me!

@thisisnic

Copy link
Copy Markdown
MemberAuthor

@nealrichardson Waiting on the CI, but otherwise I think this is ready for a final round of reviewing :D

@thisisnic
thisisnic merged commit 2cbf489 into apache:masterOct 14, 2022
@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = bd785c9 and contender = 2cbf489. 2cbf489 is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed ⬇️53.33% ⬆️0.0%] test-mac-arm
[Failed ⬇️28.22% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.39% ⬆️0.04%] ursa-thinkcentre-m75q
Buildkite builds:
[Finished] 2cbf4891 ec2-t3-xlarge-us-east-2
[Failed] 2cbf4891 test-mac-arm
[Failed] 2cbf4891 ursa-i9-9960x
[Finished] 2cbf4891 ursa-thinkcentre-m75q
[Finished] bd785c99 ec2-t3-xlarge-us-east-2
[Failed] bd785c99 test-mac-arm
[Failed] bd785c99 ursa-i9-9960x
[Finished] bd785c99 ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

@ursabot

Copy link
Copy Markdown

['Python', 'R'] benchmarks have high level of regressions.
test-mac-arm
ursa-i9-9960x

@thisisnic

Copy link
Copy Markdown
MemberAuthor

@nealrichardson I think I may need to refactor this in a follow-up PR - check out the regressions :\ Any suggestion for what I can do instead?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@thisisnic@nealrichardson@ursabot@krlmlr
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

ARROW-12105: [R] Replace vars_select, vars_rename with eval_select, eval_rename - #14371

Merged
thisisnic merged 22 commits into
apache:masterfrom
thisisnic:ARROW-12105_eval_select
Oct 14, 2022
Merged

ARROW-12105: [R] Replace vars_select, vars_rename with eval_select, eval_rename#14371
thisisnic merged 22 commits into
apache:masterfrom
thisisnic:ARROW-12105_eval_select

Conversation

@thisisnic

Copy link
Copy Markdown
Member

No description provided.

)
})

test_that("multiple select/rename and group_by", {

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I added these in because they are needed to make sure the implementation of the utility function column_select() is working properly.

@nealrichardsonnealrichardson left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for taking this on!

Comment threadr/R/arrow-package.R Outdated
#' @importFrom rlang quo_set_env quo_get_env is_formula quo_is_call f_rhs parse_expr f_env new_quosure
#' @importFrom rlang new_quosures expr_text
#' @importFrom tidyselect vars_pull vars_rename vars_select eval_select
#' @importFrom tidyselect vars_pull vars_select eval_select eval_rename

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we need to update the other usages of vars_select in the package too?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep - this is still on my to-do list, but I think all other feedback has been addressed.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done now

Comment threadr/R/util.R Outdated
Comment threadr/data-raw/docgen.R
# TODO(ARROW-17384): implement where
"Use of `where()` selection helper not yet supported"
)
docs[["dplyr::across"]] <- character(0)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎉

Comment threadr/R/dplyr-select.R Outdated
Comment threadr/R/util.R Outdated
}

simulate_data_frame <- function(schema) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Add a comment explaining why we need this function.

And do we need to export this? I think I saw some discussion between @paleolimbot and @krlmlr about needing this function in DBI or adbc.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In DBI, we need to infer the SQL data types from a RecordBatchReader, even if the DBI backend is not aware of Arrow (yet).

I wonder if there should be a way to convert a schema to a zero-row Arrow table. One way to do this is via schema -> zero-row table -> data frame. For DBI, schema -> data frame is sufficient, but perhaps schema -> zero-row table is easier to implemented in C++, and the table -> data frame operation is already efficient enough for zero-row tables.

Comment threadr/R/util.R Outdated
Comment threadr/R/util.R Outdated
abort(msg, call = call)
}

simulate_data_frame <- function(schema) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Since this is going to be called every time someone does select/rename/relocate, I'd like for this function to be cheaper. In other PRs I've been noticing the overhead of creating R6 objects, which generally is not terrible (~150 microseconds on my machine) but it adds up. And here, we're creating lots of objects we're throwing away: for each column, we create a Field, then a DataType from that, then in concat_arrays, we create a null DataType, an Array with that, and a new Array that is cast to the correct DataType. That adds up to around 1ms per column, every time this function is called. That's enough to get noticed.

Can we move this to C++? Should be a simple enough switch statement to map Arrow type ids to the corresponding R length-0 vector.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just to confirm, I benchmarked this function on the schema of the taxi dataset (20 columns), and the median time was 15ms, so a little under 1ms per column.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for confirming! Currently working on a C++ simplification though may ask for help to get it over the line ahead of the release if I can't figure it out by the end of tomorrow.

Comment threadr/src/table.cpp
Co-authored-by: Neal Richardson <neal.p.richardson@gmail.com>
@thisisnic

thisisnic commented Oct 13, 2022

Copy link
Copy Markdown
MemberAuthor

Currently getting this error message in the unit tests, due to weirdness with trying to create 0-length arrays using extension types 😬

Will try to create a nice reprex/solution/workaround, but if anyone has any suggestions in the meantime, let me know.

Error (test-extension.R:340:3): Dataset/arrow_dplyr_query can roundtrip extension types
Error: NotImplemented: MakeBuilder: cannot construct builder for type <arrow_custom_vctr[0]>
/home/nic2/arrow/cpp/src/arrow/builder.cc:289 VisitTypeInline(*type, &impl)
Backtrace:
1. ... %>% dplyr::collect()
at test-extension.R:340:2
4. arrow::select.Dataset(., number, letter, extension)
5. arrow::column_select(.data, enquos(...), op = "select")
at r/R/dplyr-select.R:24:2
6. arrow::simulate_data_frame(implicit_schema(.data))
at r/R/dplyr-select.R:97:2
8. arrow::Table__from_schema(schema)

It can be reproduced here via:

extension_vec <- vctrs::new_vctr(letters[1:10], class = "arrow_custom_vctr")
df <- tibble::tibble(x = extension_vec)
arrow_df <- arrow_table(df)
simulate_data_frame(arrow_df$schema)
Error: NotImplemented: MakeBuilder: cannot construct builder for type <arrow_custom_vctr[0]>
/home/nic2/arrow/cpp/src/arrow/builder.cc:289 VisitTypeInline(*type, &impl)

@nealrichardson

Copy link
Copy Markdown
Member

I'm sure it's solvable since we can roundtrip extension type data with >0 elements. Dewey may be best positioned to advise. Can you defer that to a followup, and find some workaround here, perhaps swap in null() type for extension types?

@thisisnic

Copy link
Copy Markdown
MemberAuthor

defer that to a followup

My favourite 5 words. Opened ARROW-18043

@thisisnic

Copy link
Copy Markdown
MemberAuthor

@nealrichardson Mind giving this another look over? I'm going to get to the "swap vars_select for eval_select" remaining bits tomorrow morning, but it'd be good to have you look at the rest of it in case there's any changes I need to make there too.

@nealrichardson

Copy link
Copy Markdown
Member

Looks great to me!

@thisisnic

Copy link
Copy Markdown
MemberAuthor

@nealrichardson Waiting on the CI, but otherwise I think this is ready for a final round of reviewing :D

@thisisnic
thisisnic merged commit 2cbf489 into apache:masterOct 14, 2022
@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = bd785c9 and contender = 2cbf489. 2cbf489 is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed ⬇️53.33% ⬆️0.0%] test-mac-arm
[Failed ⬇️28.22% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.39% ⬆️0.04%] ursa-thinkcentre-m75q
Buildkite builds:
[Finished] 2cbf4891 ec2-t3-xlarge-us-east-2
[Failed] 2cbf4891 test-mac-arm
[Failed] 2cbf4891 ursa-i9-9960x
[Finished] 2cbf4891 ursa-thinkcentre-m75q
[Finished] bd785c99 ec2-t3-xlarge-us-east-2
[Failed] bd785c99 test-mac-arm
[Failed] bd785c99 ursa-i9-9960x
[Finished] bd785c99 ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

@ursabot

Copy link
Copy Markdown

['Python', 'R'] benchmarks have high level of regressions.
test-mac-arm
ursa-i9-9960x

@thisisnic

Copy link
Copy Markdown
MemberAuthor

@nealrichardson I think I may need to refactor this in a follow-up PR - check out the regressions :\ Any suggestion for what I can do instead?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@thisisnic@nealrichardson@ursabot@krlmlr