Skip to content

ARROW-17252: [R] Intermittent valgrind failure - #13773

Merged
nealrichardson merged 8 commits into
apache:masterfrom
paleolimbot:r-valgrind-2
Aug 9, 2022
Merged

ARROW-17252: [R] Intermittent valgrind failure#13773
nealrichardson merged 8 commits into
apache:masterfrom
paleolimbot:r-valgrind-2

Conversation

@paleolimbot

@paleolimbotpaleolimbot commented Aug 2, 2022

Copy link
Copy Markdown
Member

This PR fixes intermittent leaks that occur after one of the changes from ARROW-16444: when we drain the RecordBatchReader that is emitted from the plan too quickly, it seems, some parts of the plan can leak (I don't know why this happens).

I tried removing various pieces of the RunWithCapturedR() changes (see #13746) but the only thing that removes the errors completely is draining the resulting RecordBatchReader from R (i.e., reader$read_table()) instead of in C++ (i.e., reader->ToTable()). Unfortunately, for user-defined functions to work in a plan we need a C++ level reader->ToTable(). I took the approach here of disabling the C++ level read by default, requiring a user to opt in to the version of collect() that works with a UDF. It's not ideal, but definitely safer (and more clearly marks the user-defined function behaviour as experimental).

I was able to replicate the original leaks but they are few and far between...our tests just happen to create and destroy many, many exec plans and something about the CI environment seems to trigger these more reliably (although the errors don't always occur there, either). Most of the leaks are small but there were some instances where an entire Table leaked.

@github-actions

Copy link
Copy Markdown

@github-actions

Copy link
Copy Markdown

⚠️ Ticket has not been started in JIRA, please click 'Start Progress'.

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

@github-actions crossbow submit test-r-linux-valgrind

@github-actions

Copy link
Copy Markdown

Revision: 3647aa3

Submitted crossbow builds: ursacomputing/crossbow @ actions-de212b2c40

TaskStatus
test-r-linux-valgrindAzure

@paleolimbot
paleolimbot marked this pull request as ready for review August 2, 2022 11:43
Comment threadr/R/compute.R
#' head()
#' Sys.unsetenv("R_ARROW_COLLECT_WITH_UDF")
#'
register_scalar_function <- function(name, fun, in_type, out_type,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What if you put Sys.setenv(R_ARROW_COLLECT_WITH_UDF = "true") inside of register_scalar_function()? You're already opting-in to UDFs by calling this function, and there's no reason you'd want to call this but not have COLLECT_WITH_UDF working.

Comment threadr/tests/testthat/test-compute.R Outdated
dplyr::mutate(fun_result = times_32(value)) %>%
dplyr::collect() %>%
dplyr::arrange(row_num)
withr::with_envvar(list(R_ARROW_COLLECT_WITH_UDF = "true"), {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You should be able to remove this now

Comment threadr/tests/testthat/test-compute.R Outdated
skip_if_not(CanRunWithCapturedR())
skip_if_not_available("dataset")
# Snappy has a UBSan issue: https://github.com/google/snappy/pull/148
# TODO(ARROW-17178): remove when user-defined function execution is stabilized

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note that the skip is for snappy ubsan too, so both need to be resolved before we can unskip

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

(See also ARROW-17283)

Comment threadr/tests/testthat/test-compute.R Outdated
dplyr::collect(),
tibble::tibble(a = 1L, b = 32.0)
)
withr::with_envvar(list(R_ARROW_COLLECT_WITH_UDF = "true"), {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can revert this

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we? To avoid the leaks for any subsequent tests we need the variable to not be "true".

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I was thinking you can revert this because you've already called register_scalar_function() so the env var will be set.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I see! with_envvar() was definitely not doing the right thing here (which is hopefully why the last valgrind check failed). I fixed it (I think) and requested another check!

Comment threadr/R/compute.R Outdated
@paleolimbot

Copy link
Copy Markdown
MemberAuthor

(See #13779 and #13780 for some other potential fixes, both have their checks pending)

paleolimbotand others added 2 commits August 2, 2022 14:40
Co-authored-by: Neal Richardson <neal.p.richardson@gmail.com>

@westonpacewestonpace left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

On the one hand, I can see why there might be some cases an extra collect is needed, to avoid valgrind errors. On the other hand:

  • The valgrind errors you are receiving are not the errors I would expect (I would expect direct leaks of scan-related objects).
  • I'm not entirely sure how this would be related to UDFs.

If this workaround fixes things for a little while then so be it. However, I'd really like to know exactly what is going on. You mentioned you had gotten better at reproducing. Could you share some kind of minimal(ish?) reproducer?

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

If this workaround fixes things for a little while then so be it.

I agree that this is a temporary workaround...I don't think anything all that insidious is going on that affects normal usage, but I also don't want us to get an angry CRAN note about valgrind that attracts attention to anything else we're doing.

I'd really like to know exactly what is going on.

Slight progress there: I tried waiting for the thread pools to finish before unloading the package (#13779) and that seems to remove the errors as well, although it's a bit of a hack in its own way (getting the IO thread pool is not exported in the public headers). That seems consistent with the shutting down of the thread pools leaking something?

Could you share some kind of minimal(ish?) reproducer?

From the arrow/r directory, echo 'devtools::test()' | R --no-save -d "valgrind --tool=memcheck --leak-check=full". Not very minimal, but an improvement over the 5 hours it takes the crossbow job. You can run fewer test files using devtools::test(filter = "some_regex") but I was never able to get an error using a few obvious filters ("dplyr", for example). I wonder if the CI just runs threads really really slowly which is why leaks show up more frequently then.

I'm not entirely sure how this would be related to UDFs.

The error isn't, I don't think, but was exposed by a change that we needed to make UDFs work (we need to evaluate the entire exec plan within one RunWithCapturedR(), which meant calling reader->ToTable() instead of sending the RecordBatchReader to R, then converting it to a table there). Before this PR, all exec plan results were routed through a C++-level reader->ToTable() unless an R-level RecordBatchReader was explicitly requested; after this PR, all exec plans go through an R-level RecordBatchReader unless somebody actually registers a user-defined function.

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

@github-actions crossbow submit test-r-linux-valgrind

@github-actions

Copy link
Copy Markdown

Revision: 3b71868

Submitted crossbow builds: ursacomputing/crossbow @ actions-511c82c07f

TaskStatus
test-r-linux-valgrindAzure

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

The last failure here is disappointing and leaves me stumped as to which change introduced the problem. I will spend some time tomorrow trying to reproduce this so that it's maybe possible to diagnose.

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

@github-actions crossbow submit test-r-linux-valgrind

@github-actions

Copy link
Copy Markdown

Revision: 2efade7

Submitted crossbow builds: ursacomputing/crossbow @ actions-d22fe5d8c4

TaskStatus
test-r-linux-valgrindAzure

Comment threadr/tests/testthat/test-compute.R Outdated
Sys.unsetenv("R_ARROW_COLLECT_WITH_UDF")
})

test_that("register_user_defined_function() can register multiple kernels", {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
test_that("register_user_defined_function() can register multiple kernels", {
test_that("register_scalar_function() can register multiple kernels", {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also below

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed!

Comment threadr/tests/testthat/test-compute.R Outdated
tibble::tibble(a = 1L, b = 32.0)
)

Sys.unsetenv("R_ARROW_COLLECT_WITH_UDF")

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

NBD but should this more properly go inside of on.exit above? like

on.exit({
unregister_binding("times_32", update_cache = TRUE)
# TODO(ARROW-17178) remove the need for this!
Sys.unsetenv("R_ARROW_COLLECT_WITH_UDF")
})

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed!

nealrichardson added a commit to nealrichardson/arrow that referenced this pull request Aug 7, 2022
@paleolimbot

Copy link
Copy Markdown
MemberAuthor

@github-actions crossbow submit test-r-linux-valgrind

@github-actions

Copy link
Copy Markdown

Revision: 37e6e2e

Submitted crossbow builds: ursacomputing/crossbow @ actions-d72df52d33

TaskStatus
test-r-linux-valgrindAzure

@nealrichardson

Copy link
Copy Markdown
Member

@paleolimbot is this good to merge now?

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

Yes! Sorry, forgot to check the valgrind result.

@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = 3148884 and contender = 7448322. 7448322 is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed ⬇️0.54% ⬆️0.0%] test-mac-arm
[Finished ⬇️1.09% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.18% ⬆️0.04%] ursa-thinkcentre-m75q
Buildkite builds:
[Finished] 7448322e ec2-t3-xlarge-us-east-2
[Finished] 7448322e test-mac-arm
[Finished] 7448322e ursa-i9-9960x
[Finished] 7448322e ursa-thinkcentre-m75q
[Finished] 3148884e ec2-t3-xlarge-us-east-2
[Failed] 3148884e test-mac-arm
[Finished] 3148884e ursa-i9-9960x
[Finished] 3148884e ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

@ursabot

Copy link
Copy Markdown

['Python', 'R'] benchmarks have high level of regressions.
ursa-i9-9960x

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@paleolimbot@nealrichardson@ursabot@westonpace
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
ARROW-17252: [R] Intermittent valgrind failure by paleolimbot · Pull Request #13773 · apache/arrow · GitHub
Skip to content

ARROW-17252: [R] Intermittent valgrind failure - #13773

Merged
nealrichardson merged 8 commits into
apache:masterfrom
paleolimbot:r-valgrind-2
Aug 9, 2022
Merged

ARROW-17252: [R] Intermittent valgrind failure#13773
nealrichardson merged 8 commits into
apache:masterfrom
paleolimbot:r-valgrind-2

Conversation

@paleolimbot

@paleolimbotpaleolimbot commented Aug 2, 2022

Copy link
Copy Markdown
Member

This PR fixes intermittent leaks that occur after one of the changes from ARROW-16444: when we drain the RecordBatchReader that is emitted from the plan too quickly, it seems, some parts of the plan can leak (I don't know why this happens).

I tried removing various pieces of the RunWithCapturedR() changes (see #13746) but the only thing that removes the errors completely is draining the resulting RecordBatchReader from R (i.e., reader$read_table()) instead of in C++ (i.e., reader->ToTable()). Unfortunately, for user-defined functions to work in a plan we need a C++ level reader->ToTable(). I took the approach here of disabling the C++ level read by default, requiring a user to opt in to the version of collect() that works with a UDF. It's not ideal, but definitely safer (and more clearly marks the user-defined function behaviour as experimental).

I was able to replicate the original leaks but they are few and far between...our tests just happen to create and destroy many, many exec plans and something about the CI environment seems to trigger these more reliably (although the errors don't always occur there, either). Most of the leaks are small but there were some instances where an entire Table leaked.

@github-actions

Copy link
Copy Markdown

@github-actions

Copy link
Copy Markdown

⚠️ Ticket has not been started in JIRA, please click 'Start Progress'.

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

@github-actions crossbow submit test-r-linux-valgrind

@github-actions

Copy link
Copy Markdown

Revision: 3647aa3

Submitted crossbow builds: ursacomputing/crossbow @ actions-de212b2c40

TaskStatus
test-r-linux-valgrindAzure

@paleolimbot
paleolimbot marked this pull request as ready for review August 2, 2022 11:43
Comment threadr/R/compute.R
#' head()
#' Sys.unsetenv("R_ARROW_COLLECT_WITH_UDF")
#'
register_scalar_function <- function(name, fun, in_type, out_type,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What if you put Sys.setenv(R_ARROW_COLLECT_WITH_UDF = "true") inside of register_scalar_function()? You're already opting-in to UDFs by calling this function, and there's no reason you'd want to call this but not have COLLECT_WITH_UDF working.

Comment threadr/tests/testthat/test-compute.R Outdated
dplyr::mutate(fun_result = times_32(value)) %>%
dplyr::collect() %>%
dplyr::arrange(row_num)
withr::with_envvar(list(R_ARROW_COLLECT_WITH_UDF = "true"), {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You should be able to remove this now

Comment threadr/tests/testthat/test-compute.R Outdated
skip_if_not(CanRunWithCapturedR())
skip_if_not_available("dataset")
# Snappy has a UBSan issue: https://github.com/google/snappy/pull/148
# TODO(ARROW-17178): remove when user-defined function execution is stabilized

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note that the skip is for snappy ubsan too, so both need to be resolved before we can unskip

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

(See also ARROW-17283)

Comment threadr/tests/testthat/test-compute.R Outdated
dplyr::collect(),
tibble::tibble(a = 1L, b = 32.0)
)
withr::with_envvar(list(R_ARROW_COLLECT_WITH_UDF = "true"), {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can revert this

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we? To avoid the leaks for any subsequent tests we need the variable to not be "true".

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I was thinking you can revert this because you've already called register_scalar_function() so the env var will be set.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I see! with_envvar() was definitely not doing the right thing here (which is hopefully why the last valgrind check failed). I fixed it (I think) and requested another check!

Comment threadr/R/compute.R Outdated
@paleolimbot

Copy link
Copy Markdown
MemberAuthor

(See #13779 and #13780 for some other potential fixes, both have their checks pending)

paleolimbotand others added 2 commits August 2, 2022 14:40
Co-authored-by: Neal Richardson <neal.p.richardson@gmail.com>

@westonpacewestonpace left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

On the one hand, I can see why there might be some cases an extra collect is needed, to avoid valgrind errors. On the other hand:

  • The valgrind errors you are receiving are not the errors I would expect (I would expect direct leaks of scan-related objects).
  • I'm not entirely sure how this would be related to UDFs.

If this workaround fixes things for a little while then so be it. However, I'd really like to know exactly what is going on. You mentioned you had gotten better at reproducing. Could you share some kind of minimal(ish?) reproducer?

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

If this workaround fixes things for a little while then so be it.

I agree that this is a temporary workaround...I don't think anything all that insidious is going on that affects normal usage, but I also don't want us to get an angry CRAN note about valgrind that attracts attention to anything else we're doing.

I'd really like to know exactly what is going on.

Slight progress there: I tried waiting for the thread pools to finish before unloading the package (#13779) and that seems to remove the errors as well, although it's a bit of a hack in its own way (getting the IO thread pool is not exported in the public headers). That seems consistent with the shutting down of the thread pools leaking something?

Could you share some kind of minimal(ish?) reproducer?

From the arrow/r directory, echo 'devtools::test()' | R --no-save -d "valgrind --tool=memcheck --leak-check=full". Not very minimal, but an improvement over the 5 hours it takes the crossbow job. You can run fewer test files using devtools::test(filter = "some_regex") but I was never able to get an error using a few obvious filters ("dplyr", for example). I wonder if the CI just runs threads really really slowly which is why leaks show up more frequently then.

I'm not entirely sure how this would be related to UDFs.

The error isn't, I don't think, but was exposed by a change that we needed to make UDFs work (we need to evaluate the entire exec plan within one RunWithCapturedR(), which meant calling reader->ToTable() instead of sending the RecordBatchReader to R, then converting it to a table there). Before this PR, all exec plan results were routed through a C++-level reader->ToTable() unless an R-level RecordBatchReader was explicitly requested; after this PR, all exec plans go through an R-level RecordBatchReader unless somebody actually registers a user-defined function.

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

@github-actions crossbow submit test-r-linux-valgrind

@github-actions

Copy link
Copy Markdown

Revision: 3b71868

Submitted crossbow builds: ursacomputing/crossbow @ actions-511c82c07f

TaskStatus
test-r-linux-valgrindAzure

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

The last failure here is disappointing and leaves me stumped as to which change introduced the problem. I will spend some time tomorrow trying to reproduce this so that it's maybe possible to diagnose.

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

@github-actions crossbow submit test-r-linux-valgrind

@github-actions

Copy link
Copy Markdown

Revision: 2efade7

Submitted crossbow builds: ursacomputing/crossbow @ actions-d22fe5d8c4

TaskStatus
test-r-linux-valgrindAzure

Comment threadr/tests/testthat/test-compute.R Outdated
Sys.unsetenv("R_ARROW_COLLECT_WITH_UDF")
})

test_that("register_user_defined_function() can register multiple kernels", {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
test_that("register_user_defined_function() can register multiple kernels", {
test_that("register_scalar_function() can register multiple kernels", {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also below

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed!

Comment threadr/tests/testthat/test-compute.R Outdated
tibble::tibble(a = 1L, b = 32.0)
)

Sys.unsetenv("R_ARROW_COLLECT_WITH_UDF")

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

NBD but should this more properly go inside of on.exit above? like

on.exit({
unregister_binding("times_32", update_cache = TRUE)
# TODO(ARROW-17178) remove the need for this!
Sys.unsetenv("R_ARROW_COLLECT_WITH_UDF")
})

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed!

nealrichardson added a commit to nealrichardson/arrow that referenced this pull request Aug 7, 2022
@paleolimbot

Copy link
Copy Markdown
MemberAuthor

@github-actions crossbow submit test-r-linux-valgrind

@github-actions

Copy link
Copy Markdown

Revision: 37e6e2e

Submitted crossbow builds: ursacomputing/crossbow @ actions-d72df52d33

TaskStatus
test-r-linux-valgrindAzure

@nealrichardson

Copy link
Copy Markdown
Member

@paleolimbot is this good to merge now?

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

Yes! Sorry, forgot to check the valgrind result.

@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = 3148884 and contender = 7448322. 7448322 is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed ⬇️0.54% ⬆️0.0%] test-mac-arm
[Finished ⬇️1.09% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.18% ⬆️0.04%] ursa-thinkcentre-m75q
Buildkite builds:
[Finished] 7448322e ec2-t3-xlarge-us-east-2
[Finished] 7448322e test-mac-arm
[Finished] 7448322e ursa-i9-9960x
[Finished] 7448322e ursa-thinkcentre-m75q
[Finished] 3148884e ec2-t3-xlarge-us-east-2
[Failed] 3148884e test-mac-arm
[Finished] 3148884e ursa-i9-9960x
[Finished] 3148884e ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

@ursabot

Copy link
Copy Markdown

['Python', 'R'] benchmarks have high level of regressions.
ursa-i9-9960x

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@paleolimbot@nealrichardson@ursabot@westonpace
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' ARROW-17252: [R] Intermittent valgrind failure by paleolimbot · Pull Request #13773 · apache/arrow · GitHub
Skip to content

ARROW-17252: [R] Intermittent valgrind failure - #13773

Merged
nealrichardson merged 8 commits into
apache:masterfrom
paleolimbot:r-valgrind-2
Aug 9, 2022
Merged

ARROW-17252: [R] Intermittent valgrind failure#13773
nealrichardson merged 8 commits into
apache:masterfrom
paleolimbot:r-valgrind-2

Conversation

@paleolimbot

@paleolimbotpaleolimbot commented Aug 2, 2022

Copy link
Copy Markdown
Member

This PR fixes intermittent leaks that occur after one of the changes from ARROW-16444: when we drain the RecordBatchReader that is emitted from the plan too quickly, it seems, some parts of the plan can leak (I don't know why this happens).

I tried removing various pieces of the RunWithCapturedR() changes (see #13746) but the only thing that removes the errors completely is draining the resulting RecordBatchReader from R (i.e., reader$read_table()) instead of in C++ (i.e., reader->ToTable()). Unfortunately, for user-defined functions to work in a plan we need a C++ level reader->ToTable(). I took the approach here of disabling the C++ level read by default, requiring a user to opt in to the version of collect() that works with a UDF. It's not ideal, but definitely safer (and more clearly marks the user-defined function behaviour as experimental).

I was able to replicate the original leaks but they are few and far between...our tests just happen to create and destroy many, many exec plans and something about the CI environment seems to trigger these more reliably (although the errors don't always occur there, either). Most of the leaks are small but there were some instances where an entire Table leaked.

@github-actions

Copy link
Copy Markdown

@github-actions

Copy link
Copy Markdown

⚠️ Ticket has not been started in JIRA, please click 'Start Progress'.

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

@github-actions crossbow submit test-r-linux-valgrind

@github-actions

Copy link
Copy Markdown

Revision: 3647aa3

Submitted crossbow builds: ursacomputing/crossbow @ actions-de212b2c40

TaskStatus
test-r-linux-valgrindAzure

@paleolimbot
paleolimbot marked this pull request as ready for review August 2, 2022 11:43
Comment threadr/R/compute.R
#' head()
#' Sys.unsetenv("R_ARROW_COLLECT_WITH_UDF")
#'
register_scalar_function <- function(name, fun, in_type, out_type,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What if you put Sys.setenv(R_ARROW_COLLECT_WITH_UDF = "true") inside of register_scalar_function()? You're already opting-in to UDFs by calling this function, and there's no reason you'd want to call this but not have COLLECT_WITH_UDF working.

Comment threadr/tests/testthat/test-compute.R Outdated
dplyr::mutate(fun_result = times_32(value)) %>%
dplyr::collect() %>%
dplyr::arrange(row_num)
withr::with_envvar(list(R_ARROW_COLLECT_WITH_UDF = "true"), {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You should be able to remove this now

Comment threadr/tests/testthat/test-compute.R Outdated
skip_if_not(CanRunWithCapturedR())
skip_if_not_available("dataset")
# Snappy has a UBSan issue: https://github.com/google/snappy/pull/148
# TODO(ARROW-17178): remove when user-defined function execution is stabilized

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note that the skip is for snappy ubsan too, so both need to be resolved before we can unskip

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

(See also ARROW-17283)

Comment threadr/tests/testthat/test-compute.R Outdated
dplyr::collect(),
tibble::tibble(a = 1L, b = 32.0)
)
withr::with_envvar(list(R_ARROW_COLLECT_WITH_UDF = "true"), {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can revert this

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we? To avoid the leaks for any subsequent tests we need the variable to not be "true".

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I was thinking you can revert this because you've already called register_scalar_function() so the env var will be set.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I see! with_envvar() was definitely not doing the right thing here (which is hopefully why the last valgrind check failed). I fixed it (I think) and requested another check!

Comment threadr/R/compute.R Outdated
@paleolimbot

Copy link
Copy Markdown
MemberAuthor

(See #13779 and #13780 for some other potential fixes, both have their checks pending)

paleolimbotand others added 2 commits August 2, 2022 14:40
Co-authored-by: Neal Richardson <neal.p.richardson@gmail.com>

@westonpacewestonpace left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

On the one hand, I can see why there might be some cases an extra collect is needed, to avoid valgrind errors. On the other hand:

  • The valgrind errors you are receiving are not the errors I would expect (I would expect direct leaks of scan-related objects).
  • I'm not entirely sure how this would be related to UDFs.

If this workaround fixes things for a little while then so be it. However, I'd really like to know exactly what is going on. You mentioned you had gotten better at reproducing. Could you share some kind of minimal(ish?) reproducer?

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

If this workaround fixes things for a little while then so be it.

I agree that this is a temporary workaround...I don't think anything all that insidious is going on that affects normal usage, but I also don't want us to get an angry CRAN note about valgrind that attracts attention to anything else we're doing.

I'd really like to know exactly what is going on.

Slight progress there: I tried waiting for the thread pools to finish before unloading the package (#13779) and that seems to remove the errors as well, although it's a bit of a hack in its own way (getting the IO thread pool is not exported in the public headers). That seems consistent with the shutting down of the thread pools leaking something?

Could you share some kind of minimal(ish?) reproducer?

From the arrow/r directory, echo 'devtools::test()' | R --no-save -d "valgrind --tool=memcheck --leak-check=full". Not very minimal, but an improvement over the 5 hours it takes the crossbow job. You can run fewer test files using devtools::test(filter = "some_regex") but I was never able to get an error using a few obvious filters ("dplyr", for example). I wonder if the CI just runs threads really really slowly which is why leaks show up more frequently then.

I'm not entirely sure how this would be related to UDFs.

The error isn't, I don't think, but was exposed by a change that we needed to make UDFs work (we need to evaluate the entire exec plan within one RunWithCapturedR(), which meant calling reader->ToTable() instead of sending the RecordBatchReader to R, then converting it to a table there). Before this PR, all exec plan results were routed through a C++-level reader->ToTable() unless an R-level RecordBatchReader was explicitly requested; after this PR, all exec plans go through an R-level RecordBatchReader unless somebody actually registers a user-defined function.

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

@github-actions crossbow submit test-r-linux-valgrind

@github-actions

Copy link
Copy Markdown

Revision: 3b71868

Submitted crossbow builds: ursacomputing/crossbow @ actions-511c82c07f

TaskStatus
test-r-linux-valgrindAzure

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

The last failure here is disappointing and leaves me stumped as to which change introduced the problem. I will spend some time tomorrow trying to reproduce this so that it's maybe possible to diagnose.

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

@github-actions crossbow submit test-r-linux-valgrind

@github-actions

Copy link
Copy Markdown

Revision: 2efade7

Submitted crossbow builds: ursacomputing/crossbow @ actions-d22fe5d8c4

TaskStatus
test-r-linux-valgrindAzure

Comment threadr/tests/testthat/test-compute.R Outdated
Sys.unsetenv("R_ARROW_COLLECT_WITH_UDF")
})

test_that("register_user_defined_function() can register multiple kernels", {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
test_that("register_user_defined_function() can register multiple kernels", {
test_that("register_scalar_function() can register multiple kernels", {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also below

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed!

Comment threadr/tests/testthat/test-compute.R Outdated
tibble::tibble(a = 1L, b = 32.0)
)

Sys.unsetenv("R_ARROW_COLLECT_WITH_UDF")

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

NBD but should this more properly go inside of on.exit above? like

on.exit({
unregister_binding("times_32", update_cache = TRUE)
# TODO(ARROW-17178) remove the need for this!
Sys.unsetenv("R_ARROW_COLLECT_WITH_UDF")
})

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed!

nealrichardson added a commit to nealrichardson/arrow that referenced this pull request Aug 7, 2022
@paleolimbot

Copy link
Copy Markdown
MemberAuthor

@github-actions crossbow submit test-r-linux-valgrind

@github-actions

Copy link
Copy Markdown

Revision: 37e6e2e

Submitted crossbow builds: ursacomputing/crossbow @ actions-d72df52d33

TaskStatus
test-r-linux-valgrindAzure

@nealrichardson

Copy link
Copy Markdown
Member

@paleolimbot is this good to merge now?

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

Yes! Sorry, forgot to check the valgrind result.

@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = 3148884 and contender = 7448322. 7448322 is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed ⬇️0.54% ⬆️0.0%] test-mac-arm
[Finished ⬇️1.09% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.18% ⬆️0.04%] ursa-thinkcentre-m75q
Buildkite builds:
[Finished] 7448322e ec2-t3-xlarge-us-east-2
[Finished] 7448322e test-mac-arm
[Finished] 7448322e ursa-i9-9960x
[Finished] 7448322e ursa-thinkcentre-m75q
[Finished] 3148884e ec2-t3-xlarge-us-east-2
[Failed] 3148884e test-mac-arm
[Finished] 3148884e ursa-i9-9960x
[Finished] 3148884e ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

@ursabot

Copy link
Copy Markdown

['Python', 'R'] benchmarks have high level of regressions.
ursa-i9-9960x

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@paleolimbot@nealrichardson@ursabot@westonpace
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' ARROW-17252: [R] Intermittent valgrind failure by paleolimbot · Pull Request #13773 · apache/arrow · GitHub
Skip to content

ARROW-17252: [R] Intermittent valgrind failure - #13773

Merged
nealrichardson merged 8 commits into
apache:masterfrom
paleolimbot:r-valgrind-2
Aug 9, 2022
Merged

ARROW-17252: [R] Intermittent valgrind failure#13773
nealrichardson merged 8 commits into
apache:masterfrom
paleolimbot:r-valgrind-2

Conversation

@paleolimbot

@paleolimbotpaleolimbot commented Aug 2, 2022

Copy link
Copy Markdown
Member

This PR fixes intermittent leaks that occur after one of the changes from ARROW-16444: when we drain the RecordBatchReader that is emitted from the plan too quickly, it seems, some parts of the plan can leak (I don't know why this happens).

I tried removing various pieces of the RunWithCapturedR() changes (see #13746) but the only thing that removes the errors completely is draining the resulting RecordBatchReader from R (i.e., reader$read_table()) instead of in C++ (i.e., reader->ToTable()). Unfortunately, for user-defined functions to work in a plan we need a C++ level reader->ToTable(). I took the approach here of disabling the C++ level read by default, requiring a user to opt in to the version of collect() that works with a UDF. It's not ideal, but definitely safer (and more clearly marks the user-defined function behaviour as experimental).

I was able to replicate the original leaks but they are few and far between...our tests just happen to create and destroy many, many exec plans and something about the CI environment seems to trigger these more reliably (although the errors don't always occur there, either). Most of the leaks are small but there were some instances where an entire Table leaked.

@github-actions

Copy link
Copy Markdown

@github-actions

Copy link
Copy Markdown

⚠️ Ticket has not been started in JIRA, please click 'Start Progress'.

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

@github-actions crossbow submit test-r-linux-valgrind

@github-actions

Copy link
Copy Markdown

Revision: 3647aa3

Submitted crossbow builds: ursacomputing/crossbow @ actions-de212b2c40

TaskStatus
test-r-linux-valgrindAzure

@paleolimbot
paleolimbot marked this pull request as ready for review August 2, 2022 11:43
Comment threadr/R/compute.R
#' head()
#' Sys.unsetenv("R_ARROW_COLLECT_WITH_UDF")
#'
register_scalar_function <- function(name, fun, in_type, out_type,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What if you put Sys.setenv(R_ARROW_COLLECT_WITH_UDF = "true") inside of register_scalar_function()? You're already opting-in to UDFs by calling this function, and there's no reason you'd want to call this but not have COLLECT_WITH_UDF working.

Comment threadr/tests/testthat/test-compute.R Outdated
dplyr::mutate(fun_result = times_32(value)) %>%
dplyr::collect() %>%
dplyr::arrange(row_num)
withr::with_envvar(list(R_ARROW_COLLECT_WITH_UDF = "true"), {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You should be able to remove this now

Comment threadr/tests/testthat/test-compute.R Outdated
skip_if_not(CanRunWithCapturedR())
skip_if_not_available("dataset")
# Snappy has a UBSan issue: https://github.com/google/snappy/pull/148
# TODO(ARROW-17178): remove when user-defined function execution is stabilized

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note that the skip is for snappy ubsan too, so both need to be resolved before we can unskip

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

(See also ARROW-17283)

Comment threadr/tests/testthat/test-compute.R Outdated
dplyr::collect(),
tibble::tibble(a = 1L, b = 32.0)
)
withr::with_envvar(list(R_ARROW_COLLECT_WITH_UDF = "true"), {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can revert this

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we? To avoid the leaks for any subsequent tests we need the variable to not be "true".

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I was thinking you can revert this because you've already called register_scalar_function() so the env var will be set.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I see! with_envvar() was definitely not doing the right thing here (which is hopefully why the last valgrind check failed). I fixed it (I think) and requested another check!

Comment threadr/R/compute.R Outdated
@paleolimbot

Copy link
Copy Markdown
MemberAuthor

(See #13779 and #13780 for some other potential fixes, both have their checks pending)

paleolimbotand others added 2 commits August 2, 2022 14:40
Co-authored-by: Neal Richardson <neal.p.richardson@gmail.com>

@westonpacewestonpace left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

On the one hand, I can see why there might be some cases an extra collect is needed, to avoid valgrind errors. On the other hand:

  • The valgrind errors you are receiving are not the errors I would expect (I would expect direct leaks of scan-related objects).
  • I'm not entirely sure how this would be related to UDFs.

If this workaround fixes things for a little while then so be it. However, I'd really like to know exactly what is going on. You mentioned you had gotten better at reproducing. Could you share some kind of minimal(ish?) reproducer?

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

If this workaround fixes things for a little while then so be it.

I agree that this is a temporary workaround...I don't think anything all that insidious is going on that affects normal usage, but I also don't want us to get an angry CRAN note about valgrind that attracts attention to anything else we're doing.

I'd really like to know exactly what is going on.

Slight progress there: I tried waiting for the thread pools to finish before unloading the package (#13779) and that seems to remove the errors as well, although it's a bit of a hack in its own way (getting the IO thread pool is not exported in the public headers). That seems consistent with the shutting down of the thread pools leaking something?

Could you share some kind of minimal(ish?) reproducer?

From the arrow/r directory, echo 'devtools::test()' | R --no-save -d "valgrind --tool=memcheck --leak-check=full". Not very minimal, but an improvement over the 5 hours it takes the crossbow job. You can run fewer test files using devtools::test(filter = "some_regex") but I was never able to get an error using a few obvious filters ("dplyr", for example). I wonder if the CI just runs threads really really slowly which is why leaks show up more frequently then.

I'm not entirely sure how this would be related to UDFs.

The error isn't, I don't think, but was exposed by a change that we needed to make UDFs work (we need to evaluate the entire exec plan within one RunWithCapturedR(), which meant calling reader->ToTable() instead of sending the RecordBatchReader to R, then converting it to a table there). Before this PR, all exec plan results were routed through a C++-level reader->ToTable() unless an R-level RecordBatchReader was explicitly requested; after this PR, all exec plans go through an R-level RecordBatchReader unless somebody actually registers a user-defined function.

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

@github-actions crossbow submit test-r-linux-valgrind

@github-actions

Copy link
Copy Markdown

Revision: 3b71868

Submitted crossbow builds: ursacomputing/crossbow @ actions-511c82c07f

TaskStatus
test-r-linux-valgrindAzure

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

The last failure here is disappointing and leaves me stumped as to which change introduced the problem. I will spend some time tomorrow trying to reproduce this so that it's maybe possible to diagnose.

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

@github-actions crossbow submit test-r-linux-valgrind

@github-actions

Copy link
Copy Markdown

Revision: 2efade7

Submitted crossbow builds: ursacomputing/crossbow @ actions-d22fe5d8c4

TaskStatus
test-r-linux-valgrindAzure

Comment threadr/tests/testthat/test-compute.R Outdated
Sys.unsetenv("R_ARROW_COLLECT_WITH_UDF")
})

test_that("register_user_defined_function() can register multiple kernels", {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
test_that("register_user_defined_function() can register multiple kernels", {
test_that("register_scalar_function() can register multiple kernels", {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also below

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed!

Comment threadr/tests/testthat/test-compute.R Outdated
tibble::tibble(a = 1L, b = 32.0)
)

Sys.unsetenv("R_ARROW_COLLECT_WITH_UDF")

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

NBD but should this more properly go inside of on.exit above? like

on.exit({
unregister_binding("times_32", update_cache = TRUE)
# TODO(ARROW-17178) remove the need for this!
Sys.unsetenv("R_ARROW_COLLECT_WITH_UDF")
})

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed!

nealrichardson added a commit to nealrichardson/arrow that referenced this pull request Aug 7, 2022
@paleolimbot

Copy link
Copy Markdown
MemberAuthor

@github-actions crossbow submit test-r-linux-valgrind

@github-actions

Copy link
Copy Markdown

Revision: 37e6e2e

Submitted crossbow builds: ursacomputing/crossbow @ actions-d72df52d33

TaskStatus
test-r-linux-valgrindAzure

@nealrichardson

Copy link
Copy Markdown
Member

@paleolimbot is this good to merge now?

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

Yes! Sorry, forgot to check the valgrind result.

@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = 3148884 and contender = 7448322. 7448322 is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed ⬇️0.54% ⬆️0.0%] test-mac-arm
[Finished ⬇️1.09% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.18% ⬆️0.04%] ursa-thinkcentre-m75q
Buildkite builds:
[Finished] 7448322e ec2-t3-xlarge-us-east-2
[Finished] 7448322e test-mac-arm
[Finished] 7448322e ursa-i9-9960x
[Finished] 7448322e ursa-thinkcentre-m75q
[Finished] 3148884e ec2-t3-xlarge-us-east-2
[Failed] 3148884e test-mac-arm
[Finished] 3148884e ursa-i9-9960x
[Finished] 3148884e ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

@ursabot

Copy link
Copy Markdown

['Python', 'R'] benchmarks have high level of regressions.
ursa-i9-9960x

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@paleolimbot@nealrichardson@ursabot@westonpace
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' ARROW-17252: [R] Intermittent valgrind failure by paleolimbot · Pull Request #13773 · apache/arrow · GitHub
Skip to content

ARROW-17252: [R] Intermittent valgrind failure - #13773

Merged
nealrichardson merged 8 commits into
apache:masterfrom
paleolimbot:r-valgrind-2
Aug 9, 2022
Merged

ARROW-17252: [R] Intermittent valgrind failure#13773
nealrichardson merged 8 commits into
apache:masterfrom
paleolimbot:r-valgrind-2

Conversation

@paleolimbot

@paleolimbotpaleolimbot commented Aug 2, 2022

Copy link
Copy Markdown
Member

This PR fixes intermittent leaks that occur after one of the changes from ARROW-16444: when we drain the RecordBatchReader that is emitted from the plan too quickly, it seems, some parts of the plan can leak (I don't know why this happens).

I tried removing various pieces of the RunWithCapturedR() changes (see #13746) but the only thing that removes the errors completely is draining the resulting RecordBatchReader from R (i.e., reader$read_table()) instead of in C++ (i.e., reader->ToTable()). Unfortunately, for user-defined functions to work in a plan we need a C++ level reader->ToTable(). I took the approach here of disabling the C++ level read by default, requiring a user to opt in to the version of collect() that works with a UDF. It's not ideal, but definitely safer (and more clearly marks the user-defined function behaviour as experimental).

I was able to replicate the original leaks but they are few and far between...our tests just happen to create and destroy many, many exec plans and something about the CI environment seems to trigger these more reliably (although the errors don't always occur there, either). Most of the leaks are small but there were some instances where an entire Table leaked.

@github-actions

Copy link
Copy Markdown

@github-actions

Copy link
Copy Markdown

⚠️ Ticket has not been started in JIRA, please click 'Start Progress'.

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

@github-actions crossbow submit test-r-linux-valgrind

@github-actions

Copy link
Copy Markdown

Revision: 3647aa3

Submitted crossbow builds: ursacomputing/crossbow @ actions-de212b2c40

TaskStatus
test-r-linux-valgrindAzure

@paleolimbot
paleolimbot marked this pull request as ready for review August 2, 2022 11:43
Comment threadr/R/compute.R
#' head()
#' Sys.unsetenv("R_ARROW_COLLECT_WITH_UDF")
#'
register_scalar_function <- function(name, fun, in_type, out_type,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What if you put Sys.setenv(R_ARROW_COLLECT_WITH_UDF = "true") inside of register_scalar_function()? You're already opting-in to UDFs by calling this function, and there's no reason you'd want to call this but not have COLLECT_WITH_UDF working.

Comment threadr/tests/testthat/test-compute.R Outdated
dplyr::mutate(fun_result = times_32(value)) %>%
dplyr::collect() %>%
dplyr::arrange(row_num)
withr::with_envvar(list(R_ARROW_COLLECT_WITH_UDF = "true"), {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You should be able to remove this now

Comment threadr/tests/testthat/test-compute.R Outdated
skip_if_not(CanRunWithCapturedR())
skip_if_not_available("dataset")
# Snappy has a UBSan issue: https://github.com/google/snappy/pull/148
# TODO(ARROW-17178): remove when user-defined function execution is stabilized

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note that the skip is for snappy ubsan too, so both need to be resolved before we can unskip

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

(See also ARROW-17283)

Comment threadr/tests/testthat/test-compute.R Outdated
dplyr::collect(),
tibble::tibble(a = 1L, b = 32.0)
)
withr::with_envvar(list(R_ARROW_COLLECT_WITH_UDF = "true"), {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can revert this

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we? To avoid the leaks for any subsequent tests we need the variable to not be "true".

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I was thinking you can revert this because you've already called register_scalar_function() so the env var will be set.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I see! with_envvar() was definitely not doing the right thing here (which is hopefully why the last valgrind check failed). I fixed it (I think) and requested another check!

Comment threadr/R/compute.R Outdated
@paleolimbot

Copy link
Copy Markdown
MemberAuthor

(See #13779 and #13780 for some other potential fixes, both have their checks pending)

paleolimbotand others added 2 commits August 2, 2022 14:40
Co-authored-by: Neal Richardson <neal.p.richardson@gmail.com>

@westonpacewestonpace left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

On the one hand, I can see why there might be some cases an extra collect is needed, to avoid valgrind errors. On the other hand:

  • The valgrind errors you are receiving are not the errors I would expect (I would expect direct leaks of scan-related objects).
  • I'm not entirely sure how this would be related to UDFs.

If this workaround fixes things for a little while then so be it. However, I'd really like to know exactly what is going on. You mentioned you had gotten better at reproducing. Could you share some kind of minimal(ish?) reproducer?

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

If this workaround fixes things for a little while then so be it.

I agree that this is a temporary workaround...I don't think anything all that insidious is going on that affects normal usage, but I also don't want us to get an angry CRAN note about valgrind that attracts attention to anything else we're doing.

I'd really like to know exactly what is going on.

Slight progress there: I tried waiting for the thread pools to finish before unloading the package (#13779) and that seems to remove the errors as well, although it's a bit of a hack in its own way (getting the IO thread pool is not exported in the public headers). That seems consistent with the shutting down of the thread pools leaking something?

Could you share some kind of minimal(ish?) reproducer?

From the arrow/r directory, echo 'devtools::test()' | R --no-save -d "valgrind --tool=memcheck --leak-check=full". Not very minimal, but an improvement over the 5 hours it takes the crossbow job. You can run fewer test files using devtools::test(filter = "some_regex") but I was never able to get an error using a few obvious filters ("dplyr", for example). I wonder if the CI just runs threads really really slowly which is why leaks show up more frequently then.

I'm not entirely sure how this would be related to UDFs.

The error isn't, I don't think, but was exposed by a change that we needed to make UDFs work (we need to evaluate the entire exec plan within one RunWithCapturedR(), which meant calling reader->ToTable() instead of sending the RecordBatchReader to R, then converting it to a table there). Before this PR, all exec plan results were routed through a C++-level reader->ToTable() unless an R-level RecordBatchReader was explicitly requested; after this PR, all exec plans go through an R-level RecordBatchReader unless somebody actually registers a user-defined function.

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

@github-actions crossbow submit test-r-linux-valgrind

@github-actions

Copy link
Copy Markdown

Revision: 3b71868

Submitted crossbow builds: ursacomputing/crossbow @ actions-511c82c07f

TaskStatus
test-r-linux-valgrindAzure

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

The last failure here is disappointing and leaves me stumped as to which change introduced the problem. I will spend some time tomorrow trying to reproduce this so that it's maybe possible to diagnose.

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

@github-actions crossbow submit test-r-linux-valgrind

@github-actions

Copy link
Copy Markdown

Revision: 2efade7

Submitted crossbow builds: ursacomputing/crossbow @ actions-d22fe5d8c4

TaskStatus
test-r-linux-valgrindAzure

Comment threadr/tests/testthat/test-compute.R Outdated
Sys.unsetenv("R_ARROW_COLLECT_WITH_UDF")
})

test_that("register_user_defined_function() can register multiple kernels", {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
test_that("register_user_defined_function() can register multiple kernels", {
test_that("register_scalar_function() can register multiple kernels", {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also below

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed!

Comment threadr/tests/testthat/test-compute.R Outdated
tibble::tibble(a = 1L, b = 32.0)
)

Sys.unsetenv("R_ARROW_COLLECT_WITH_UDF")

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

NBD but should this more properly go inside of on.exit above? like

on.exit({
unregister_binding("times_32", update_cache = TRUE)
# TODO(ARROW-17178) remove the need for this!
Sys.unsetenv("R_ARROW_COLLECT_WITH_UDF")
})

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed!

nealrichardson added a commit to nealrichardson/arrow that referenced this pull request Aug 7, 2022
@paleolimbot

Copy link
Copy Markdown
MemberAuthor

@github-actions crossbow submit test-r-linux-valgrind

@github-actions

Copy link
Copy Markdown

Revision: 37e6e2e

Submitted crossbow builds: ursacomputing/crossbow @ actions-d72df52d33

TaskStatus
test-r-linux-valgrindAzure

@nealrichardson

Copy link
Copy Markdown
Member

@paleolimbot is this good to merge now?

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

Yes! Sorry, forgot to check the valgrind result.

@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = 3148884 and contender = 7448322. 7448322 is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed ⬇️0.54% ⬆️0.0%] test-mac-arm
[Finished ⬇️1.09% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.18% ⬆️0.04%] ursa-thinkcentre-m75q
Buildkite builds:
[Finished] 7448322e ec2-t3-xlarge-us-east-2
[Finished] 7448322e test-mac-arm
[Finished] 7448322e ursa-i9-9960x
[Finished] 7448322e ursa-thinkcentre-m75q
[Finished] 3148884e ec2-t3-xlarge-us-east-2
[Failed] 3148884e test-mac-arm
[Finished] 3148884e ursa-i9-9960x
[Finished] 3148884e ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

@ursabot

Copy link
Copy Markdown

['Python', 'R'] benchmarks have high level of regressions.
ursa-i9-9960x

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@paleolimbot@nealrichardson@ursabot@westonpace
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' ARROW-17252: [R] Intermittent valgrind failure by paleolimbot · Pull Request #13773 · apache/arrow · GitHub
Skip to content

ARROW-17252: [R] Intermittent valgrind failure - #13773

Merged
nealrichardson merged 8 commits into
apache:masterfrom
paleolimbot:r-valgrind-2
Aug 9, 2022
Merged

ARROW-17252: [R] Intermittent valgrind failure#13773
nealrichardson merged 8 commits into
apache:masterfrom
paleolimbot:r-valgrind-2

Conversation

@paleolimbot

@paleolimbotpaleolimbot commented Aug 2, 2022

Copy link
Copy Markdown
Member

This PR fixes intermittent leaks that occur after one of the changes from ARROW-16444: when we drain the RecordBatchReader that is emitted from the plan too quickly, it seems, some parts of the plan can leak (I don't know why this happens).

I tried removing various pieces of the RunWithCapturedR() changes (see #13746) but the only thing that removes the errors completely is draining the resulting RecordBatchReader from R (i.e., reader$read_table()) instead of in C++ (i.e., reader->ToTable()). Unfortunately, for user-defined functions to work in a plan we need a C++ level reader->ToTable(). I took the approach here of disabling the C++ level read by default, requiring a user to opt in to the version of collect() that works with a UDF. It's not ideal, but definitely safer (and more clearly marks the user-defined function behaviour as experimental).

I was able to replicate the original leaks but they are few and far between...our tests just happen to create and destroy many, many exec plans and something about the CI environment seems to trigger these more reliably (although the errors don't always occur there, either). Most of the leaks are small but there were some instances where an entire Table leaked.

@github-actions

Copy link
Copy Markdown

@github-actions

Copy link
Copy Markdown

⚠️ Ticket has not been started in JIRA, please click 'Start Progress'.

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

@github-actions crossbow submit test-r-linux-valgrind

@github-actions

Copy link
Copy Markdown

Revision: 3647aa3

Submitted crossbow builds: ursacomputing/crossbow @ actions-de212b2c40

TaskStatus
test-r-linux-valgrindAzure

@paleolimbot
paleolimbot marked this pull request as ready for review August 2, 2022 11:43
Comment threadr/R/compute.R
#' head()
#' Sys.unsetenv("R_ARROW_COLLECT_WITH_UDF")
#'
register_scalar_function <- function(name, fun, in_type, out_type,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What if you put Sys.setenv(R_ARROW_COLLECT_WITH_UDF = "true") inside of register_scalar_function()? You're already opting-in to UDFs by calling this function, and there's no reason you'd want to call this but not have COLLECT_WITH_UDF working.

Comment threadr/tests/testthat/test-compute.R Outdated
dplyr::mutate(fun_result = times_32(value)) %>%
dplyr::collect() %>%
dplyr::arrange(row_num)
withr::with_envvar(list(R_ARROW_COLLECT_WITH_UDF = "true"), {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You should be able to remove this now

Comment threadr/tests/testthat/test-compute.R Outdated
skip_if_not(CanRunWithCapturedR())
skip_if_not_available("dataset")
# Snappy has a UBSan issue: https://github.com/google/snappy/pull/148
# TODO(ARROW-17178): remove when user-defined function execution is stabilized

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note that the skip is for snappy ubsan too, so both need to be resolved before we can unskip

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

(See also ARROW-17283)

Comment threadr/tests/testthat/test-compute.R Outdated
dplyr::collect(),
tibble::tibble(a = 1L, b = 32.0)
)
withr::with_envvar(list(R_ARROW_COLLECT_WITH_UDF = "true"), {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can revert this

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we? To avoid the leaks for any subsequent tests we need the variable to not be "true".

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I was thinking you can revert this because you've already called register_scalar_function() so the env var will be set.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I see! with_envvar() was definitely not doing the right thing here (which is hopefully why the last valgrind check failed). I fixed it (I think) and requested another check!

Comment threadr/R/compute.R Outdated
@paleolimbot

Copy link
Copy Markdown
MemberAuthor

(See #13779 and #13780 for some other potential fixes, both have their checks pending)

paleolimbotand others added 2 commits August 2, 2022 14:40
Co-authored-by: Neal Richardson <neal.p.richardson@gmail.com>

@westonpacewestonpace left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

On the one hand, I can see why there might be some cases an extra collect is needed, to avoid valgrind errors. On the other hand:

  • The valgrind errors you are receiving are not the errors I would expect (I would expect direct leaks of scan-related objects).
  • I'm not entirely sure how this would be related to UDFs.

If this workaround fixes things for a little while then so be it. However, I'd really like to know exactly what is going on. You mentioned you had gotten better at reproducing. Could you share some kind of minimal(ish?) reproducer?

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

If this workaround fixes things for a little while then so be it.

I agree that this is a temporary workaround...I don't think anything all that insidious is going on that affects normal usage, but I also don't want us to get an angry CRAN note about valgrind that attracts attention to anything else we're doing.

I'd really like to know exactly what is going on.

Slight progress there: I tried waiting for the thread pools to finish before unloading the package (#13779) and that seems to remove the errors as well, although it's a bit of a hack in its own way (getting the IO thread pool is not exported in the public headers). That seems consistent with the shutting down of the thread pools leaking something?

Could you share some kind of minimal(ish?) reproducer?

From the arrow/r directory, echo 'devtools::test()' | R --no-save -d "valgrind --tool=memcheck --leak-check=full". Not very minimal, but an improvement over the 5 hours it takes the crossbow job. You can run fewer test files using devtools::test(filter = "some_regex") but I was never able to get an error using a few obvious filters ("dplyr", for example). I wonder if the CI just runs threads really really slowly which is why leaks show up more frequently then.

I'm not entirely sure how this would be related to UDFs.

The error isn't, I don't think, but was exposed by a change that we needed to make UDFs work (we need to evaluate the entire exec plan within one RunWithCapturedR(), which meant calling reader->ToTable() instead of sending the RecordBatchReader to R, then converting it to a table there). Before this PR, all exec plan results were routed through a C++-level reader->ToTable() unless an R-level RecordBatchReader was explicitly requested; after this PR, all exec plans go through an R-level RecordBatchReader unless somebody actually registers a user-defined function.

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

@github-actions crossbow submit test-r-linux-valgrind

@github-actions

Copy link
Copy Markdown

Revision: 3b71868

Submitted crossbow builds: ursacomputing/crossbow @ actions-511c82c07f

TaskStatus
test-r-linux-valgrindAzure

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

The last failure here is disappointing and leaves me stumped as to which change introduced the problem. I will spend some time tomorrow trying to reproduce this so that it's maybe possible to diagnose.

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

@github-actions crossbow submit test-r-linux-valgrind

@github-actions

Copy link
Copy Markdown

Revision: 2efade7

Submitted crossbow builds: ursacomputing/crossbow @ actions-d22fe5d8c4

TaskStatus
test-r-linux-valgrindAzure

Comment threadr/tests/testthat/test-compute.R Outdated
Sys.unsetenv("R_ARROW_COLLECT_WITH_UDF")
})

test_that("register_user_defined_function() can register multiple kernels", {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
test_that("register_user_defined_function() can register multiple kernels", {
test_that("register_scalar_function() can register multiple kernels", {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also below

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed!

Comment threadr/tests/testthat/test-compute.R Outdated
tibble::tibble(a = 1L, b = 32.0)
)

Sys.unsetenv("R_ARROW_COLLECT_WITH_UDF")

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

NBD but should this more properly go inside of on.exit above? like

on.exit({
unregister_binding("times_32", update_cache = TRUE)
# TODO(ARROW-17178) remove the need for this!
Sys.unsetenv("R_ARROW_COLLECT_WITH_UDF")
})

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed!

nealrichardson added a commit to nealrichardson/arrow that referenced this pull request Aug 7, 2022
@paleolimbot

Copy link
Copy Markdown
MemberAuthor

@github-actions crossbow submit test-r-linux-valgrind

@github-actions

Copy link
Copy Markdown

Revision: 37e6e2e

Submitted crossbow builds: ursacomputing/crossbow @ actions-d72df52d33

TaskStatus
test-r-linux-valgrindAzure

@nealrichardson

Copy link
Copy Markdown
Member

@paleolimbot is this good to merge now?

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

Yes! Sorry, forgot to check the valgrind result.

@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = 3148884 and contender = 7448322. 7448322 is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed ⬇️0.54% ⬆️0.0%] test-mac-arm
[Finished ⬇️1.09% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.18% ⬆️0.04%] ursa-thinkcentre-m75q
Buildkite builds:
[Finished] 7448322e ec2-t3-xlarge-us-east-2
[Finished] 7448322e test-mac-arm
[Finished] 7448322e ursa-i9-9960x
[Finished] 7448322e ursa-thinkcentre-m75q
[Finished] 3148884e ec2-t3-xlarge-us-east-2
[Failed] 3148884e test-mac-arm
[Finished] 3148884e ursa-i9-9960x
[Finished] 3148884e ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

@ursabot

Copy link
Copy Markdown

['Python', 'R'] benchmarks have high level of regressions.
ursa-i9-9960x

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@paleolimbot@nealrichardson@ursabot@westonpace
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' ARROW-17252: [R] Intermittent valgrind failure by paleolimbot · Pull Request #13773 · apache/arrow · GitHub
Skip to content

ARROW-17252: [R] Intermittent valgrind failure - #13773

Merged
nealrichardson merged 8 commits into
apache:masterfrom
paleolimbot:r-valgrind-2
Aug 9, 2022
Merged

ARROW-17252: [R] Intermittent valgrind failure#13773
nealrichardson merged 8 commits into
apache:masterfrom
paleolimbot:r-valgrind-2

Conversation

@paleolimbot

@paleolimbotpaleolimbot commented Aug 2, 2022

Copy link
Copy Markdown
Member

This PR fixes intermittent leaks that occur after one of the changes from ARROW-16444: when we drain the RecordBatchReader that is emitted from the plan too quickly, it seems, some parts of the plan can leak (I don't know why this happens).

I tried removing various pieces of the RunWithCapturedR() changes (see #13746) but the only thing that removes the errors completely is draining the resulting RecordBatchReader from R (i.e., reader$read_table()) instead of in C++ (i.e., reader->ToTable()). Unfortunately, for user-defined functions to work in a plan we need a C++ level reader->ToTable(). I took the approach here of disabling the C++ level read by default, requiring a user to opt in to the version of collect() that works with a UDF. It's not ideal, but definitely safer (and more clearly marks the user-defined function behaviour as experimental).

I was able to replicate the original leaks but they are few and far between...our tests just happen to create and destroy many, many exec plans and something about the CI environment seems to trigger these more reliably (although the errors don't always occur there, either). Most of the leaks are small but there were some instances where an entire Table leaked.

@github-actions

Copy link
Copy Markdown

@github-actions

Copy link
Copy Markdown

⚠️ Ticket has not been started in JIRA, please click 'Start Progress'.

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

@github-actions crossbow submit test-r-linux-valgrind

@github-actions

Copy link
Copy Markdown

Revision: 3647aa3

Submitted crossbow builds: ursacomputing/crossbow @ actions-de212b2c40

TaskStatus
test-r-linux-valgrindAzure

@paleolimbot
paleolimbot marked this pull request as ready for review August 2, 2022 11:43
Comment threadr/R/compute.R
#' head()
#' Sys.unsetenv("R_ARROW_COLLECT_WITH_UDF")
#'
register_scalar_function <- function(name, fun, in_type, out_type,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What if you put Sys.setenv(R_ARROW_COLLECT_WITH_UDF = "true") inside of register_scalar_function()? You're already opting-in to UDFs by calling this function, and there's no reason you'd want to call this but not have COLLECT_WITH_UDF working.

Comment threadr/tests/testthat/test-compute.R Outdated
dplyr::mutate(fun_result = times_32(value)) %>%
dplyr::collect() %>%
dplyr::arrange(row_num)
withr::with_envvar(list(R_ARROW_COLLECT_WITH_UDF = "true"), {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You should be able to remove this now

Comment threadr/tests/testthat/test-compute.R Outdated
skip_if_not(CanRunWithCapturedR())
skip_if_not_available("dataset")
# Snappy has a UBSan issue: https://github.com/google/snappy/pull/148
# TODO(ARROW-17178): remove when user-defined function execution is stabilized

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note that the skip is for snappy ubsan too, so both need to be resolved before we can unskip

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

(See also ARROW-17283)

Comment threadr/tests/testthat/test-compute.R Outdated
dplyr::collect(),
tibble::tibble(a = 1L, b = 32.0)
)
withr::with_envvar(list(R_ARROW_COLLECT_WITH_UDF = "true"), {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can revert this

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we? To avoid the leaks for any subsequent tests we need the variable to not be "true".

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I was thinking you can revert this because you've already called register_scalar_function() so the env var will be set.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I see! with_envvar() was definitely not doing the right thing here (which is hopefully why the last valgrind check failed). I fixed it (I think) and requested another check!

Comment threadr/R/compute.R Outdated
@paleolimbot

Copy link
Copy Markdown
MemberAuthor

(See #13779 and #13780 for some other potential fixes, both have their checks pending)

paleolimbotand others added 2 commits August 2, 2022 14:40
Co-authored-by: Neal Richardson <neal.p.richardson@gmail.com>

@westonpacewestonpace left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

On the one hand, I can see why there might be some cases an extra collect is needed, to avoid valgrind errors. On the other hand:

  • The valgrind errors you are receiving are not the errors I would expect (I would expect direct leaks of scan-related objects).
  • I'm not entirely sure how this would be related to UDFs.

If this workaround fixes things for a little while then so be it. However, I'd really like to know exactly what is going on. You mentioned you had gotten better at reproducing. Could you share some kind of minimal(ish?) reproducer?

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

If this workaround fixes things for a little while then so be it.

I agree that this is a temporary workaround...I don't think anything all that insidious is going on that affects normal usage, but I also don't want us to get an angry CRAN note about valgrind that attracts attention to anything else we're doing.

I'd really like to know exactly what is going on.

Slight progress there: I tried waiting for the thread pools to finish before unloading the package (#13779) and that seems to remove the errors as well, although it's a bit of a hack in its own way (getting the IO thread pool is not exported in the public headers). That seems consistent with the shutting down of the thread pools leaking something?

Could you share some kind of minimal(ish?) reproducer?

From the arrow/r directory, echo 'devtools::test()' | R --no-save -d "valgrind --tool=memcheck --leak-check=full". Not very minimal, but an improvement over the 5 hours it takes the crossbow job. You can run fewer test files using devtools::test(filter = "some_regex") but I was never able to get an error using a few obvious filters ("dplyr", for example). I wonder if the CI just runs threads really really slowly which is why leaks show up more frequently then.

I'm not entirely sure how this would be related to UDFs.

The error isn't, I don't think, but was exposed by a change that we needed to make UDFs work (we need to evaluate the entire exec plan within one RunWithCapturedR(), which meant calling reader->ToTable() instead of sending the RecordBatchReader to R, then converting it to a table there). Before this PR, all exec plan results were routed through a C++-level reader->ToTable() unless an R-level RecordBatchReader was explicitly requested; after this PR, all exec plans go through an R-level RecordBatchReader unless somebody actually registers a user-defined function.

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

@github-actions crossbow submit test-r-linux-valgrind

@github-actions

Copy link
Copy Markdown

Revision: 3b71868

Submitted crossbow builds: ursacomputing/crossbow @ actions-511c82c07f

TaskStatus
test-r-linux-valgrindAzure

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

The last failure here is disappointing and leaves me stumped as to which change introduced the problem. I will spend some time tomorrow trying to reproduce this so that it's maybe possible to diagnose.

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

@github-actions crossbow submit test-r-linux-valgrind

@github-actions

Copy link
Copy Markdown

Revision: 2efade7

Submitted crossbow builds: ursacomputing/crossbow @ actions-d22fe5d8c4

TaskStatus
test-r-linux-valgrindAzure

Comment threadr/tests/testthat/test-compute.R Outdated
Sys.unsetenv("R_ARROW_COLLECT_WITH_UDF")
})

test_that("register_user_defined_function() can register multiple kernels", {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
test_that("register_user_defined_function() can register multiple kernels", {
test_that("register_scalar_function() can register multiple kernels", {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also below

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed!

Comment threadr/tests/testthat/test-compute.R Outdated
tibble::tibble(a = 1L, b = 32.0)
)

Sys.unsetenv("R_ARROW_COLLECT_WITH_UDF")

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

NBD but should this more properly go inside of on.exit above? like

on.exit({
unregister_binding("times_32", update_cache = TRUE)
# TODO(ARROW-17178) remove the need for this!
Sys.unsetenv("R_ARROW_COLLECT_WITH_UDF")
})

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed!

nealrichardson added a commit to nealrichardson/arrow that referenced this pull request Aug 7, 2022
@paleolimbot

Copy link
Copy Markdown
MemberAuthor

@github-actions crossbow submit test-r-linux-valgrind

@github-actions

Copy link
Copy Markdown

Revision: 37e6e2e

Submitted crossbow builds: ursacomputing/crossbow @ actions-d72df52d33

TaskStatus
test-r-linux-valgrindAzure

@nealrichardson

Copy link
Copy Markdown
Member

@paleolimbot is this good to merge now?

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

Yes! Sorry, forgot to check the valgrind result.

@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = 3148884 and contender = 7448322. 7448322 is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed ⬇️0.54% ⬆️0.0%] test-mac-arm
[Finished ⬇️1.09% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.18% ⬆️0.04%] ursa-thinkcentre-m75q
Buildkite builds:
[Finished] 7448322e ec2-t3-xlarge-us-east-2
[Finished] 7448322e test-mac-arm
[Finished] 7448322e ursa-i9-9960x
[Finished] 7448322e ursa-thinkcentre-m75q
[Finished] 3148884e ec2-t3-xlarge-us-east-2
[Failed] 3148884e test-mac-arm
[Finished] 3148884e ursa-i9-9960x
[Finished] 3148884e ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

@ursabot

Copy link
Copy Markdown

['Python', 'R'] benchmarks have high level of regressions.
ursa-i9-9960x

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@paleolimbot@nealrichardson@ursabot@westonpace
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); ARROW-17252: [R] Intermittent valgrind failure by paleolimbot · Pull Request #13773 · apache/arrow · GitHub
Skip to content

ARROW-17252: [R] Intermittent valgrind failure - #13773

Merged
nealrichardson merged 8 commits into
apache:masterfrom
paleolimbot:r-valgrind-2
Aug 9, 2022
Merged

ARROW-17252: [R] Intermittent valgrind failure#13773
nealrichardson merged 8 commits into
apache:masterfrom
paleolimbot:r-valgrind-2

Conversation

@paleolimbot

@paleolimbotpaleolimbot commented Aug 2, 2022

Copy link
Copy Markdown
Member

This PR fixes intermittent leaks that occur after one of the changes from ARROW-16444: when we drain the RecordBatchReader that is emitted from the plan too quickly, it seems, some parts of the plan can leak (I don't know why this happens).

I tried removing various pieces of the RunWithCapturedR() changes (see #13746) but the only thing that removes the errors completely is draining the resulting RecordBatchReader from R (i.e., reader$read_table()) instead of in C++ (i.e., reader->ToTable()). Unfortunately, for user-defined functions to work in a plan we need a C++ level reader->ToTable(). I took the approach here of disabling the C++ level read by default, requiring a user to opt in to the version of collect() that works with a UDF. It's not ideal, but definitely safer (and more clearly marks the user-defined function behaviour as experimental).

I was able to replicate the original leaks but they are few and far between...our tests just happen to create and destroy many, many exec plans and something about the CI environment seems to trigger these more reliably (although the errors don't always occur there, either). Most of the leaks are small but there were some instances where an entire Table leaked.

@github-actions

Copy link
Copy Markdown

@github-actions

Copy link
Copy Markdown

⚠️ Ticket has not been started in JIRA, please click 'Start Progress'.

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

@github-actions crossbow submit test-r-linux-valgrind

@github-actions

Copy link
Copy Markdown

Revision: 3647aa3

Submitted crossbow builds: ursacomputing/crossbow @ actions-de212b2c40

TaskStatus
test-r-linux-valgrindAzure

@paleolimbot
paleolimbot marked this pull request as ready for review August 2, 2022 11:43
Comment threadr/R/compute.R
#' head()
#' Sys.unsetenv("R_ARROW_COLLECT_WITH_UDF")
#'
register_scalar_function <- function(name, fun, in_type, out_type,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What if you put Sys.setenv(R_ARROW_COLLECT_WITH_UDF = "true") inside of register_scalar_function()? You're already opting-in to UDFs by calling this function, and there's no reason you'd want to call this but not have COLLECT_WITH_UDF working.

Comment threadr/tests/testthat/test-compute.R Outdated
dplyr::mutate(fun_result = times_32(value)) %>%
dplyr::collect() %>%
dplyr::arrange(row_num)
withr::with_envvar(list(R_ARROW_COLLECT_WITH_UDF = "true"), {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You should be able to remove this now

Comment threadr/tests/testthat/test-compute.R Outdated
skip_if_not(CanRunWithCapturedR())
skip_if_not_available("dataset")
# Snappy has a UBSan issue: https://github.com/google/snappy/pull/148
# TODO(ARROW-17178): remove when user-defined function execution is stabilized

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note that the skip is for snappy ubsan too, so both need to be resolved before we can unskip

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

(See also ARROW-17283)

Comment threadr/tests/testthat/test-compute.R Outdated
dplyr::collect(),
tibble::tibble(a = 1L, b = 32.0)
)
withr::with_envvar(list(R_ARROW_COLLECT_WITH_UDF = "true"), {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can revert this

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we? To avoid the leaks for any subsequent tests we need the variable to not be "true".

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I was thinking you can revert this because you've already called register_scalar_function() so the env var will be set.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I see! with_envvar() was definitely not doing the right thing here (which is hopefully why the last valgrind check failed). I fixed it (I think) and requested another check!

Comment threadr/R/compute.R Outdated
@paleolimbot

Copy link
Copy Markdown
MemberAuthor

(See #13779 and #13780 for some other potential fixes, both have their checks pending)

paleolimbotand others added 2 commits August 2, 2022 14:40
Co-authored-by: Neal Richardson <neal.p.richardson@gmail.com>

@westonpacewestonpace left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

On the one hand, I can see why there might be some cases an extra collect is needed, to avoid valgrind errors. On the other hand:

  • The valgrind errors you are receiving are not the errors I would expect (I would expect direct leaks of scan-related objects).
  • I'm not entirely sure how this would be related to UDFs.

If this workaround fixes things for a little while then so be it. However, I'd really like to know exactly what is going on. You mentioned you had gotten better at reproducing. Could you share some kind of minimal(ish?) reproducer?

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

If this workaround fixes things for a little while then so be it.

I agree that this is a temporary workaround...I don't think anything all that insidious is going on that affects normal usage, but I also don't want us to get an angry CRAN note about valgrind that attracts attention to anything else we're doing.

I'd really like to know exactly what is going on.

Slight progress there: I tried waiting for the thread pools to finish before unloading the package (#13779) and that seems to remove the errors as well, although it's a bit of a hack in its own way (getting the IO thread pool is not exported in the public headers). That seems consistent with the shutting down of the thread pools leaking something?

Could you share some kind of minimal(ish?) reproducer?

From the arrow/r directory, echo 'devtools::test()' | R --no-save -d "valgrind --tool=memcheck --leak-check=full". Not very minimal, but an improvement over the 5 hours it takes the crossbow job. You can run fewer test files using devtools::test(filter = "some_regex") but I was never able to get an error using a few obvious filters ("dplyr", for example). I wonder if the CI just runs threads really really slowly which is why leaks show up more frequently then.

I'm not entirely sure how this would be related to UDFs.

The error isn't, I don't think, but was exposed by a change that we needed to make UDFs work (we need to evaluate the entire exec plan within one RunWithCapturedR(), which meant calling reader->ToTable() instead of sending the RecordBatchReader to R, then converting it to a table there). Before this PR, all exec plan results were routed through a C++-level reader->ToTable() unless an R-level RecordBatchReader was explicitly requested; after this PR, all exec plans go through an R-level RecordBatchReader unless somebody actually registers a user-defined function.

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

@github-actions crossbow submit test-r-linux-valgrind

@github-actions

Copy link
Copy Markdown

Revision: 3b71868

Submitted crossbow builds: ursacomputing/crossbow @ actions-511c82c07f

TaskStatus
test-r-linux-valgrindAzure

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

The last failure here is disappointing and leaves me stumped as to which change introduced the problem. I will spend some time tomorrow trying to reproduce this so that it's maybe possible to diagnose.

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

@github-actions crossbow submit test-r-linux-valgrind

@github-actions

Copy link
Copy Markdown

Revision: 2efade7

Submitted crossbow builds: ursacomputing/crossbow @ actions-d22fe5d8c4

TaskStatus
test-r-linux-valgrindAzure

Comment threadr/tests/testthat/test-compute.R Outdated
Sys.unsetenv("R_ARROW_COLLECT_WITH_UDF")
})

test_that("register_user_defined_function() can register multiple kernels", {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
test_that("register_user_defined_function() can register multiple kernels", {
test_that("register_scalar_function() can register multiple kernels", {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also below

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed!

Comment threadr/tests/testthat/test-compute.R Outdated
tibble::tibble(a = 1L, b = 32.0)
)

Sys.unsetenv("R_ARROW_COLLECT_WITH_UDF")

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

NBD but should this more properly go inside of on.exit above? like

on.exit({
unregister_binding("times_32", update_cache = TRUE)
# TODO(ARROW-17178) remove the need for this!
Sys.unsetenv("R_ARROW_COLLECT_WITH_UDF")
})

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed!

nealrichardson added a commit to nealrichardson/arrow that referenced this pull request Aug 7, 2022
@paleolimbot

Copy link
Copy Markdown
MemberAuthor

@github-actions crossbow submit test-r-linux-valgrind

@github-actions

Copy link
Copy Markdown

Revision: 37e6e2e

Submitted crossbow builds: ursacomputing/crossbow @ actions-d72df52d33

TaskStatus
test-r-linux-valgrindAzure

@nealrichardson

Copy link
Copy Markdown
Member

@paleolimbot is this good to merge now?

@paleolimbot

Copy link
Copy Markdown
MemberAuthor

Yes! Sorry, forgot to check the valgrind result.

@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = 3148884 and contender = 7448322. 7448322 is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed ⬇️0.54% ⬆️0.0%] test-mac-arm
[Finished ⬇️1.09% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.18% ⬆️0.04%] ursa-thinkcentre-m75q
Buildkite builds:
[Finished] 7448322e ec2-t3-xlarge-us-east-2
[Finished] 7448322e test-mac-arm
[Finished] 7448322e ursa-i9-9960x
[Finished] 7448322e ursa-thinkcentre-m75q
[Finished] 3148884e ec2-t3-xlarge-us-east-2
[Failed] 3148884e test-mac-arm
[Finished] 3148884e ursa-i9-9960x
[Finished] 3148884e ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

@ursabot

Copy link
Copy Markdown

['Python', 'R'] benchmarks have high level of regressions.
ursa-i9-9960x

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@paleolimbot@nealrichardson@ursabot@westonpace