Use caller CUDA stream for D2H and H2D copies (#20498) - #20498

Merged
meta-codesync[bot] merged 1 commit into
pytorch:mainfrom
Conarnar:export-D109590531
Jun 30, 2026
Merged

Use caller CUDA stream for D2H and H2D copies (#20498)#20498
meta-codesync[bot] merged 1 commit into
pytorch:mainfrom
Conarnar:export-D109590531

Conversation

@Conarnar

@ConarnarConarnar commented Jun 24, 2026

Copy link
Copy Markdown
Contributor

Summary:

CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via getCallerStream()), copy_host_to_device and copy_device_to_host use cudaMemcpyAsync. When no caller stream is set, the synchronous cudaMemcpy path is used as before.

Additionally:

  • Added null pointer and zero-byte validation — null dst/src return Error::InvalidArgument instead of aborting in cudaMemcpy, and zero-byte copies return Error::Ok early.
  • Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
  • Wired //executorch/extension/cuda:caller_stream dependency in TARGETS.
  • Added extension_cuda dependencies to CMakeLists.txt.
  • Added test_cuda_allocator with coverage for sync/async paths and error handling.
  • Added CIs for unit tests.

Reviewed By: Gasoonjia

Differential Revision: D109590531

CopilotAI review requested due to automatic review settings June 24, 2026 22:51
@pytorch-bot

pytorch-botBot commented Jun 24, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20498

Note: Links to docs will display an error until the docs builds have been completed.

✅ No Failures

As of commit d3f5462 with merge base b331ebd (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 24, 2026
@meta-codesync

Copy link
Copy Markdown
Contributor

@Conarnar has exported this pull request. If you are a Meta employee, you can view the originating Diff in D109590531.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot encountered an error and was unable to review this pull request. You can try again by re-requesting a review.

@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@meta-codesyncmeta-codesyncBot changed the title Use caller CUDA stream for D2H and H2D copiesUse caller CUDA stream for D2H and H2D copies (#20498)Jun 24, 2026
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 24, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync` and synchronize the stream before returning — preserving the blocking API contract while allowing work to be issued on the caller's stream. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
@Conarnar
Conarnarforce-pushed the export-D109590531 branch 2 times, most recently from 3d8da75 to 07765c3CompareJune 25, 2026 17:10
CopilotAI review requested due to automatic review settings June 25, 2026 17:10
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync` and synchronize the stream before returning — preserving the blocking API contract while allowing work to be issued on the caller's stream. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 4 out of 4 changed files in this pull request and generated 5 comments.

Comment on lines +161 to +168
cudaError_t err = cudaSuccess;
const auto caller_stream = executorch::extension::cuda::getCallerStream();
if (caller_stream) {
err = cudaMemcpyAsync(
dst, src, nbytes, cudaMemcpyHostToDevice, *caller_stream);
} else {
err = cudaMemcpy(dst, src, nbytes, cudaMemcpyHostToDevice);
}
Comment on lines +144 to +150
// TODO: validate caller stream device matches index.
// For now assert single-GPU case.
ET_CHECK_OR_RETURN_ERROR(
index == -1 || index == 0,
InvalidArgument,
"CudaAllocator::copy_host_to_device only supports device 0, got %d",
static_cast<int>(index));
Comment on lines +202 to +208
// TODO: validate caller stream device matches index.
// For now assert single-GPU case.
ET_CHECK_OR_RETURN_ERROR(
index == -1 || index == 0,
InvalidArgument,
"CudaAllocator::copy_device_to_host only supports device 0, got %d",
static_cast<int>(index));
Comment on lines +78 to +90
cudaStream_t s;
ASSERT_EQ(cudaStreamCreate(&s), cudaSuccess);
executorch::extension::cuda::CallerStreamGuard g(s);

CudaAllocator& a = CudaAllocator::instance();
auto res = a.allocate(256, 0);
ASSERT_TRUE(res.ok());
void* d = res.get();
std::vector<uint8_t> h(256, 7);
// should take async branch internally, still return Ok
EXPECT_EQ(a.copy_host_to_device(d, h.data(), 256, 0), Error::Ok);
a.deallocate(d, 0);
cudaStreamDestroy(s);
Comment on lines +103 to +117
cudaStream_t s;
ASSERT_EQ(cudaStreamCreate(&s), cudaSuccess);
executorch::extension::cuda::CallerStreamGuard g(s);

CudaAllocator& a = CudaAllocator::instance();
auto res = a.allocate(256, 0);
ASSERT_TRUE(res.ok());
void* d = res.get();
std::vector<uint8_t> h_src(256, 5), h_dst(256, 0);
ASSERT_EQ(a.copy_host_to_device(d, h_src.data(), 256, 0), Error::Ok);
EXPECT_EQ(a.copy_device_to_host(h_dst.data(), d, 256, 0), Error::Ok);
EXPECT_EQ(h_src, h_dst);

a.deallocate(d, 0);
cudaStreamDestroy(s);
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 25, 2026 18:57

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 4 out of 4 changed files in this pull request and generated 5 comments.

Comment on lines +144 to +150
// TODO: validate caller stream device matches index.
// For now assert index is -1 or 0.
ET_CHECK_OR_RETURN_ERROR(
index == -1 || index == 0,
InvalidArgument,
"CudaAllocator::copy_host_to_device only supports device 0 or -1 (current), got %d",
static_cast<int>(index));
Comment on lines +204 to +210
// TODO: validate caller stream device matches index.
// For now assert index is -1 or 0.
ET_CHECK_OR_RETURN_ERROR(
index == -1 || index == 0,
InvalidArgument,
"CudaAllocator::copy_device_to_host only supports device 0 or -1 (current), got %d",
static_cast<int>(index));
Comment on lines +161 to +166
cudaError_t err = cudaSuccess;
const auto caller_stream = executorch::extension::cuda::getCallerStream();
if (caller_stream) {
err = cudaMemcpyAsync(
dst, src, nbytes, cudaMemcpyHostToDevice, *caller_stream);
// We don't synchronize the stream here because the caller is expected to
Comment on lines +223 to +228
if (caller_stream) {
err = cudaMemcpyAsync(
dst, src, nbytes, cudaMemcpyDeviceToHost, *caller_stream);
if (err == cudaSuccess) {
err = cudaStreamSynchronize(*caller_stream);
}
Comment on lines +116 to +118
ASSERT_EQ(a.copy_host_to_device(d, h_src.data(), 256, 0), Error::Ok);
EXPECT_EQ(a.copy_device_to_host(h_dst.data(), d, 256, 0), Error::Ok);
EXPECT_EQ(h_src, h_dst);
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 27, 2026 10:24
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
@Conarnar
Conarnarforce-pushed the export-D109590531 branch 2 times, most recently from 32968c0 to 056c25cCompareJune 27, 2026 10:24

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
Pull Request resolved: pytorch#20498
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 27, 2026 19:34
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 28, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 28, 2026 00:10

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 28, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 28, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
- Added CIs for unit tests.
Reviewed By: Gasoonjia
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 28, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
- Added CIs for unit tests.
Reviewed By: Gasoonjia
Differential Revision: D109590531
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
- Added CIs for unit tests.
Reviewed By: Gasoonjia
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 29, 2026 18:40

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ciflow/cudaCLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@Conarnar@Gasoonjia@shoumikhin
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Use caller CUDA stream for D2H and H2D copies (#20498) - #20498

Merged
meta-codesync[bot] merged 1 commit into
pytorch:mainfrom
Conarnar:export-D109590531
Jun 30, 2026
Merged

Use caller CUDA stream for D2H and H2D copies (#20498)#20498
meta-codesync[bot] merged 1 commit into
pytorch:mainfrom
Conarnar:export-D109590531

Conversation

@Conarnar

@ConarnarConarnar commented Jun 24, 2026

Copy link
Copy Markdown
Contributor

Summary:

CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via getCallerStream()), copy_host_to_device and copy_device_to_host use cudaMemcpyAsync. When no caller stream is set, the synchronous cudaMemcpy path is used as before.

Additionally:

  • Added null pointer and zero-byte validation — null dst/src return Error::InvalidArgument instead of aborting in cudaMemcpy, and zero-byte copies return Error::Ok early.
  • Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
  • Wired //executorch/extension/cuda:caller_stream dependency in TARGETS.
  • Added extension_cuda dependencies to CMakeLists.txt.
  • Added test_cuda_allocator with coverage for sync/async paths and error handling.
  • Added CIs for unit tests.

Reviewed By: Gasoonjia

Differential Revision: D109590531

CopilotAI review requested due to automatic review settings June 24, 2026 22:51
@pytorch-bot

pytorch-botBot commented Jun 24, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20498

Note: Links to docs will display an error until the docs builds have been completed.

✅ No Failures

As of commit d3f5462 with merge base b331ebd (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 24, 2026
@meta-codesync

Copy link
Copy Markdown
Contributor

@Conarnar has exported this pull request. If you are a Meta employee, you can view the originating Diff in D109590531.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot encountered an error and was unable to review this pull request. You can try again by re-requesting a review.

@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@meta-codesyncmeta-codesyncBot changed the title Use caller CUDA stream for D2H and H2D copiesUse caller CUDA stream for D2H and H2D copies (#20498)Jun 24, 2026
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 24, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync` and synchronize the stream before returning — preserving the blocking API contract while allowing work to be issued on the caller's stream. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
@Conarnar
Conarnarforce-pushed the export-D109590531 branch 2 times, most recently from 3d8da75 to 07765c3CompareJune 25, 2026 17:10
CopilotAI review requested due to automatic review settings June 25, 2026 17:10
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync` and synchronize the stream before returning — preserving the blocking API contract while allowing work to be issued on the caller's stream. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 4 out of 4 changed files in this pull request and generated 5 comments.

Comment on lines +161 to +168
cudaError_t err = cudaSuccess;
const auto caller_stream = executorch::extension::cuda::getCallerStream();
if (caller_stream) {
err = cudaMemcpyAsync(
dst, src, nbytes, cudaMemcpyHostToDevice, *caller_stream);
} else {
err = cudaMemcpy(dst, src, nbytes, cudaMemcpyHostToDevice);
}
Comment on lines +144 to +150
// TODO: validate caller stream device matches index.
// For now assert single-GPU case.
ET_CHECK_OR_RETURN_ERROR(
index == -1 || index == 0,
InvalidArgument,
"CudaAllocator::copy_host_to_device only supports device 0, got %d",
static_cast<int>(index));
Comment on lines +202 to +208
// TODO: validate caller stream device matches index.
// For now assert single-GPU case.
ET_CHECK_OR_RETURN_ERROR(
index == -1 || index == 0,
InvalidArgument,
"CudaAllocator::copy_device_to_host only supports device 0, got %d",
static_cast<int>(index));
Comment on lines +78 to +90
cudaStream_t s;
ASSERT_EQ(cudaStreamCreate(&s), cudaSuccess);
executorch::extension::cuda::CallerStreamGuard g(s);

CudaAllocator& a = CudaAllocator::instance();
auto res = a.allocate(256, 0);
ASSERT_TRUE(res.ok());
void* d = res.get();
std::vector<uint8_t> h(256, 7);
// should take async branch internally, still return Ok
EXPECT_EQ(a.copy_host_to_device(d, h.data(), 256, 0), Error::Ok);
a.deallocate(d, 0);
cudaStreamDestroy(s);
Comment on lines +103 to +117
cudaStream_t s;
ASSERT_EQ(cudaStreamCreate(&s), cudaSuccess);
executorch::extension::cuda::CallerStreamGuard g(s);

CudaAllocator& a = CudaAllocator::instance();
auto res = a.allocate(256, 0);
ASSERT_TRUE(res.ok());
void* d = res.get();
std::vector<uint8_t> h_src(256, 5), h_dst(256, 0);
ASSERT_EQ(a.copy_host_to_device(d, h_src.data(), 256, 0), Error::Ok);
EXPECT_EQ(a.copy_device_to_host(h_dst.data(), d, 256, 0), Error::Ok);
EXPECT_EQ(h_src, h_dst);

a.deallocate(d, 0);
cudaStreamDestroy(s);
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 25, 2026 18:57

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 4 out of 4 changed files in this pull request and generated 5 comments.

Comment on lines +144 to +150
// TODO: validate caller stream device matches index.
// For now assert index is -1 or 0.
ET_CHECK_OR_RETURN_ERROR(
index == -1 || index == 0,
InvalidArgument,
"CudaAllocator::copy_host_to_device only supports device 0 or -1 (current), got %d",
static_cast<int>(index));
Comment on lines +204 to +210
// TODO: validate caller stream device matches index.
// For now assert index is -1 or 0.
ET_CHECK_OR_RETURN_ERROR(
index == -1 || index == 0,
InvalidArgument,
"CudaAllocator::copy_device_to_host only supports device 0 or -1 (current), got %d",
static_cast<int>(index));
Comment on lines +161 to +166
cudaError_t err = cudaSuccess;
const auto caller_stream = executorch::extension::cuda::getCallerStream();
if (caller_stream) {
err = cudaMemcpyAsync(
dst, src, nbytes, cudaMemcpyHostToDevice, *caller_stream);
// We don't synchronize the stream here because the caller is expected to
Comment on lines +223 to +228
if (caller_stream) {
err = cudaMemcpyAsync(
dst, src, nbytes, cudaMemcpyDeviceToHost, *caller_stream);
if (err == cudaSuccess) {
err = cudaStreamSynchronize(*caller_stream);
}
Comment on lines +116 to +118
ASSERT_EQ(a.copy_host_to_device(d, h_src.data(), 256, 0), Error::Ok);
EXPECT_EQ(a.copy_device_to_host(h_dst.data(), d, 256, 0), Error::Ok);
EXPECT_EQ(h_src, h_dst);
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 27, 2026 10:24
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
@Conarnar
Conarnarforce-pushed the export-D109590531 branch 2 times, most recently from 32968c0 to 056c25cCompareJune 27, 2026 10:24

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
Pull Request resolved: pytorch#20498
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 27, 2026 19:34
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 28, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 28, 2026 00:10

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 28, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 28, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
- Added CIs for unit tests.
Reviewed By: Gasoonjia
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 28, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
- Added CIs for unit tests.
Reviewed By: Gasoonjia
Differential Revision: D109590531
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
- Added CIs for unit tests.
Reviewed By: Gasoonjia
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 29, 2026 18:40

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ciflow/cudaCLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@Conarnar@Gasoonjia@shoumikhin
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Use caller CUDA stream for D2H and H2D copies (#20498) - #20498

Merged
meta-codesync[bot] merged 1 commit into
pytorch:mainfrom
Conarnar:export-D109590531
Jun 30, 2026
Merged

Use caller CUDA stream for D2H and H2D copies (#20498)#20498
meta-codesync[bot] merged 1 commit into
pytorch:mainfrom
Conarnar:export-D109590531

Conversation

@Conarnar

@ConarnarConarnar commented Jun 24, 2026

Copy link
Copy Markdown
Contributor

Summary:

CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via getCallerStream()), copy_host_to_device and copy_device_to_host use cudaMemcpyAsync. When no caller stream is set, the synchronous cudaMemcpy path is used as before.

Additionally:

  • Added null pointer and zero-byte validation — null dst/src return Error::InvalidArgument instead of aborting in cudaMemcpy, and zero-byte copies return Error::Ok early.
  • Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
  • Wired //executorch/extension/cuda:caller_stream dependency in TARGETS.
  • Added extension_cuda dependencies to CMakeLists.txt.
  • Added test_cuda_allocator with coverage for sync/async paths and error handling.
  • Added CIs for unit tests.

Reviewed By: Gasoonjia

Differential Revision: D109590531

CopilotAI review requested due to automatic review settings June 24, 2026 22:51
@pytorch-bot

pytorch-botBot commented Jun 24, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20498

Note: Links to docs will display an error until the docs builds have been completed.

✅ No Failures

As of commit d3f5462 with merge base b331ebd (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 24, 2026
@meta-codesync

Copy link
Copy Markdown
Contributor

@Conarnar has exported this pull request. If you are a Meta employee, you can view the originating Diff in D109590531.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot encountered an error and was unable to review this pull request. You can try again by re-requesting a review.

@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@meta-codesyncmeta-codesyncBot changed the title Use caller CUDA stream for D2H and H2D copiesUse caller CUDA stream for D2H and H2D copies (#20498)Jun 24, 2026
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 24, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync` and synchronize the stream before returning — preserving the blocking API contract while allowing work to be issued on the caller's stream. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
@Conarnar
Conarnarforce-pushed the export-D109590531 branch 2 times, most recently from 3d8da75 to 07765c3CompareJune 25, 2026 17:10
CopilotAI review requested due to automatic review settings June 25, 2026 17:10
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync` and synchronize the stream before returning — preserving the blocking API contract while allowing work to be issued on the caller's stream. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 4 out of 4 changed files in this pull request and generated 5 comments.

Comment on lines +161 to +168
cudaError_t err = cudaSuccess;
const auto caller_stream = executorch::extension::cuda::getCallerStream();
if (caller_stream) {
err = cudaMemcpyAsync(
dst, src, nbytes, cudaMemcpyHostToDevice, *caller_stream);
} else {
err = cudaMemcpy(dst, src, nbytes, cudaMemcpyHostToDevice);
}
Comment on lines +144 to +150
// TODO: validate caller stream device matches index.
// For now assert single-GPU case.
ET_CHECK_OR_RETURN_ERROR(
index == -1 || index == 0,
InvalidArgument,
"CudaAllocator::copy_host_to_device only supports device 0, got %d",
static_cast<int>(index));
Comment on lines +202 to +208
// TODO: validate caller stream device matches index.
// For now assert single-GPU case.
ET_CHECK_OR_RETURN_ERROR(
index == -1 || index == 0,
InvalidArgument,
"CudaAllocator::copy_device_to_host only supports device 0, got %d",
static_cast<int>(index));
Comment on lines +78 to +90
cudaStream_t s;
ASSERT_EQ(cudaStreamCreate(&s), cudaSuccess);
executorch::extension::cuda::CallerStreamGuard g(s);

CudaAllocator& a = CudaAllocator::instance();
auto res = a.allocate(256, 0);
ASSERT_TRUE(res.ok());
void* d = res.get();
std::vector<uint8_t> h(256, 7);
// should take async branch internally, still return Ok
EXPECT_EQ(a.copy_host_to_device(d, h.data(), 256, 0), Error::Ok);
a.deallocate(d, 0);
cudaStreamDestroy(s);
Comment on lines +103 to +117
cudaStream_t s;
ASSERT_EQ(cudaStreamCreate(&s), cudaSuccess);
executorch::extension::cuda::CallerStreamGuard g(s);

CudaAllocator& a = CudaAllocator::instance();
auto res = a.allocate(256, 0);
ASSERT_TRUE(res.ok());
void* d = res.get();
std::vector<uint8_t> h_src(256, 5), h_dst(256, 0);
ASSERT_EQ(a.copy_host_to_device(d, h_src.data(), 256, 0), Error::Ok);
EXPECT_EQ(a.copy_device_to_host(h_dst.data(), d, 256, 0), Error::Ok);
EXPECT_EQ(h_src, h_dst);

a.deallocate(d, 0);
cudaStreamDestroy(s);
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 25, 2026 18:57

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 4 out of 4 changed files in this pull request and generated 5 comments.

Comment on lines +144 to +150
// TODO: validate caller stream device matches index.
// For now assert index is -1 or 0.
ET_CHECK_OR_RETURN_ERROR(
index == -1 || index == 0,
InvalidArgument,
"CudaAllocator::copy_host_to_device only supports device 0 or -1 (current), got %d",
static_cast<int>(index));
Comment on lines +204 to +210
// TODO: validate caller stream device matches index.
// For now assert index is -1 or 0.
ET_CHECK_OR_RETURN_ERROR(
index == -1 || index == 0,
InvalidArgument,
"CudaAllocator::copy_device_to_host only supports device 0 or -1 (current), got %d",
static_cast<int>(index));
Comment on lines +161 to +166
cudaError_t err = cudaSuccess;
const auto caller_stream = executorch::extension::cuda::getCallerStream();
if (caller_stream) {
err = cudaMemcpyAsync(
dst, src, nbytes, cudaMemcpyHostToDevice, *caller_stream);
// We don't synchronize the stream here because the caller is expected to
Comment on lines +223 to +228
if (caller_stream) {
err = cudaMemcpyAsync(
dst, src, nbytes, cudaMemcpyDeviceToHost, *caller_stream);
if (err == cudaSuccess) {
err = cudaStreamSynchronize(*caller_stream);
}
Comment on lines +116 to +118
ASSERT_EQ(a.copy_host_to_device(d, h_src.data(), 256, 0), Error::Ok);
EXPECT_EQ(a.copy_device_to_host(h_dst.data(), d, 256, 0), Error::Ok);
EXPECT_EQ(h_src, h_dst);
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 27, 2026 10:24
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
@Conarnar
Conarnarforce-pushed the export-D109590531 branch 2 times, most recently from 32968c0 to 056c25cCompareJune 27, 2026 10:24

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
Pull Request resolved: pytorch#20498
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 27, 2026 19:34
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 28, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 28, 2026 00:10

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 28, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 28, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
- Added CIs for unit tests.
Reviewed By: Gasoonjia
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 28, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
- Added CIs for unit tests.
Reviewed By: Gasoonjia
Differential Revision: D109590531
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
- Added CIs for unit tests.
Reviewed By: Gasoonjia
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 29, 2026 18:40

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ciflow/cudaCLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@Conarnar@Gasoonjia@shoumikhin
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Use caller CUDA stream for D2H and H2D copies (#20498) - #20498

Merged
meta-codesync[bot] merged 1 commit into
pytorch:mainfrom
Conarnar:export-D109590531
Jun 30, 2026
Merged

Use caller CUDA stream for D2H and H2D copies (#20498)#20498
meta-codesync[bot] merged 1 commit into
pytorch:mainfrom
Conarnar:export-D109590531

Conversation

@Conarnar

@ConarnarConarnar commented Jun 24, 2026

Copy link
Copy Markdown
Contributor

Summary:

CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via getCallerStream()), copy_host_to_device and copy_device_to_host use cudaMemcpyAsync. When no caller stream is set, the synchronous cudaMemcpy path is used as before.

Additionally:

  • Added null pointer and zero-byte validation — null dst/src return Error::InvalidArgument instead of aborting in cudaMemcpy, and zero-byte copies return Error::Ok early.
  • Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
  • Wired //executorch/extension/cuda:caller_stream dependency in TARGETS.
  • Added extension_cuda dependencies to CMakeLists.txt.
  • Added test_cuda_allocator with coverage for sync/async paths and error handling.
  • Added CIs for unit tests.

Reviewed By: Gasoonjia

Differential Revision: D109590531

CopilotAI review requested due to automatic review settings June 24, 2026 22:51
@pytorch-bot

pytorch-botBot commented Jun 24, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20498

Note: Links to docs will display an error until the docs builds have been completed.

✅ No Failures

As of commit d3f5462 with merge base b331ebd (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 24, 2026
@meta-codesync

Copy link
Copy Markdown
Contributor

@Conarnar has exported this pull request. If you are a Meta employee, you can view the originating Diff in D109590531.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot encountered an error and was unable to review this pull request. You can try again by re-requesting a review.

@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@meta-codesyncmeta-codesyncBot changed the title Use caller CUDA stream for D2H and H2D copiesUse caller CUDA stream for D2H and H2D copies (#20498)Jun 24, 2026
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 24, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync` and synchronize the stream before returning — preserving the blocking API contract while allowing work to be issued on the caller's stream. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
@Conarnar
Conarnarforce-pushed the export-D109590531 branch 2 times, most recently from 3d8da75 to 07765c3CompareJune 25, 2026 17:10
CopilotAI review requested due to automatic review settings June 25, 2026 17:10
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync` and synchronize the stream before returning — preserving the blocking API contract while allowing work to be issued on the caller's stream. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 4 out of 4 changed files in this pull request and generated 5 comments.

Comment on lines +161 to +168
cudaError_t err = cudaSuccess;
const auto caller_stream = executorch::extension::cuda::getCallerStream();
if (caller_stream) {
err = cudaMemcpyAsync(
dst, src, nbytes, cudaMemcpyHostToDevice, *caller_stream);
} else {
err = cudaMemcpy(dst, src, nbytes, cudaMemcpyHostToDevice);
}
Comment on lines +144 to +150
// TODO: validate caller stream device matches index.
// For now assert single-GPU case.
ET_CHECK_OR_RETURN_ERROR(
index == -1 || index == 0,
InvalidArgument,
"CudaAllocator::copy_host_to_device only supports device 0, got %d",
static_cast<int>(index));
Comment on lines +202 to +208
// TODO: validate caller stream device matches index.
// For now assert single-GPU case.
ET_CHECK_OR_RETURN_ERROR(
index == -1 || index == 0,
InvalidArgument,
"CudaAllocator::copy_device_to_host only supports device 0, got %d",
static_cast<int>(index));
Comment on lines +78 to +90
cudaStream_t s;
ASSERT_EQ(cudaStreamCreate(&s), cudaSuccess);
executorch::extension::cuda::CallerStreamGuard g(s);

CudaAllocator& a = CudaAllocator::instance();
auto res = a.allocate(256, 0);
ASSERT_TRUE(res.ok());
void* d = res.get();
std::vector<uint8_t> h(256, 7);
// should take async branch internally, still return Ok
EXPECT_EQ(a.copy_host_to_device(d, h.data(), 256, 0), Error::Ok);
a.deallocate(d, 0);
cudaStreamDestroy(s);
Comment on lines +103 to +117
cudaStream_t s;
ASSERT_EQ(cudaStreamCreate(&s), cudaSuccess);
executorch::extension::cuda::CallerStreamGuard g(s);

CudaAllocator& a = CudaAllocator::instance();
auto res = a.allocate(256, 0);
ASSERT_TRUE(res.ok());
void* d = res.get();
std::vector<uint8_t> h_src(256, 5), h_dst(256, 0);
ASSERT_EQ(a.copy_host_to_device(d, h_src.data(), 256, 0), Error::Ok);
EXPECT_EQ(a.copy_device_to_host(h_dst.data(), d, 256, 0), Error::Ok);
EXPECT_EQ(h_src, h_dst);

a.deallocate(d, 0);
cudaStreamDestroy(s);
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 25, 2026 18:57

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 4 out of 4 changed files in this pull request and generated 5 comments.

Comment on lines +144 to +150
// TODO: validate caller stream device matches index.
// For now assert index is -1 or 0.
ET_CHECK_OR_RETURN_ERROR(
index == -1 || index == 0,
InvalidArgument,
"CudaAllocator::copy_host_to_device only supports device 0 or -1 (current), got %d",
static_cast<int>(index));
Comment on lines +204 to +210
// TODO: validate caller stream device matches index.
// For now assert index is -1 or 0.
ET_CHECK_OR_RETURN_ERROR(
index == -1 || index == 0,
InvalidArgument,
"CudaAllocator::copy_device_to_host only supports device 0 or -1 (current), got %d",
static_cast<int>(index));
Comment on lines +161 to +166
cudaError_t err = cudaSuccess;
const auto caller_stream = executorch::extension::cuda::getCallerStream();
if (caller_stream) {
err = cudaMemcpyAsync(
dst, src, nbytes, cudaMemcpyHostToDevice, *caller_stream);
// We don't synchronize the stream here because the caller is expected to
Comment on lines +223 to +228
if (caller_stream) {
err = cudaMemcpyAsync(
dst, src, nbytes, cudaMemcpyDeviceToHost, *caller_stream);
if (err == cudaSuccess) {
err = cudaStreamSynchronize(*caller_stream);
}
Comment on lines +116 to +118
ASSERT_EQ(a.copy_host_to_device(d, h_src.data(), 256, 0), Error::Ok);
EXPECT_EQ(a.copy_device_to_host(h_dst.data(), d, 256, 0), Error::Ok);
EXPECT_EQ(h_src, h_dst);
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 27, 2026 10:24
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
@Conarnar
Conarnarforce-pushed the export-D109590531 branch 2 times, most recently from 32968c0 to 056c25cCompareJune 27, 2026 10:24

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
Pull Request resolved: pytorch#20498
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 27, 2026 19:34
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 28, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 28, 2026 00:10

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 28, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 28, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
- Added CIs for unit tests.
Reviewed By: Gasoonjia
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 28, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
- Added CIs for unit tests.
Reviewed By: Gasoonjia
Differential Revision: D109590531
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
- Added CIs for unit tests.
Reviewed By: Gasoonjia
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 29, 2026 18:40

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ciflow/cudaCLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@Conarnar@Gasoonjia@shoumikhin
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Use caller CUDA stream for D2H and H2D copies (#20498) - #20498

Merged
meta-codesync[bot] merged 1 commit into
pytorch:mainfrom
Conarnar:export-D109590531
Jun 30, 2026
Merged

Use caller CUDA stream for D2H and H2D copies (#20498)#20498
meta-codesync[bot] merged 1 commit into
pytorch:mainfrom
Conarnar:export-D109590531

Conversation

@Conarnar

@ConarnarConarnar commented Jun 24, 2026

Copy link
Copy Markdown
Contributor

Summary:

CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via getCallerStream()), copy_host_to_device and copy_device_to_host use cudaMemcpyAsync. When no caller stream is set, the synchronous cudaMemcpy path is used as before.

Additionally:

  • Added null pointer and zero-byte validation — null dst/src return Error::InvalidArgument instead of aborting in cudaMemcpy, and zero-byte copies return Error::Ok early.
  • Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
  • Wired //executorch/extension/cuda:caller_stream dependency in TARGETS.
  • Added extension_cuda dependencies to CMakeLists.txt.
  • Added test_cuda_allocator with coverage for sync/async paths and error handling.
  • Added CIs for unit tests.

Reviewed By: Gasoonjia

Differential Revision: D109590531

CopilotAI review requested due to automatic review settings June 24, 2026 22:51
@pytorch-bot

pytorch-botBot commented Jun 24, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20498

Note: Links to docs will display an error until the docs builds have been completed.

✅ No Failures

As of commit d3f5462 with merge base b331ebd (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 24, 2026
@meta-codesync

Copy link
Copy Markdown
Contributor

@Conarnar has exported this pull request. If you are a Meta employee, you can view the originating Diff in D109590531.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot encountered an error and was unable to review this pull request. You can try again by re-requesting a review.

@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@meta-codesyncmeta-codesyncBot changed the title Use caller CUDA stream for D2H and H2D copiesUse caller CUDA stream for D2H and H2D copies (#20498)Jun 24, 2026
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 24, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync` and synchronize the stream before returning — preserving the blocking API contract while allowing work to be issued on the caller's stream. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
@Conarnar
Conarnarforce-pushed the export-D109590531 branch 2 times, most recently from 3d8da75 to 07765c3CompareJune 25, 2026 17:10
CopilotAI review requested due to automatic review settings June 25, 2026 17:10
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync` and synchronize the stream before returning — preserving the blocking API contract while allowing work to be issued on the caller's stream. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 4 out of 4 changed files in this pull request and generated 5 comments.

Comment on lines +161 to +168
cudaError_t err = cudaSuccess;
const auto caller_stream = executorch::extension::cuda::getCallerStream();
if (caller_stream) {
err = cudaMemcpyAsync(
dst, src, nbytes, cudaMemcpyHostToDevice, *caller_stream);
} else {
err = cudaMemcpy(dst, src, nbytes, cudaMemcpyHostToDevice);
}
Comment on lines +144 to +150
// TODO: validate caller stream device matches index.
// For now assert single-GPU case.
ET_CHECK_OR_RETURN_ERROR(
index == -1 || index == 0,
InvalidArgument,
"CudaAllocator::copy_host_to_device only supports device 0, got %d",
static_cast<int>(index));
Comment on lines +202 to +208
// TODO: validate caller stream device matches index.
// For now assert single-GPU case.
ET_CHECK_OR_RETURN_ERROR(
index == -1 || index == 0,
InvalidArgument,
"CudaAllocator::copy_device_to_host only supports device 0, got %d",
static_cast<int>(index));
Comment on lines +78 to +90
cudaStream_t s;
ASSERT_EQ(cudaStreamCreate(&s), cudaSuccess);
executorch::extension::cuda::CallerStreamGuard g(s);

CudaAllocator& a = CudaAllocator::instance();
auto res = a.allocate(256, 0);
ASSERT_TRUE(res.ok());
void* d = res.get();
std::vector<uint8_t> h(256, 7);
// should take async branch internally, still return Ok
EXPECT_EQ(a.copy_host_to_device(d, h.data(), 256, 0), Error::Ok);
a.deallocate(d, 0);
cudaStreamDestroy(s);
Comment on lines +103 to +117
cudaStream_t s;
ASSERT_EQ(cudaStreamCreate(&s), cudaSuccess);
executorch::extension::cuda::CallerStreamGuard g(s);

CudaAllocator& a = CudaAllocator::instance();
auto res = a.allocate(256, 0);
ASSERT_TRUE(res.ok());
void* d = res.get();
std::vector<uint8_t> h_src(256, 5), h_dst(256, 0);
ASSERT_EQ(a.copy_host_to_device(d, h_src.data(), 256, 0), Error::Ok);
EXPECT_EQ(a.copy_device_to_host(h_dst.data(), d, 256, 0), Error::Ok);
EXPECT_EQ(h_src, h_dst);

a.deallocate(d, 0);
cudaStreamDestroy(s);
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 25, 2026 18:57

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 4 out of 4 changed files in this pull request and generated 5 comments.

Comment on lines +144 to +150
// TODO: validate caller stream device matches index.
// For now assert index is -1 or 0.
ET_CHECK_OR_RETURN_ERROR(
index == -1 || index == 0,
InvalidArgument,
"CudaAllocator::copy_host_to_device only supports device 0 or -1 (current), got %d",
static_cast<int>(index));
Comment on lines +204 to +210
// TODO: validate caller stream device matches index.
// For now assert index is -1 or 0.
ET_CHECK_OR_RETURN_ERROR(
index == -1 || index == 0,
InvalidArgument,
"CudaAllocator::copy_device_to_host only supports device 0 or -1 (current), got %d",
static_cast<int>(index));
Comment on lines +161 to +166
cudaError_t err = cudaSuccess;
const auto caller_stream = executorch::extension::cuda::getCallerStream();
if (caller_stream) {
err = cudaMemcpyAsync(
dst, src, nbytes, cudaMemcpyHostToDevice, *caller_stream);
// We don't synchronize the stream here because the caller is expected to
Comment on lines +223 to +228
if (caller_stream) {
err = cudaMemcpyAsync(
dst, src, nbytes, cudaMemcpyDeviceToHost, *caller_stream);
if (err == cudaSuccess) {
err = cudaStreamSynchronize(*caller_stream);
}
Comment on lines +116 to +118
ASSERT_EQ(a.copy_host_to_device(d, h_src.data(), 256, 0), Error::Ok);
EXPECT_EQ(a.copy_device_to_host(h_dst.data(), d, 256, 0), Error::Ok);
EXPECT_EQ(h_src, h_dst);
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 27, 2026 10:24
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
@Conarnar
Conarnarforce-pushed the export-D109590531 branch 2 times, most recently from 32968c0 to 056c25cCompareJune 27, 2026 10:24

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
Pull Request resolved: pytorch#20498
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 27, 2026 19:34
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 28, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 28, 2026 00:10

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 28, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 28, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
- Added CIs for unit tests.
Reviewed By: Gasoonjia
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 28, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
- Added CIs for unit tests.
Reviewed By: Gasoonjia
Differential Revision: D109590531
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
- Added CIs for unit tests.
Reviewed By: Gasoonjia
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 29, 2026 18:40

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ciflow/cudaCLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@Conarnar@Gasoonjia@shoumikhin
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Use caller CUDA stream for D2H and H2D copies (#20498) - #20498

Merged
meta-codesync[bot] merged 1 commit into
pytorch:mainfrom
Conarnar:export-D109590531
Jun 30, 2026
Merged

Use caller CUDA stream for D2H and H2D copies (#20498)#20498
meta-codesync[bot] merged 1 commit into
pytorch:mainfrom
Conarnar:export-D109590531

Conversation

@Conarnar

@ConarnarConarnar commented Jun 24, 2026

Copy link
Copy Markdown
Contributor

Summary:

CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via getCallerStream()), copy_host_to_device and copy_device_to_host use cudaMemcpyAsync. When no caller stream is set, the synchronous cudaMemcpy path is used as before.

Additionally:

  • Added null pointer and zero-byte validation — null dst/src return Error::InvalidArgument instead of aborting in cudaMemcpy, and zero-byte copies return Error::Ok early.
  • Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
  • Wired //executorch/extension/cuda:caller_stream dependency in TARGETS.
  • Added extension_cuda dependencies to CMakeLists.txt.
  • Added test_cuda_allocator with coverage for sync/async paths and error handling.
  • Added CIs for unit tests.

Reviewed By: Gasoonjia

Differential Revision: D109590531

CopilotAI review requested due to automatic review settings June 24, 2026 22:51
@pytorch-bot

pytorch-botBot commented Jun 24, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20498

Note: Links to docs will display an error until the docs builds have been completed.

✅ No Failures

As of commit d3f5462 with merge base b331ebd (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 24, 2026
@meta-codesync

Copy link
Copy Markdown
Contributor

@Conarnar has exported this pull request. If you are a Meta employee, you can view the originating Diff in D109590531.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot encountered an error and was unable to review this pull request. You can try again by re-requesting a review.

@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@meta-codesyncmeta-codesyncBot changed the title Use caller CUDA stream for D2H and H2D copiesUse caller CUDA stream for D2H and H2D copies (#20498)Jun 24, 2026
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 24, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync` and synchronize the stream before returning — preserving the blocking API contract while allowing work to be issued on the caller's stream. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
@Conarnar
Conarnarforce-pushed the export-D109590531 branch 2 times, most recently from 3d8da75 to 07765c3CompareJune 25, 2026 17:10
CopilotAI review requested due to automatic review settings June 25, 2026 17:10
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync` and synchronize the stream before returning — preserving the blocking API contract while allowing work to be issued on the caller's stream. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 4 out of 4 changed files in this pull request and generated 5 comments.

Comment on lines +161 to +168
cudaError_t err = cudaSuccess;
const auto caller_stream = executorch::extension::cuda::getCallerStream();
if (caller_stream) {
err = cudaMemcpyAsync(
dst, src, nbytes, cudaMemcpyHostToDevice, *caller_stream);
} else {
err = cudaMemcpy(dst, src, nbytes, cudaMemcpyHostToDevice);
}
Comment on lines +144 to +150
// TODO: validate caller stream device matches index.
// For now assert single-GPU case.
ET_CHECK_OR_RETURN_ERROR(
index == -1 || index == 0,
InvalidArgument,
"CudaAllocator::copy_host_to_device only supports device 0, got %d",
static_cast<int>(index));
Comment on lines +202 to +208
// TODO: validate caller stream device matches index.
// For now assert single-GPU case.
ET_CHECK_OR_RETURN_ERROR(
index == -1 || index == 0,
InvalidArgument,
"CudaAllocator::copy_device_to_host only supports device 0, got %d",
static_cast<int>(index));
Comment on lines +78 to +90
cudaStream_t s;
ASSERT_EQ(cudaStreamCreate(&s), cudaSuccess);
executorch::extension::cuda::CallerStreamGuard g(s);

CudaAllocator& a = CudaAllocator::instance();
auto res = a.allocate(256, 0);
ASSERT_TRUE(res.ok());
void* d = res.get();
std::vector<uint8_t> h(256, 7);
// should take async branch internally, still return Ok
EXPECT_EQ(a.copy_host_to_device(d, h.data(), 256, 0), Error::Ok);
a.deallocate(d, 0);
cudaStreamDestroy(s);
Comment on lines +103 to +117
cudaStream_t s;
ASSERT_EQ(cudaStreamCreate(&s), cudaSuccess);
executorch::extension::cuda::CallerStreamGuard g(s);

CudaAllocator& a = CudaAllocator::instance();
auto res = a.allocate(256, 0);
ASSERT_TRUE(res.ok());
void* d = res.get();
std::vector<uint8_t> h_src(256, 5), h_dst(256, 0);
ASSERT_EQ(a.copy_host_to_device(d, h_src.data(), 256, 0), Error::Ok);
EXPECT_EQ(a.copy_device_to_host(h_dst.data(), d, 256, 0), Error::Ok);
EXPECT_EQ(h_src, h_dst);

a.deallocate(d, 0);
cudaStreamDestroy(s);
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 25, 2026 18:57

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 4 out of 4 changed files in this pull request and generated 5 comments.

Comment on lines +144 to +150
// TODO: validate caller stream device matches index.
// For now assert index is -1 or 0.
ET_CHECK_OR_RETURN_ERROR(
index == -1 || index == 0,
InvalidArgument,
"CudaAllocator::copy_host_to_device only supports device 0 or -1 (current), got %d",
static_cast<int>(index));
Comment on lines +204 to +210
// TODO: validate caller stream device matches index.
// For now assert index is -1 or 0.
ET_CHECK_OR_RETURN_ERROR(
index == -1 || index == 0,
InvalidArgument,
"CudaAllocator::copy_device_to_host only supports device 0 or -1 (current), got %d",
static_cast<int>(index));
Comment on lines +161 to +166
cudaError_t err = cudaSuccess;
const auto caller_stream = executorch::extension::cuda::getCallerStream();
if (caller_stream) {
err = cudaMemcpyAsync(
dst, src, nbytes, cudaMemcpyHostToDevice, *caller_stream);
// We don't synchronize the stream here because the caller is expected to
Comment on lines +223 to +228
if (caller_stream) {
err = cudaMemcpyAsync(
dst, src, nbytes, cudaMemcpyDeviceToHost, *caller_stream);
if (err == cudaSuccess) {
err = cudaStreamSynchronize(*caller_stream);
}
Comment on lines +116 to +118
ASSERT_EQ(a.copy_host_to_device(d, h_src.data(), 256, 0), Error::Ok);
EXPECT_EQ(a.copy_device_to_host(h_dst.data(), d, 256, 0), Error::Ok);
EXPECT_EQ(h_src, h_dst);
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 27, 2026 10:24
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
@Conarnar
Conarnarforce-pushed the export-D109590531 branch 2 times, most recently from 32968c0 to 056c25cCompareJune 27, 2026 10:24

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
Pull Request resolved: pytorch#20498
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 27, 2026 19:34
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 28, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 28, 2026 00:10

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 28, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 28, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
- Added CIs for unit tests.
Reviewed By: Gasoonjia
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 28, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
- Added CIs for unit tests.
Reviewed By: Gasoonjia
Differential Revision: D109590531
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
- Added CIs for unit tests.
Reviewed By: Gasoonjia
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 29, 2026 18:40

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ciflow/cudaCLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@Conarnar@Gasoonjia@shoumikhin
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Use caller CUDA stream for D2H and H2D copies (#20498) - #20498

Merged
meta-codesync[bot] merged 1 commit into
pytorch:mainfrom
Conarnar:export-D109590531
Jun 30, 2026
Merged

Use caller CUDA stream for D2H and H2D copies (#20498)#20498
meta-codesync[bot] merged 1 commit into
pytorch:mainfrom
Conarnar:export-D109590531

Conversation

@Conarnar

@ConarnarConarnar commented Jun 24, 2026

Copy link
Copy Markdown
Contributor

Summary:

CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via getCallerStream()), copy_host_to_device and copy_device_to_host use cudaMemcpyAsync. When no caller stream is set, the synchronous cudaMemcpy path is used as before.

Additionally:

  • Added null pointer and zero-byte validation — null dst/src return Error::InvalidArgument instead of aborting in cudaMemcpy, and zero-byte copies return Error::Ok early.
  • Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
  • Wired //executorch/extension/cuda:caller_stream dependency in TARGETS.
  • Added extension_cuda dependencies to CMakeLists.txt.
  • Added test_cuda_allocator with coverage for sync/async paths and error handling.
  • Added CIs for unit tests.

Reviewed By: Gasoonjia

Differential Revision: D109590531

CopilotAI review requested due to automatic review settings June 24, 2026 22:51
@pytorch-bot

pytorch-botBot commented Jun 24, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20498

Note: Links to docs will display an error until the docs builds have been completed.

✅ No Failures

As of commit d3f5462 with merge base b331ebd (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 24, 2026
@meta-codesync

Copy link
Copy Markdown
Contributor

@Conarnar has exported this pull request. If you are a Meta employee, you can view the originating Diff in D109590531.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot encountered an error and was unable to review this pull request. You can try again by re-requesting a review.

@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@meta-codesyncmeta-codesyncBot changed the title Use caller CUDA stream for D2H and H2D copiesUse caller CUDA stream for D2H and H2D copies (#20498)Jun 24, 2026
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 24, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync` and synchronize the stream before returning — preserving the blocking API contract while allowing work to be issued on the caller's stream. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
@Conarnar
Conarnarforce-pushed the export-D109590531 branch 2 times, most recently from 3d8da75 to 07765c3CompareJune 25, 2026 17:10
CopilotAI review requested due to automatic review settings June 25, 2026 17:10
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync` and synchronize the stream before returning — preserving the blocking API contract while allowing work to be issued on the caller's stream. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 4 out of 4 changed files in this pull request and generated 5 comments.

Comment on lines +161 to +168
cudaError_t err = cudaSuccess;
const auto caller_stream = executorch::extension::cuda::getCallerStream();
if (caller_stream) {
err = cudaMemcpyAsync(
dst, src, nbytes, cudaMemcpyHostToDevice, *caller_stream);
} else {
err = cudaMemcpy(dst, src, nbytes, cudaMemcpyHostToDevice);
}
Comment on lines +144 to +150
// TODO: validate caller stream device matches index.
// For now assert single-GPU case.
ET_CHECK_OR_RETURN_ERROR(
index == -1 || index == 0,
InvalidArgument,
"CudaAllocator::copy_host_to_device only supports device 0, got %d",
static_cast<int>(index));
Comment on lines +202 to +208
// TODO: validate caller stream device matches index.
// For now assert single-GPU case.
ET_CHECK_OR_RETURN_ERROR(
index == -1 || index == 0,
InvalidArgument,
"CudaAllocator::copy_device_to_host only supports device 0, got %d",
static_cast<int>(index));
Comment on lines +78 to +90
cudaStream_t s;
ASSERT_EQ(cudaStreamCreate(&s), cudaSuccess);
executorch::extension::cuda::CallerStreamGuard g(s);

CudaAllocator& a = CudaAllocator::instance();
auto res = a.allocate(256, 0);
ASSERT_TRUE(res.ok());
void* d = res.get();
std::vector<uint8_t> h(256, 7);
// should take async branch internally, still return Ok
EXPECT_EQ(a.copy_host_to_device(d, h.data(), 256, 0), Error::Ok);
a.deallocate(d, 0);
cudaStreamDestroy(s);
Comment on lines +103 to +117
cudaStream_t s;
ASSERT_EQ(cudaStreamCreate(&s), cudaSuccess);
executorch::extension::cuda::CallerStreamGuard g(s);

CudaAllocator& a = CudaAllocator::instance();
auto res = a.allocate(256, 0);
ASSERT_TRUE(res.ok());
void* d = res.get();
std::vector<uint8_t> h_src(256, 5), h_dst(256, 0);
ASSERT_EQ(a.copy_host_to_device(d, h_src.data(), 256, 0), Error::Ok);
EXPECT_EQ(a.copy_device_to_host(h_dst.data(), d, 256, 0), Error::Ok);
EXPECT_EQ(h_src, h_dst);

a.deallocate(d, 0);
cudaStreamDestroy(s);
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 25, 2026 18:57

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 4 out of 4 changed files in this pull request and generated 5 comments.

Comment on lines +144 to +150
// TODO: validate caller stream device matches index.
// For now assert index is -1 or 0.
ET_CHECK_OR_RETURN_ERROR(
index == -1 || index == 0,
InvalidArgument,
"CudaAllocator::copy_host_to_device only supports device 0 or -1 (current), got %d",
static_cast<int>(index));
Comment on lines +204 to +210
// TODO: validate caller stream device matches index.
// For now assert index is -1 or 0.
ET_CHECK_OR_RETURN_ERROR(
index == -1 || index == 0,
InvalidArgument,
"CudaAllocator::copy_device_to_host only supports device 0 or -1 (current), got %d",
static_cast<int>(index));
Comment on lines +161 to +166
cudaError_t err = cudaSuccess;
const auto caller_stream = executorch::extension::cuda::getCallerStream();
if (caller_stream) {
err = cudaMemcpyAsync(
dst, src, nbytes, cudaMemcpyHostToDevice, *caller_stream);
// We don't synchronize the stream here because the caller is expected to
Comment on lines +223 to +228
if (caller_stream) {
err = cudaMemcpyAsync(
dst, src, nbytes, cudaMemcpyDeviceToHost, *caller_stream);
if (err == cudaSuccess) {
err = cudaStreamSynchronize(*caller_stream);
}
Comment on lines +116 to +118
ASSERT_EQ(a.copy_host_to_device(d, h_src.data(), 256, 0), Error::Ok);
EXPECT_EQ(a.copy_device_to_host(h_dst.data(), d, 256, 0), Error::Ok);
EXPECT_EQ(h_src, h_dst);
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 27, 2026 10:24
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
@Conarnar
Conarnarforce-pushed the export-D109590531 branch 2 times, most recently from 32968c0 to 056c25cCompareJune 27, 2026 10:24

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
Pull Request resolved: pytorch#20498
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 27, 2026 19:34
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 28, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 28, 2026 00:10

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 28, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 28, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
- Added CIs for unit tests.
Reviewed By: Gasoonjia
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 28, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
- Added CIs for unit tests.
Reviewed By: Gasoonjia
Differential Revision: D109590531
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
- Added CIs for unit tests.
Reviewed By: Gasoonjia
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 29, 2026 18:40

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ciflow/cudaCLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@Conarnar@Gasoonjia@shoumikhin
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Use caller CUDA stream for D2H and H2D copies (#20498) - #20498

Merged
meta-codesync[bot] merged 1 commit into
pytorch:mainfrom
Conarnar:export-D109590531
Jun 30, 2026
Merged

Use caller CUDA stream for D2H and H2D copies (#20498)#20498
meta-codesync[bot] merged 1 commit into
pytorch:mainfrom
Conarnar:export-D109590531

Conversation

@Conarnar

@ConarnarConarnar commented Jun 24, 2026

Copy link
Copy Markdown
Contributor

Summary:

CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via getCallerStream()), copy_host_to_device and copy_device_to_host use cudaMemcpyAsync. When no caller stream is set, the synchronous cudaMemcpy path is used as before.

Additionally:

  • Added null pointer and zero-byte validation — null dst/src return Error::InvalidArgument instead of aborting in cudaMemcpy, and zero-byte copies return Error::Ok early.
  • Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
  • Wired //executorch/extension/cuda:caller_stream dependency in TARGETS.
  • Added extension_cuda dependencies to CMakeLists.txt.
  • Added test_cuda_allocator with coverage for sync/async paths and error handling.
  • Added CIs for unit tests.

Reviewed By: Gasoonjia

Differential Revision: D109590531

CopilotAI review requested due to automatic review settings June 24, 2026 22:51
@pytorch-bot

pytorch-botBot commented Jun 24, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20498

Note: Links to docs will display an error until the docs builds have been completed.

✅ No Failures

As of commit d3f5462 with merge base b331ebd (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 24, 2026
@meta-codesync

Copy link
Copy Markdown
Contributor

@Conarnar has exported this pull request. If you are a Meta employee, you can view the originating Diff in D109590531.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot encountered an error and was unable to review this pull request. You can try again by re-requesting a review.

@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@meta-codesyncmeta-codesyncBot changed the title Use caller CUDA stream for D2H and H2D copiesUse caller CUDA stream for D2H and H2D copies (#20498)Jun 24, 2026
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 24, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync` and synchronize the stream before returning — preserving the blocking API contract while allowing work to be issued on the caller's stream. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
@Conarnar
Conarnarforce-pushed the export-D109590531 branch 2 times, most recently from 3d8da75 to 07765c3CompareJune 25, 2026 17:10
CopilotAI review requested due to automatic review settings June 25, 2026 17:10
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync` and synchronize the stream before returning — preserving the blocking API contract while allowing work to be issued on the caller's stream. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 4 out of 4 changed files in this pull request and generated 5 comments.

Comment on lines +161 to +168
cudaError_t err = cudaSuccess;
const auto caller_stream = executorch::extension::cuda::getCallerStream();
if (caller_stream) {
err = cudaMemcpyAsync(
dst, src, nbytes, cudaMemcpyHostToDevice, *caller_stream);
} else {
err = cudaMemcpy(dst, src, nbytes, cudaMemcpyHostToDevice);
}
Comment on lines +144 to +150
// TODO: validate caller stream device matches index.
// For now assert single-GPU case.
ET_CHECK_OR_RETURN_ERROR(
index == -1 || index == 0,
InvalidArgument,
"CudaAllocator::copy_host_to_device only supports device 0, got %d",
static_cast<int>(index));
Comment on lines +202 to +208
// TODO: validate caller stream device matches index.
// For now assert single-GPU case.
ET_CHECK_OR_RETURN_ERROR(
index == -1 || index == 0,
InvalidArgument,
"CudaAllocator::copy_device_to_host only supports device 0, got %d",
static_cast<int>(index));
Comment on lines +78 to +90
cudaStream_t s;
ASSERT_EQ(cudaStreamCreate(&s), cudaSuccess);
executorch::extension::cuda::CallerStreamGuard g(s);

CudaAllocator& a = CudaAllocator::instance();
auto res = a.allocate(256, 0);
ASSERT_TRUE(res.ok());
void* d = res.get();
std::vector<uint8_t> h(256, 7);
// should take async branch internally, still return Ok
EXPECT_EQ(a.copy_host_to_device(d, h.data(), 256, 0), Error::Ok);
a.deallocate(d, 0);
cudaStreamDestroy(s);
Comment on lines +103 to +117
cudaStream_t s;
ASSERT_EQ(cudaStreamCreate(&s), cudaSuccess);
executorch::extension::cuda::CallerStreamGuard g(s);

CudaAllocator& a = CudaAllocator::instance();
auto res = a.allocate(256, 0);
ASSERT_TRUE(res.ok());
void* d = res.get();
std::vector<uint8_t> h_src(256, 5), h_dst(256, 0);
ASSERT_EQ(a.copy_host_to_device(d, h_src.data(), 256, 0), Error::Ok);
EXPECT_EQ(a.copy_device_to_host(h_dst.data(), d, 256, 0), Error::Ok);
EXPECT_EQ(h_src, h_dst);

a.deallocate(d, 0);
cudaStreamDestroy(s);
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 25, 2026 18:57

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 4 out of 4 changed files in this pull request and generated 5 comments.

Comment on lines +144 to +150
// TODO: validate caller stream device matches index.
// For now assert index is -1 or 0.
ET_CHECK_OR_RETURN_ERROR(
index == -1 || index == 0,
InvalidArgument,
"CudaAllocator::copy_host_to_device only supports device 0 or -1 (current), got %d",
static_cast<int>(index));
Comment on lines +204 to +210
// TODO: validate caller stream device matches index.
// For now assert index is -1 or 0.
ET_CHECK_OR_RETURN_ERROR(
index == -1 || index == 0,
InvalidArgument,
"CudaAllocator::copy_device_to_host only supports device 0 or -1 (current), got %d",
static_cast<int>(index));
Comment on lines +161 to +166
cudaError_t err = cudaSuccess;
const auto caller_stream = executorch::extension::cuda::getCallerStream();
if (caller_stream) {
err = cudaMemcpyAsync(
dst, src, nbytes, cudaMemcpyHostToDevice, *caller_stream);
// We don't synchronize the stream here because the caller is expected to
Comment on lines +223 to +228
if (caller_stream) {
err = cudaMemcpyAsync(
dst, src, nbytes, cudaMemcpyDeviceToHost, *caller_stream);
if (err == cudaSuccess) {
err = cudaStreamSynchronize(*caller_stream);
}
Comment on lines +116 to +118
ASSERT_EQ(a.copy_host_to_device(d, h_src.data(), 256, 0), Error::Ok);
EXPECT_EQ(a.copy_device_to_host(h_dst.data(), d, 256, 0), Error::Ok);
EXPECT_EQ(h_src, h_dst);
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 25, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 27, 2026 10:24
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
@Conarnar
Conarnarforce-pushed the export-D109590531 branch 2 times, most recently from 32968c0 to 056c25cCompareJune 27, 2026 10:24

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
Pull Request resolved: pytorch#20498
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 27, 2026 19:34
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 27, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 28, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 28, 2026 00:10

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 28, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
Reviewed By: Gasoonjia
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 28, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
- Added CIs for unit tests.
Reviewed By: Gasoonjia
Differential Revision: D109590531
Conarnar added a commit to Conarnar/executorch that referenced this pull request Jun 28, 2026
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
- Added CIs for unit tests.
Reviewed By: Gasoonjia
Differential Revision: D109590531
Summary:
CudaAllocator memory copies now support async copy on a caller-provided CUDA stream. When a caller stream is available (via `getCallerStream()`), `copy_host_to_device` and `copy_device_to_host` use `cudaMemcpyAsync`. When no caller stream is set, the synchronous `cudaMemcpy` path is used as before.
Additionally:
- Added null pointer and zero-byte validation — null `dst`/`src` return `Error::InvalidArgument` instead of aborting in `cudaMemcpy`, and zero-byte copies return `Error::Ok` early.
- Assert single-GPU case (index 0 or -1) until multi-GPU stream validation is added.
- Wired `//executorch/extension/cuda:caller_stream` dependency in TARGETS.
- Added `extension_cuda` dependencies to CMakeLists.txt.
- Added `test_cuda_allocator` with coverage for sync/async paths and error handling.
- Added CIs for unit tests.
Reviewed By: Gasoonjia
Differential Revision: D109590531
CopilotAI review requested due to automatic review settings June 29, 2026 18:40

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ciflow/cudaCLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@Conarnar@Gasoonjia@shoumikhin