Skip to content

Adopt CUDA stream compatibility accessors - #3136

Open
bdice wants to merge 11 commits into
NVIDIA:mainfrom
bdice:cuda-stream-ref-prep
Open

Adopt CUDA stream compatibility accessors#3136
bdice wants to merge 11 commits into
NVIDIA:mainfrom
bdice:cuda-stream-ref-prep

Conversation

@bdice

@bdicebdice commented Sep 5, 2026

Copy link
Copy Markdown
Contributor

Summary

This PR helps RAFT adopt cuda::stream_ref. It is a non-breaking subset of the changes in #3129 to ease migration.

Use the get() and sync() compatibility aliases added in RMM #2537. These spellings are shared by rmm::cuda_stream_view and cuda::stream_ref.

This preserves existing stream types and public APIs while extracting mechanical accessor updates from the broader stream migration. It is independently buildable without RMM #2372 and leaves the migration PR focused on actual type and signature changes.

This updates raw CUDA, library, kernel-launch, and legacy API boundaries throughout RAFT and uses the CUDA stream umbrella header where stream definitions are required. Existing rmm::cuda_stream_view return types remain unchanged; their migration stays in RAFT #3129.

@copy-pr-bot

Copy link
Copy Markdown

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

@bdice
bdiceforce-pushed the cuda-stream-ref-prep branch from 72b7fbd to 5fbe443CompareSeptember 5, 2026 02:43
@bdice
bdice marked this pull request as ready for review September 5, 2026 20:02
@bdice
bdice requested a review from a team as a code ownerSeptember 5, 2026 20:02
@bdicebdice added the improvement Improvement / enhancement to an existing function label Sep 5, 2026
@coderabbitai

coderabbitaiBot commented Sep 5, 2026

Copy link
Copy Markdown

Review Change Stack

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: ed911b60-4c07-4b19-a216-6049bdc42f5e

📥 Commits

Reviewing files that changed from the base of the PR and between 51a2375 and e40580b.

📒 Files selected for processing (1)
  • cpp/bench/prims/util/fast_int_div.cu

Included review availability: Your plan provides up to 12 included reviews per hour; 9 remain after this review.


📝 Summary

Summary by CodeRabbit

  • Bug Fixes

    • Improved CUDA stream interoperability across communication, linear algebra, matrix, random, sparse, spectral, and statistics operations.
    • Ensured asynchronous GPU operations consistently receive valid native stream handles.
    • Improved stream synchronization reliability across GPU operations and benchmarks.
  • Chores

    • Replaced deprecated CUDA stream headers with the current stream interface.
    • Updated validation, test, and benchmark coverage without changing public APIs or numerical behavior.

Walkthrough

Changes

CUDA stream handle migration

Layer / File(s)Summary
Production stream handle conversion
cpp/include/raft/**, cpp/src/**
Production code now passes raw cudaStream_t values from stream wrappers through .get(). Deprecated cuda/stream_ref includes and .value() access are replaced.
Test and benchmark stream conversion
cpp/tests/**, cpp/bench/**
Tests and benchmarks now pass raw CUDA stream handles to CUDA, RAFT, RMM, and library APIs. Updated synchronization calls use sync().

Estimated code review effort: 3 (Moderate) | ~20 minutes

Merge Risk:⚪ Minimal · up to e4058

This updates benchmark stream accessor calls without changing public APIs or stream ordering. No current merge-blocking risk is identified.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check nameStatusExplanationResolution
Docstring Coverage⚠️ WarningDocstring coverage is 26.73% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 101 functions across 66 files.Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check nameStatusExplanation
Linked Issues check✅ PassedCheck skipped because no linked issues were found for this pull request.
Out of Scope Changes check✅ PassedCheck skipped because no linked issues were found for this pull request.
Title check✅ PassedThe title clearly summarizes the main change: adopting CUDA stream compatibility accessors.
Description check✅ PassedThe description directly explains the stream accessor and synchronization updates, compatibility goals, preserved APIs, and migration scope.
  • Fix all pre-merge checks with AI
✨ Finishing Touches 💡 1
🛠️ Fix failing CI checks 💡
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Comment @coderabbitai help to get the list of available commands.

@bdicebdice added the non-breaking Non-breaking change label Sep 5, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

improvementImprovement / enhancement to an existing functionnon-breakingNon-breaking change

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@bdice