Uh oh!
There was an error while loading. Please reload this page.
feat: add runtime backend API foundation - #14
Conversation
31b8520 to
faef619Comparefaef619 to
4a1d37cCompareUh oh!
There was an error while loading. Please reload this page.
* feat!: align runtime API and add runtime dispatch (#11) * Align runtime API with generated wrappers * Add default runtime dispatch specialization * Refactor runtime dispatch namespace * Use Abseil status for runtime device API * Revert "Use Abseil status for runtime device API" This reverts commit a26ddff. * Address runtime dispatch review feedback * Keep runtime API list in generator * Add TensorView constructor guard test * Align runtime memcpy kind constants with CUDA API * Use CUDA-style runtime memcpy constants * Use CUDA-style runtime memcpy constants * Move TensorView tests back into core test * Remove standalone TensorView test target * Remove standalone TensorView test file * Use fully qualified runtime API names in README * style: format runtime dispatch test * feat: refactor InfiniCore CPU runtime to InfiniRT (#8) Co-authored-by: Jiacheng Huang <huangjiacheng0709@outlook.com> * feat: add platform-adaptive runtime tests (#15) * feat: add runtime backend API foundation (#14) --------- Co-authored-by: spike-zhu <74974704+spike-zhu@users.noreply.github.com>
* feat!: align runtime API and add runtime dispatch (#11) * Align runtime API with generated wrappers * Add default runtime dispatch specialization * Refactor runtime dispatch namespace * Use Abseil status for runtime device API * Revert "Use Abseil status for runtime device API" This reverts commit a26ddff. * Address runtime dispatch review feedback * Keep runtime API list in generator * Add TensorView constructor guard test * Align runtime memcpy kind constants with CUDA API * Use CUDA-style runtime memcpy constants * Use CUDA-style runtime memcpy constants * Move TensorView tests back into core test * Remove standalone TensorView test target * Remove standalone TensorView test file * Use fully qualified runtime API names in README * style: format runtime dispatch test * feat: refactor InfiniCore CPU runtime to InfiniRT (#8) Co-authored-by: Jiacheng Huang <huangjiacheng0709@outlook.com> * feat: add platform-adaptive runtime tests (#15) * feat: add runtime backend API foundation (#14) --------- Co-authored-by: spike-zhu <74974704+spike-zhu@users.noreply.github.com>
* feat!: align runtime API and add runtime dispatch (#11) * Align runtime API with generated wrappers * Add default runtime dispatch specialization * Refactor runtime dispatch namespace * Use Abseil status for runtime device API * Revert "Use Abseil status for runtime device API" This reverts commit a26ddff. * Address runtime dispatch review feedback * Keep runtime API list in generator * Add TensorView constructor guard test * Align runtime memcpy kind constants with CUDA API * Use CUDA-style runtime memcpy constants * Use CUDA-style runtime memcpy constants * Move TensorView tests back into core test * Remove standalone TensorView test target * Remove standalone TensorView test file * Use fully qualified runtime API names in README * style: format runtime dispatch test * feat: refactor InfiniCore CPU runtime to InfiniRT (#8) Co-authored-by: Jiacheng Huang <huangjiacheng0709@outlook.com> * feat: add platform-adaptive runtime tests (#15) * feat: add runtime backend API foundation (#14) --------- Co-authored-by: spike-zhu <74974704+spike-zhu@users.noreply.github.com>
* feat: add graph runtime api * feat: add graph c api * fix: align graph runtime API with runtime namespace * feat!: align runtime API and add runtime dispatch (#11) * Align runtime API with generated wrappers * Add default runtime dispatch specialization * Refactor runtime dispatch namespace * Use Abseil status for runtime device API * Revert "Use Abseil status for runtime device API" This reverts commit a26ddff. * Address runtime dispatch review feedback * Keep runtime API list in generator * Add TensorView constructor guard test * Align runtime memcpy kind constants with CUDA API * Use CUDA-style runtime memcpy constants * Use CUDA-style runtime memcpy constants * Move TensorView tests back into core test * Remove standalone TensorView test target * Remove standalone TensorView test file * Use fully qualified runtime API names in README * style: format runtime dispatch test * feat: refactor InfiniCore CPU runtime to InfiniRT (#8) Co-authored-by: Jiacheng Huang <huangjiacheng0709@outlook.com> * feat: add platform-adaptive runtime tests (#15) * feat: add runtime backend API foundation (#14) --------- Co-authored-by: spike-zhu <74974704+spike-zhu@users.noreply.github.com> * refactor: use generated C++ graph runtime API * chore: drop unrelated gitignore change * feat: add Ascend graph runtime support Map the Ascend graph lifecycle to aclmdlRI capture and replay symbols loaded from AscendCL at runtime. Expose Ascend graph capture and replay through the generated C++ runtime API and link dl only for WITH_ASCEND builds. * fix: address graph runtime review comments Move graph runtime validation into the common Runtime contract and keep StreamCreate aligned with CUDA semantics. Hide Ascend RI symbol probing helpers, use the Runtime Error alias consistently, and drop unverified Moore/Metax graph placeholders. * fix: support Ascend async memset Route the Ascend runtime MemsetAsync API to aclrtMemsetAsync so Core can clear device buffers through the InfiniRT C++ runtime path. --------- Co-authored-by: Jiacheng Huang <45955067+voltjia@users.noreply.github.com> Co-authored-by: spike-zhu <74974704+spike-zhu@users.noreply.github.com> Co-authored-by: Jiacheng Huang <huangjiacheng0709@outlook.com>
Summary
Motivation
This splits the backend runtime API foundation out of #9 so that #9 can focus on graph runtime APIs only. It follows the surface already added for CPU in #8 and now uses the shared platform-adaptive test mechanism from #15.
Closes N/A
Type of Change
feat- new feature / new operator / new platformfix- bug fixperf- performance improvement (no behavioral change)refactor- code restructuring without behavior changetest- adding or fixing tests onlydocs- documentation onlybuild/ci- build system or CI configurationchore- tooling, formatting, or other non-code changes!in the Conventional Commits prefix or aBREAKING CHANGE:footer)Platforms Affected
WITH_CPU)WITH_NVIDIA)WITH_ILUVATAR)WITH_METAX)WITH_CAMBRICON)WITH_MOORE)WITH_ASCEND)WITH_TORCH)Smoke Test Result
Test Results on Supported Platforms
ssh nvidiawithaccelerator-dev/nvidia:latestssh iluvatarwithaccelerator-dev/iluvatar:latestssh metaxwithaccelerator-dev/metax:latest; installed CMake in-container forctestssh cambriconwithaccelerator-dev/cambricon:latestssh moorewithaccelerator-dev/moore:latest; async alloc/free are expected unsupported on current SDKssh ascendwithaccelerator-dev/ascend:latest; SSH wrapper returned non-zero due known host issueRepresentative NVIDIA `ctest` output
Benchmark / Performance Impact
N/A
Notes for Reviewers
musaFreeAsync, so this PR exposes the API but reports async allocation/free as unsupported for Moore.