Add Neon Inner product u4*u4 kernel - #1353
Conversation
There was a problem hiding this comment.
Pull request overview
This PR adds an AArch64-specific Neon+dotprod inner-product kernel for 4-bit (USlice<4> × USlice<4>) paths in spherical quantization, and wires it into the spherical quantizer’s architecture dispatch so the Neon implementation can be selected where available.
Changes:
- Add
aarch64spherical__codegeninstantiations for the 4-bit Neon inner-product paths. - Enable Neon dispatch for spherical quantization
AsData<4>andAsQuery<4>without downcasting to Scalar. - Implement an AArch64 Neon
InnerProductkernel forUSlice<4> × USlice<4>and adjust retargeting to avoid overlap.
Reviewed changes
Copilot reviewed 4 out of 4 changed files in this pull request and generated 2 comments.
| File | Description |
|---|---|
diskann-quantization/src/spherical/iface.rs |
Updates dispatch mapping so 4-bit spherical paths can use Neon directly (no downcast). |
diskann-quantization/src/spherical/__codegen/mod.rs |
Adds an AArch64 codegen module behind cfg(target_arch = "aarch64"). |
diskann-quantization/src/spherical/__codegen/aarch64.rs |
New AArch64 instantiation helpers for the 4-bit Neon inner-product distance computer. |
diskann-quantization/src/bits/distances.rs |
Adds the Neon USlice<4> × USlice<4> inner-product implementation and updates retargeting accordingly. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
|
@microsoft-github-policy-service agree company="Arm" |
Mark Hildebrand (hildebrandmw)
left a comment
There was a problem hiding this comment.
Thanks! Looks good to me - great to start having Neon kernels!
Outside of the small tweak to the test bounds, please start a new aarch64.rs file in diskann-quantization/src/__codegen.
5a885d5 to
a3506ac
Compare
a3506ac to
a654681
Compare
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #1353 +/- ##
==========================================
- Coverage 91.55% 91.55% -0.01%
==========================================
Files 521 521
Lines 100371 100371
==========================================
- Hits 91898 91895 -3
- Misses 8473 8476 +3
Flags with carried forward coverage won't be shown. Click here to find out more.
🚀 New features to boost your workflow:
|
fabcb9b
into
microsoft:main
# Breaking Changes Flat search visitors are now query-aware (microsoft#1359) * `flat::FlatIndex` has been removed. The search entry point is now the free function `flat::knn_search`, and the `DistancesUnordered` visitor is constructed per-query rather than reused. The `ElementRef`, `QueryComputer`, and `QueryComputerError` associated types and the visitor GAT have also been removed. `DistancesUnordered` now yields `(id, distance)` pairs directly. Migration: Callers of the old `FlatIndex`/visitor API should: 1. Drop `FlatIndex` and construct your `DistancesUnordered` visitor directly for the query being searched. 2. Replace calls into the removed wrapper with `flat::knn_search(&mut visitor, k, processor, query, &mut output)`. 3. Remove any `ElementRef`/`QueryComputer` implementation, fuse scanning and distance computation directly in your `DistancesUnordered::distances_unordered` implementation. # All Changes * Make `UnalignedSlice` Send and Sync. by @hildebrandmw in microsoft#1348 * Bump actions/checkout from 4.4.0 to 7.0.1 in the github-actions group by @dependabot[bot] in microsoft#1344 * Deduplicate virtual start point edges during disk serialization by @partychen in microsoft#1350 * Fix alpha pruning documentation by @xinyuwen2 in microsoft#1351 * Bump the github-actions group with 4 updates by @dependabot[bot] in microsoft#1356 * Add Neon Inner product u4*u4 kernel by @pfoxARM in microsoft#1353 * Strengthen arguments to `robust_prune`. by @hildebrandmw in microsoft#1358 * Allow inspection of the paged search accessor. by @hildebrandmw in microsoft#1364 * Make flat search visitors query-aware by @partychen in microsoft#1359 * Add Neon inner-product kernel for USlice<2> with spherical wiring by @pfoxARM in microsoft#1363 ## New Contributors * @pfoxARM made their first contribution in microsoft#1353 **Full Changelog**: microsoft/DiskANN@v0.56.0...v0.57.0 Co-authored-by: Mark Hildebrand <mhildebrand@microsoft.com>
What does this implement?
Exclusive aarch64 USlice4 * USlice 4 Inner Product kernel using Neon and dotprod. We also add quantization instantiation for spherical quantization.
Any other comments?
We use dot_simd() heavily in this kernel, for performance we rely on dotprod feature.