Uh oh!
There was an error while loading. Please reload this page.
lakefile: drop Blake3 from the native-decide dynlib - #607
Draft
samuelburnham wants to merge 5 commits into
Draft
Conversation
Update the multi-stark dependency and Rust toolchain for Plonky3 0.6, along with the Rust 1.98 lint migrations required to keep the workspace warning-free. Refresh the Rust-compatible Blake3.lean pin in both root and compile-package manifests. Adapt recursive Aiur verification to Plonky3's pruned FRI multiproofs. Native proofs retain their compact serialized representation and native verification path; the FFI expands authenticated Merkle frontiers into per-query advice only when entering the existing recursive verifier circuit. Preserve the packed claim-digest convention in the recursion diagnostic and exercise the advice boundary in the end-to-end test and benchmark paths. CPU and CUDA recursive q1 runs produce identical 823,485-byte inner proofs and 331,273-byte outer proofs. The q50 Vector.extract_append workload retains identical CPU/CUDA proof sizes. Inner plus outer STARK proving measures 65.87s on CPU and 8.81s with CUDA on the RTX PRO 6000, a 7.48x speedup.
PR benchmark runs execute trusted workflow YAML from the default branch while loading composite actions from the PR checkout. When Bencher data and binary caches are unavailable, the workflow checks out main under base/ and asks Lake to rebuild it without first installing the Rust channel pinned by that checkout. Teach the existing CPU provenance action to install the base checkout's validated Rust channel and profile immediately before an uncached base build. The step is a no-op when the toolchain is already available and leaves cached benchmark comparisons unchanged.
Consume the Plonky3 0.6 batch-opening layout directly in Aiur instead of expanding every pruned Merkle frontier into one authentication path per FRI query. Sample all query indices from the unchanged transcript, sort and deduplicate them with an O(q log q) merge sort, authenticate each input and commit-phase commitment once, then retain the existing per-query reduced-opening and FRI arithmetic. Bind every frontier to transcript-derived indices, consume boundary digests in Plonky3's level/parent/child order, reject trailing frontier elements and inconsistent duplicate leaves, and assert all native opening dimensions and sibling counts. Explicitly constrain the digest-bound protocol specialization to cap height 0, binary FRI, and a constant final polynomial. Move memo_u32_less_than into IxVM Core so both substitution and multiproof sorting share its constrained rows. Strengthen the recursive negative test to mutate a structurally valid stage-1 commitment. Regenerate both checked-in Aiur Rust executors and retain interpreter/codegen query-count parity. On Vector.extract_append q50, recursive-verifier FFT cost falls from 204.073B to 201.166B. CPU outer proving improves from 50.09s to 45.03s and the full CPU pipeline from 90.64s to 82.90s. GPU outer proving improves from 15.85s to 13.72s and the full GPU pipeline from 28.86s to 26.69s. The outer proof grows from 3.92 MB to 4.17 MB. Validated with the MultiStark primitive suite, recursive honest/tamper/parity tests, codegen --check, release workspace clippy, release CUDA clippy, rustfmt, and diff checks.
The Warp x64 runner pool mixes Intel Granite Rapids and AMD Zen 5, and a build job may land on one vendor while the job that runs its binaries lands on the other. Neither vendor's feature set contains the other's, so `-Ctarget-cpu=native` does not produce a portable binary: Zen 5 enables SSE4A, and LLVM emits it. Disassembling the workspace built for znver5 finds 31 SSE4A instructions, all INSERTQ, in `ix-ffi` and in `aiur_ixvm_witness::add_entries_parallel`. Granite Rapids has no SSE4A, so the first one executed raises #UD, killing the process with SIGILL during witness generation. That is what turned every row of #605's benchmark into a crash. Pin the measured intersection of the two CPUs instead. x86-64-v4 covers every AVX-512 subset Plonky3 uses; +avx512vbmi2 preserves its VPSHRDQ interleave and +gfni preserves LLVM's byte-shift lowering. A workspace built with these flags contains no instruction absent from either vendor and has an instruction vocabulary identical to a graniterapids build. blake3 dispatches on CPUID at runtime and is unaffected either way. `.cargo/config.toml` keeps `-Ctarget-cpu=native`: a developer builds and runs on one machine, and x86-64-v4 would exclude every host without AVX-512. Only CI has the split, so only CI pins the ISA. The new guard fails the job when a runner lacks a required feature, so the assumption is enforced rather than assumed, and the shared `warp-x64` cargo cache key becomes sound now that codegen no longer varies by host. RUSTFLAGS is hashed into that key, so the flag change rotates it on its own. Pinning also removes a benchmarking hazard that never crashed: LLVM sets prefer-256-bit for Granite Rapids but not for Zen 5, so the same source vectorized 3.2x more widely depending on the build host, and main-vs-PR timings were not comparable across a vendor split.
Blake3 now precompiles its libraries, so Lake loads their shared objects -- which bundle the C and Rust FFI objects -- into any process elaborating a module that imports them. The Blake3 half of `ix_native_decide_dynlib` was assembling that by hand from a `blake3_rs_shared` cdylib, and that target no longer exists upstream. The target keeps Ix's own externs, which nothing else supplies. Precompiling `Ix.Unsigned` instead would work, but only as its own library declared after `Ix`: both would claim the module, `Package.findModule?` resolves with `findSomeRev?`, and losing that race silently stops precompiling it -- with the symptom appearing as a missing native implementation inside a proof file rather than as a configuration error. A local dynlib naming its modules outright is worth more than the lines it costs. The Blake3 pin moves to the revision that turned precompilation on, and moves in lakefile.lean, lake-manifest.json, flake.nix and flake.lock together. The lakefile drops the cdylib that the older revision still provides, so a pin left behind in any one of them pairs the new lakefile with a Blake3 that does not precompile -- and that mismatch surfaces as a missing native implementation inside a proof file rather than as a build error.
samuelburnhamforce-pushed
the
sb/blake3-precompile
branch
from
September 1, 2026 21:46
53a4803 to
b3d6044Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for freeto join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Blake3 now precompiles its libraries, so Lake loads their shared objects -- which bundle the C and Rust FFI objects -- into any process elaborating a module that imports them. The Blake3 half of
ix_native_decide_dynlibwas assembling that by hand from ablake3_rs_sharedcdylib, and that target no longer exists upstream.The target keeps Ix's own externs, which nothing else supplies. Precompiling
Ix.Unsignedinstead would work, but only as its own library declared afterIx: both would claim the module,Package.findModule?resolves withfindSomeRev?, and losing that race silently stops precompiling it -- with the symptom appearing as a missing native implementation inside a proof file rather than as a configuration error. A local dynlib naming its modules outright is worth more than the lines it costs.The Blake3 pin moves to the revision that turned precompilation on, and moves in lakefile.lean, lake-manifest.json, flake.nix and flake.lock together. The lakefile drops the cdylib that the older revision still provides, so a pin left behind in any one of them pairs the new lakefile with a Blake3 that does not precompile -- and that mismatch surfaces as a missing native implementation inside a proof file rather than as a build error.