Update Aiur to Plonky3 0.6 - #605

Draft
arthurpaulino wants to merge 5 commits into
mainfrom
ap/bump-p3
Draft

Update Aiur to Plonky3 0.6#605
arthurpaulino wants to merge 5 commits into
mainfrom
ap/bump-p3

Conversation

@arthurpaulino

Copy link
Copy Markdown
Member

Update the multi-stark dependency and Rust toolchain for Plonky3 0.6, along with the Rust 1.98 lint migrations required to keep the workspace warning-free. Refresh the Rust-compatible Blake3.lean pin in both root and compile-package manifests.

Adapt recursive Aiur verification to Plonky3's pruned FRI multiproofs. Native proofs retain their compact serialized representation and native verification path; the FFI expands authenticated Merkle frontiers into per-query advice only when entering the existing recursive verifier circuit.

Preserve the packed claim-digest convention in the recursion diagnostic and exercise the advice boundary in the end-to-end test and benchmark paths. CPU and CUDA recursive q1 runs produce identical 823,485-byte inner proofs and 331,273-byte outer proofs.

The q50 Vector.extract_append workload retains identical CPU/CUDA proof sizes. Inner plus outer STARK proving measures 65.87s on CPU and 8.81s with CUDA on the RTX PRO 6000, a 7.48x speedup.

Update the multi-stark dependency and Rust toolchain for Plonky3 0.6, along with the Rust 1.98 lint migrations required to keep the workspace warning-free. Refresh the Rust-compatible Blake3.lean pin in both root and compile-package manifests.
Adapt recursive Aiur verification to Plonky3's pruned FRI multiproofs. Native proofs retain their compact serialized representation and native verification path; the FFI expands authenticated Merkle frontiers into per-query advice only when entering the existing recursive verifier circuit.
Preserve the packed claim-digest convention in the recursion diagnostic and exercise the advice boundary in the end-to-end test and benchmark paths. CPU and CUDA recursive q1 runs produce identical 823,485-byte inner proofs and 331,273-byte outer proofs.
The q50 Vector.extract_append workload retains identical CPU/CUDA proof sizes. Inner plus outer STARK proving measures 65.87s on CPU and 8.81s with CUDA on the RTX PRO 6000, a 7.48x speedup.
PR benchmark runs execute trusted workflow YAML from the default branch while loading composite actions from the PR checkout. When Bencher data and binary caches are unavailable, the workflow checks out main under base/ and asks Lake to rebuild it without first installing the Rust channel pinned by that checkout.
Teach the existing CPU provenance action to install the base checkout's validated Rust channel and profile immediately before an uncached base build. The step is a no-op when the toolchain is already available and leaves cached benchmark comparisons unchanged.
@arthurpaulino

Copy link
Copy Markdown
MemberAuthor

!benchmark fresh

@argument-ci-bot

argument-ci-botBot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

!benchmark — main vs eab18e4

backends: aiur=prove · envs: InitStd · baseline: fresh (benchmark products rebuilt, base-SHA run, bencher bypassed)

aiur · InitStd · prove — main from: base run @ e1ca8e2 (fresh — bencher bypassed)

7 constants · 1 with regressions · 7 with improvements (|Δ| > 3.0% on any metric).

IxVM on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append8.816 s8.990 s+2.0%30.095 s29.094 s-3.3% 🟢92.21095.380+3.4% 🟢70.72 GiB72.35 GiB+2.3%11.33 MiB11.07 MiB-2.3%68.7 ms60.4 ms-12.1% (1.14× faster) 🟢134.35B134.35B+0.0%
Char.ofOrdinal_le_of_le6.689 s6.790 s+1.5%25.631 s25.209 s-1.6%107.800109.600+1.7%63.87 GiB65.41 GiB+2.4%11.33 MiB11.07 MiB-2.3%73.7 ms60.3 ms-18.2% (1.22× faster) 🟢102.60B102.60B+0.0%
Array.extract_append6.679 s6.415 s-4.0% 🟢23.076 s22.072 s-4.4% 🟢69.60072.760+4.5% 🟢52.03 GiB52.79 GiB+1.4%11.24 MiB10.99 MiB-2.2%72.3 ms64.6 ms-10.6% (1.12× faster) 🟢97.08B97.08B+0.0%
Std.HashMap3.964 s4.002 s+1.0%15.519 s15.133 s-2.5%131.580134.940+2.6%36.34 GiB37.08 GiB+2.1%11.26 MiB11.00 MiB-2.3%75.1 ms65.5 ms-12.8% (1.15× faster) 🟢61.88B61.88B+0.0%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq3.598 s3.528 s-1.9%14.327 s13.806 s-3.6% 🟢130.310135.230+3.8% 🟢33.96 GiB34.69 GiB+2.1%11.26 MiB11.01 MiB-2.2%69.9 ms59.5 ms-15.0% (1.18× faster) 🟢55.68B55.68B+0.0%
String.append424.3 ms423.0 ms-0.3%2.278 s2.124 s-6.7% (1.07× faster) 🟢143.540153.930+7.2% (1.07× faster) 🟢4.89 GiB5.00 GiB+2.4%9.94 MiB9.74 MiB-2.0%64.2 ms52.0 ms-19.0% (1.24× faster) 🟢3.37B3.37B+0.0%
Nat.add_comm260.1 ms259.8 ms-0.1%1.068 s970.5 ms-9.1% (1.10× faster) 🟢43.07047.400+10.1% (1.10× faster) 🟢3.99 GiB4.64 GiB+16.4% (1.16× larger) ⚠️9.09 MiB8.91 MiB-2.0%53.9 ms41.9 ms-22.3% (1.29× faster) 🟢308.40M308.40M+0.0%
FRI verifier on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append5.488 s5.490 s+0.0%31.664 s30.699 s-3.0% 🟢87.64090.400+3.1% 🟢101.37 GiB101.07 GiB-0.3%3.97 MiB3.72 MiB-6.3% (1.07× smaller) 🟢29.2 ms23.8 ms-18.5% (1.23× faster) 🟢210.23B210.23B+0.0%
Char.ofOrdinal_le_of_le5.386 s5.489 s+1.9%30.846 s30.914 s+0.2%89.57089.380-0.2%100.32 GiB100.32 GiB+0.0%3.97 MiB3.72 MiB-6.3% (1.07× smaller) 🟢24.8 ms23.1 ms-6.6% (1.07× faster) 🟢207.18B207.18B+0.0%
Array.extract_append5.209 s5.235 s+0.5%29.948 s28.844 s-3.7% 🟢53.63055.680+3.8% 🟢94.59 GiB95.37 GiB+0.8%3.97 MiB3.73 MiB-6.1% (1.07× smaller) 🟢25.7 ms21.3 ms-17.1% (1.21× faster) 🟢200.65B200.65B+0.0%
Std.HashMap5.314 s5.306 s-0.1%29.652 s29.584 s-0.2%68.87069.020+0.2%94.56 GiB95.34 GiB+0.8%3.97 MiB3.72 MiB-6.3% (1.07× smaller) 🟢25.2 ms21.2 ms-15.7% (1.19× faster) 🟢203.35B203.35B+0.0%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq5.400 s5.333 s-1.2%31.051 s30.174 s-2.8%60.13061.870+2.9%98.68 GiB99.48 GiB+0.8%3.97 MiB3.73 MiB-6.2% (1.07× smaller) 🟢30.2 ms31.0 ms+2.5%205.59B205.59B+0.0%
String.append4.334 s4.419 s+2.0%27.489 s26.612 s-3.2% 🟢11.90012.290+3.3% 🟢87.87 GiB88.58 GiB+0.8%3.97 MiB3.73 MiB-6.1% (1.07× smaller) 🟢34.5 ms22.7 ms-34.3% (1.52× faster) 🟢168.67B168.67B+0.0%
Nat.add_comm3.518 s3.499 s-0.5%18.993 s17.942 s-5.5% (1.06× faster) 🟢2.4202.560+5.8% (1.06× faster) 🟢58.70 GiB59.42 GiB+1.2%3.97 MiB3.72 MiB-6.2% (1.07× smaller) 🟢24.8 ms20.8 ms-16.1% (1.19× faster) 🟢130.84B130.84B+0.0%
Pipeline total (7 constants)
constanttotal-time (main)total-time (PR)Δ%pipeline-throughput (const/s) (main)pipeline-throughput (const/s) (PR)Δ%pipeline-peak-ram (main)pipeline-peak-ram (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append1m 1.8s59.792 s-3.2% 🟢44.93046.410+3.3% 🟢101.37 GiB101.07 GiB-0.3%
Char.ofOrdinal_le_of_le56.477 s56.124 s-0.6%48.92049.230+0.6%100.32 GiB100.32 GiB+0.0%
Array.extract_append53.024 s50.916 s-4.0% 🟢30.29031.540+4.1% 🟢94.59 GiB95.37 GiB+0.8%
Std.HashMap45.172 s44.717 s-1.0%45.21045.660+1.0%94.56 GiB95.34 GiB+0.8%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq45.377 s43.980 s-3.1% 🟢41.14042.450+3.2% 🟢98.68 GiB99.48 GiB+0.8%
String.append29.767 s28.737 s-3.5% 🟢10.99011.380+3.5% 🟢87.87 GiB88.58 GiB+0.8%
Nat.add_comm20.061 s18.912 s-5.7% (1.06× faster) 🟢2.2902.430+6.1% (1.06× faster) 🟢58.70 GiB59.42 GiB+1.2%

Workflow logs

Consume the Plonky3 0.6 batch-opening layout directly in Aiur instead of expanding every pruned Merkle frontier into one authentication path per FRI query. Sample all query indices from the unchanged transcript, sort and deduplicate them with an O(q log q) merge sort, authenticate each input and commit-phase commitment once, then retain the existing per-query reduced-opening and FRI arithmetic.
Bind every frontier to transcript-derived indices, consume boundary digests in Plonky3's level/parent/child order, reject trailing frontier elements and inconsistent duplicate leaves, and assert all native opening dimensions and sibling counts. Explicitly constrain the digest-bound protocol specialization to cap height 0, binary FRI, and a constant final polynomial. Move memo_u32_less_than into IxVM Core so both substitution and multiproof sorting share its constrained rows.
Strengthen the recursive negative test to mutate a structurally valid stage-1 commitment. Regenerate both checked-in Aiur Rust executors and retain interpreter/codegen query-count parity.
On Vector.extract_append q50, recursive-verifier FFT cost falls from 204.073B to 201.166B. CPU outer proving improves from 50.09s to 45.03s and the full CPU pipeline from 90.64s to 82.90s. GPU outer proving improves from 15.85s to 13.72s and the full GPU pipeline from 28.86s to 26.69s. The outer proof grows from 3.92 MB to 4.17 MB.
Validated with the MultiStark primitive suite, recursive honest/tamper/parity tests, codegen --check, release workspace clippy, release CUDA clippy, rustfmt, and diff checks.
@arthurpaulino

Copy link
Copy Markdown
MemberAuthor

!benchmark fresh

@argument-ci-bot

argument-ci-botBot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

!benchmark — main vs 840932a

backends: aiur=prove · envs: InitStd · baseline: fresh (benchmark products rebuilt, base-SHA run, bencher bypassed)

Warning

  • CPU model mismatch for PR benchmark binaries in this job: built on AMD EPYC 9R45; measured on Intel(R) Xeon(R) 6975P-C. Native Rust code uses -Ctarget-cpu=native.

aiur · InitStd · prove — main from: base run @ e1ca8e2 (fresh — bencher bypassed)

7 constants · 0 with regressions · 0 with improvements (|Δ| > 3.0% on any metric).

IxVM on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append8.874 s💥 CRASHn/a44.291 s💥 CRASHn/a62.650💥 CRASHn/a70.80 GiB💥 CRASHn/a11.33 MiB💥 CRASHn/a77.4 ms💥 CRASHn/a134.35B💥 CRASHn/a
Char.ofOrdinal_le_of_le6.603 s💥 CRASHn/a38.242 s💥 CRASHn/a72.250💥 CRASHn/a63.81 GiB💥 CRASHn/a11.33 MiB💥 CRASHn/a76.9 ms💥 CRASHn/a102.60B💥 CRASHn/a
Array.extract_append6.311 s💥 CRASHn/a33.420 s💥 CRASHn/a48.060💥 CRASHn/a51.95 GiB💥 CRASHn/a11.24 MiB💥 CRASHn/a86.5 ms💥 CRASHn/a97.08B💥 CRASHn/a
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq3.774 s💥 CRASHn/a20.628 s💥 CRASHn/a90.510💥 CRASHn/a33.97 GiB💥 CRASHn/a11.26 MiB💥 CRASHn/a82.6 ms💥 CRASHn/a55.68B💥 CRASHn/a
Std.HashMap4.060 s💥 CRASHn/a22.501 s💥 CRASHn/a90.750💥 CRASHn/a36.36 GiB💥 CRASHn/a11.26 MiB💥 CRASHn/a84.1 ms💥 CRASHn/a61.88B💥 CRASHn/a
String.append708.3 ms💥 CRASHn/a2.863 s💥 CRASHn/a114.220💥 CRASHn/a4.99 GiB💥 CRASHn/a9.94 MiB💥 CRASHn/a70.3 ms💥 CRASHn/a3.37B💥 CRASHn/a
Nat.add_comm496.0 ms💥 CRASHn/a1.330 s💥 CRASHn/a34.580💥 CRASHn/a4.33 GiB💥 CRASHn/a9.09 MiB💥 CRASHn/a58.9 ms💥 CRASHn/a308.40M💥 CRASHn/a
FRI verifier on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append6.880 s💥 CRASHn/a54.265 s💥 CRASHn/a51.140💥 CRASHn/a101.05 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a49.5 ms💥 CRASHn/a210.23B💥 CRASHn/a
Char.ofOrdinal_le_of_le6.813 s💥 CRASHn/a53.449 s💥 CRASHn/a51.690💥 CRASHn/a99.81 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a45.1 ms💥 CRASHn/a207.18B💥 CRASHn/a
Array.extract_append6.371 s💥 CRASHn/a50.695 s💥 CRASHn/a31.680💥 CRASHn/a95.19 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a58.1 ms💥 CRASHn/a200.65B💥 CRASHn/a
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq6.748 s💥 CRASHn/a53.378 s💥 CRASHn/a34.980💥 CRASHn/a98.71 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a46.1 ms💥 CRASHn/a205.59B💥 CRASHn/a
Std.HashMap6.455 s💥 CRASHn/a50.868 s💥 CRASHn/a40.140💥 CRASHn/a95.03 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a45.7 ms💥 CRASHn/a203.35B💥 CRASHn/a
String.append5.307 s💥 CRASHn/a47.595 s💥 CRASHn/a6.870💥 CRASHn/a87.89 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a32.4 ms💥 CRASHn/a168.67B💥 CRASHn/a
Nat.add_comm4.526 s💥 CRASHn/a31.415 s💥 CRASHn/a1.460💥 CRASHn/a58.66 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a28.0 ms💥 CRASHn/a130.84B💥 CRASHn/a
Pipeline total (7 constants)
constanttotal-time (main)total-time (PR)Δ%pipeline-throughput (const/s) (main)pipeline-throughput (const/s) (PR)Δ%pipeline-peak-ram (main)pipeline-peak-ram (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append1m 38.6s💥 CRASHn/a28.160💥 CRASHn/a101.05 GiB💥 CRASHn/a
Char.ofOrdinal_le_of_le1m 31.7s💥 CRASHn/a30.130💥 CRASHn/a99.81 GiB💥 CRASHn/a
Array.extract_append1m 24.1s💥 CRASHn/a19.090💥 CRASHn/a95.19 GiB💥 CRASHn/a
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq1m 14.0s💥 CRASHn/a25.230💥 CRASHn/a98.71 GiB💥 CRASHn/a
Std.HashMap1m 13.4s💥 CRASHn/a27.830💥 CRASHn/a95.03 GiB💥 CRASHn/a
String.append50.458 s💥 CRASHn/a6.480💥 CRASHn/a87.89 GiB💥 CRASHn/a
Nat.add_comm32.745 s💥 CRASHn/a1.400💥 CRASHn/a58.66 GiB💥 CRASHn/a

Workflow logs

samuelburnham added a commit that referenced this pull request Sep 1, 2026
The Warp x64 runner pool mixes Intel Granite Rapids and AMD Zen 5, and a
build job may land on one vendor while the job that runs its binaries
lands on the other. Neither vendor's feature set contains the other's, so
`-Ctarget-cpu=native` does not produce a portable binary: Zen 5 enables
SSE4A, and LLVM emits it. Disassembling the workspace built for znver5
finds 31 SSE4A instructions, all INSERTQ, in `ix-ffi` and in
`aiur_ixvm_witness::add_entries_parallel`. Granite Rapids has no SSE4A,
so the first one executed raises #UD, killing the process with SIGILL
during witness generation. That is what turned every row of #605's
benchmark into a crash.
Pin the measured intersection of the two CPUs instead. x86-64-v4 covers
every AVX-512 subset Plonky3 uses; +avx512vbmi2 preserves its VPSHRDQ
interleave and +gfni preserves LLVM's byte-shift lowering. A workspace
built with these flags contains no instruction absent from either vendor
and has an instruction vocabulary identical to a graniterapids build.
blake3 dispatches on CPUID at runtime and is unaffected either way.
`.cargo/config.toml` keeps `-Ctarget-cpu=native`: a developer builds and
runs on one machine, and x86-64-v4 would exclude every host without
AVX-512. Only CI has the split, so only CI pins the ISA. The new guard
fails the job when a runner lacks a required feature, so the assumption
is enforced rather than assumed, and the shared `warp-x64` cargo cache
key becomes sound now that codegen no longer varies by host.
Pinning also removes a benchmarking hazard that never crashed: LLVM sets
prefer-256-bit for Granite Rapids but not for Zen 5, so the same source
vectorized 3.2x more widely depending on the build host, and main-vs-PR
timings were not comparable across a vendor split.
Drop the CPU-mismatch warning the benchmark comment used to carry. It
detected a real problem, but the flags above prevent that problem, and
computing it in one job to render it in another cost a Markdown file
threaded through cache entries, artifacts, and a `--warning-file` flag on
`ix bench compare`. Warnings belong to the run that finds them.
@arthurpaulino

Copy link
Copy Markdown
MemberAuthor

!benchmark fresh

@argument-ci-bot

argument-ci-botBot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

⚠️ Some benchmark jobs failed — results may be partial.

!benchmark — main vs d7fc2f1

backends: aiur=prove · envs: InitStd · baseline: fresh (benchmark products rebuilt, base-SHA run, bencher bypassed)

No result tables were produced — see the run logs.

Workflow logs

The Warp x64 runner pool mixes Intel Granite Rapids and AMD Zen 5, and a
build job may land on one vendor while the job that runs its binaries
lands on the other. Neither vendor's feature set contains the other's, so
`-Ctarget-cpu=native` does not produce a portable binary: Zen 5 enables
SSE4A, and LLVM emits it. Disassembling the workspace built for znver5
finds 31 SSE4A instructions, all INSERTQ, in `ix-ffi` and in
`aiur_ixvm_witness::add_entries_parallel`. Granite Rapids has no SSE4A,
so the first one executed raises #UD, killing the process with SIGILL
during witness generation. That is what turned every row of #605's
benchmark into a crash.
Pin the measured intersection of the two CPUs instead. x86-64-v4 covers
every AVX-512 subset Plonky3 uses; +avx512vbmi2 preserves its VPSHRDQ
interleave and +gfni preserves LLVM's byte-shift lowering. A workspace
built with these flags contains no instruction absent from either vendor
and has an instruction vocabulary identical to a graniterapids build.
blake3 dispatches on CPUID at runtime and is unaffected either way.
`.cargo/config.toml` keeps `-Ctarget-cpu=native`: a developer builds and
runs on one machine, and x86-64-v4 would exclude every host without
AVX-512. Only CI has the split, so only CI pins the ISA. The new guard
fails the job when a runner lacks a required feature, so the assumption
is enforced rather than assumed, and the shared `warp-x64` cargo cache
key becomes sound now that codegen no longer varies by host. RUSTFLAGS
is hashed into that key, so the flag change rotates it on its own.
Pinning also removes a benchmarking hazard that never crashed: LLVM sets
prefer-256-bit for Granite Rapids but not for Zen 5, so the same source
vectorized 3.2x more widely depending on the build host, and main-vs-PR
timings were not comparable across a vendor split.
@samuelburnham

Copy link
Copy Markdown
Member

!benchmark fresh

@argument-ci-bot

argument-ci-botBot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

!benchmark — main vs feb014d

backends: aiur=prove · envs: InitStd · baseline: fresh (benchmark products rebuilt, base-SHA run, bencher bypassed)

aiur · InitStd · prove — main from: base run @ e1ca8e2 (fresh — bencher bypassed)

7 constants · 3 with regressions · 7 with improvements (|Δ| > 3.0% on any metric).

IxVM on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append9.269 s9.073 s-2.1%31.250 s30.269 s-3.1% 🟢88.80091.680+3.2% 🟢70.72 GiB72.28 GiB+2.2%11.33 MiB11.06 MiB-2.4%68.5 ms62.7 ms-8.5% (1.09× faster) 🟢134.35B134.35B+0.0%
Char.ofOrdinal_le_of_le6.964 s6.879 s-1.2%26.884 s25.793 s-4.1% 🟢102.780107.120+4.2% 🟢63.83 GiB65.31 GiB+2.3%11.33 MiB11.07 MiB-2.3%77.9 ms59.3 ms-23.8% (1.31× faster) 🟢102.60B102.60B+0.0%
Array.extract_append6.676 s6.751 s+1.1%23.925 s23.179 s-3.1% 🟢67.13069.290+3.2% 🟢51.97 GiB52.72 GiB+1.4%11.24 MiB10.99 MiB-2.2%71.1 ms57.2 ms-19.5% (1.24× faster) 🟢97.08B97.08B+0.0%
Std.HashMap4.171 s4.079 s-2.2%16.222 s15.628 s-3.7% 🟢125.880130.670+3.8% 🟢36.29 GiB37.10 GiB+2.2%11.26 MiB11.01 MiB-2.2%74.8 ms66.6 ms-11.0% (1.12× faster) 🟢61.88B61.88B+0.0%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq3.697 s3.648 s-1.3%14.858 s14.287 s-3.8% 🟢125.650130.680+4.0% 🟢34.02 GiB34.75 GiB+2.1%11.26 MiB11.00 MiB-2.2%75.5 ms57.7 ms-23.6% (1.31× faster) 🟢55.68B55.68B+0.0%
String.append435.5 ms434.0 ms-0.4%2.307 s2.129 s-7.7% (1.08× faster) 🟢141.750153.600+8.4% (1.08× faster) 🟢5.74 GiB5.52 GiB-3.9% 🟢9.94 MiB9.74 MiB-2.1%64.4 ms50.2 ms-22.1% (1.28× faster) 🟢3.37B3.37B+0.0%
Nat.add_comm267.1 ms264.3 ms-1.1%1.068 s984.9 ms-7.8% (1.08× faster) 🟢43.06046.710+8.5% (1.08× faster) 🟢4.51 GiB3.99 GiB-11.5% (1.13× smaller) 🟢9.09 MiB8.90 MiB-2.1%53.3 ms47.9 ms-10.2% (1.11× faster) 🟢308.40M308.40M+0.0%
FRI verifier on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append5.680 s4.975 s-12.4% (1.14× faster) 🟢33.109 s30.692 s-7.3% (1.08× faster) 🟢83.82090.410+7.9% (1.08× faster) 🟢101.44 GiB99.74 GiB-1.7%3.97 MiB3.98 MiB+0.2%26.3 ms22.3 ms-15.4% (1.18× faster) 🟢210.23B203.74B-3.1% 🟢
Char.ofOrdinal_le_of_le5.599 s5.014 s-10.5% (1.12× faster) 🟢32.280 s31.379 s-2.8%85.59088.050+2.9%99.81 GiB101.09 GiB+1.3%3.97 MiB3.98 MiB+0.2%27.6 ms28.6 ms+3.6% ⚠️207.18B208.08B+0.4%
Array.extract_append5.315 s4.858 s-8.6% (1.09× faster) 🟢30.962 s29.672 s-4.2% 🟢51.87054.120+4.3% 🟢95.00 GiB95.36 GiB+0.4%3.97 MiB3.99 MiB+0.4%26.0 ms39.4 ms+51.3% (1.51× slower) ⚠️200.65B200.40B-0.1%
Std.HashMap5.445 s4.864 s-10.7% (1.12× faster) 🟢31.021 s29.949 s-3.5% 🟢65.83068.180+3.6% 🟢94.50 GiB95.90 GiB+1.5%3.97 MiB3.98 MiB+0.2%25.4 ms25.6 ms+0.8%203.35B204.11B+0.4%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq5.496 s4.898 s-10.9% (1.12× faster) 🟢32.335 s29.501 s-8.8% (1.10× faster) 🟢57.74063.290+9.6% (1.10× faster) 🟢98.72 GiB95.35 GiB-3.4% 🟢3.97 MiB3.98 MiB+0.3%25.9 ms31.9 ms+23.0% (1.23× slower) ⚠️205.59B199.65B-2.9%
String.append4.487 s3.955 s-11.8% (1.13× faster) 🟢28.710 s27.141 s-5.5% (1.06× faster) 🟢11.39012.050+5.8% (1.06× faster) 🟢87.89 GiB88.65 GiB+0.9%3.97 MiB3.97 MiB+0.1%37.6 ms27.4 ms-27.2% (1.37× faster) 🟢168.67B164.80B-2.3%
Nat.add_comm3.619 s3.077 s-15.0% (1.18× faster) 🟢19.611 s17.954 s-8.5% (1.09× faster) 🟢2.3502.560+8.9% (1.09× faster) 🟢58.66 GiB58.03 GiB-1.1%3.97 MiB3.98 MiB+0.2%27.6 ms22.0 ms-20.4% (1.26× faster) 🟢130.84B124.20B-5.1% (1.05× fewer) 🟢
Pipeline total (7 constants)
constanttotal-time (main)total-time (PR)Δ%pipeline-throughput (const/s) (main)pipeline-throughput (const/s) (PR)Δ%pipeline-peak-ram (main)pipeline-peak-ram (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append1m 4.4s1m 1.0s-5.3% (1.06× faster) 🟢43.12045.520+5.6% (1.06× faster) 🟢101.44 GiB99.74 GiB-1.7%
Char.ofOrdinal_le_of_le59.164 s57.172 s-3.4% 🟢46.70048.330+3.5% 🟢99.81 GiB101.09 GiB+1.3%
Array.extract_append54.887 s52.851 s-3.7% 🟢29.26030.390+3.9% 🟢95.00 GiB95.36 GiB+0.4%
Std.HashMap47.243 s45.577 s-3.5% 🟢43.22044.800+3.7% 🟢94.50 GiB95.90 GiB+1.5%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq47.193 s43.788 s-7.2% (1.08× faster) 🟢39.56042.640+7.8% (1.08× faster) 🟢98.72 GiB95.35 GiB-3.4% 🟢
String.append31.017 s29.270 s-5.6% (1.06× faster) 🟢10.54011.170+6.0% (1.06× faster) 🟢87.89 GiB88.65 GiB+0.9%
Nat.add_comm20.680 s18.939 s-8.4% (1.09× faster) 🟢2.2202.430+9.5% (1.09× faster) 🟢58.66 GiB58.03 GiB-1.1%

Workflow logs

#606)
* Authenticate all frontier group members in the recursive verifier
The in-circuit pruned-multiproof walk (mmcs_verify_multi / frontier_level)
collapses queries that share a parent to a single lead node and hashes
only the lead's rows via inject_maybe(ar, ...). Non-lead members' opened
rows for the shorter (injected) matrices were still consumed in their own
per-query FRI arithmetic (batch_views_at) but never authenticated against
any commitment — the leaf hash covers only the tallest matrices, and
shorter ones are bound solely through injection. A prover could therefore
forge a non-lead member's shorter-matrix opening. Plonky3's reference
verify_batch_pruned guards exactly this with InconsistentGroupOpening
(and InconsistentDuplicateOpenings for equal-index queries); the port had
neither. The prior per-query walk did not have the gap, so it was
introduced with the direct multiproof consumption.
- frontier_level: on a group merge, assert the lead and member agree on
every not-yet-injected matrix (height <= next_lh) via select_rows_le +
pointer equality. Transitive across pairwise merges, so the whole group
is pinned; matches InconsistentGroupOpening.
- frontier_merge: duplicate transcript indices must open the SAME full
rows, not merely the same tallest-matrix leaf digest; matches
InconsistentDuplicateOpenings.
Pointer equality is admissible inside assert_eq! (equal pointers imply
equal content; a spurious mismatch costs only completeness — see
IxVM.Core). select_rows_le selects rows of matrices at height <= target,
mirroring select_rows.
Validated: the group-merge branch is genuinely reached by the factorial
recursion proof (an always-false variant of the new assert fails the
honest test), the honest proof still verifies with the real assert
(completeness preserved), the existing tamper tests still reject, and
the full lake test suite is green (2717 checks). aiur_multi_stark.rs
regenerated; kernel executor unchanged.
* Drop the multi_stark::advice dependency
Companion to multi-stark removing its unused per-query advice module.
ix consumed native pruned multiproofs directly and referenced only
advice::AdviceError, whose two arms (verification failed, serialization
failed) were immediately string-formatted by the FFI. Replace it with a
plain Result<Vec<u8>, String>: AiurSystem::proof_to_advice_bytes maps
both failures to a message, and the FFI passes the string straight to
LeanExcept::error_string. No behavior change; the Lean binding
(Except String ByteArray) is unaffected.
Bump the multi-stark pin to the advice-removed revision. Requires that
multi-stark's ap/bump-p3-drop-advice be pushed first, exactly as with
every other pin in this series.
* Bump multi-stark audit revision
---------
Co-authored-by: Arthur Paulino <arthurleonardo.ap@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@arthurpaulino@samuelburnham
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Update Aiur to Plonky3 0.6 - #605

Draft
arthurpaulino wants to merge 5 commits into
mainfrom
ap/bump-p3
Draft

Update Aiur to Plonky3 0.6#605
arthurpaulino wants to merge 5 commits into
mainfrom
ap/bump-p3

Conversation

@arthurpaulino

Copy link
Copy Markdown
Member

Update the multi-stark dependency and Rust toolchain for Plonky3 0.6, along with the Rust 1.98 lint migrations required to keep the workspace warning-free. Refresh the Rust-compatible Blake3.lean pin in both root and compile-package manifests.

Adapt recursive Aiur verification to Plonky3's pruned FRI multiproofs. Native proofs retain their compact serialized representation and native verification path; the FFI expands authenticated Merkle frontiers into per-query advice only when entering the existing recursive verifier circuit.

Preserve the packed claim-digest convention in the recursion diagnostic and exercise the advice boundary in the end-to-end test and benchmark paths. CPU and CUDA recursive q1 runs produce identical 823,485-byte inner proofs and 331,273-byte outer proofs.

The q50 Vector.extract_append workload retains identical CPU/CUDA proof sizes. Inner plus outer STARK proving measures 65.87s on CPU and 8.81s with CUDA on the RTX PRO 6000, a 7.48x speedup.

Update the multi-stark dependency and Rust toolchain for Plonky3 0.6, along with the Rust 1.98 lint migrations required to keep the workspace warning-free. Refresh the Rust-compatible Blake3.lean pin in both root and compile-package manifests.
Adapt recursive Aiur verification to Plonky3's pruned FRI multiproofs. Native proofs retain their compact serialized representation and native verification path; the FFI expands authenticated Merkle frontiers into per-query advice only when entering the existing recursive verifier circuit.
Preserve the packed claim-digest convention in the recursion diagnostic and exercise the advice boundary in the end-to-end test and benchmark paths. CPU and CUDA recursive q1 runs produce identical 823,485-byte inner proofs and 331,273-byte outer proofs.
The q50 Vector.extract_append workload retains identical CPU/CUDA proof sizes. Inner plus outer STARK proving measures 65.87s on CPU and 8.81s with CUDA on the RTX PRO 6000, a 7.48x speedup.
PR benchmark runs execute trusted workflow YAML from the default branch while loading composite actions from the PR checkout. When Bencher data and binary caches are unavailable, the workflow checks out main under base/ and asks Lake to rebuild it without first installing the Rust channel pinned by that checkout.
Teach the existing CPU provenance action to install the base checkout's validated Rust channel and profile immediately before an uncached base build. The step is a no-op when the toolchain is already available and leaves cached benchmark comparisons unchanged.
@arthurpaulino

Copy link
Copy Markdown
MemberAuthor

!benchmark fresh

@argument-ci-bot

argument-ci-botBot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

!benchmark — main vs eab18e4

backends: aiur=prove · envs: InitStd · baseline: fresh (benchmark products rebuilt, base-SHA run, bencher bypassed)

aiur · InitStd · prove — main from: base run @ e1ca8e2 (fresh — bencher bypassed)

7 constants · 1 with regressions · 7 with improvements (|Δ| > 3.0% on any metric).

IxVM on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append8.816 s8.990 s+2.0%30.095 s29.094 s-3.3% 🟢92.21095.380+3.4% 🟢70.72 GiB72.35 GiB+2.3%11.33 MiB11.07 MiB-2.3%68.7 ms60.4 ms-12.1% (1.14× faster) 🟢134.35B134.35B+0.0%
Char.ofOrdinal_le_of_le6.689 s6.790 s+1.5%25.631 s25.209 s-1.6%107.800109.600+1.7%63.87 GiB65.41 GiB+2.4%11.33 MiB11.07 MiB-2.3%73.7 ms60.3 ms-18.2% (1.22× faster) 🟢102.60B102.60B+0.0%
Array.extract_append6.679 s6.415 s-4.0% 🟢23.076 s22.072 s-4.4% 🟢69.60072.760+4.5% 🟢52.03 GiB52.79 GiB+1.4%11.24 MiB10.99 MiB-2.2%72.3 ms64.6 ms-10.6% (1.12× faster) 🟢97.08B97.08B+0.0%
Std.HashMap3.964 s4.002 s+1.0%15.519 s15.133 s-2.5%131.580134.940+2.6%36.34 GiB37.08 GiB+2.1%11.26 MiB11.00 MiB-2.3%75.1 ms65.5 ms-12.8% (1.15× faster) 🟢61.88B61.88B+0.0%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq3.598 s3.528 s-1.9%14.327 s13.806 s-3.6% 🟢130.310135.230+3.8% 🟢33.96 GiB34.69 GiB+2.1%11.26 MiB11.01 MiB-2.2%69.9 ms59.5 ms-15.0% (1.18× faster) 🟢55.68B55.68B+0.0%
String.append424.3 ms423.0 ms-0.3%2.278 s2.124 s-6.7% (1.07× faster) 🟢143.540153.930+7.2% (1.07× faster) 🟢4.89 GiB5.00 GiB+2.4%9.94 MiB9.74 MiB-2.0%64.2 ms52.0 ms-19.0% (1.24× faster) 🟢3.37B3.37B+0.0%
Nat.add_comm260.1 ms259.8 ms-0.1%1.068 s970.5 ms-9.1% (1.10× faster) 🟢43.07047.400+10.1% (1.10× faster) 🟢3.99 GiB4.64 GiB+16.4% (1.16× larger) ⚠️9.09 MiB8.91 MiB-2.0%53.9 ms41.9 ms-22.3% (1.29× faster) 🟢308.40M308.40M+0.0%
FRI verifier on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append5.488 s5.490 s+0.0%31.664 s30.699 s-3.0% 🟢87.64090.400+3.1% 🟢101.37 GiB101.07 GiB-0.3%3.97 MiB3.72 MiB-6.3% (1.07× smaller) 🟢29.2 ms23.8 ms-18.5% (1.23× faster) 🟢210.23B210.23B+0.0%
Char.ofOrdinal_le_of_le5.386 s5.489 s+1.9%30.846 s30.914 s+0.2%89.57089.380-0.2%100.32 GiB100.32 GiB+0.0%3.97 MiB3.72 MiB-6.3% (1.07× smaller) 🟢24.8 ms23.1 ms-6.6% (1.07× faster) 🟢207.18B207.18B+0.0%
Array.extract_append5.209 s5.235 s+0.5%29.948 s28.844 s-3.7% 🟢53.63055.680+3.8% 🟢94.59 GiB95.37 GiB+0.8%3.97 MiB3.73 MiB-6.1% (1.07× smaller) 🟢25.7 ms21.3 ms-17.1% (1.21× faster) 🟢200.65B200.65B+0.0%
Std.HashMap5.314 s5.306 s-0.1%29.652 s29.584 s-0.2%68.87069.020+0.2%94.56 GiB95.34 GiB+0.8%3.97 MiB3.72 MiB-6.3% (1.07× smaller) 🟢25.2 ms21.2 ms-15.7% (1.19× faster) 🟢203.35B203.35B+0.0%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq5.400 s5.333 s-1.2%31.051 s30.174 s-2.8%60.13061.870+2.9%98.68 GiB99.48 GiB+0.8%3.97 MiB3.73 MiB-6.2% (1.07× smaller) 🟢30.2 ms31.0 ms+2.5%205.59B205.59B+0.0%
String.append4.334 s4.419 s+2.0%27.489 s26.612 s-3.2% 🟢11.90012.290+3.3% 🟢87.87 GiB88.58 GiB+0.8%3.97 MiB3.73 MiB-6.1% (1.07× smaller) 🟢34.5 ms22.7 ms-34.3% (1.52× faster) 🟢168.67B168.67B+0.0%
Nat.add_comm3.518 s3.499 s-0.5%18.993 s17.942 s-5.5% (1.06× faster) 🟢2.4202.560+5.8% (1.06× faster) 🟢58.70 GiB59.42 GiB+1.2%3.97 MiB3.72 MiB-6.2% (1.07× smaller) 🟢24.8 ms20.8 ms-16.1% (1.19× faster) 🟢130.84B130.84B+0.0%
Pipeline total (7 constants)
constanttotal-time (main)total-time (PR)Δ%pipeline-throughput (const/s) (main)pipeline-throughput (const/s) (PR)Δ%pipeline-peak-ram (main)pipeline-peak-ram (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append1m 1.8s59.792 s-3.2% 🟢44.93046.410+3.3% 🟢101.37 GiB101.07 GiB-0.3%
Char.ofOrdinal_le_of_le56.477 s56.124 s-0.6%48.92049.230+0.6%100.32 GiB100.32 GiB+0.0%
Array.extract_append53.024 s50.916 s-4.0% 🟢30.29031.540+4.1% 🟢94.59 GiB95.37 GiB+0.8%
Std.HashMap45.172 s44.717 s-1.0%45.21045.660+1.0%94.56 GiB95.34 GiB+0.8%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq45.377 s43.980 s-3.1% 🟢41.14042.450+3.2% 🟢98.68 GiB99.48 GiB+0.8%
String.append29.767 s28.737 s-3.5% 🟢10.99011.380+3.5% 🟢87.87 GiB88.58 GiB+0.8%
Nat.add_comm20.061 s18.912 s-5.7% (1.06× faster) 🟢2.2902.430+6.1% (1.06× faster) 🟢58.70 GiB59.42 GiB+1.2%

Workflow logs

Consume the Plonky3 0.6 batch-opening layout directly in Aiur instead of expanding every pruned Merkle frontier into one authentication path per FRI query. Sample all query indices from the unchanged transcript, sort and deduplicate them with an O(q log q) merge sort, authenticate each input and commit-phase commitment once, then retain the existing per-query reduced-opening and FRI arithmetic.
Bind every frontier to transcript-derived indices, consume boundary digests in Plonky3's level/parent/child order, reject trailing frontier elements and inconsistent duplicate leaves, and assert all native opening dimensions and sibling counts. Explicitly constrain the digest-bound protocol specialization to cap height 0, binary FRI, and a constant final polynomial. Move memo_u32_less_than into IxVM Core so both substitution and multiproof sorting share its constrained rows.
Strengthen the recursive negative test to mutate a structurally valid stage-1 commitment. Regenerate both checked-in Aiur Rust executors and retain interpreter/codegen query-count parity.
On Vector.extract_append q50, recursive-verifier FFT cost falls from 204.073B to 201.166B. CPU outer proving improves from 50.09s to 45.03s and the full CPU pipeline from 90.64s to 82.90s. GPU outer proving improves from 15.85s to 13.72s and the full GPU pipeline from 28.86s to 26.69s. The outer proof grows from 3.92 MB to 4.17 MB.
Validated with the MultiStark primitive suite, recursive honest/tamper/parity tests, codegen --check, release workspace clippy, release CUDA clippy, rustfmt, and diff checks.
@arthurpaulino

Copy link
Copy Markdown
MemberAuthor

!benchmark fresh

@argument-ci-bot

argument-ci-botBot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

!benchmark — main vs 840932a

backends: aiur=prove · envs: InitStd · baseline: fresh (benchmark products rebuilt, base-SHA run, bencher bypassed)

Warning

  • CPU model mismatch for PR benchmark binaries in this job: built on AMD EPYC 9R45; measured on Intel(R) Xeon(R) 6975P-C. Native Rust code uses -Ctarget-cpu=native.

aiur · InitStd · prove — main from: base run @ e1ca8e2 (fresh — bencher bypassed)

7 constants · 0 with regressions · 0 with improvements (|Δ| > 3.0% on any metric).

IxVM on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append8.874 s💥 CRASHn/a44.291 s💥 CRASHn/a62.650💥 CRASHn/a70.80 GiB💥 CRASHn/a11.33 MiB💥 CRASHn/a77.4 ms💥 CRASHn/a134.35B💥 CRASHn/a
Char.ofOrdinal_le_of_le6.603 s💥 CRASHn/a38.242 s💥 CRASHn/a72.250💥 CRASHn/a63.81 GiB💥 CRASHn/a11.33 MiB💥 CRASHn/a76.9 ms💥 CRASHn/a102.60B💥 CRASHn/a
Array.extract_append6.311 s💥 CRASHn/a33.420 s💥 CRASHn/a48.060💥 CRASHn/a51.95 GiB💥 CRASHn/a11.24 MiB💥 CRASHn/a86.5 ms💥 CRASHn/a97.08B💥 CRASHn/a
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq3.774 s💥 CRASHn/a20.628 s💥 CRASHn/a90.510💥 CRASHn/a33.97 GiB💥 CRASHn/a11.26 MiB💥 CRASHn/a82.6 ms💥 CRASHn/a55.68B💥 CRASHn/a
Std.HashMap4.060 s💥 CRASHn/a22.501 s💥 CRASHn/a90.750💥 CRASHn/a36.36 GiB💥 CRASHn/a11.26 MiB💥 CRASHn/a84.1 ms💥 CRASHn/a61.88B💥 CRASHn/a
String.append708.3 ms💥 CRASHn/a2.863 s💥 CRASHn/a114.220💥 CRASHn/a4.99 GiB💥 CRASHn/a9.94 MiB💥 CRASHn/a70.3 ms💥 CRASHn/a3.37B💥 CRASHn/a
Nat.add_comm496.0 ms💥 CRASHn/a1.330 s💥 CRASHn/a34.580💥 CRASHn/a4.33 GiB💥 CRASHn/a9.09 MiB💥 CRASHn/a58.9 ms💥 CRASHn/a308.40M💥 CRASHn/a
FRI verifier on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append6.880 s💥 CRASHn/a54.265 s💥 CRASHn/a51.140💥 CRASHn/a101.05 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a49.5 ms💥 CRASHn/a210.23B💥 CRASHn/a
Char.ofOrdinal_le_of_le6.813 s💥 CRASHn/a53.449 s💥 CRASHn/a51.690💥 CRASHn/a99.81 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a45.1 ms💥 CRASHn/a207.18B💥 CRASHn/a
Array.extract_append6.371 s💥 CRASHn/a50.695 s💥 CRASHn/a31.680💥 CRASHn/a95.19 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a58.1 ms💥 CRASHn/a200.65B💥 CRASHn/a
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq6.748 s💥 CRASHn/a53.378 s💥 CRASHn/a34.980💥 CRASHn/a98.71 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a46.1 ms💥 CRASHn/a205.59B💥 CRASHn/a
Std.HashMap6.455 s💥 CRASHn/a50.868 s💥 CRASHn/a40.140💥 CRASHn/a95.03 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a45.7 ms💥 CRASHn/a203.35B💥 CRASHn/a
String.append5.307 s💥 CRASHn/a47.595 s💥 CRASHn/a6.870💥 CRASHn/a87.89 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a32.4 ms💥 CRASHn/a168.67B💥 CRASHn/a
Nat.add_comm4.526 s💥 CRASHn/a31.415 s💥 CRASHn/a1.460💥 CRASHn/a58.66 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a28.0 ms💥 CRASHn/a130.84B💥 CRASHn/a
Pipeline total (7 constants)
constanttotal-time (main)total-time (PR)Δ%pipeline-throughput (const/s) (main)pipeline-throughput (const/s) (PR)Δ%pipeline-peak-ram (main)pipeline-peak-ram (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append1m 38.6s💥 CRASHn/a28.160💥 CRASHn/a101.05 GiB💥 CRASHn/a
Char.ofOrdinal_le_of_le1m 31.7s💥 CRASHn/a30.130💥 CRASHn/a99.81 GiB💥 CRASHn/a
Array.extract_append1m 24.1s💥 CRASHn/a19.090💥 CRASHn/a95.19 GiB💥 CRASHn/a
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq1m 14.0s💥 CRASHn/a25.230💥 CRASHn/a98.71 GiB💥 CRASHn/a
Std.HashMap1m 13.4s💥 CRASHn/a27.830💥 CRASHn/a95.03 GiB💥 CRASHn/a
String.append50.458 s💥 CRASHn/a6.480💥 CRASHn/a87.89 GiB💥 CRASHn/a
Nat.add_comm32.745 s💥 CRASHn/a1.400💥 CRASHn/a58.66 GiB💥 CRASHn/a

Workflow logs

samuelburnham added a commit that referenced this pull request Sep 1, 2026
The Warp x64 runner pool mixes Intel Granite Rapids and AMD Zen 5, and a
build job may land on one vendor while the job that runs its binaries
lands on the other. Neither vendor's feature set contains the other's, so
`-Ctarget-cpu=native` does not produce a portable binary: Zen 5 enables
SSE4A, and LLVM emits it. Disassembling the workspace built for znver5
finds 31 SSE4A instructions, all INSERTQ, in `ix-ffi` and in
`aiur_ixvm_witness::add_entries_parallel`. Granite Rapids has no SSE4A,
so the first one executed raises #UD, killing the process with SIGILL
during witness generation. That is what turned every row of #605's
benchmark into a crash.
Pin the measured intersection of the two CPUs instead. x86-64-v4 covers
every AVX-512 subset Plonky3 uses; +avx512vbmi2 preserves its VPSHRDQ
interleave and +gfni preserves LLVM's byte-shift lowering. A workspace
built with these flags contains no instruction absent from either vendor
and has an instruction vocabulary identical to a graniterapids build.
blake3 dispatches on CPUID at runtime and is unaffected either way.
`.cargo/config.toml` keeps `-Ctarget-cpu=native`: a developer builds and
runs on one machine, and x86-64-v4 would exclude every host without
AVX-512. Only CI has the split, so only CI pins the ISA. The new guard
fails the job when a runner lacks a required feature, so the assumption
is enforced rather than assumed, and the shared `warp-x64` cargo cache
key becomes sound now that codegen no longer varies by host.
Pinning also removes a benchmarking hazard that never crashed: LLVM sets
prefer-256-bit for Granite Rapids but not for Zen 5, so the same source
vectorized 3.2x more widely depending on the build host, and main-vs-PR
timings were not comparable across a vendor split.
Drop the CPU-mismatch warning the benchmark comment used to carry. It
detected a real problem, but the flags above prevent that problem, and
computing it in one job to render it in another cost a Markdown file
threaded through cache entries, artifacts, and a `--warning-file` flag on
`ix bench compare`. Warnings belong to the run that finds them.
@arthurpaulino

Copy link
Copy Markdown
MemberAuthor

!benchmark fresh

@argument-ci-bot

argument-ci-botBot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

⚠️ Some benchmark jobs failed — results may be partial.

!benchmark — main vs d7fc2f1

backends: aiur=prove · envs: InitStd · baseline: fresh (benchmark products rebuilt, base-SHA run, bencher bypassed)

No result tables were produced — see the run logs.

Workflow logs

The Warp x64 runner pool mixes Intel Granite Rapids and AMD Zen 5, and a
build job may land on one vendor while the job that runs its binaries
lands on the other. Neither vendor's feature set contains the other's, so
`-Ctarget-cpu=native` does not produce a portable binary: Zen 5 enables
SSE4A, and LLVM emits it. Disassembling the workspace built for znver5
finds 31 SSE4A instructions, all INSERTQ, in `ix-ffi` and in
`aiur_ixvm_witness::add_entries_parallel`. Granite Rapids has no SSE4A,
so the first one executed raises #UD, killing the process with SIGILL
during witness generation. That is what turned every row of #605's
benchmark into a crash.
Pin the measured intersection of the two CPUs instead. x86-64-v4 covers
every AVX-512 subset Plonky3 uses; +avx512vbmi2 preserves its VPSHRDQ
interleave and +gfni preserves LLVM's byte-shift lowering. A workspace
built with these flags contains no instruction absent from either vendor
and has an instruction vocabulary identical to a graniterapids build.
blake3 dispatches on CPUID at runtime and is unaffected either way.
`.cargo/config.toml` keeps `-Ctarget-cpu=native`: a developer builds and
runs on one machine, and x86-64-v4 would exclude every host without
AVX-512. Only CI has the split, so only CI pins the ISA. The new guard
fails the job when a runner lacks a required feature, so the assumption
is enforced rather than assumed, and the shared `warp-x64` cargo cache
key becomes sound now that codegen no longer varies by host. RUSTFLAGS
is hashed into that key, so the flag change rotates it on its own.
Pinning also removes a benchmarking hazard that never crashed: LLVM sets
prefer-256-bit for Granite Rapids but not for Zen 5, so the same source
vectorized 3.2x more widely depending on the build host, and main-vs-PR
timings were not comparable across a vendor split.
@samuelburnham

Copy link
Copy Markdown
Member

!benchmark fresh

@argument-ci-bot

argument-ci-botBot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

!benchmark — main vs feb014d

backends: aiur=prove · envs: InitStd · baseline: fresh (benchmark products rebuilt, base-SHA run, bencher bypassed)

aiur · InitStd · prove — main from: base run @ e1ca8e2 (fresh — bencher bypassed)

7 constants · 3 with regressions · 7 with improvements (|Δ| > 3.0% on any metric).

IxVM on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append9.269 s9.073 s-2.1%31.250 s30.269 s-3.1% 🟢88.80091.680+3.2% 🟢70.72 GiB72.28 GiB+2.2%11.33 MiB11.06 MiB-2.4%68.5 ms62.7 ms-8.5% (1.09× faster) 🟢134.35B134.35B+0.0%
Char.ofOrdinal_le_of_le6.964 s6.879 s-1.2%26.884 s25.793 s-4.1% 🟢102.780107.120+4.2% 🟢63.83 GiB65.31 GiB+2.3%11.33 MiB11.07 MiB-2.3%77.9 ms59.3 ms-23.8% (1.31× faster) 🟢102.60B102.60B+0.0%
Array.extract_append6.676 s6.751 s+1.1%23.925 s23.179 s-3.1% 🟢67.13069.290+3.2% 🟢51.97 GiB52.72 GiB+1.4%11.24 MiB10.99 MiB-2.2%71.1 ms57.2 ms-19.5% (1.24× faster) 🟢97.08B97.08B+0.0%
Std.HashMap4.171 s4.079 s-2.2%16.222 s15.628 s-3.7% 🟢125.880130.670+3.8% 🟢36.29 GiB37.10 GiB+2.2%11.26 MiB11.01 MiB-2.2%74.8 ms66.6 ms-11.0% (1.12× faster) 🟢61.88B61.88B+0.0%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq3.697 s3.648 s-1.3%14.858 s14.287 s-3.8% 🟢125.650130.680+4.0% 🟢34.02 GiB34.75 GiB+2.1%11.26 MiB11.00 MiB-2.2%75.5 ms57.7 ms-23.6% (1.31× faster) 🟢55.68B55.68B+0.0%
String.append435.5 ms434.0 ms-0.4%2.307 s2.129 s-7.7% (1.08× faster) 🟢141.750153.600+8.4% (1.08× faster) 🟢5.74 GiB5.52 GiB-3.9% 🟢9.94 MiB9.74 MiB-2.1%64.4 ms50.2 ms-22.1% (1.28× faster) 🟢3.37B3.37B+0.0%
Nat.add_comm267.1 ms264.3 ms-1.1%1.068 s984.9 ms-7.8% (1.08× faster) 🟢43.06046.710+8.5% (1.08× faster) 🟢4.51 GiB3.99 GiB-11.5% (1.13× smaller) 🟢9.09 MiB8.90 MiB-2.1%53.3 ms47.9 ms-10.2% (1.11× faster) 🟢308.40M308.40M+0.0%
FRI verifier on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append5.680 s4.975 s-12.4% (1.14× faster) 🟢33.109 s30.692 s-7.3% (1.08× faster) 🟢83.82090.410+7.9% (1.08× faster) 🟢101.44 GiB99.74 GiB-1.7%3.97 MiB3.98 MiB+0.2%26.3 ms22.3 ms-15.4% (1.18× faster) 🟢210.23B203.74B-3.1% 🟢
Char.ofOrdinal_le_of_le5.599 s5.014 s-10.5% (1.12× faster) 🟢32.280 s31.379 s-2.8%85.59088.050+2.9%99.81 GiB101.09 GiB+1.3%3.97 MiB3.98 MiB+0.2%27.6 ms28.6 ms+3.6% ⚠️207.18B208.08B+0.4%
Array.extract_append5.315 s4.858 s-8.6% (1.09× faster) 🟢30.962 s29.672 s-4.2% 🟢51.87054.120+4.3% 🟢95.00 GiB95.36 GiB+0.4%3.97 MiB3.99 MiB+0.4%26.0 ms39.4 ms+51.3% (1.51× slower) ⚠️200.65B200.40B-0.1%
Std.HashMap5.445 s4.864 s-10.7% (1.12× faster) 🟢31.021 s29.949 s-3.5% 🟢65.83068.180+3.6% 🟢94.50 GiB95.90 GiB+1.5%3.97 MiB3.98 MiB+0.2%25.4 ms25.6 ms+0.8%203.35B204.11B+0.4%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq5.496 s4.898 s-10.9% (1.12× faster) 🟢32.335 s29.501 s-8.8% (1.10× faster) 🟢57.74063.290+9.6% (1.10× faster) 🟢98.72 GiB95.35 GiB-3.4% 🟢3.97 MiB3.98 MiB+0.3%25.9 ms31.9 ms+23.0% (1.23× slower) ⚠️205.59B199.65B-2.9%
String.append4.487 s3.955 s-11.8% (1.13× faster) 🟢28.710 s27.141 s-5.5% (1.06× faster) 🟢11.39012.050+5.8% (1.06× faster) 🟢87.89 GiB88.65 GiB+0.9%3.97 MiB3.97 MiB+0.1%37.6 ms27.4 ms-27.2% (1.37× faster) 🟢168.67B164.80B-2.3%
Nat.add_comm3.619 s3.077 s-15.0% (1.18× faster) 🟢19.611 s17.954 s-8.5% (1.09× faster) 🟢2.3502.560+8.9% (1.09× faster) 🟢58.66 GiB58.03 GiB-1.1%3.97 MiB3.98 MiB+0.2%27.6 ms22.0 ms-20.4% (1.26× faster) 🟢130.84B124.20B-5.1% (1.05× fewer) 🟢
Pipeline total (7 constants)
constanttotal-time (main)total-time (PR)Δ%pipeline-throughput (const/s) (main)pipeline-throughput (const/s) (PR)Δ%pipeline-peak-ram (main)pipeline-peak-ram (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append1m 4.4s1m 1.0s-5.3% (1.06× faster) 🟢43.12045.520+5.6% (1.06× faster) 🟢101.44 GiB99.74 GiB-1.7%
Char.ofOrdinal_le_of_le59.164 s57.172 s-3.4% 🟢46.70048.330+3.5% 🟢99.81 GiB101.09 GiB+1.3%
Array.extract_append54.887 s52.851 s-3.7% 🟢29.26030.390+3.9% 🟢95.00 GiB95.36 GiB+0.4%
Std.HashMap47.243 s45.577 s-3.5% 🟢43.22044.800+3.7% 🟢94.50 GiB95.90 GiB+1.5%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq47.193 s43.788 s-7.2% (1.08× faster) 🟢39.56042.640+7.8% (1.08× faster) 🟢98.72 GiB95.35 GiB-3.4% 🟢
String.append31.017 s29.270 s-5.6% (1.06× faster) 🟢10.54011.170+6.0% (1.06× faster) 🟢87.89 GiB88.65 GiB+0.9%
Nat.add_comm20.680 s18.939 s-8.4% (1.09× faster) 🟢2.2202.430+9.5% (1.09× faster) 🟢58.66 GiB58.03 GiB-1.1%

Workflow logs

#606)
* Authenticate all frontier group members in the recursive verifier
The in-circuit pruned-multiproof walk (mmcs_verify_multi / frontier_level)
collapses queries that share a parent to a single lead node and hashes
only the lead's rows via inject_maybe(ar, ...). Non-lead members' opened
rows for the shorter (injected) matrices were still consumed in their own
per-query FRI arithmetic (batch_views_at) but never authenticated against
any commitment — the leaf hash covers only the tallest matrices, and
shorter ones are bound solely through injection. A prover could therefore
forge a non-lead member's shorter-matrix opening. Plonky3's reference
verify_batch_pruned guards exactly this with InconsistentGroupOpening
(and InconsistentDuplicateOpenings for equal-index queries); the port had
neither. The prior per-query walk did not have the gap, so it was
introduced with the direct multiproof consumption.
- frontier_level: on a group merge, assert the lead and member agree on
every not-yet-injected matrix (height <= next_lh) via select_rows_le +
pointer equality. Transitive across pairwise merges, so the whole group
is pinned; matches InconsistentGroupOpening.
- frontier_merge: duplicate transcript indices must open the SAME full
rows, not merely the same tallest-matrix leaf digest; matches
InconsistentDuplicateOpenings.
Pointer equality is admissible inside assert_eq! (equal pointers imply
equal content; a spurious mismatch costs only completeness — see
IxVM.Core). select_rows_le selects rows of matrices at height <= target,
mirroring select_rows.
Validated: the group-merge branch is genuinely reached by the factorial
recursion proof (an always-false variant of the new assert fails the
honest test), the honest proof still verifies with the real assert
(completeness preserved), the existing tamper tests still reject, and
the full lake test suite is green (2717 checks). aiur_multi_stark.rs
regenerated; kernel executor unchanged.
* Drop the multi_stark::advice dependency
Companion to multi-stark removing its unused per-query advice module.
ix consumed native pruned multiproofs directly and referenced only
advice::AdviceError, whose two arms (verification failed, serialization
failed) were immediately string-formatted by the FFI. Replace it with a
plain Result<Vec<u8>, String>: AiurSystem::proof_to_advice_bytes maps
both failures to a message, and the FFI passes the string straight to
LeanExcept::error_string. No behavior change; the Lean binding
(Except String ByteArray) is unaffected.
Bump the multi-stark pin to the advice-removed revision. Requires that
multi-stark's ap/bump-p3-drop-advice be pushed first, exactly as with
every other pin in this series.
* Bump multi-stark audit revision
---------
Co-authored-by: Arthur Paulino <arthurleonardo.ap@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@arthurpaulino@samuelburnham
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Update Aiur to Plonky3 0.6 - #605

Draft
arthurpaulino wants to merge 5 commits into
mainfrom
ap/bump-p3
Draft

Update Aiur to Plonky3 0.6#605
arthurpaulino wants to merge 5 commits into
mainfrom
ap/bump-p3

Conversation

@arthurpaulino

Copy link
Copy Markdown
Member

Update the multi-stark dependency and Rust toolchain for Plonky3 0.6, along with the Rust 1.98 lint migrations required to keep the workspace warning-free. Refresh the Rust-compatible Blake3.lean pin in both root and compile-package manifests.

Adapt recursive Aiur verification to Plonky3's pruned FRI multiproofs. Native proofs retain their compact serialized representation and native verification path; the FFI expands authenticated Merkle frontiers into per-query advice only when entering the existing recursive verifier circuit.

Preserve the packed claim-digest convention in the recursion diagnostic and exercise the advice boundary in the end-to-end test and benchmark paths. CPU and CUDA recursive q1 runs produce identical 823,485-byte inner proofs and 331,273-byte outer proofs.

The q50 Vector.extract_append workload retains identical CPU/CUDA proof sizes. Inner plus outer STARK proving measures 65.87s on CPU and 8.81s with CUDA on the RTX PRO 6000, a 7.48x speedup.

Update the multi-stark dependency and Rust toolchain for Plonky3 0.6, along with the Rust 1.98 lint migrations required to keep the workspace warning-free. Refresh the Rust-compatible Blake3.lean pin in both root and compile-package manifests.
Adapt recursive Aiur verification to Plonky3's pruned FRI multiproofs. Native proofs retain their compact serialized representation and native verification path; the FFI expands authenticated Merkle frontiers into per-query advice only when entering the existing recursive verifier circuit.
Preserve the packed claim-digest convention in the recursion diagnostic and exercise the advice boundary in the end-to-end test and benchmark paths. CPU and CUDA recursive q1 runs produce identical 823,485-byte inner proofs and 331,273-byte outer proofs.
The q50 Vector.extract_append workload retains identical CPU/CUDA proof sizes. Inner plus outer STARK proving measures 65.87s on CPU and 8.81s with CUDA on the RTX PRO 6000, a 7.48x speedup.
PR benchmark runs execute trusted workflow YAML from the default branch while loading composite actions from the PR checkout. When Bencher data and binary caches are unavailable, the workflow checks out main under base/ and asks Lake to rebuild it without first installing the Rust channel pinned by that checkout.
Teach the existing CPU provenance action to install the base checkout's validated Rust channel and profile immediately before an uncached base build. The step is a no-op when the toolchain is already available and leaves cached benchmark comparisons unchanged.
@arthurpaulino

Copy link
Copy Markdown
MemberAuthor

!benchmark fresh

@argument-ci-bot

argument-ci-botBot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

!benchmark — main vs eab18e4

backends: aiur=prove · envs: InitStd · baseline: fresh (benchmark products rebuilt, base-SHA run, bencher bypassed)

aiur · InitStd · prove — main from: base run @ e1ca8e2 (fresh — bencher bypassed)

7 constants · 1 with regressions · 7 with improvements (|Δ| > 3.0% on any metric).

IxVM on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append8.816 s8.990 s+2.0%30.095 s29.094 s-3.3% 🟢92.21095.380+3.4% 🟢70.72 GiB72.35 GiB+2.3%11.33 MiB11.07 MiB-2.3%68.7 ms60.4 ms-12.1% (1.14× faster) 🟢134.35B134.35B+0.0%
Char.ofOrdinal_le_of_le6.689 s6.790 s+1.5%25.631 s25.209 s-1.6%107.800109.600+1.7%63.87 GiB65.41 GiB+2.4%11.33 MiB11.07 MiB-2.3%73.7 ms60.3 ms-18.2% (1.22× faster) 🟢102.60B102.60B+0.0%
Array.extract_append6.679 s6.415 s-4.0% 🟢23.076 s22.072 s-4.4% 🟢69.60072.760+4.5% 🟢52.03 GiB52.79 GiB+1.4%11.24 MiB10.99 MiB-2.2%72.3 ms64.6 ms-10.6% (1.12× faster) 🟢97.08B97.08B+0.0%
Std.HashMap3.964 s4.002 s+1.0%15.519 s15.133 s-2.5%131.580134.940+2.6%36.34 GiB37.08 GiB+2.1%11.26 MiB11.00 MiB-2.3%75.1 ms65.5 ms-12.8% (1.15× faster) 🟢61.88B61.88B+0.0%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq3.598 s3.528 s-1.9%14.327 s13.806 s-3.6% 🟢130.310135.230+3.8% 🟢33.96 GiB34.69 GiB+2.1%11.26 MiB11.01 MiB-2.2%69.9 ms59.5 ms-15.0% (1.18× faster) 🟢55.68B55.68B+0.0%
String.append424.3 ms423.0 ms-0.3%2.278 s2.124 s-6.7% (1.07× faster) 🟢143.540153.930+7.2% (1.07× faster) 🟢4.89 GiB5.00 GiB+2.4%9.94 MiB9.74 MiB-2.0%64.2 ms52.0 ms-19.0% (1.24× faster) 🟢3.37B3.37B+0.0%
Nat.add_comm260.1 ms259.8 ms-0.1%1.068 s970.5 ms-9.1% (1.10× faster) 🟢43.07047.400+10.1% (1.10× faster) 🟢3.99 GiB4.64 GiB+16.4% (1.16× larger) ⚠️9.09 MiB8.91 MiB-2.0%53.9 ms41.9 ms-22.3% (1.29× faster) 🟢308.40M308.40M+0.0%
FRI verifier on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append5.488 s5.490 s+0.0%31.664 s30.699 s-3.0% 🟢87.64090.400+3.1% 🟢101.37 GiB101.07 GiB-0.3%3.97 MiB3.72 MiB-6.3% (1.07× smaller) 🟢29.2 ms23.8 ms-18.5% (1.23× faster) 🟢210.23B210.23B+0.0%
Char.ofOrdinal_le_of_le5.386 s5.489 s+1.9%30.846 s30.914 s+0.2%89.57089.380-0.2%100.32 GiB100.32 GiB+0.0%3.97 MiB3.72 MiB-6.3% (1.07× smaller) 🟢24.8 ms23.1 ms-6.6% (1.07× faster) 🟢207.18B207.18B+0.0%
Array.extract_append5.209 s5.235 s+0.5%29.948 s28.844 s-3.7% 🟢53.63055.680+3.8% 🟢94.59 GiB95.37 GiB+0.8%3.97 MiB3.73 MiB-6.1% (1.07× smaller) 🟢25.7 ms21.3 ms-17.1% (1.21× faster) 🟢200.65B200.65B+0.0%
Std.HashMap5.314 s5.306 s-0.1%29.652 s29.584 s-0.2%68.87069.020+0.2%94.56 GiB95.34 GiB+0.8%3.97 MiB3.72 MiB-6.3% (1.07× smaller) 🟢25.2 ms21.2 ms-15.7% (1.19× faster) 🟢203.35B203.35B+0.0%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq5.400 s5.333 s-1.2%31.051 s30.174 s-2.8%60.13061.870+2.9%98.68 GiB99.48 GiB+0.8%3.97 MiB3.73 MiB-6.2% (1.07× smaller) 🟢30.2 ms31.0 ms+2.5%205.59B205.59B+0.0%
String.append4.334 s4.419 s+2.0%27.489 s26.612 s-3.2% 🟢11.90012.290+3.3% 🟢87.87 GiB88.58 GiB+0.8%3.97 MiB3.73 MiB-6.1% (1.07× smaller) 🟢34.5 ms22.7 ms-34.3% (1.52× faster) 🟢168.67B168.67B+0.0%
Nat.add_comm3.518 s3.499 s-0.5%18.993 s17.942 s-5.5% (1.06× faster) 🟢2.4202.560+5.8% (1.06× faster) 🟢58.70 GiB59.42 GiB+1.2%3.97 MiB3.72 MiB-6.2% (1.07× smaller) 🟢24.8 ms20.8 ms-16.1% (1.19× faster) 🟢130.84B130.84B+0.0%
Pipeline total (7 constants)
constanttotal-time (main)total-time (PR)Δ%pipeline-throughput (const/s) (main)pipeline-throughput (const/s) (PR)Δ%pipeline-peak-ram (main)pipeline-peak-ram (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append1m 1.8s59.792 s-3.2% 🟢44.93046.410+3.3% 🟢101.37 GiB101.07 GiB-0.3%
Char.ofOrdinal_le_of_le56.477 s56.124 s-0.6%48.92049.230+0.6%100.32 GiB100.32 GiB+0.0%
Array.extract_append53.024 s50.916 s-4.0% 🟢30.29031.540+4.1% 🟢94.59 GiB95.37 GiB+0.8%
Std.HashMap45.172 s44.717 s-1.0%45.21045.660+1.0%94.56 GiB95.34 GiB+0.8%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq45.377 s43.980 s-3.1% 🟢41.14042.450+3.2% 🟢98.68 GiB99.48 GiB+0.8%
String.append29.767 s28.737 s-3.5% 🟢10.99011.380+3.5% 🟢87.87 GiB88.58 GiB+0.8%
Nat.add_comm20.061 s18.912 s-5.7% (1.06× faster) 🟢2.2902.430+6.1% (1.06× faster) 🟢58.70 GiB59.42 GiB+1.2%

Workflow logs

Consume the Plonky3 0.6 batch-opening layout directly in Aiur instead of expanding every pruned Merkle frontier into one authentication path per FRI query. Sample all query indices from the unchanged transcript, sort and deduplicate them with an O(q log q) merge sort, authenticate each input and commit-phase commitment once, then retain the existing per-query reduced-opening and FRI arithmetic.
Bind every frontier to transcript-derived indices, consume boundary digests in Plonky3's level/parent/child order, reject trailing frontier elements and inconsistent duplicate leaves, and assert all native opening dimensions and sibling counts. Explicitly constrain the digest-bound protocol specialization to cap height 0, binary FRI, and a constant final polynomial. Move memo_u32_less_than into IxVM Core so both substitution and multiproof sorting share its constrained rows.
Strengthen the recursive negative test to mutate a structurally valid stage-1 commitment. Regenerate both checked-in Aiur Rust executors and retain interpreter/codegen query-count parity.
On Vector.extract_append q50, recursive-verifier FFT cost falls from 204.073B to 201.166B. CPU outer proving improves from 50.09s to 45.03s and the full CPU pipeline from 90.64s to 82.90s. GPU outer proving improves from 15.85s to 13.72s and the full GPU pipeline from 28.86s to 26.69s. The outer proof grows from 3.92 MB to 4.17 MB.
Validated with the MultiStark primitive suite, recursive honest/tamper/parity tests, codegen --check, release workspace clippy, release CUDA clippy, rustfmt, and diff checks.
@arthurpaulino

Copy link
Copy Markdown
MemberAuthor

!benchmark fresh

@argument-ci-bot

argument-ci-botBot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

!benchmark — main vs 840932a

backends: aiur=prove · envs: InitStd · baseline: fresh (benchmark products rebuilt, base-SHA run, bencher bypassed)

Warning

  • CPU model mismatch for PR benchmark binaries in this job: built on AMD EPYC 9R45; measured on Intel(R) Xeon(R) 6975P-C. Native Rust code uses -Ctarget-cpu=native.

aiur · InitStd · prove — main from: base run @ e1ca8e2 (fresh — bencher bypassed)

7 constants · 0 with regressions · 0 with improvements (|Δ| > 3.0% on any metric).

IxVM on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append8.874 s💥 CRASHn/a44.291 s💥 CRASHn/a62.650💥 CRASHn/a70.80 GiB💥 CRASHn/a11.33 MiB💥 CRASHn/a77.4 ms💥 CRASHn/a134.35B💥 CRASHn/a
Char.ofOrdinal_le_of_le6.603 s💥 CRASHn/a38.242 s💥 CRASHn/a72.250💥 CRASHn/a63.81 GiB💥 CRASHn/a11.33 MiB💥 CRASHn/a76.9 ms💥 CRASHn/a102.60B💥 CRASHn/a
Array.extract_append6.311 s💥 CRASHn/a33.420 s💥 CRASHn/a48.060💥 CRASHn/a51.95 GiB💥 CRASHn/a11.24 MiB💥 CRASHn/a86.5 ms💥 CRASHn/a97.08B💥 CRASHn/a
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq3.774 s💥 CRASHn/a20.628 s💥 CRASHn/a90.510💥 CRASHn/a33.97 GiB💥 CRASHn/a11.26 MiB💥 CRASHn/a82.6 ms💥 CRASHn/a55.68B💥 CRASHn/a
Std.HashMap4.060 s💥 CRASHn/a22.501 s💥 CRASHn/a90.750💥 CRASHn/a36.36 GiB💥 CRASHn/a11.26 MiB💥 CRASHn/a84.1 ms💥 CRASHn/a61.88B💥 CRASHn/a
String.append708.3 ms💥 CRASHn/a2.863 s💥 CRASHn/a114.220💥 CRASHn/a4.99 GiB💥 CRASHn/a9.94 MiB💥 CRASHn/a70.3 ms💥 CRASHn/a3.37B💥 CRASHn/a
Nat.add_comm496.0 ms💥 CRASHn/a1.330 s💥 CRASHn/a34.580💥 CRASHn/a4.33 GiB💥 CRASHn/a9.09 MiB💥 CRASHn/a58.9 ms💥 CRASHn/a308.40M💥 CRASHn/a
FRI verifier on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append6.880 s💥 CRASHn/a54.265 s💥 CRASHn/a51.140💥 CRASHn/a101.05 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a49.5 ms💥 CRASHn/a210.23B💥 CRASHn/a
Char.ofOrdinal_le_of_le6.813 s💥 CRASHn/a53.449 s💥 CRASHn/a51.690💥 CRASHn/a99.81 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a45.1 ms💥 CRASHn/a207.18B💥 CRASHn/a
Array.extract_append6.371 s💥 CRASHn/a50.695 s💥 CRASHn/a31.680💥 CRASHn/a95.19 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a58.1 ms💥 CRASHn/a200.65B💥 CRASHn/a
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq6.748 s💥 CRASHn/a53.378 s💥 CRASHn/a34.980💥 CRASHn/a98.71 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a46.1 ms💥 CRASHn/a205.59B💥 CRASHn/a
Std.HashMap6.455 s💥 CRASHn/a50.868 s💥 CRASHn/a40.140💥 CRASHn/a95.03 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a45.7 ms💥 CRASHn/a203.35B💥 CRASHn/a
String.append5.307 s💥 CRASHn/a47.595 s💥 CRASHn/a6.870💥 CRASHn/a87.89 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a32.4 ms💥 CRASHn/a168.67B💥 CRASHn/a
Nat.add_comm4.526 s💥 CRASHn/a31.415 s💥 CRASHn/a1.460💥 CRASHn/a58.66 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a28.0 ms💥 CRASHn/a130.84B💥 CRASHn/a
Pipeline total (7 constants)
constanttotal-time (main)total-time (PR)Δ%pipeline-throughput (const/s) (main)pipeline-throughput (const/s) (PR)Δ%pipeline-peak-ram (main)pipeline-peak-ram (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append1m 38.6s💥 CRASHn/a28.160💥 CRASHn/a101.05 GiB💥 CRASHn/a
Char.ofOrdinal_le_of_le1m 31.7s💥 CRASHn/a30.130💥 CRASHn/a99.81 GiB💥 CRASHn/a
Array.extract_append1m 24.1s💥 CRASHn/a19.090💥 CRASHn/a95.19 GiB💥 CRASHn/a
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq1m 14.0s💥 CRASHn/a25.230💥 CRASHn/a98.71 GiB💥 CRASHn/a
Std.HashMap1m 13.4s💥 CRASHn/a27.830💥 CRASHn/a95.03 GiB💥 CRASHn/a
String.append50.458 s💥 CRASHn/a6.480💥 CRASHn/a87.89 GiB💥 CRASHn/a
Nat.add_comm32.745 s💥 CRASHn/a1.400💥 CRASHn/a58.66 GiB💥 CRASHn/a

Workflow logs

samuelburnham added a commit that referenced this pull request Sep 1, 2026
The Warp x64 runner pool mixes Intel Granite Rapids and AMD Zen 5, and a
build job may land on one vendor while the job that runs its binaries
lands on the other. Neither vendor's feature set contains the other's, so
`-Ctarget-cpu=native` does not produce a portable binary: Zen 5 enables
SSE4A, and LLVM emits it. Disassembling the workspace built for znver5
finds 31 SSE4A instructions, all INSERTQ, in `ix-ffi` and in
`aiur_ixvm_witness::add_entries_parallel`. Granite Rapids has no SSE4A,
so the first one executed raises #UD, killing the process with SIGILL
during witness generation. That is what turned every row of #605's
benchmark into a crash.
Pin the measured intersection of the two CPUs instead. x86-64-v4 covers
every AVX-512 subset Plonky3 uses; +avx512vbmi2 preserves its VPSHRDQ
interleave and +gfni preserves LLVM's byte-shift lowering. A workspace
built with these flags contains no instruction absent from either vendor
and has an instruction vocabulary identical to a graniterapids build.
blake3 dispatches on CPUID at runtime and is unaffected either way.
`.cargo/config.toml` keeps `-Ctarget-cpu=native`: a developer builds and
runs on one machine, and x86-64-v4 would exclude every host without
AVX-512. Only CI has the split, so only CI pins the ISA. The new guard
fails the job when a runner lacks a required feature, so the assumption
is enforced rather than assumed, and the shared `warp-x64` cargo cache
key becomes sound now that codegen no longer varies by host.
Pinning also removes a benchmarking hazard that never crashed: LLVM sets
prefer-256-bit for Granite Rapids but not for Zen 5, so the same source
vectorized 3.2x more widely depending on the build host, and main-vs-PR
timings were not comparable across a vendor split.
Drop the CPU-mismatch warning the benchmark comment used to carry. It
detected a real problem, but the flags above prevent that problem, and
computing it in one job to render it in another cost a Markdown file
threaded through cache entries, artifacts, and a `--warning-file` flag on
`ix bench compare`. Warnings belong to the run that finds them.
@arthurpaulino

Copy link
Copy Markdown
MemberAuthor

!benchmark fresh

@argument-ci-bot

argument-ci-botBot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

⚠️ Some benchmark jobs failed — results may be partial.

!benchmark — main vs d7fc2f1

backends: aiur=prove · envs: InitStd · baseline: fresh (benchmark products rebuilt, base-SHA run, bencher bypassed)

No result tables were produced — see the run logs.

Workflow logs

The Warp x64 runner pool mixes Intel Granite Rapids and AMD Zen 5, and a
build job may land on one vendor while the job that runs its binaries
lands on the other. Neither vendor's feature set contains the other's, so
`-Ctarget-cpu=native` does not produce a portable binary: Zen 5 enables
SSE4A, and LLVM emits it. Disassembling the workspace built for znver5
finds 31 SSE4A instructions, all INSERTQ, in `ix-ffi` and in
`aiur_ixvm_witness::add_entries_parallel`. Granite Rapids has no SSE4A,
so the first one executed raises #UD, killing the process with SIGILL
during witness generation. That is what turned every row of #605's
benchmark into a crash.
Pin the measured intersection of the two CPUs instead. x86-64-v4 covers
every AVX-512 subset Plonky3 uses; +avx512vbmi2 preserves its VPSHRDQ
interleave and +gfni preserves LLVM's byte-shift lowering. A workspace
built with these flags contains no instruction absent from either vendor
and has an instruction vocabulary identical to a graniterapids build.
blake3 dispatches on CPUID at runtime and is unaffected either way.
`.cargo/config.toml` keeps `-Ctarget-cpu=native`: a developer builds and
runs on one machine, and x86-64-v4 would exclude every host without
AVX-512. Only CI has the split, so only CI pins the ISA. The new guard
fails the job when a runner lacks a required feature, so the assumption
is enforced rather than assumed, and the shared `warp-x64` cargo cache
key becomes sound now that codegen no longer varies by host. RUSTFLAGS
is hashed into that key, so the flag change rotates it on its own.
Pinning also removes a benchmarking hazard that never crashed: LLVM sets
prefer-256-bit for Granite Rapids but not for Zen 5, so the same source
vectorized 3.2x more widely depending on the build host, and main-vs-PR
timings were not comparable across a vendor split.
@samuelburnham

Copy link
Copy Markdown
Member

!benchmark fresh

@argument-ci-bot

argument-ci-botBot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

!benchmark — main vs feb014d

backends: aiur=prove · envs: InitStd · baseline: fresh (benchmark products rebuilt, base-SHA run, bencher bypassed)

aiur · InitStd · prove — main from: base run @ e1ca8e2 (fresh — bencher bypassed)

7 constants · 3 with regressions · 7 with improvements (|Δ| > 3.0% on any metric).

IxVM on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append9.269 s9.073 s-2.1%31.250 s30.269 s-3.1% 🟢88.80091.680+3.2% 🟢70.72 GiB72.28 GiB+2.2%11.33 MiB11.06 MiB-2.4%68.5 ms62.7 ms-8.5% (1.09× faster) 🟢134.35B134.35B+0.0%
Char.ofOrdinal_le_of_le6.964 s6.879 s-1.2%26.884 s25.793 s-4.1% 🟢102.780107.120+4.2% 🟢63.83 GiB65.31 GiB+2.3%11.33 MiB11.07 MiB-2.3%77.9 ms59.3 ms-23.8% (1.31× faster) 🟢102.60B102.60B+0.0%
Array.extract_append6.676 s6.751 s+1.1%23.925 s23.179 s-3.1% 🟢67.13069.290+3.2% 🟢51.97 GiB52.72 GiB+1.4%11.24 MiB10.99 MiB-2.2%71.1 ms57.2 ms-19.5% (1.24× faster) 🟢97.08B97.08B+0.0%
Std.HashMap4.171 s4.079 s-2.2%16.222 s15.628 s-3.7% 🟢125.880130.670+3.8% 🟢36.29 GiB37.10 GiB+2.2%11.26 MiB11.01 MiB-2.2%74.8 ms66.6 ms-11.0% (1.12× faster) 🟢61.88B61.88B+0.0%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq3.697 s3.648 s-1.3%14.858 s14.287 s-3.8% 🟢125.650130.680+4.0% 🟢34.02 GiB34.75 GiB+2.1%11.26 MiB11.00 MiB-2.2%75.5 ms57.7 ms-23.6% (1.31× faster) 🟢55.68B55.68B+0.0%
String.append435.5 ms434.0 ms-0.4%2.307 s2.129 s-7.7% (1.08× faster) 🟢141.750153.600+8.4% (1.08× faster) 🟢5.74 GiB5.52 GiB-3.9% 🟢9.94 MiB9.74 MiB-2.1%64.4 ms50.2 ms-22.1% (1.28× faster) 🟢3.37B3.37B+0.0%
Nat.add_comm267.1 ms264.3 ms-1.1%1.068 s984.9 ms-7.8% (1.08× faster) 🟢43.06046.710+8.5% (1.08× faster) 🟢4.51 GiB3.99 GiB-11.5% (1.13× smaller) 🟢9.09 MiB8.90 MiB-2.1%53.3 ms47.9 ms-10.2% (1.11× faster) 🟢308.40M308.40M+0.0%
FRI verifier on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append5.680 s4.975 s-12.4% (1.14× faster) 🟢33.109 s30.692 s-7.3% (1.08× faster) 🟢83.82090.410+7.9% (1.08× faster) 🟢101.44 GiB99.74 GiB-1.7%3.97 MiB3.98 MiB+0.2%26.3 ms22.3 ms-15.4% (1.18× faster) 🟢210.23B203.74B-3.1% 🟢
Char.ofOrdinal_le_of_le5.599 s5.014 s-10.5% (1.12× faster) 🟢32.280 s31.379 s-2.8%85.59088.050+2.9%99.81 GiB101.09 GiB+1.3%3.97 MiB3.98 MiB+0.2%27.6 ms28.6 ms+3.6% ⚠️207.18B208.08B+0.4%
Array.extract_append5.315 s4.858 s-8.6% (1.09× faster) 🟢30.962 s29.672 s-4.2% 🟢51.87054.120+4.3% 🟢95.00 GiB95.36 GiB+0.4%3.97 MiB3.99 MiB+0.4%26.0 ms39.4 ms+51.3% (1.51× slower) ⚠️200.65B200.40B-0.1%
Std.HashMap5.445 s4.864 s-10.7% (1.12× faster) 🟢31.021 s29.949 s-3.5% 🟢65.83068.180+3.6% 🟢94.50 GiB95.90 GiB+1.5%3.97 MiB3.98 MiB+0.2%25.4 ms25.6 ms+0.8%203.35B204.11B+0.4%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq5.496 s4.898 s-10.9% (1.12× faster) 🟢32.335 s29.501 s-8.8% (1.10× faster) 🟢57.74063.290+9.6% (1.10× faster) 🟢98.72 GiB95.35 GiB-3.4% 🟢3.97 MiB3.98 MiB+0.3%25.9 ms31.9 ms+23.0% (1.23× slower) ⚠️205.59B199.65B-2.9%
String.append4.487 s3.955 s-11.8% (1.13× faster) 🟢28.710 s27.141 s-5.5% (1.06× faster) 🟢11.39012.050+5.8% (1.06× faster) 🟢87.89 GiB88.65 GiB+0.9%3.97 MiB3.97 MiB+0.1%37.6 ms27.4 ms-27.2% (1.37× faster) 🟢168.67B164.80B-2.3%
Nat.add_comm3.619 s3.077 s-15.0% (1.18× faster) 🟢19.611 s17.954 s-8.5% (1.09× faster) 🟢2.3502.560+8.9% (1.09× faster) 🟢58.66 GiB58.03 GiB-1.1%3.97 MiB3.98 MiB+0.2%27.6 ms22.0 ms-20.4% (1.26× faster) 🟢130.84B124.20B-5.1% (1.05× fewer) 🟢
Pipeline total (7 constants)
constanttotal-time (main)total-time (PR)Δ%pipeline-throughput (const/s) (main)pipeline-throughput (const/s) (PR)Δ%pipeline-peak-ram (main)pipeline-peak-ram (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append1m 4.4s1m 1.0s-5.3% (1.06× faster) 🟢43.12045.520+5.6% (1.06× faster) 🟢101.44 GiB99.74 GiB-1.7%
Char.ofOrdinal_le_of_le59.164 s57.172 s-3.4% 🟢46.70048.330+3.5% 🟢99.81 GiB101.09 GiB+1.3%
Array.extract_append54.887 s52.851 s-3.7% 🟢29.26030.390+3.9% 🟢95.00 GiB95.36 GiB+0.4%
Std.HashMap47.243 s45.577 s-3.5% 🟢43.22044.800+3.7% 🟢94.50 GiB95.90 GiB+1.5%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq47.193 s43.788 s-7.2% (1.08× faster) 🟢39.56042.640+7.8% (1.08× faster) 🟢98.72 GiB95.35 GiB-3.4% 🟢
String.append31.017 s29.270 s-5.6% (1.06× faster) 🟢10.54011.170+6.0% (1.06× faster) 🟢87.89 GiB88.65 GiB+0.9%
Nat.add_comm20.680 s18.939 s-8.4% (1.09× faster) 🟢2.2202.430+9.5% (1.09× faster) 🟢58.66 GiB58.03 GiB-1.1%

Workflow logs

#606)
* Authenticate all frontier group members in the recursive verifier
The in-circuit pruned-multiproof walk (mmcs_verify_multi / frontier_level)
collapses queries that share a parent to a single lead node and hashes
only the lead's rows via inject_maybe(ar, ...). Non-lead members' opened
rows for the shorter (injected) matrices were still consumed in their own
per-query FRI arithmetic (batch_views_at) but never authenticated against
any commitment — the leaf hash covers only the tallest matrices, and
shorter ones are bound solely through injection. A prover could therefore
forge a non-lead member's shorter-matrix opening. Plonky3's reference
verify_batch_pruned guards exactly this with InconsistentGroupOpening
(and InconsistentDuplicateOpenings for equal-index queries); the port had
neither. The prior per-query walk did not have the gap, so it was
introduced with the direct multiproof consumption.
- frontier_level: on a group merge, assert the lead and member agree on
every not-yet-injected matrix (height <= next_lh) via select_rows_le +
pointer equality. Transitive across pairwise merges, so the whole group
is pinned; matches InconsistentGroupOpening.
- frontier_merge: duplicate transcript indices must open the SAME full
rows, not merely the same tallest-matrix leaf digest; matches
InconsistentDuplicateOpenings.
Pointer equality is admissible inside assert_eq! (equal pointers imply
equal content; a spurious mismatch costs only completeness — see
IxVM.Core). select_rows_le selects rows of matrices at height <= target,
mirroring select_rows.
Validated: the group-merge branch is genuinely reached by the factorial
recursion proof (an always-false variant of the new assert fails the
honest test), the honest proof still verifies with the real assert
(completeness preserved), the existing tamper tests still reject, and
the full lake test suite is green (2717 checks). aiur_multi_stark.rs
regenerated; kernel executor unchanged.
* Drop the multi_stark::advice dependency
Companion to multi-stark removing its unused per-query advice module.
ix consumed native pruned multiproofs directly and referenced only
advice::AdviceError, whose two arms (verification failed, serialization
failed) were immediately string-formatted by the FFI. Replace it with a
plain Result<Vec<u8>, String>: AiurSystem::proof_to_advice_bytes maps
both failures to a message, and the FFI passes the string straight to
LeanExcept::error_string. No behavior change; the Lean binding
(Except String ByteArray) is unaffected.
Bump the multi-stark pin to the advice-removed revision. Requires that
multi-stark's ap/bump-p3-drop-advice be pushed first, exactly as with
every other pin in this series.
* Bump multi-stark audit revision
---------
Co-authored-by: Arthur Paulino <arthurleonardo.ap@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@arthurpaulino@samuelburnham
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Update Aiur to Plonky3 0.6 - #605

Draft
arthurpaulino wants to merge 5 commits into
mainfrom
ap/bump-p3
Draft

Update Aiur to Plonky3 0.6#605
arthurpaulino wants to merge 5 commits into
mainfrom
ap/bump-p3

Conversation

@arthurpaulino

Copy link
Copy Markdown
Member

Update the multi-stark dependency and Rust toolchain for Plonky3 0.6, along with the Rust 1.98 lint migrations required to keep the workspace warning-free. Refresh the Rust-compatible Blake3.lean pin in both root and compile-package manifests.

Adapt recursive Aiur verification to Plonky3's pruned FRI multiproofs. Native proofs retain their compact serialized representation and native verification path; the FFI expands authenticated Merkle frontiers into per-query advice only when entering the existing recursive verifier circuit.

Preserve the packed claim-digest convention in the recursion diagnostic and exercise the advice boundary in the end-to-end test and benchmark paths. CPU and CUDA recursive q1 runs produce identical 823,485-byte inner proofs and 331,273-byte outer proofs.

The q50 Vector.extract_append workload retains identical CPU/CUDA proof sizes. Inner plus outer STARK proving measures 65.87s on CPU and 8.81s with CUDA on the RTX PRO 6000, a 7.48x speedup.

Update the multi-stark dependency and Rust toolchain for Plonky3 0.6, along with the Rust 1.98 lint migrations required to keep the workspace warning-free. Refresh the Rust-compatible Blake3.lean pin in both root and compile-package manifests.
Adapt recursive Aiur verification to Plonky3's pruned FRI multiproofs. Native proofs retain their compact serialized representation and native verification path; the FFI expands authenticated Merkle frontiers into per-query advice only when entering the existing recursive verifier circuit.
Preserve the packed claim-digest convention in the recursion diagnostic and exercise the advice boundary in the end-to-end test and benchmark paths. CPU and CUDA recursive q1 runs produce identical 823,485-byte inner proofs and 331,273-byte outer proofs.
The q50 Vector.extract_append workload retains identical CPU/CUDA proof sizes. Inner plus outer STARK proving measures 65.87s on CPU and 8.81s with CUDA on the RTX PRO 6000, a 7.48x speedup.
PR benchmark runs execute trusted workflow YAML from the default branch while loading composite actions from the PR checkout. When Bencher data and binary caches are unavailable, the workflow checks out main under base/ and asks Lake to rebuild it without first installing the Rust channel pinned by that checkout.
Teach the existing CPU provenance action to install the base checkout's validated Rust channel and profile immediately before an uncached base build. The step is a no-op when the toolchain is already available and leaves cached benchmark comparisons unchanged.
@arthurpaulino

Copy link
Copy Markdown
MemberAuthor

!benchmark fresh

@argument-ci-bot

argument-ci-botBot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

!benchmark — main vs eab18e4

backends: aiur=prove · envs: InitStd · baseline: fresh (benchmark products rebuilt, base-SHA run, bencher bypassed)

aiur · InitStd · prove — main from: base run @ e1ca8e2 (fresh — bencher bypassed)

7 constants · 1 with regressions · 7 with improvements (|Δ| > 3.0% on any metric).

IxVM on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append8.816 s8.990 s+2.0%30.095 s29.094 s-3.3% 🟢92.21095.380+3.4% 🟢70.72 GiB72.35 GiB+2.3%11.33 MiB11.07 MiB-2.3%68.7 ms60.4 ms-12.1% (1.14× faster) 🟢134.35B134.35B+0.0%
Char.ofOrdinal_le_of_le6.689 s6.790 s+1.5%25.631 s25.209 s-1.6%107.800109.600+1.7%63.87 GiB65.41 GiB+2.4%11.33 MiB11.07 MiB-2.3%73.7 ms60.3 ms-18.2% (1.22× faster) 🟢102.60B102.60B+0.0%
Array.extract_append6.679 s6.415 s-4.0% 🟢23.076 s22.072 s-4.4% 🟢69.60072.760+4.5% 🟢52.03 GiB52.79 GiB+1.4%11.24 MiB10.99 MiB-2.2%72.3 ms64.6 ms-10.6% (1.12× faster) 🟢97.08B97.08B+0.0%
Std.HashMap3.964 s4.002 s+1.0%15.519 s15.133 s-2.5%131.580134.940+2.6%36.34 GiB37.08 GiB+2.1%11.26 MiB11.00 MiB-2.3%75.1 ms65.5 ms-12.8% (1.15× faster) 🟢61.88B61.88B+0.0%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq3.598 s3.528 s-1.9%14.327 s13.806 s-3.6% 🟢130.310135.230+3.8% 🟢33.96 GiB34.69 GiB+2.1%11.26 MiB11.01 MiB-2.2%69.9 ms59.5 ms-15.0% (1.18× faster) 🟢55.68B55.68B+0.0%
String.append424.3 ms423.0 ms-0.3%2.278 s2.124 s-6.7% (1.07× faster) 🟢143.540153.930+7.2% (1.07× faster) 🟢4.89 GiB5.00 GiB+2.4%9.94 MiB9.74 MiB-2.0%64.2 ms52.0 ms-19.0% (1.24× faster) 🟢3.37B3.37B+0.0%
Nat.add_comm260.1 ms259.8 ms-0.1%1.068 s970.5 ms-9.1% (1.10× faster) 🟢43.07047.400+10.1% (1.10× faster) 🟢3.99 GiB4.64 GiB+16.4% (1.16× larger) ⚠️9.09 MiB8.91 MiB-2.0%53.9 ms41.9 ms-22.3% (1.29× faster) 🟢308.40M308.40M+0.0%
FRI verifier on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append5.488 s5.490 s+0.0%31.664 s30.699 s-3.0% 🟢87.64090.400+3.1% 🟢101.37 GiB101.07 GiB-0.3%3.97 MiB3.72 MiB-6.3% (1.07× smaller) 🟢29.2 ms23.8 ms-18.5% (1.23× faster) 🟢210.23B210.23B+0.0%
Char.ofOrdinal_le_of_le5.386 s5.489 s+1.9%30.846 s30.914 s+0.2%89.57089.380-0.2%100.32 GiB100.32 GiB+0.0%3.97 MiB3.72 MiB-6.3% (1.07× smaller) 🟢24.8 ms23.1 ms-6.6% (1.07× faster) 🟢207.18B207.18B+0.0%
Array.extract_append5.209 s5.235 s+0.5%29.948 s28.844 s-3.7% 🟢53.63055.680+3.8% 🟢94.59 GiB95.37 GiB+0.8%3.97 MiB3.73 MiB-6.1% (1.07× smaller) 🟢25.7 ms21.3 ms-17.1% (1.21× faster) 🟢200.65B200.65B+0.0%
Std.HashMap5.314 s5.306 s-0.1%29.652 s29.584 s-0.2%68.87069.020+0.2%94.56 GiB95.34 GiB+0.8%3.97 MiB3.72 MiB-6.3% (1.07× smaller) 🟢25.2 ms21.2 ms-15.7% (1.19× faster) 🟢203.35B203.35B+0.0%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq5.400 s5.333 s-1.2%31.051 s30.174 s-2.8%60.13061.870+2.9%98.68 GiB99.48 GiB+0.8%3.97 MiB3.73 MiB-6.2% (1.07× smaller) 🟢30.2 ms31.0 ms+2.5%205.59B205.59B+0.0%
String.append4.334 s4.419 s+2.0%27.489 s26.612 s-3.2% 🟢11.90012.290+3.3% 🟢87.87 GiB88.58 GiB+0.8%3.97 MiB3.73 MiB-6.1% (1.07× smaller) 🟢34.5 ms22.7 ms-34.3% (1.52× faster) 🟢168.67B168.67B+0.0%
Nat.add_comm3.518 s3.499 s-0.5%18.993 s17.942 s-5.5% (1.06× faster) 🟢2.4202.560+5.8% (1.06× faster) 🟢58.70 GiB59.42 GiB+1.2%3.97 MiB3.72 MiB-6.2% (1.07× smaller) 🟢24.8 ms20.8 ms-16.1% (1.19× faster) 🟢130.84B130.84B+0.0%
Pipeline total (7 constants)
constanttotal-time (main)total-time (PR)Δ%pipeline-throughput (const/s) (main)pipeline-throughput (const/s) (PR)Δ%pipeline-peak-ram (main)pipeline-peak-ram (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append1m 1.8s59.792 s-3.2% 🟢44.93046.410+3.3% 🟢101.37 GiB101.07 GiB-0.3%
Char.ofOrdinal_le_of_le56.477 s56.124 s-0.6%48.92049.230+0.6%100.32 GiB100.32 GiB+0.0%
Array.extract_append53.024 s50.916 s-4.0% 🟢30.29031.540+4.1% 🟢94.59 GiB95.37 GiB+0.8%
Std.HashMap45.172 s44.717 s-1.0%45.21045.660+1.0%94.56 GiB95.34 GiB+0.8%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq45.377 s43.980 s-3.1% 🟢41.14042.450+3.2% 🟢98.68 GiB99.48 GiB+0.8%
String.append29.767 s28.737 s-3.5% 🟢10.99011.380+3.5% 🟢87.87 GiB88.58 GiB+0.8%
Nat.add_comm20.061 s18.912 s-5.7% (1.06× faster) 🟢2.2902.430+6.1% (1.06× faster) 🟢58.70 GiB59.42 GiB+1.2%

Workflow logs

Consume the Plonky3 0.6 batch-opening layout directly in Aiur instead of expanding every pruned Merkle frontier into one authentication path per FRI query. Sample all query indices from the unchanged transcript, sort and deduplicate them with an O(q log q) merge sort, authenticate each input and commit-phase commitment once, then retain the existing per-query reduced-opening and FRI arithmetic.
Bind every frontier to transcript-derived indices, consume boundary digests in Plonky3's level/parent/child order, reject trailing frontier elements and inconsistent duplicate leaves, and assert all native opening dimensions and sibling counts. Explicitly constrain the digest-bound protocol specialization to cap height 0, binary FRI, and a constant final polynomial. Move memo_u32_less_than into IxVM Core so both substitution and multiproof sorting share its constrained rows.
Strengthen the recursive negative test to mutate a structurally valid stage-1 commitment. Regenerate both checked-in Aiur Rust executors and retain interpreter/codegen query-count parity.
On Vector.extract_append q50, recursive-verifier FFT cost falls from 204.073B to 201.166B. CPU outer proving improves from 50.09s to 45.03s and the full CPU pipeline from 90.64s to 82.90s. GPU outer proving improves from 15.85s to 13.72s and the full GPU pipeline from 28.86s to 26.69s. The outer proof grows from 3.92 MB to 4.17 MB.
Validated with the MultiStark primitive suite, recursive honest/tamper/parity tests, codegen --check, release workspace clippy, release CUDA clippy, rustfmt, and diff checks.
@arthurpaulino

Copy link
Copy Markdown
MemberAuthor

!benchmark fresh

@argument-ci-bot

argument-ci-botBot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

!benchmark — main vs 840932a

backends: aiur=prove · envs: InitStd · baseline: fresh (benchmark products rebuilt, base-SHA run, bencher bypassed)

Warning

  • CPU model mismatch for PR benchmark binaries in this job: built on AMD EPYC 9R45; measured on Intel(R) Xeon(R) 6975P-C. Native Rust code uses -Ctarget-cpu=native.

aiur · InitStd · prove — main from: base run @ e1ca8e2 (fresh — bencher bypassed)

7 constants · 0 with regressions · 0 with improvements (|Δ| > 3.0% on any metric).

IxVM on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append8.874 s💥 CRASHn/a44.291 s💥 CRASHn/a62.650💥 CRASHn/a70.80 GiB💥 CRASHn/a11.33 MiB💥 CRASHn/a77.4 ms💥 CRASHn/a134.35B💥 CRASHn/a
Char.ofOrdinal_le_of_le6.603 s💥 CRASHn/a38.242 s💥 CRASHn/a72.250💥 CRASHn/a63.81 GiB💥 CRASHn/a11.33 MiB💥 CRASHn/a76.9 ms💥 CRASHn/a102.60B💥 CRASHn/a
Array.extract_append6.311 s💥 CRASHn/a33.420 s💥 CRASHn/a48.060💥 CRASHn/a51.95 GiB💥 CRASHn/a11.24 MiB💥 CRASHn/a86.5 ms💥 CRASHn/a97.08B💥 CRASHn/a
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq3.774 s💥 CRASHn/a20.628 s💥 CRASHn/a90.510💥 CRASHn/a33.97 GiB💥 CRASHn/a11.26 MiB💥 CRASHn/a82.6 ms💥 CRASHn/a55.68B💥 CRASHn/a
Std.HashMap4.060 s💥 CRASHn/a22.501 s💥 CRASHn/a90.750💥 CRASHn/a36.36 GiB💥 CRASHn/a11.26 MiB💥 CRASHn/a84.1 ms💥 CRASHn/a61.88B💥 CRASHn/a
String.append708.3 ms💥 CRASHn/a2.863 s💥 CRASHn/a114.220💥 CRASHn/a4.99 GiB💥 CRASHn/a9.94 MiB💥 CRASHn/a70.3 ms💥 CRASHn/a3.37B💥 CRASHn/a
Nat.add_comm496.0 ms💥 CRASHn/a1.330 s💥 CRASHn/a34.580💥 CRASHn/a4.33 GiB💥 CRASHn/a9.09 MiB💥 CRASHn/a58.9 ms💥 CRASHn/a308.40M💥 CRASHn/a
FRI verifier on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append6.880 s💥 CRASHn/a54.265 s💥 CRASHn/a51.140💥 CRASHn/a101.05 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a49.5 ms💥 CRASHn/a210.23B💥 CRASHn/a
Char.ofOrdinal_le_of_le6.813 s💥 CRASHn/a53.449 s💥 CRASHn/a51.690💥 CRASHn/a99.81 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a45.1 ms💥 CRASHn/a207.18B💥 CRASHn/a
Array.extract_append6.371 s💥 CRASHn/a50.695 s💥 CRASHn/a31.680💥 CRASHn/a95.19 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a58.1 ms💥 CRASHn/a200.65B💥 CRASHn/a
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq6.748 s💥 CRASHn/a53.378 s💥 CRASHn/a34.980💥 CRASHn/a98.71 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a46.1 ms💥 CRASHn/a205.59B💥 CRASHn/a
Std.HashMap6.455 s💥 CRASHn/a50.868 s💥 CRASHn/a40.140💥 CRASHn/a95.03 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a45.7 ms💥 CRASHn/a203.35B💥 CRASHn/a
String.append5.307 s💥 CRASHn/a47.595 s💥 CRASHn/a6.870💥 CRASHn/a87.89 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a32.4 ms💥 CRASHn/a168.67B💥 CRASHn/a
Nat.add_comm4.526 s💥 CRASHn/a31.415 s💥 CRASHn/a1.460💥 CRASHn/a58.66 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a28.0 ms💥 CRASHn/a130.84B💥 CRASHn/a
Pipeline total (7 constants)
constanttotal-time (main)total-time (PR)Δ%pipeline-throughput (const/s) (main)pipeline-throughput (const/s) (PR)Δ%pipeline-peak-ram (main)pipeline-peak-ram (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append1m 38.6s💥 CRASHn/a28.160💥 CRASHn/a101.05 GiB💥 CRASHn/a
Char.ofOrdinal_le_of_le1m 31.7s💥 CRASHn/a30.130💥 CRASHn/a99.81 GiB💥 CRASHn/a
Array.extract_append1m 24.1s💥 CRASHn/a19.090💥 CRASHn/a95.19 GiB💥 CRASHn/a
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq1m 14.0s💥 CRASHn/a25.230💥 CRASHn/a98.71 GiB💥 CRASHn/a
Std.HashMap1m 13.4s💥 CRASHn/a27.830💥 CRASHn/a95.03 GiB💥 CRASHn/a
String.append50.458 s💥 CRASHn/a6.480💥 CRASHn/a87.89 GiB💥 CRASHn/a
Nat.add_comm32.745 s💥 CRASHn/a1.400💥 CRASHn/a58.66 GiB💥 CRASHn/a

Workflow logs

samuelburnham added a commit that referenced this pull request Sep 1, 2026
The Warp x64 runner pool mixes Intel Granite Rapids and AMD Zen 5, and a
build job may land on one vendor while the job that runs its binaries
lands on the other. Neither vendor's feature set contains the other's, so
`-Ctarget-cpu=native` does not produce a portable binary: Zen 5 enables
SSE4A, and LLVM emits it. Disassembling the workspace built for znver5
finds 31 SSE4A instructions, all INSERTQ, in `ix-ffi` and in
`aiur_ixvm_witness::add_entries_parallel`. Granite Rapids has no SSE4A,
so the first one executed raises #UD, killing the process with SIGILL
during witness generation. That is what turned every row of #605's
benchmark into a crash.
Pin the measured intersection of the two CPUs instead. x86-64-v4 covers
every AVX-512 subset Plonky3 uses; +avx512vbmi2 preserves its VPSHRDQ
interleave and +gfni preserves LLVM's byte-shift lowering. A workspace
built with these flags contains no instruction absent from either vendor
and has an instruction vocabulary identical to a graniterapids build.
blake3 dispatches on CPUID at runtime and is unaffected either way.
`.cargo/config.toml` keeps `-Ctarget-cpu=native`: a developer builds and
runs on one machine, and x86-64-v4 would exclude every host without
AVX-512. Only CI has the split, so only CI pins the ISA. The new guard
fails the job when a runner lacks a required feature, so the assumption
is enforced rather than assumed, and the shared `warp-x64` cargo cache
key becomes sound now that codegen no longer varies by host.
Pinning also removes a benchmarking hazard that never crashed: LLVM sets
prefer-256-bit for Granite Rapids but not for Zen 5, so the same source
vectorized 3.2x more widely depending on the build host, and main-vs-PR
timings were not comparable across a vendor split.
Drop the CPU-mismatch warning the benchmark comment used to carry. It
detected a real problem, but the flags above prevent that problem, and
computing it in one job to render it in another cost a Markdown file
threaded through cache entries, artifacts, and a `--warning-file` flag on
`ix bench compare`. Warnings belong to the run that finds them.
@arthurpaulino

Copy link
Copy Markdown
MemberAuthor

!benchmark fresh

@argument-ci-bot

argument-ci-botBot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

⚠️ Some benchmark jobs failed — results may be partial.

!benchmark — main vs d7fc2f1

backends: aiur=prove · envs: InitStd · baseline: fresh (benchmark products rebuilt, base-SHA run, bencher bypassed)

No result tables were produced — see the run logs.

Workflow logs

The Warp x64 runner pool mixes Intel Granite Rapids and AMD Zen 5, and a
build job may land on one vendor while the job that runs its binaries
lands on the other. Neither vendor's feature set contains the other's, so
`-Ctarget-cpu=native` does not produce a portable binary: Zen 5 enables
SSE4A, and LLVM emits it. Disassembling the workspace built for znver5
finds 31 SSE4A instructions, all INSERTQ, in `ix-ffi` and in
`aiur_ixvm_witness::add_entries_parallel`. Granite Rapids has no SSE4A,
so the first one executed raises #UD, killing the process with SIGILL
during witness generation. That is what turned every row of #605's
benchmark into a crash.
Pin the measured intersection of the two CPUs instead. x86-64-v4 covers
every AVX-512 subset Plonky3 uses; +avx512vbmi2 preserves its VPSHRDQ
interleave and +gfni preserves LLVM's byte-shift lowering. A workspace
built with these flags contains no instruction absent from either vendor
and has an instruction vocabulary identical to a graniterapids build.
blake3 dispatches on CPUID at runtime and is unaffected either way.
`.cargo/config.toml` keeps `-Ctarget-cpu=native`: a developer builds and
runs on one machine, and x86-64-v4 would exclude every host without
AVX-512. Only CI has the split, so only CI pins the ISA. The new guard
fails the job when a runner lacks a required feature, so the assumption
is enforced rather than assumed, and the shared `warp-x64` cargo cache
key becomes sound now that codegen no longer varies by host. RUSTFLAGS
is hashed into that key, so the flag change rotates it on its own.
Pinning also removes a benchmarking hazard that never crashed: LLVM sets
prefer-256-bit for Granite Rapids but not for Zen 5, so the same source
vectorized 3.2x more widely depending on the build host, and main-vs-PR
timings were not comparable across a vendor split.
@samuelburnham

Copy link
Copy Markdown
Member

!benchmark fresh

@argument-ci-bot

argument-ci-botBot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

!benchmark — main vs feb014d

backends: aiur=prove · envs: InitStd · baseline: fresh (benchmark products rebuilt, base-SHA run, bencher bypassed)

aiur · InitStd · prove — main from: base run @ e1ca8e2 (fresh — bencher bypassed)

7 constants · 3 with regressions · 7 with improvements (|Δ| > 3.0% on any metric).

IxVM on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append9.269 s9.073 s-2.1%31.250 s30.269 s-3.1% 🟢88.80091.680+3.2% 🟢70.72 GiB72.28 GiB+2.2%11.33 MiB11.06 MiB-2.4%68.5 ms62.7 ms-8.5% (1.09× faster) 🟢134.35B134.35B+0.0%
Char.ofOrdinal_le_of_le6.964 s6.879 s-1.2%26.884 s25.793 s-4.1% 🟢102.780107.120+4.2% 🟢63.83 GiB65.31 GiB+2.3%11.33 MiB11.07 MiB-2.3%77.9 ms59.3 ms-23.8% (1.31× faster) 🟢102.60B102.60B+0.0%
Array.extract_append6.676 s6.751 s+1.1%23.925 s23.179 s-3.1% 🟢67.13069.290+3.2% 🟢51.97 GiB52.72 GiB+1.4%11.24 MiB10.99 MiB-2.2%71.1 ms57.2 ms-19.5% (1.24× faster) 🟢97.08B97.08B+0.0%
Std.HashMap4.171 s4.079 s-2.2%16.222 s15.628 s-3.7% 🟢125.880130.670+3.8% 🟢36.29 GiB37.10 GiB+2.2%11.26 MiB11.01 MiB-2.2%74.8 ms66.6 ms-11.0% (1.12× faster) 🟢61.88B61.88B+0.0%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq3.697 s3.648 s-1.3%14.858 s14.287 s-3.8% 🟢125.650130.680+4.0% 🟢34.02 GiB34.75 GiB+2.1%11.26 MiB11.00 MiB-2.2%75.5 ms57.7 ms-23.6% (1.31× faster) 🟢55.68B55.68B+0.0%
String.append435.5 ms434.0 ms-0.4%2.307 s2.129 s-7.7% (1.08× faster) 🟢141.750153.600+8.4% (1.08× faster) 🟢5.74 GiB5.52 GiB-3.9% 🟢9.94 MiB9.74 MiB-2.1%64.4 ms50.2 ms-22.1% (1.28× faster) 🟢3.37B3.37B+0.0%
Nat.add_comm267.1 ms264.3 ms-1.1%1.068 s984.9 ms-7.8% (1.08× faster) 🟢43.06046.710+8.5% (1.08× faster) 🟢4.51 GiB3.99 GiB-11.5% (1.13× smaller) 🟢9.09 MiB8.90 MiB-2.1%53.3 ms47.9 ms-10.2% (1.11× faster) 🟢308.40M308.40M+0.0%
FRI verifier on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append5.680 s4.975 s-12.4% (1.14× faster) 🟢33.109 s30.692 s-7.3% (1.08× faster) 🟢83.82090.410+7.9% (1.08× faster) 🟢101.44 GiB99.74 GiB-1.7%3.97 MiB3.98 MiB+0.2%26.3 ms22.3 ms-15.4% (1.18× faster) 🟢210.23B203.74B-3.1% 🟢
Char.ofOrdinal_le_of_le5.599 s5.014 s-10.5% (1.12× faster) 🟢32.280 s31.379 s-2.8%85.59088.050+2.9%99.81 GiB101.09 GiB+1.3%3.97 MiB3.98 MiB+0.2%27.6 ms28.6 ms+3.6% ⚠️207.18B208.08B+0.4%
Array.extract_append5.315 s4.858 s-8.6% (1.09× faster) 🟢30.962 s29.672 s-4.2% 🟢51.87054.120+4.3% 🟢95.00 GiB95.36 GiB+0.4%3.97 MiB3.99 MiB+0.4%26.0 ms39.4 ms+51.3% (1.51× slower) ⚠️200.65B200.40B-0.1%
Std.HashMap5.445 s4.864 s-10.7% (1.12× faster) 🟢31.021 s29.949 s-3.5% 🟢65.83068.180+3.6% 🟢94.50 GiB95.90 GiB+1.5%3.97 MiB3.98 MiB+0.2%25.4 ms25.6 ms+0.8%203.35B204.11B+0.4%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq5.496 s4.898 s-10.9% (1.12× faster) 🟢32.335 s29.501 s-8.8% (1.10× faster) 🟢57.74063.290+9.6% (1.10× faster) 🟢98.72 GiB95.35 GiB-3.4% 🟢3.97 MiB3.98 MiB+0.3%25.9 ms31.9 ms+23.0% (1.23× slower) ⚠️205.59B199.65B-2.9%
String.append4.487 s3.955 s-11.8% (1.13× faster) 🟢28.710 s27.141 s-5.5% (1.06× faster) 🟢11.39012.050+5.8% (1.06× faster) 🟢87.89 GiB88.65 GiB+0.9%3.97 MiB3.97 MiB+0.1%37.6 ms27.4 ms-27.2% (1.37× faster) 🟢168.67B164.80B-2.3%
Nat.add_comm3.619 s3.077 s-15.0% (1.18× faster) 🟢19.611 s17.954 s-8.5% (1.09× faster) 🟢2.3502.560+8.9% (1.09× faster) 🟢58.66 GiB58.03 GiB-1.1%3.97 MiB3.98 MiB+0.2%27.6 ms22.0 ms-20.4% (1.26× faster) 🟢130.84B124.20B-5.1% (1.05× fewer) 🟢
Pipeline total (7 constants)
constanttotal-time (main)total-time (PR)Δ%pipeline-throughput (const/s) (main)pipeline-throughput (const/s) (PR)Δ%pipeline-peak-ram (main)pipeline-peak-ram (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append1m 4.4s1m 1.0s-5.3% (1.06× faster) 🟢43.12045.520+5.6% (1.06× faster) 🟢101.44 GiB99.74 GiB-1.7%
Char.ofOrdinal_le_of_le59.164 s57.172 s-3.4% 🟢46.70048.330+3.5% 🟢99.81 GiB101.09 GiB+1.3%
Array.extract_append54.887 s52.851 s-3.7% 🟢29.26030.390+3.9% 🟢95.00 GiB95.36 GiB+0.4%
Std.HashMap47.243 s45.577 s-3.5% 🟢43.22044.800+3.7% 🟢94.50 GiB95.90 GiB+1.5%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq47.193 s43.788 s-7.2% (1.08× faster) 🟢39.56042.640+7.8% (1.08× faster) 🟢98.72 GiB95.35 GiB-3.4% 🟢
String.append31.017 s29.270 s-5.6% (1.06× faster) 🟢10.54011.170+6.0% (1.06× faster) 🟢87.89 GiB88.65 GiB+0.9%
Nat.add_comm20.680 s18.939 s-8.4% (1.09× faster) 🟢2.2202.430+9.5% (1.09× faster) 🟢58.66 GiB58.03 GiB-1.1%

Workflow logs

#606)
* Authenticate all frontier group members in the recursive verifier
The in-circuit pruned-multiproof walk (mmcs_verify_multi / frontier_level)
collapses queries that share a parent to a single lead node and hashes
only the lead's rows via inject_maybe(ar, ...). Non-lead members' opened
rows for the shorter (injected) matrices were still consumed in their own
per-query FRI arithmetic (batch_views_at) but never authenticated against
any commitment — the leaf hash covers only the tallest matrices, and
shorter ones are bound solely through injection. A prover could therefore
forge a non-lead member's shorter-matrix opening. Plonky3's reference
verify_batch_pruned guards exactly this with InconsistentGroupOpening
(and InconsistentDuplicateOpenings for equal-index queries); the port had
neither. The prior per-query walk did not have the gap, so it was
introduced with the direct multiproof consumption.
- frontier_level: on a group merge, assert the lead and member agree on
every not-yet-injected matrix (height <= next_lh) via select_rows_le +
pointer equality. Transitive across pairwise merges, so the whole group
is pinned; matches InconsistentGroupOpening.
- frontier_merge: duplicate transcript indices must open the SAME full
rows, not merely the same tallest-matrix leaf digest; matches
InconsistentDuplicateOpenings.
Pointer equality is admissible inside assert_eq! (equal pointers imply
equal content; a spurious mismatch costs only completeness — see
IxVM.Core). select_rows_le selects rows of matrices at height <= target,
mirroring select_rows.
Validated: the group-merge branch is genuinely reached by the factorial
recursion proof (an always-false variant of the new assert fails the
honest test), the honest proof still verifies with the real assert
(completeness preserved), the existing tamper tests still reject, and
the full lake test suite is green (2717 checks). aiur_multi_stark.rs
regenerated; kernel executor unchanged.
* Drop the multi_stark::advice dependency
Companion to multi-stark removing its unused per-query advice module.
ix consumed native pruned multiproofs directly and referenced only
advice::AdviceError, whose two arms (verification failed, serialization
failed) were immediately string-formatted by the FFI. Replace it with a
plain Result<Vec<u8>, String>: AiurSystem::proof_to_advice_bytes maps
both failures to a message, and the FFI passes the string straight to
LeanExcept::error_string. No behavior change; the Lean binding
(Except String ByteArray) is unaffected.
Bump the multi-stark pin to the advice-removed revision. Requires that
multi-stark's ap/bump-p3-drop-advice be pushed first, exactly as with
every other pin in this series.
* Bump multi-stark audit revision
---------
Co-authored-by: Arthur Paulino <arthurleonardo.ap@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@arthurpaulino@samuelburnham
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Update Aiur to Plonky3 0.6 - #605

Draft
arthurpaulino wants to merge 5 commits into
mainfrom
ap/bump-p3
Draft

Update Aiur to Plonky3 0.6#605
arthurpaulino wants to merge 5 commits into
mainfrom
ap/bump-p3

Conversation

@arthurpaulino

Copy link
Copy Markdown
Member

Update the multi-stark dependency and Rust toolchain for Plonky3 0.6, along with the Rust 1.98 lint migrations required to keep the workspace warning-free. Refresh the Rust-compatible Blake3.lean pin in both root and compile-package manifests.

Adapt recursive Aiur verification to Plonky3's pruned FRI multiproofs. Native proofs retain their compact serialized representation and native verification path; the FFI expands authenticated Merkle frontiers into per-query advice only when entering the existing recursive verifier circuit.

Preserve the packed claim-digest convention in the recursion diagnostic and exercise the advice boundary in the end-to-end test and benchmark paths. CPU and CUDA recursive q1 runs produce identical 823,485-byte inner proofs and 331,273-byte outer proofs.

The q50 Vector.extract_append workload retains identical CPU/CUDA proof sizes. Inner plus outer STARK proving measures 65.87s on CPU and 8.81s with CUDA on the RTX PRO 6000, a 7.48x speedup.

Update the multi-stark dependency and Rust toolchain for Plonky3 0.6, along with the Rust 1.98 lint migrations required to keep the workspace warning-free. Refresh the Rust-compatible Blake3.lean pin in both root and compile-package manifests.
Adapt recursive Aiur verification to Plonky3's pruned FRI multiproofs. Native proofs retain their compact serialized representation and native verification path; the FFI expands authenticated Merkle frontiers into per-query advice only when entering the existing recursive verifier circuit.
Preserve the packed claim-digest convention in the recursion diagnostic and exercise the advice boundary in the end-to-end test and benchmark paths. CPU and CUDA recursive q1 runs produce identical 823,485-byte inner proofs and 331,273-byte outer proofs.
The q50 Vector.extract_append workload retains identical CPU/CUDA proof sizes. Inner plus outer STARK proving measures 65.87s on CPU and 8.81s with CUDA on the RTX PRO 6000, a 7.48x speedup.
PR benchmark runs execute trusted workflow YAML from the default branch while loading composite actions from the PR checkout. When Bencher data and binary caches are unavailable, the workflow checks out main under base/ and asks Lake to rebuild it without first installing the Rust channel pinned by that checkout.
Teach the existing CPU provenance action to install the base checkout's validated Rust channel and profile immediately before an uncached base build. The step is a no-op when the toolchain is already available and leaves cached benchmark comparisons unchanged.
@arthurpaulino

Copy link
Copy Markdown
MemberAuthor

!benchmark fresh

@argument-ci-bot

argument-ci-botBot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

!benchmark — main vs eab18e4

backends: aiur=prove · envs: InitStd · baseline: fresh (benchmark products rebuilt, base-SHA run, bencher bypassed)

aiur · InitStd · prove — main from: base run @ e1ca8e2 (fresh — bencher bypassed)

7 constants · 1 with regressions · 7 with improvements (|Δ| > 3.0% on any metric).

IxVM on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append8.816 s8.990 s+2.0%30.095 s29.094 s-3.3% 🟢92.21095.380+3.4% 🟢70.72 GiB72.35 GiB+2.3%11.33 MiB11.07 MiB-2.3%68.7 ms60.4 ms-12.1% (1.14× faster) 🟢134.35B134.35B+0.0%
Char.ofOrdinal_le_of_le6.689 s6.790 s+1.5%25.631 s25.209 s-1.6%107.800109.600+1.7%63.87 GiB65.41 GiB+2.4%11.33 MiB11.07 MiB-2.3%73.7 ms60.3 ms-18.2% (1.22× faster) 🟢102.60B102.60B+0.0%
Array.extract_append6.679 s6.415 s-4.0% 🟢23.076 s22.072 s-4.4% 🟢69.60072.760+4.5% 🟢52.03 GiB52.79 GiB+1.4%11.24 MiB10.99 MiB-2.2%72.3 ms64.6 ms-10.6% (1.12× faster) 🟢97.08B97.08B+0.0%
Std.HashMap3.964 s4.002 s+1.0%15.519 s15.133 s-2.5%131.580134.940+2.6%36.34 GiB37.08 GiB+2.1%11.26 MiB11.00 MiB-2.3%75.1 ms65.5 ms-12.8% (1.15× faster) 🟢61.88B61.88B+0.0%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq3.598 s3.528 s-1.9%14.327 s13.806 s-3.6% 🟢130.310135.230+3.8% 🟢33.96 GiB34.69 GiB+2.1%11.26 MiB11.01 MiB-2.2%69.9 ms59.5 ms-15.0% (1.18× faster) 🟢55.68B55.68B+0.0%
String.append424.3 ms423.0 ms-0.3%2.278 s2.124 s-6.7% (1.07× faster) 🟢143.540153.930+7.2% (1.07× faster) 🟢4.89 GiB5.00 GiB+2.4%9.94 MiB9.74 MiB-2.0%64.2 ms52.0 ms-19.0% (1.24× faster) 🟢3.37B3.37B+0.0%
Nat.add_comm260.1 ms259.8 ms-0.1%1.068 s970.5 ms-9.1% (1.10× faster) 🟢43.07047.400+10.1% (1.10× faster) 🟢3.99 GiB4.64 GiB+16.4% (1.16× larger) ⚠️9.09 MiB8.91 MiB-2.0%53.9 ms41.9 ms-22.3% (1.29× faster) 🟢308.40M308.40M+0.0%
FRI verifier on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append5.488 s5.490 s+0.0%31.664 s30.699 s-3.0% 🟢87.64090.400+3.1% 🟢101.37 GiB101.07 GiB-0.3%3.97 MiB3.72 MiB-6.3% (1.07× smaller) 🟢29.2 ms23.8 ms-18.5% (1.23× faster) 🟢210.23B210.23B+0.0%
Char.ofOrdinal_le_of_le5.386 s5.489 s+1.9%30.846 s30.914 s+0.2%89.57089.380-0.2%100.32 GiB100.32 GiB+0.0%3.97 MiB3.72 MiB-6.3% (1.07× smaller) 🟢24.8 ms23.1 ms-6.6% (1.07× faster) 🟢207.18B207.18B+0.0%
Array.extract_append5.209 s5.235 s+0.5%29.948 s28.844 s-3.7% 🟢53.63055.680+3.8% 🟢94.59 GiB95.37 GiB+0.8%3.97 MiB3.73 MiB-6.1% (1.07× smaller) 🟢25.7 ms21.3 ms-17.1% (1.21× faster) 🟢200.65B200.65B+0.0%
Std.HashMap5.314 s5.306 s-0.1%29.652 s29.584 s-0.2%68.87069.020+0.2%94.56 GiB95.34 GiB+0.8%3.97 MiB3.72 MiB-6.3% (1.07× smaller) 🟢25.2 ms21.2 ms-15.7% (1.19× faster) 🟢203.35B203.35B+0.0%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq5.400 s5.333 s-1.2%31.051 s30.174 s-2.8%60.13061.870+2.9%98.68 GiB99.48 GiB+0.8%3.97 MiB3.73 MiB-6.2% (1.07× smaller) 🟢30.2 ms31.0 ms+2.5%205.59B205.59B+0.0%
String.append4.334 s4.419 s+2.0%27.489 s26.612 s-3.2% 🟢11.90012.290+3.3% 🟢87.87 GiB88.58 GiB+0.8%3.97 MiB3.73 MiB-6.1% (1.07× smaller) 🟢34.5 ms22.7 ms-34.3% (1.52× faster) 🟢168.67B168.67B+0.0%
Nat.add_comm3.518 s3.499 s-0.5%18.993 s17.942 s-5.5% (1.06× faster) 🟢2.4202.560+5.8% (1.06× faster) 🟢58.70 GiB59.42 GiB+1.2%3.97 MiB3.72 MiB-6.2% (1.07× smaller) 🟢24.8 ms20.8 ms-16.1% (1.19× faster) 🟢130.84B130.84B+0.0%
Pipeline total (7 constants)
constanttotal-time (main)total-time (PR)Δ%pipeline-throughput (const/s) (main)pipeline-throughput (const/s) (PR)Δ%pipeline-peak-ram (main)pipeline-peak-ram (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append1m 1.8s59.792 s-3.2% 🟢44.93046.410+3.3% 🟢101.37 GiB101.07 GiB-0.3%
Char.ofOrdinal_le_of_le56.477 s56.124 s-0.6%48.92049.230+0.6%100.32 GiB100.32 GiB+0.0%
Array.extract_append53.024 s50.916 s-4.0% 🟢30.29031.540+4.1% 🟢94.59 GiB95.37 GiB+0.8%
Std.HashMap45.172 s44.717 s-1.0%45.21045.660+1.0%94.56 GiB95.34 GiB+0.8%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq45.377 s43.980 s-3.1% 🟢41.14042.450+3.2% 🟢98.68 GiB99.48 GiB+0.8%
String.append29.767 s28.737 s-3.5% 🟢10.99011.380+3.5% 🟢87.87 GiB88.58 GiB+0.8%
Nat.add_comm20.061 s18.912 s-5.7% (1.06× faster) 🟢2.2902.430+6.1% (1.06× faster) 🟢58.70 GiB59.42 GiB+1.2%

Workflow logs

Consume the Plonky3 0.6 batch-opening layout directly in Aiur instead of expanding every pruned Merkle frontier into one authentication path per FRI query. Sample all query indices from the unchanged transcript, sort and deduplicate them with an O(q log q) merge sort, authenticate each input and commit-phase commitment once, then retain the existing per-query reduced-opening and FRI arithmetic.
Bind every frontier to transcript-derived indices, consume boundary digests in Plonky3's level/parent/child order, reject trailing frontier elements and inconsistent duplicate leaves, and assert all native opening dimensions and sibling counts. Explicitly constrain the digest-bound protocol specialization to cap height 0, binary FRI, and a constant final polynomial. Move memo_u32_less_than into IxVM Core so both substitution and multiproof sorting share its constrained rows.
Strengthen the recursive negative test to mutate a structurally valid stage-1 commitment. Regenerate both checked-in Aiur Rust executors and retain interpreter/codegen query-count parity.
On Vector.extract_append q50, recursive-verifier FFT cost falls from 204.073B to 201.166B. CPU outer proving improves from 50.09s to 45.03s and the full CPU pipeline from 90.64s to 82.90s. GPU outer proving improves from 15.85s to 13.72s and the full GPU pipeline from 28.86s to 26.69s. The outer proof grows from 3.92 MB to 4.17 MB.
Validated with the MultiStark primitive suite, recursive honest/tamper/parity tests, codegen --check, release workspace clippy, release CUDA clippy, rustfmt, and diff checks.
@arthurpaulino

Copy link
Copy Markdown
MemberAuthor

!benchmark fresh

@argument-ci-bot

argument-ci-botBot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

!benchmark — main vs 840932a

backends: aiur=prove · envs: InitStd · baseline: fresh (benchmark products rebuilt, base-SHA run, bencher bypassed)

Warning

  • CPU model mismatch for PR benchmark binaries in this job: built on AMD EPYC 9R45; measured on Intel(R) Xeon(R) 6975P-C. Native Rust code uses -Ctarget-cpu=native.

aiur · InitStd · prove — main from: base run @ e1ca8e2 (fresh — bencher bypassed)

7 constants · 0 with regressions · 0 with improvements (|Δ| > 3.0% on any metric).

IxVM on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append8.874 s💥 CRASHn/a44.291 s💥 CRASHn/a62.650💥 CRASHn/a70.80 GiB💥 CRASHn/a11.33 MiB💥 CRASHn/a77.4 ms💥 CRASHn/a134.35B💥 CRASHn/a
Char.ofOrdinal_le_of_le6.603 s💥 CRASHn/a38.242 s💥 CRASHn/a72.250💥 CRASHn/a63.81 GiB💥 CRASHn/a11.33 MiB💥 CRASHn/a76.9 ms💥 CRASHn/a102.60B💥 CRASHn/a
Array.extract_append6.311 s💥 CRASHn/a33.420 s💥 CRASHn/a48.060💥 CRASHn/a51.95 GiB💥 CRASHn/a11.24 MiB💥 CRASHn/a86.5 ms💥 CRASHn/a97.08B💥 CRASHn/a
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq3.774 s💥 CRASHn/a20.628 s💥 CRASHn/a90.510💥 CRASHn/a33.97 GiB💥 CRASHn/a11.26 MiB💥 CRASHn/a82.6 ms💥 CRASHn/a55.68B💥 CRASHn/a
Std.HashMap4.060 s💥 CRASHn/a22.501 s💥 CRASHn/a90.750💥 CRASHn/a36.36 GiB💥 CRASHn/a11.26 MiB💥 CRASHn/a84.1 ms💥 CRASHn/a61.88B💥 CRASHn/a
String.append708.3 ms💥 CRASHn/a2.863 s💥 CRASHn/a114.220💥 CRASHn/a4.99 GiB💥 CRASHn/a9.94 MiB💥 CRASHn/a70.3 ms💥 CRASHn/a3.37B💥 CRASHn/a
Nat.add_comm496.0 ms💥 CRASHn/a1.330 s💥 CRASHn/a34.580💥 CRASHn/a4.33 GiB💥 CRASHn/a9.09 MiB💥 CRASHn/a58.9 ms💥 CRASHn/a308.40M💥 CRASHn/a
FRI verifier on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append6.880 s💥 CRASHn/a54.265 s💥 CRASHn/a51.140💥 CRASHn/a101.05 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a49.5 ms💥 CRASHn/a210.23B💥 CRASHn/a
Char.ofOrdinal_le_of_le6.813 s💥 CRASHn/a53.449 s💥 CRASHn/a51.690💥 CRASHn/a99.81 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a45.1 ms💥 CRASHn/a207.18B💥 CRASHn/a
Array.extract_append6.371 s💥 CRASHn/a50.695 s💥 CRASHn/a31.680💥 CRASHn/a95.19 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a58.1 ms💥 CRASHn/a200.65B💥 CRASHn/a
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq6.748 s💥 CRASHn/a53.378 s💥 CRASHn/a34.980💥 CRASHn/a98.71 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a46.1 ms💥 CRASHn/a205.59B💥 CRASHn/a
Std.HashMap6.455 s💥 CRASHn/a50.868 s💥 CRASHn/a40.140💥 CRASHn/a95.03 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a45.7 ms💥 CRASHn/a203.35B💥 CRASHn/a
String.append5.307 s💥 CRASHn/a47.595 s💥 CRASHn/a6.870💥 CRASHn/a87.89 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a32.4 ms💥 CRASHn/a168.67B💥 CRASHn/a
Nat.add_comm4.526 s💥 CRASHn/a31.415 s💥 CRASHn/a1.460💥 CRASHn/a58.66 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a28.0 ms💥 CRASHn/a130.84B💥 CRASHn/a
Pipeline total (7 constants)
constanttotal-time (main)total-time (PR)Δ%pipeline-throughput (const/s) (main)pipeline-throughput (const/s) (PR)Δ%pipeline-peak-ram (main)pipeline-peak-ram (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append1m 38.6s💥 CRASHn/a28.160💥 CRASHn/a101.05 GiB💥 CRASHn/a
Char.ofOrdinal_le_of_le1m 31.7s💥 CRASHn/a30.130💥 CRASHn/a99.81 GiB💥 CRASHn/a
Array.extract_append1m 24.1s💥 CRASHn/a19.090💥 CRASHn/a95.19 GiB💥 CRASHn/a
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq1m 14.0s💥 CRASHn/a25.230💥 CRASHn/a98.71 GiB💥 CRASHn/a
Std.HashMap1m 13.4s💥 CRASHn/a27.830💥 CRASHn/a95.03 GiB💥 CRASHn/a
String.append50.458 s💥 CRASHn/a6.480💥 CRASHn/a87.89 GiB💥 CRASHn/a
Nat.add_comm32.745 s💥 CRASHn/a1.400💥 CRASHn/a58.66 GiB💥 CRASHn/a

Workflow logs

samuelburnham added a commit that referenced this pull request Sep 1, 2026
The Warp x64 runner pool mixes Intel Granite Rapids and AMD Zen 5, and a
build job may land on one vendor while the job that runs its binaries
lands on the other. Neither vendor's feature set contains the other's, so
`-Ctarget-cpu=native` does not produce a portable binary: Zen 5 enables
SSE4A, and LLVM emits it. Disassembling the workspace built for znver5
finds 31 SSE4A instructions, all INSERTQ, in `ix-ffi` and in
`aiur_ixvm_witness::add_entries_parallel`. Granite Rapids has no SSE4A,
so the first one executed raises #UD, killing the process with SIGILL
during witness generation. That is what turned every row of #605's
benchmark into a crash.
Pin the measured intersection of the two CPUs instead. x86-64-v4 covers
every AVX-512 subset Plonky3 uses; +avx512vbmi2 preserves its VPSHRDQ
interleave and +gfni preserves LLVM's byte-shift lowering. A workspace
built with these flags contains no instruction absent from either vendor
and has an instruction vocabulary identical to a graniterapids build.
blake3 dispatches on CPUID at runtime and is unaffected either way.
`.cargo/config.toml` keeps `-Ctarget-cpu=native`: a developer builds and
runs on one machine, and x86-64-v4 would exclude every host without
AVX-512. Only CI has the split, so only CI pins the ISA. The new guard
fails the job when a runner lacks a required feature, so the assumption
is enforced rather than assumed, and the shared `warp-x64` cargo cache
key becomes sound now that codegen no longer varies by host.
Pinning also removes a benchmarking hazard that never crashed: LLVM sets
prefer-256-bit for Granite Rapids but not for Zen 5, so the same source
vectorized 3.2x more widely depending on the build host, and main-vs-PR
timings were not comparable across a vendor split.
Drop the CPU-mismatch warning the benchmark comment used to carry. It
detected a real problem, but the flags above prevent that problem, and
computing it in one job to render it in another cost a Markdown file
threaded through cache entries, artifacts, and a `--warning-file` flag on
`ix bench compare`. Warnings belong to the run that finds them.
@arthurpaulino

Copy link
Copy Markdown
MemberAuthor

!benchmark fresh

@argument-ci-bot

argument-ci-botBot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

⚠️ Some benchmark jobs failed — results may be partial.

!benchmark — main vs d7fc2f1

backends: aiur=prove · envs: InitStd · baseline: fresh (benchmark products rebuilt, base-SHA run, bencher bypassed)

No result tables were produced — see the run logs.

Workflow logs

The Warp x64 runner pool mixes Intel Granite Rapids and AMD Zen 5, and a
build job may land on one vendor while the job that runs its binaries
lands on the other. Neither vendor's feature set contains the other's, so
`-Ctarget-cpu=native` does not produce a portable binary: Zen 5 enables
SSE4A, and LLVM emits it. Disassembling the workspace built for znver5
finds 31 SSE4A instructions, all INSERTQ, in `ix-ffi` and in
`aiur_ixvm_witness::add_entries_parallel`. Granite Rapids has no SSE4A,
so the first one executed raises #UD, killing the process with SIGILL
during witness generation. That is what turned every row of #605's
benchmark into a crash.
Pin the measured intersection of the two CPUs instead. x86-64-v4 covers
every AVX-512 subset Plonky3 uses; +avx512vbmi2 preserves its VPSHRDQ
interleave and +gfni preserves LLVM's byte-shift lowering. A workspace
built with these flags contains no instruction absent from either vendor
and has an instruction vocabulary identical to a graniterapids build.
blake3 dispatches on CPUID at runtime and is unaffected either way.
`.cargo/config.toml` keeps `-Ctarget-cpu=native`: a developer builds and
runs on one machine, and x86-64-v4 would exclude every host without
AVX-512. Only CI has the split, so only CI pins the ISA. The new guard
fails the job when a runner lacks a required feature, so the assumption
is enforced rather than assumed, and the shared `warp-x64` cargo cache
key becomes sound now that codegen no longer varies by host. RUSTFLAGS
is hashed into that key, so the flag change rotates it on its own.
Pinning also removes a benchmarking hazard that never crashed: LLVM sets
prefer-256-bit for Granite Rapids but not for Zen 5, so the same source
vectorized 3.2x more widely depending on the build host, and main-vs-PR
timings were not comparable across a vendor split.
@samuelburnham

Copy link
Copy Markdown
Member

!benchmark fresh

@argument-ci-bot

argument-ci-botBot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

!benchmark — main vs feb014d

backends: aiur=prove · envs: InitStd · baseline: fresh (benchmark products rebuilt, base-SHA run, bencher bypassed)

aiur · InitStd · prove — main from: base run @ e1ca8e2 (fresh — bencher bypassed)

7 constants · 3 with regressions · 7 with improvements (|Δ| > 3.0% on any metric).

IxVM on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append9.269 s9.073 s-2.1%31.250 s30.269 s-3.1% 🟢88.80091.680+3.2% 🟢70.72 GiB72.28 GiB+2.2%11.33 MiB11.06 MiB-2.4%68.5 ms62.7 ms-8.5% (1.09× faster) 🟢134.35B134.35B+0.0%
Char.ofOrdinal_le_of_le6.964 s6.879 s-1.2%26.884 s25.793 s-4.1% 🟢102.780107.120+4.2% 🟢63.83 GiB65.31 GiB+2.3%11.33 MiB11.07 MiB-2.3%77.9 ms59.3 ms-23.8% (1.31× faster) 🟢102.60B102.60B+0.0%
Array.extract_append6.676 s6.751 s+1.1%23.925 s23.179 s-3.1% 🟢67.13069.290+3.2% 🟢51.97 GiB52.72 GiB+1.4%11.24 MiB10.99 MiB-2.2%71.1 ms57.2 ms-19.5% (1.24× faster) 🟢97.08B97.08B+0.0%
Std.HashMap4.171 s4.079 s-2.2%16.222 s15.628 s-3.7% 🟢125.880130.670+3.8% 🟢36.29 GiB37.10 GiB+2.2%11.26 MiB11.01 MiB-2.2%74.8 ms66.6 ms-11.0% (1.12× faster) 🟢61.88B61.88B+0.0%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq3.697 s3.648 s-1.3%14.858 s14.287 s-3.8% 🟢125.650130.680+4.0% 🟢34.02 GiB34.75 GiB+2.1%11.26 MiB11.00 MiB-2.2%75.5 ms57.7 ms-23.6% (1.31× faster) 🟢55.68B55.68B+0.0%
String.append435.5 ms434.0 ms-0.4%2.307 s2.129 s-7.7% (1.08× faster) 🟢141.750153.600+8.4% (1.08× faster) 🟢5.74 GiB5.52 GiB-3.9% 🟢9.94 MiB9.74 MiB-2.1%64.4 ms50.2 ms-22.1% (1.28× faster) 🟢3.37B3.37B+0.0%
Nat.add_comm267.1 ms264.3 ms-1.1%1.068 s984.9 ms-7.8% (1.08× faster) 🟢43.06046.710+8.5% (1.08× faster) 🟢4.51 GiB3.99 GiB-11.5% (1.13× smaller) 🟢9.09 MiB8.90 MiB-2.1%53.3 ms47.9 ms-10.2% (1.11× faster) 🟢308.40M308.40M+0.0%
FRI verifier on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append5.680 s4.975 s-12.4% (1.14× faster) 🟢33.109 s30.692 s-7.3% (1.08× faster) 🟢83.82090.410+7.9% (1.08× faster) 🟢101.44 GiB99.74 GiB-1.7%3.97 MiB3.98 MiB+0.2%26.3 ms22.3 ms-15.4% (1.18× faster) 🟢210.23B203.74B-3.1% 🟢
Char.ofOrdinal_le_of_le5.599 s5.014 s-10.5% (1.12× faster) 🟢32.280 s31.379 s-2.8%85.59088.050+2.9%99.81 GiB101.09 GiB+1.3%3.97 MiB3.98 MiB+0.2%27.6 ms28.6 ms+3.6% ⚠️207.18B208.08B+0.4%
Array.extract_append5.315 s4.858 s-8.6% (1.09× faster) 🟢30.962 s29.672 s-4.2% 🟢51.87054.120+4.3% 🟢95.00 GiB95.36 GiB+0.4%3.97 MiB3.99 MiB+0.4%26.0 ms39.4 ms+51.3% (1.51× slower) ⚠️200.65B200.40B-0.1%
Std.HashMap5.445 s4.864 s-10.7% (1.12× faster) 🟢31.021 s29.949 s-3.5% 🟢65.83068.180+3.6% 🟢94.50 GiB95.90 GiB+1.5%3.97 MiB3.98 MiB+0.2%25.4 ms25.6 ms+0.8%203.35B204.11B+0.4%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq5.496 s4.898 s-10.9% (1.12× faster) 🟢32.335 s29.501 s-8.8% (1.10× faster) 🟢57.74063.290+9.6% (1.10× faster) 🟢98.72 GiB95.35 GiB-3.4% 🟢3.97 MiB3.98 MiB+0.3%25.9 ms31.9 ms+23.0% (1.23× slower) ⚠️205.59B199.65B-2.9%
String.append4.487 s3.955 s-11.8% (1.13× faster) 🟢28.710 s27.141 s-5.5% (1.06× faster) 🟢11.39012.050+5.8% (1.06× faster) 🟢87.89 GiB88.65 GiB+0.9%3.97 MiB3.97 MiB+0.1%37.6 ms27.4 ms-27.2% (1.37× faster) 🟢168.67B164.80B-2.3%
Nat.add_comm3.619 s3.077 s-15.0% (1.18× faster) 🟢19.611 s17.954 s-8.5% (1.09× faster) 🟢2.3502.560+8.9% (1.09× faster) 🟢58.66 GiB58.03 GiB-1.1%3.97 MiB3.98 MiB+0.2%27.6 ms22.0 ms-20.4% (1.26× faster) 🟢130.84B124.20B-5.1% (1.05× fewer) 🟢
Pipeline total (7 constants)
constanttotal-time (main)total-time (PR)Δ%pipeline-throughput (const/s) (main)pipeline-throughput (const/s) (PR)Δ%pipeline-peak-ram (main)pipeline-peak-ram (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append1m 4.4s1m 1.0s-5.3% (1.06× faster) 🟢43.12045.520+5.6% (1.06× faster) 🟢101.44 GiB99.74 GiB-1.7%
Char.ofOrdinal_le_of_le59.164 s57.172 s-3.4% 🟢46.70048.330+3.5% 🟢99.81 GiB101.09 GiB+1.3%
Array.extract_append54.887 s52.851 s-3.7% 🟢29.26030.390+3.9% 🟢95.00 GiB95.36 GiB+0.4%
Std.HashMap47.243 s45.577 s-3.5% 🟢43.22044.800+3.7% 🟢94.50 GiB95.90 GiB+1.5%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq47.193 s43.788 s-7.2% (1.08× faster) 🟢39.56042.640+7.8% (1.08× faster) 🟢98.72 GiB95.35 GiB-3.4% 🟢
String.append31.017 s29.270 s-5.6% (1.06× faster) 🟢10.54011.170+6.0% (1.06× faster) 🟢87.89 GiB88.65 GiB+0.9%
Nat.add_comm20.680 s18.939 s-8.4% (1.09× faster) 🟢2.2202.430+9.5% (1.09× faster) 🟢58.66 GiB58.03 GiB-1.1%

Workflow logs

#606)
* Authenticate all frontier group members in the recursive verifier
The in-circuit pruned-multiproof walk (mmcs_verify_multi / frontier_level)
collapses queries that share a parent to a single lead node and hashes
only the lead's rows via inject_maybe(ar, ...). Non-lead members' opened
rows for the shorter (injected) matrices were still consumed in their own
per-query FRI arithmetic (batch_views_at) but never authenticated against
any commitment — the leaf hash covers only the tallest matrices, and
shorter ones are bound solely through injection. A prover could therefore
forge a non-lead member's shorter-matrix opening. Plonky3's reference
verify_batch_pruned guards exactly this with InconsistentGroupOpening
(and InconsistentDuplicateOpenings for equal-index queries); the port had
neither. The prior per-query walk did not have the gap, so it was
introduced with the direct multiproof consumption.
- frontier_level: on a group merge, assert the lead and member agree on
every not-yet-injected matrix (height <= next_lh) via select_rows_le +
pointer equality. Transitive across pairwise merges, so the whole group
is pinned; matches InconsistentGroupOpening.
- frontier_merge: duplicate transcript indices must open the SAME full
rows, not merely the same tallest-matrix leaf digest; matches
InconsistentDuplicateOpenings.
Pointer equality is admissible inside assert_eq! (equal pointers imply
equal content; a spurious mismatch costs only completeness — see
IxVM.Core). select_rows_le selects rows of matrices at height <= target,
mirroring select_rows.
Validated: the group-merge branch is genuinely reached by the factorial
recursion proof (an always-false variant of the new assert fails the
honest test), the honest proof still verifies with the real assert
(completeness preserved), the existing tamper tests still reject, and
the full lake test suite is green (2717 checks). aiur_multi_stark.rs
regenerated; kernel executor unchanged.
* Drop the multi_stark::advice dependency
Companion to multi-stark removing its unused per-query advice module.
ix consumed native pruned multiproofs directly and referenced only
advice::AdviceError, whose two arms (verification failed, serialization
failed) were immediately string-formatted by the FFI. Replace it with a
plain Result<Vec<u8>, String>: AiurSystem::proof_to_advice_bytes maps
both failures to a message, and the FFI passes the string straight to
LeanExcept::error_string. No behavior change; the Lean binding
(Except String ByteArray) is unaffected.
Bump the multi-stark pin to the advice-removed revision. Requires that
multi-stark's ap/bump-p3-drop-advice be pushed first, exactly as with
every other pin in this series.
* Bump multi-stark audit revision
---------
Co-authored-by: Arthur Paulino <arthurleonardo.ap@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@arthurpaulino@samuelburnham
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Update Aiur to Plonky3 0.6 - #605

Draft
arthurpaulino wants to merge 5 commits into
mainfrom
ap/bump-p3
Draft

Update Aiur to Plonky3 0.6#605
arthurpaulino wants to merge 5 commits into
mainfrom
ap/bump-p3

Conversation

@arthurpaulino

Copy link
Copy Markdown
Member

Update the multi-stark dependency and Rust toolchain for Plonky3 0.6, along with the Rust 1.98 lint migrations required to keep the workspace warning-free. Refresh the Rust-compatible Blake3.lean pin in both root and compile-package manifests.

Adapt recursive Aiur verification to Plonky3's pruned FRI multiproofs. Native proofs retain their compact serialized representation and native verification path; the FFI expands authenticated Merkle frontiers into per-query advice only when entering the existing recursive verifier circuit.

Preserve the packed claim-digest convention in the recursion diagnostic and exercise the advice boundary in the end-to-end test and benchmark paths. CPU and CUDA recursive q1 runs produce identical 823,485-byte inner proofs and 331,273-byte outer proofs.

The q50 Vector.extract_append workload retains identical CPU/CUDA proof sizes. Inner plus outer STARK proving measures 65.87s on CPU and 8.81s with CUDA on the RTX PRO 6000, a 7.48x speedup.

Update the multi-stark dependency and Rust toolchain for Plonky3 0.6, along with the Rust 1.98 lint migrations required to keep the workspace warning-free. Refresh the Rust-compatible Blake3.lean pin in both root and compile-package manifests.
Adapt recursive Aiur verification to Plonky3's pruned FRI multiproofs. Native proofs retain their compact serialized representation and native verification path; the FFI expands authenticated Merkle frontiers into per-query advice only when entering the existing recursive verifier circuit.
Preserve the packed claim-digest convention in the recursion diagnostic and exercise the advice boundary in the end-to-end test and benchmark paths. CPU and CUDA recursive q1 runs produce identical 823,485-byte inner proofs and 331,273-byte outer proofs.
The q50 Vector.extract_append workload retains identical CPU/CUDA proof sizes. Inner plus outer STARK proving measures 65.87s on CPU and 8.81s with CUDA on the RTX PRO 6000, a 7.48x speedup.
PR benchmark runs execute trusted workflow YAML from the default branch while loading composite actions from the PR checkout. When Bencher data and binary caches are unavailable, the workflow checks out main under base/ and asks Lake to rebuild it without first installing the Rust channel pinned by that checkout.
Teach the existing CPU provenance action to install the base checkout's validated Rust channel and profile immediately before an uncached base build. The step is a no-op when the toolchain is already available and leaves cached benchmark comparisons unchanged.
@arthurpaulino

Copy link
Copy Markdown
MemberAuthor

!benchmark fresh

@argument-ci-bot

argument-ci-botBot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

!benchmark — main vs eab18e4

backends: aiur=prove · envs: InitStd · baseline: fresh (benchmark products rebuilt, base-SHA run, bencher bypassed)

aiur · InitStd · prove — main from: base run @ e1ca8e2 (fresh — bencher bypassed)

7 constants · 1 with regressions · 7 with improvements (|Δ| > 3.0% on any metric).

IxVM on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append8.816 s8.990 s+2.0%30.095 s29.094 s-3.3% 🟢92.21095.380+3.4% 🟢70.72 GiB72.35 GiB+2.3%11.33 MiB11.07 MiB-2.3%68.7 ms60.4 ms-12.1% (1.14× faster) 🟢134.35B134.35B+0.0%
Char.ofOrdinal_le_of_le6.689 s6.790 s+1.5%25.631 s25.209 s-1.6%107.800109.600+1.7%63.87 GiB65.41 GiB+2.4%11.33 MiB11.07 MiB-2.3%73.7 ms60.3 ms-18.2% (1.22× faster) 🟢102.60B102.60B+0.0%
Array.extract_append6.679 s6.415 s-4.0% 🟢23.076 s22.072 s-4.4% 🟢69.60072.760+4.5% 🟢52.03 GiB52.79 GiB+1.4%11.24 MiB10.99 MiB-2.2%72.3 ms64.6 ms-10.6% (1.12× faster) 🟢97.08B97.08B+0.0%
Std.HashMap3.964 s4.002 s+1.0%15.519 s15.133 s-2.5%131.580134.940+2.6%36.34 GiB37.08 GiB+2.1%11.26 MiB11.00 MiB-2.3%75.1 ms65.5 ms-12.8% (1.15× faster) 🟢61.88B61.88B+0.0%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq3.598 s3.528 s-1.9%14.327 s13.806 s-3.6% 🟢130.310135.230+3.8% 🟢33.96 GiB34.69 GiB+2.1%11.26 MiB11.01 MiB-2.2%69.9 ms59.5 ms-15.0% (1.18× faster) 🟢55.68B55.68B+0.0%
String.append424.3 ms423.0 ms-0.3%2.278 s2.124 s-6.7% (1.07× faster) 🟢143.540153.930+7.2% (1.07× faster) 🟢4.89 GiB5.00 GiB+2.4%9.94 MiB9.74 MiB-2.0%64.2 ms52.0 ms-19.0% (1.24× faster) 🟢3.37B3.37B+0.0%
Nat.add_comm260.1 ms259.8 ms-0.1%1.068 s970.5 ms-9.1% (1.10× faster) 🟢43.07047.400+10.1% (1.10× faster) 🟢3.99 GiB4.64 GiB+16.4% (1.16× larger) ⚠️9.09 MiB8.91 MiB-2.0%53.9 ms41.9 ms-22.3% (1.29× faster) 🟢308.40M308.40M+0.0%
FRI verifier on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append5.488 s5.490 s+0.0%31.664 s30.699 s-3.0% 🟢87.64090.400+3.1% 🟢101.37 GiB101.07 GiB-0.3%3.97 MiB3.72 MiB-6.3% (1.07× smaller) 🟢29.2 ms23.8 ms-18.5% (1.23× faster) 🟢210.23B210.23B+0.0%
Char.ofOrdinal_le_of_le5.386 s5.489 s+1.9%30.846 s30.914 s+0.2%89.57089.380-0.2%100.32 GiB100.32 GiB+0.0%3.97 MiB3.72 MiB-6.3% (1.07× smaller) 🟢24.8 ms23.1 ms-6.6% (1.07× faster) 🟢207.18B207.18B+0.0%
Array.extract_append5.209 s5.235 s+0.5%29.948 s28.844 s-3.7% 🟢53.63055.680+3.8% 🟢94.59 GiB95.37 GiB+0.8%3.97 MiB3.73 MiB-6.1% (1.07× smaller) 🟢25.7 ms21.3 ms-17.1% (1.21× faster) 🟢200.65B200.65B+0.0%
Std.HashMap5.314 s5.306 s-0.1%29.652 s29.584 s-0.2%68.87069.020+0.2%94.56 GiB95.34 GiB+0.8%3.97 MiB3.72 MiB-6.3% (1.07× smaller) 🟢25.2 ms21.2 ms-15.7% (1.19× faster) 🟢203.35B203.35B+0.0%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq5.400 s5.333 s-1.2%31.051 s30.174 s-2.8%60.13061.870+2.9%98.68 GiB99.48 GiB+0.8%3.97 MiB3.73 MiB-6.2% (1.07× smaller) 🟢30.2 ms31.0 ms+2.5%205.59B205.59B+0.0%
String.append4.334 s4.419 s+2.0%27.489 s26.612 s-3.2% 🟢11.90012.290+3.3% 🟢87.87 GiB88.58 GiB+0.8%3.97 MiB3.73 MiB-6.1% (1.07× smaller) 🟢34.5 ms22.7 ms-34.3% (1.52× faster) 🟢168.67B168.67B+0.0%
Nat.add_comm3.518 s3.499 s-0.5%18.993 s17.942 s-5.5% (1.06× faster) 🟢2.4202.560+5.8% (1.06× faster) 🟢58.70 GiB59.42 GiB+1.2%3.97 MiB3.72 MiB-6.2% (1.07× smaller) 🟢24.8 ms20.8 ms-16.1% (1.19× faster) 🟢130.84B130.84B+0.0%
Pipeline total (7 constants)
constanttotal-time (main)total-time (PR)Δ%pipeline-throughput (const/s) (main)pipeline-throughput (const/s) (PR)Δ%pipeline-peak-ram (main)pipeline-peak-ram (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append1m 1.8s59.792 s-3.2% 🟢44.93046.410+3.3% 🟢101.37 GiB101.07 GiB-0.3%
Char.ofOrdinal_le_of_le56.477 s56.124 s-0.6%48.92049.230+0.6%100.32 GiB100.32 GiB+0.0%
Array.extract_append53.024 s50.916 s-4.0% 🟢30.29031.540+4.1% 🟢94.59 GiB95.37 GiB+0.8%
Std.HashMap45.172 s44.717 s-1.0%45.21045.660+1.0%94.56 GiB95.34 GiB+0.8%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq45.377 s43.980 s-3.1% 🟢41.14042.450+3.2% 🟢98.68 GiB99.48 GiB+0.8%
String.append29.767 s28.737 s-3.5% 🟢10.99011.380+3.5% 🟢87.87 GiB88.58 GiB+0.8%
Nat.add_comm20.061 s18.912 s-5.7% (1.06× faster) 🟢2.2902.430+6.1% (1.06× faster) 🟢58.70 GiB59.42 GiB+1.2%

Workflow logs

Consume the Plonky3 0.6 batch-opening layout directly in Aiur instead of expanding every pruned Merkle frontier into one authentication path per FRI query. Sample all query indices from the unchanged transcript, sort and deduplicate them with an O(q log q) merge sort, authenticate each input and commit-phase commitment once, then retain the existing per-query reduced-opening and FRI arithmetic.
Bind every frontier to transcript-derived indices, consume boundary digests in Plonky3's level/parent/child order, reject trailing frontier elements and inconsistent duplicate leaves, and assert all native opening dimensions and sibling counts. Explicitly constrain the digest-bound protocol specialization to cap height 0, binary FRI, and a constant final polynomial. Move memo_u32_less_than into IxVM Core so both substitution and multiproof sorting share its constrained rows.
Strengthen the recursive negative test to mutate a structurally valid stage-1 commitment. Regenerate both checked-in Aiur Rust executors and retain interpreter/codegen query-count parity.
On Vector.extract_append q50, recursive-verifier FFT cost falls from 204.073B to 201.166B. CPU outer proving improves from 50.09s to 45.03s and the full CPU pipeline from 90.64s to 82.90s. GPU outer proving improves from 15.85s to 13.72s and the full GPU pipeline from 28.86s to 26.69s. The outer proof grows from 3.92 MB to 4.17 MB.
Validated with the MultiStark primitive suite, recursive honest/tamper/parity tests, codegen --check, release workspace clippy, release CUDA clippy, rustfmt, and diff checks.
@arthurpaulino

Copy link
Copy Markdown
MemberAuthor

!benchmark fresh

@argument-ci-bot

argument-ci-botBot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

!benchmark — main vs 840932a

backends: aiur=prove · envs: InitStd · baseline: fresh (benchmark products rebuilt, base-SHA run, bencher bypassed)

Warning

  • CPU model mismatch for PR benchmark binaries in this job: built on AMD EPYC 9R45; measured on Intel(R) Xeon(R) 6975P-C. Native Rust code uses -Ctarget-cpu=native.

aiur · InitStd · prove — main from: base run @ e1ca8e2 (fresh — bencher bypassed)

7 constants · 0 with regressions · 0 with improvements (|Δ| > 3.0% on any metric).

IxVM on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append8.874 s💥 CRASHn/a44.291 s💥 CRASHn/a62.650💥 CRASHn/a70.80 GiB💥 CRASHn/a11.33 MiB💥 CRASHn/a77.4 ms💥 CRASHn/a134.35B💥 CRASHn/a
Char.ofOrdinal_le_of_le6.603 s💥 CRASHn/a38.242 s💥 CRASHn/a72.250💥 CRASHn/a63.81 GiB💥 CRASHn/a11.33 MiB💥 CRASHn/a76.9 ms💥 CRASHn/a102.60B💥 CRASHn/a
Array.extract_append6.311 s💥 CRASHn/a33.420 s💥 CRASHn/a48.060💥 CRASHn/a51.95 GiB💥 CRASHn/a11.24 MiB💥 CRASHn/a86.5 ms💥 CRASHn/a97.08B💥 CRASHn/a
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq3.774 s💥 CRASHn/a20.628 s💥 CRASHn/a90.510💥 CRASHn/a33.97 GiB💥 CRASHn/a11.26 MiB💥 CRASHn/a82.6 ms💥 CRASHn/a55.68B💥 CRASHn/a
Std.HashMap4.060 s💥 CRASHn/a22.501 s💥 CRASHn/a90.750💥 CRASHn/a36.36 GiB💥 CRASHn/a11.26 MiB💥 CRASHn/a84.1 ms💥 CRASHn/a61.88B💥 CRASHn/a
String.append708.3 ms💥 CRASHn/a2.863 s💥 CRASHn/a114.220💥 CRASHn/a4.99 GiB💥 CRASHn/a9.94 MiB💥 CRASHn/a70.3 ms💥 CRASHn/a3.37B💥 CRASHn/a
Nat.add_comm496.0 ms💥 CRASHn/a1.330 s💥 CRASHn/a34.580💥 CRASHn/a4.33 GiB💥 CRASHn/a9.09 MiB💥 CRASHn/a58.9 ms💥 CRASHn/a308.40M💥 CRASHn/a
FRI verifier on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append6.880 s💥 CRASHn/a54.265 s💥 CRASHn/a51.140💥 CRASHn/a101.05 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a49.5 ms💥 CRASHn/a210.23B💥 CRASHn/a
Char.ofOrdinal_le_of_le6.813 s💥 CRASHn/a53.449 s💥 CRASHn/a51.690💥 CRASHn/a99.81 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a45.1 ms💥 CRASHn/a207.18B💥 CRASHn/a
Array.extract_append6.371 s💥 CRASHn/a50.695 s💥 CRASHn/a31.680💥 CRASHn/a95.19 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a58.1 ms💥 CRASHn/a200.65B💥 CRASHn/a
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq6.748 s💥 CRASHn/a53.378 s💥 CRASHn/a34.980💥 CRASHn/a98.71 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a46.1 ms💥 CRASHn/a205.59B💥 CRASHn/a
Std.HashMap6.455 s💥 CRASHn/a50.868 s💥 CRASHn/a40.140💥 CRASHn/a95.03 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a45.7 ms💥 CRASHn/a203.35B💥 CRASHn/a
String.append5.307 s💥 CRASHn/a47.595 s💥 CRASHn/a6.870💥 CRASHn/a87.89 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a32.4 ms💥 CRASHn/a168.67B💥 CRASHn/a
Nat.add_comm4.526 s💥 CRASHn/a31.415 s💥 CRASHn/a1.460💥 CRASHn/a58.66 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a28.0 ms💥 CRASHn/a130.84B💥 CRASHn/a
Pipeline total (7 constants)
constanttotal-time (main)total-time (PR)Δ%pipeline-throughput (const/s) (main)pipeline-throughput (const/s) (PR)Δ%pipeline-peak-ram (main)pipeline-peak-ram (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append1m 38.6s💥 CRASHn/a28.160💥 CRASHn/a101.05 GiB💥 CRASHn/a
Char.ofOrdinal_le_of_le1m 31.7s💥 CRASHn/a30.130💥 CRASHn/a99.81 GiB💥 CRASHn/a
Array.extract_append1m 24.1s💥 CRASHn/a19.090💥 CRASHn/a95.19 GiB💥 CRASHn/a
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq1m 14.0s💥 CRASHn/a25.230💥 CRASHn/a98.71 GiB💥 CRASHn/a
Std.HashMap1m 13.4s💥 CRASHn/a27.830💥 CRASHn/a95.03 GiB💥 CRASHn/a
String.append50.458 s💥 CRASHn/a6.480💥 CRASHn/a87.89 GiB💥 CRASHn/a
Nat.add_comm32.745 s💥 CRASHn/a1.400💥 CRASHn/a58.66 GiB💥 CRASHn/a

Workflow logs

samuelburnham added a commit that referenced this pull request Sep 1, 2026
The Warp x64 runner pool mixes Intel Granite Rapids and AMD Zen 5, and a
build job may land on one vendor while the job that runs its binaries
lands on the other. Neither vendor's feature set contains the other's, so
`-Ctarget-cpu=native` does not produce a portable binary: Zen 5 enables
SSE4A, and LLVM emits it. Disassembling the workspace built for znver5
finds 31 SSE4A instructions, all INSERTQ, in `ix-ffi` and in
`aiur_ixvm_witness::add_entries_parallel`. Granite Rapids has no SSE4A,
so the first one executed raises #UD, killing the process with SIGILL
during witness generation. That is what turned every row of #605's
benchmark into a crash.
Pin the measured intersection of the two CPUs instead. x86-64-v4 covers
every AVX-512 subset Plonky3 uses; +avx512vbmi2 preserves its VPSHRDQ
interleave and +gfni preserves LLVM's byte-shift lowering. A workspace
built with these flags contains no instruction absent from either vendor
and has an instruction vocabulary identical to a graniterapids build.
blake3 dispatches on CPUID at runtime and is unaffected either way.
`.cargo/config.toml` keeps `-Ctarget-cpu=native`: a developer builds and
runs on one machine, and x86-64-v4 would exclude every host without
AVX-512. Only CI has the split, so only CI pins the ISA. The new guard
fails the job when a runner lacks a required feature, so the assumption
is enforced rather than assumed, and the shared `warp-x64` cargo cache
key becomes sound now that codegen no longer varies by host.
Pinning also removes a benchmarking hazard that never crashed: LLVM sets
prefer-256-bit for Granite Rapids but not for Zen 5, so the same source
vectorized 3.2x more widely depending on the build host, and main-vs-PR
timings were not comparable across a vendor split.
Drop the CPU-mismatch warning the benchmark comment used to carry. It
detected a real problem, but the flags above prevent that problem, and
computing it in one job to render it in another cost a Markdown file
threaded through cache entries, artifacts, and a `--warning-file` flag on
`ix bench compare`. Warnings belong to the run that finds them.
@arthurpaulino

Copy link
Copy Markdown
MemberAuthor

!benchmark fresh

@argument-ci-bot

argument-ci-botBot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

⚠️ Some benchmark jobs failed — results may be partial.

!benchmark — main vs d7fc2f1

backends: aiur=prove · envs: InitStd · baseline: fresh (benchmark products rebuilt, base-SHA run, bencher bypassed)

No result tables were produced — see the run logs.

Workflow logs

The Warp x64 runner pool mixes Intel Granite Rapids and AMD Zen 5, and a
build job may land on one vendor while the job that runs its binaries
lands on the other. Neither vendor's feature set contains the other's, so
`-Ctarget-cpu=native` does not produce a portable binary: Zen 5 enables
SSE4A, and LLVM emits it. Disassembling the workspace built for znver5
finds 31 SSE4A instructions, all INSERTQ, in `ix-ffi` and in
`aiur_ixvm_witness::add_entries_parallel`. Granite Rapids has no SSE4A,
so the first one executed raises #UD, killing the process with SIGILL
during witness generation. That is what turned every row of #605's
benchmark into a crash.
Pin the measured intersection of the two CPUs instead. x86-64-v4 covers
every AVX-512 subset Plonky3 uses; +avx512vbmi2 preserves its VPSHRDQ
interleave and +gfni preserves LLVM's byte-shift lowering. A workspace
built with these flags contains no instruction absent from either vendor
and has an instruction vocabulary identical to a graniterapids build.
blake3 dispatches on CPUID at runtime and is unaffected either way.
`.cargo/config.toml` keeps `-Ctarget-cpu=native`: a developer builds and
runs on one machine, and x86-64-v4 would exclude every host without
AVX-512. Only CI has the split, so only CI pins the ISA. The new guard
fails the job when a runner lacks a required feature, so the assumption
is enforced rather than assumed, and the shared `warp-x64` cargo cache
key becomes sound now that codegen no longer varies by host. RUSTFLAGS
is hashed into that key, so the flag change rotates it on its own.
Pinning also removes a benchmarking hazard that never crashed: LLVM sets
prefer-256-bit for Granite Rapids but not for Zen 5, so the same source
vectorized 3.2x more widely depending on the build host, and main-vs-PR
timings were not comparable across a vendor split.
@samuelburnham

Copy link
Copy Markdown
Member

!benchmark fresh

@argument-ci-bot

argument-ci-botBot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

!benchmark — main vs feb014d

backends: aiur=prove · envs: InitStd · baseline: fresh (benchmark products rebuilt, base-SHA run, bencher bypassed)

aiur · InitStd · prove — main from: base run @ e1ca8e2 (fresh — bencher bypassed)

7 constants · 3 with regressions · 7 with improvements (|Δ| > 3.0% on any metric).

IxVM on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append9.269 s9.073 s-2.1%31.250 s30.269 s-3.1% 🟢88.80091.680+3.2% 🟢70.72 GiB72.28 GiB+2.2%11.33 MiB11.06 MiB-2.4%68.5 ms62.7 ms-8.5% (1.09× faster) 🟢134.35B134.35B+0.0%
Char.ofOrdinal_le_of_le6.964 s6.879 s-1.2%26.884 s25.793 s-4.1% 🟢102.780107.120+4.2% 🟢63.83 GiB65.31 GiB+2.3%11.33 MiB11.07 MiB-2.3%77.9 ms59.3 ms-23.8% (1.31× faster) 🟢102.60B102.60B+0.0%
Array.extract_append6.676 s6.751 s+1.1%23.925 s23.179 s-3.1% 🟢67.13069.290+3.2% 🟢51.97 GiB52.72 GiB+1.4%11.24 MiB10.99 MiB-2.2%71.1 ms57.2 ms-19.5% (1.24× faster) 🟢97.08B97.08B+0.0%
Std.HashMap4.171 s4.079 s-2.2%16.222 s15.628 s-3.7% 🟢125.880130.670+3.8% 🟢36.29 GiB37.10 GiB+2.2%11.26 MiB11.01 MiB-2.2%74.8 ms66.6 ms-11.0% (1.12× faster) 🟢61.88B61.88B+0.0%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq3.697 s3.648 s-1.3%14.858 s14.287 s-3.8% 🟢125.650130.680+4.0% 🟢34.02 GiB34.75 GiB+2.1%11.26 MiB11.00 MiB-2.2%75.5 ms57.7 ms-23.6% (1.31× faster) 🟢55.68B55.68B+0.0%
String.append435.5 ms434.0 ms-0.4%2.307 s2.129 s-7.7% (1.08× faster) 🟢141.750153.600+8.4% (1.08× faster) 🟢5.74 GiB5.52 GiB-3.9% 🟢9.94 MiB9.74 MiB-2.1%64.4 ms50.2 ms-22.1% (1.28× faster) 🟢3.37B3.37B+0.0%
Nat.add_comm267.1 ms264.3 ms-1.1%1.068 s984.9 ms-7.8% (1.08× faster) 🟢43.06046.710+8.5% (1.08× faster) 🟢4.51 GiB3.99 GiB-11.5% (1.13× smaller) 🟢9.09 MiB8.90 MiB-2.1%53.3 ms47.9 ms-10.2% (1.11× faster) 🟢308.40M308.40M+0.0%
FRI verifier on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append5.680 s4.975 s-12.4% (1.14× faster) 🟢33.109 s30.692 s-7.3% (1.08× faster) 🟢83.82090.410+7.9% (1.08× faster) 🟢101.44 GiB99.74 GiB-1.7%3.97 MiB3.98 MiB+0.2%26.3 ms22.3 ms-15.4% (1.18× faster) 🟢210.23B203.74B-3.1% 🟢
Char.ofOrdinal_le_of_le5.599 s5.014 s-10.5% (1.12× faster) 🟢32.280 s31.379 s-2.8%85.59088.050+2.9%99.81 GiB101.09 GiB+1.3%3.97 MiB3.98 MiB+0.2%27.6 ms28.6 ms+3.6% ⚠️207.18B208.08B+0.4%
Array.extract_append5.315 s4.858 s-8.6% (1.09× faster) 🟢30.962 s29.672 s-4.2% 🟢51.87054.120+4.3% 🟢95.00 GiB95.36 GiB+0.4%3.97 MiB3.99 MiB+0.4%26.0 ms39.4 ms+51.3% (1.51× slower) ⚠️200.65B200.40B-0.1%
Std.HashMap5.445 s4.864 s-10.7% (1.12× faster) 🟢31.021 s29.949 s-3.5% 🟢65.83068.180+3.6% 🟢94.50 GiB95.90 GiB+1.5%3.97 MiB3.98 MiB+0.2%25.4 ms25.6 ms+0.8%203.35B204.11B+0.4%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq5.496 s4.898 s-10.9% (1.12× faster) 🟢32.335 s29.501 s-8.8% (1.10× faster) 🟢57.74063.290+9.6% (1.10× faster) 🟢98.72 GiB95.35 GiB-3.4% 🟢3.97 MiB3.98 MiB+0.3%25.9 ms31.9 ms+23.0% (1.23× slower) ⚠️205.59B199.65B-2.9%
String.append4.487 s3.955 s-11.8% (1.13× faster) 🟢28.710 s27.141 s-5.5% (1.06× faster) 🟢11.39012.050+5.8% (1.06× faster) 🟢87.89 GiB88.65 GiB+0.9%3.97 MiB3.97 MiB+0.1%37.6 ms27.4 ms-27.2% (1.37× faster) 🟢168.67B164.80B-2.3%
Nat.add_comm3.619 s3.077 s-15.0% (1.18× faster) 🟢19.611 s17.954 s-8.5% (1.09× faster) 🟢2.3502.560+8.9% (1.09× faster) 🟢58.66 GiB58.03 GiB-1.1%3.97 MiB3.98 MiB+0.2%27.6 ms22.0 ms-20.4% (1.26× faster) 🟢130.84B124.20B-5.1% (1.05× fewer) 🟢
Pipeline total (7 constants)
constanttotal-time (main)total-time (PR)Δ%pipeline-throughput (const/s) (main)pipeline-throughput (const/s) (PR)Δ%pipeline-peak-ram (main)pipeline-peak-ram (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append1m 4.4s1m 1.0s-5.3% (1.06× faster) 🟢43.12045.520+5.6% (1.06× faster) 🟢101.44 GiB99.74 GiB-1.7%
Char.ofOrdinal_le_of_le59.164 s57.172 s-3.4% 🟢46.70048.330+3.5% 🟢99.81 GiB101.09 GiB+1.3%
Array.extract_append54.887 s52.851 s-3.7% 🟢29.26030.390+3.9% 🟢95.00 GiB95.36 GiB+0.4%
Std.HashMap47.243 s45.577 s-3.5% 🟢43.22044.800+3.7% 🟢94.50 GiB95.90 GiB+1.5%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq47.193 s43.788 s-7.2% (1.08× faster) 🟢39.56042.640+7.8% (1.08× faster) 🟢98.72 GiB95.35 GiB-3.4% 🟢
String.append31.017 s29.270 s-5.6% (1.06× faster) 🟢10.54011.170+6.0% (1.06× faster) 🟢87.89 GiB88.65 GiB+0.9%
Nat.add_comm20.680 s18.939 s-8.4% (1.09× faster) 🟢2.2202.430+9.5% (1.09× faster) 🟢58.66 GiB58.03 GiB-1.1%

Workflow logs

#606)
* Authenticate all frontier group members in the recursive verifier
The in-circuit pruned-multiproof walk (mmcs_verify_multi / frontier_level)
collapses queries that share a parent to a single lead node and hashes
only the lead's rows via inject_maybe(ar, ...). Non-lead members' opened
rows for the shorter (injected) matrices were still consumed in their own
per-query FRI arithmetic (batch_views_at) but never authenticated against
any commitment — the leaf hash covers only the tallest matrices, and
shorter ones are bound solely through injection. A prover could therefore
forge a non-lead member's shorter-matrix opening. Plonky3's reference
verify_batch_pruned guards exactly this with InconsistentGroupOpening
(and InconsistentDuplicateOpenings for equal-index queries); the port had
neither. The prior per-query walk did not have the gap, so it was
introduced with the direct multiproof consumption.
- frontier_level: on a group merge, assert the lead and member agree on
every not-yet-injected matrix (height <= next_lh) via select_rows_le +
pointer equality. Transitive across pairwise merges, so the whole group
is pinned; matches InconsistentGroupOpening.
- frontier_merge: duplicate transcript indices must open the SAME full
rows, not merely the same tallest-matrix leaf digest; matches
InconsistentDuplicateOpenings.
Pointer equality is admissible inside assert_eq! (equal pointers imply
equal content; a spurious mismatch costs only completeness — see
IxVM.Core). select_rows_le selects rows of matrices at height <= target,
mirroring select_rows.
Validated: the group-merge branch is genuinely reached by the factorial
recursion proof (an always-false variant of the new assert fails the
honest test), the honest proof still verifies with the real assert
(completeness preserved), the existing tamper tests still reject, and
the full lake test suite is green (2717 checks). aiur_multi_stark.rs
regenerated; kernel executor unchanged.
* Drop the multi_stark::advice dependency
Companion to multi-stark removing its unused per-query advice module.
ix consumed native pruned multiproofs directly and referenced only
advice::AdviceError, whose two arms (verification failed, serialization
failed) were immediately string-formatted by the FFI. Replace it with a
plain Result<Vec<u8>, String>: AiurSystem::proof_to_advice_bytes maps
both failures to a message, and the FFI passes the string straight to
LeanExcept::error_string. No behavior change; the Lean binding
(Except String ByteArray) is unaffected.
Bump the multi-stark pin to the advice-removed revision. Requires that
multi-stark's ap/bump-p3-drop-advice be pushed first, exactly as with
every other pin in this series.
* Bump multi-stark audit revision
---------
Co-authored-by: Arthur Paulino <arthurleonardo.ap@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@arthurpaulino@samuelburnham
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Update Aiur to Plonky3 0.6 - #605

Draft
arthurpaulino wants to merge 5 commits into
mainfrom
ap/bump-p3
Draft

Update Aiur to Plonky3 0.6#605
arthurpaulino wants to merge 5 commits into
mainfrom
ap/bump-p3

Conversation

@arthurpaulino

Copy link
Copy Markdown
Member

Update the multi-stark dependency and Rust toolchain for Plonky3 0.6, along with the Rust 1.98 lint migrations required to keep the workspace warning-free. Refresh the Rust-compatible Blake3.lean pin in both root and compile-package manifests.

Adapt recursive Aiur verification to Plonky3's pruned FRI multiproofs. Native proofs retain their compact serialized representation and native verification path; the FFI expands authenticated Merkle frontiers into per-query advice only when entering the existing recursive verifier circuit.

Preserve the packed claim-digest convention in the recursion diagnostic and exercise the advice boundary in the end-to-end test and benchmark paths. CPU and CUDA recursive q1 runs produce identical 823,485-byte inner proofs and 331,273-byte outer proofs.

The q50 Vector.extract_append workload retains identical CPU/CUDA proof sizes. Inner plus outer STARK proving measures 65.87s on CPU and 8.81s with CUDA on the RTX PRO 6000, a 7.48x speedup.

Update the multi-stark dependency and Rust toolchain for Plonky3 0.6, along with the Rust 1.98 lint migrations required to keep the workspace warning-free. Refresh the Rust-compatible Blake3.lean pin in both root and compile-package manifests.
Adapt recursive Aiur verification to Plonky3's pruned FRI multiproofs. Native proofs retain their compact serialized representation and native verification path; the FFI expands authenticated Merkle frontiers into per-query advice only when entering the existing recursive verifier circuit.
Preserve the packed claim-digest convention in the recursion diagnostic and exercise the advice boundary in the end-to-end test and benchmark paths. CPU and CUDA recursive q1 runs produce identical 823,485-byte inner proofs and 331,273-byte outer proofs.
The q50 Vector.extract_append workload retains identical CPU/CUDA proof sizes. Inner plus outer STARK proving measures 65.87s on CPU and 8.81s with CUDA on the RTX PRO 6000, a 7.48x speedup.
PR benchmark runs execute trusted workflow YAML from the default branch while loading composite actions from the PR checkout. When Bencher data and binary caches are unavailable, the workflow checks out main under base/ and asks Lake to rebuild it without first installing the Rust channel pinned by that checkout.
Teach the existing CPU provenance action to install the base checkout's validated Rust channel and profile immediately before an uncached base build. The step is a no-op when the toolchain is already available and leaves cached benchmark comparisons unchanged.
@arthurpaulino

Copy link
Copy Markdown
MemberAuthor

!benchmark fresh

@argument-ci-bot

argument-ci-botBot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

!benchmark — main vs eab18e4

backends: aiur=prove · envs: InitStd · baseline: fresh (benchmark products rebuilt, base-SHA run, bencher bypassed)

aiur · InitStd · prove — main from: base run @ e1ca8e2 (fresh — bencher bypassed)

7 constants · 1 with regressions · 7 with improvements (|Δ| > 3.0% on any metric).

IxVM on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append8.816 s8.990 s+2.0%30.095 s29.094 s-3.3% 🟢92.21095.380+3.4% 🟢70.72 GiB72.35 GiB+2.3%11.33 MiB11.07 MiB-2.3%68.7 ms60.4 ms-12.1% (1.14× faster) 🟢134.35B134.35B+0.0%
Char.ofOrdinal_le_of_le6.689 s6.790 s+1.5%25.631 s25.209 s-1.6%107.800109.600+1.7%63.87 GiB65.41 GiB+2.4%11.33 MiB11.07 MiB-2.3%73.7 ms60.3 ms-18.2% (1.22× faster) 🟢102.60B102.60B+0.0%
Array.extract_append6.679 s6.415 s-4.0% 🟢23.076 s22.072 s-4.4% 🟢69.60072.760+4.5% 🟢52.03 GiB52.79 GiB+1.4%11.24 MiB10.99 MiB-2.2%72.3 ms64.6 ms-10.6% (1.12× faster) 🟢97.08B97.08B+0.0%
Std.HashMap3.964 s4.002 s+1.0%15.519 s15.133 s-2.5%131.580134.940+2.6%36.34 GiB37.08 GiB+2.1%11.26 MiB11.00 MiB-2.3%75.1 ms65.5 ms-12.8% (1.15× faster) 🟢61.88B61.88B+0.0%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq3.598 s3.528 s-1.9%14.327 s13.806 s-3.6% 🟢130.310135.230+3.8% 🟢33.96 GiB34.69 GiB+2.1%11.26 MiB11.01 MiB-2.2%69.9 ms59.5 ms-15.0% (1.18× faster) 🟢55.68B55.68B+0.0%
String.append424.3 ms423.0 ms-0.3%2.278 s2.124 s-6.7% (1.07× faster) 🟢143.540153.930+7.2% (1.07× faster) 🟢4.89 GiB5.00 GiB+2.4%9.94 MiB9.74 MiB-2.0%64.2 ms52.0 ms-19.0% (1.24× faster) 🟢3.37B3.37B+0.0%
Nat.add_comm260.1 ms259.8 ms-0.1%1.068 s970.5 ms-9.1% (1.10× faster) 🟢43.07047.400+10.1% (1.10× faster) 🟢3.99 GiB4.64 GiB+16.4% (1.16× larger) ⚠️9.09 MiB8.91 MiB-2.0%53.9 ms41.9 ms-22.3% (1.29× faster) 🟢308.40M308.40M+0.0%
FRI verifier on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append5.488 s5.490 s+0.0%31.664 s30.699 s-3.0% 🟢87.64090.400+3.1% 🟢101.37 GiB101.07 GiB-0.3%3.97 MiB3.72 MiB-6.3% (1.07× smaller) 🟢29.2 ms23.8 ms-18.5% (1.23× faster) 🟢210.23B210.23B+0.0%
Char.ofOrdinal_le_of_le5.386 s5.489 s+1.9%30.846 s30.914 s+0.2%89.57089.380-0.2%100.32 GiB100.32 GiB+0.0%3.97 MiB3.72 MiB-6.3% (1.07× smaller) 🟢24.8 ms23.1 ms-6.6% (1.07× faster) 🟢207.18B207.18B+0.0%
Array.extract_append5.209 s5.235 s+0.5%29.948 s28.844 s-3.7% 🟢53.63055.680+3.8% 🟢94.59 GiB95.37 GiB+0.8%3.97 MiB3.73 MiB-6.1% (1.07× smaller) 🟢25.7 ms21.3 ms-17.1% (1.21× faster) 🟢200.65B200.65B+0.0%
Std.HashMap5.314 s5.306 s-0.1%29.652 s29.584 s-0.2%68.87069.020+0.2%94.56 GiB95.34 GiB+0.8%3.97 MiB3.72 MiB-6.3% (1.07× smaller) 🟢25.2 ms21.2 ms-15.7% (1.19× faster) 🟢203.35B203.35B+0.0%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq5.400 s5.333 s-1.2%31.051 s30.174 s-2.8%60.13061.870+2.9%98.68 GiB99.48 GiB+0.8%3.97 MiB3.73 MiB-6.2% (1.07× smaller) 🟢30.2 ms31.0 ms+2.5%205.59B205.59B+0.0%
String.append4.334 s4.419 s+2.0%27.489 s26.612 s-3.2% 🟢11.90012.290+3.3% 🟢87.87 GiB88.58 GiB+0.8%3.97 MiB3.73 MiB-6.1% (1.07× smaller) 🟢34.5 ms22.7 ms-34.3% (1.52× faster) 🟢168.67B168.67B+0.0%
Nat.add_comm3.518 s3.499 s-0.5%18.993 s17.942 s-5.5% (1.06× faster) 🟢2.4202.560+5.8% (1.06× faster) 🟢58.70 GiB59.42 GiB+1.2%3.97 MiB3.72 MiB-6.2% (1.07× smaller) 🟢24.8 ms20.8 ms-16.1% (1.19× faster) 🟢130.84B130.84B+0.0%
Pipeline total (7 constants)
constanttotal-time (main)total-time (PR)Δ%pipeline-throughput (const/s) (main)pipeline-throughput (const/s) (PR)Δ%pipeline-peak-ram (main)pipeline-peak-ram (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append1m 1.8s59.792 s-3.2% 🟢44.93046.410+3.3% 🟢101.37 GiB101.07 GiB-0.3%
Char.ofOrdinal_le_of_le56.477 s56.124 s-0.6%48.92049.230+0.6%100.32 GiB100.32 GiB+0.0%
Array.extract_append53.024 s50.916 s-4.0% 🟢30.29031.540+4.1% 🟢94.59 GiB95.37 GiB+0.8%
Std.HashMap45.172 s44.717 s-1.0%45.21045.660+1.0%94.56 GiB95.34 GiB+0.8%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq45.377 s43.980 s-3.1% 🟢41.14042.450+3.2% 🟢98.68 GiB99.48 GiB+0.8%
String.append29.767 s28.737 s-3.5% 🟢10.99011.380+3.5% 🟢87.87 GiB88.58 GiB+0.8%
Nat.add_comm20.061 s18.912 s-5.7% (1.06× faster) 🟢2.2902.430+6.1% (1.06× faster) 🟢58.70 GiB59.42 GiB+1.2%

Workflow logs

Consume the Plonky3 0.6 batch-opening layout directly in Aiur instead of expanding every pruned Merkle frontier into one authentication path per FRI query. Sample all query indices from the unchanged transcript, sort and deduplicate them with an O(q log q) merge sort, authenticate each input and commit-phase commitment once, then retain the existing per-query reduced-opening and FRI arithmetic.
Bind every frontier to transcript-derived indices, consume boundary digests in Plonky3's level/parent/child order, reject trailing frontier elements and inconsistent duplicate leaves, and assert all native opening dimensions and sibling counts. Explicitly constrain the digest-bound protocol specialization to cap height 0, binary FRI, and a constant final polynomial. Move memo_u32_less_than into IxVM Core so both substitution and multiproof sorting share its constrained rows.
Strengthen the recursive negative test to mutate a structurally valid stage-1 commitment. Regenerate both checked-in Aiur Rust executors and retain interpreter/codegen query-count parity.
On Vector.extract_append q50, recursive-verifier FFT cost falls from 204.073B to 201.166B. CPU outer proving improves from 50.09s to 45.03s and the full CPU pipeline from 90.64s to 82.90s. GPU outer proving improves from 15.85s to 13.72s and the full GPU pipeline from 28.86s to 26.69s. The outer proof grows from 3.92 MB to 4.17 MB.
Validated with the MultiStark primitive suite, recursive honest/tamper/parity tests, codegen --check, release workspace clippy, release CUDA clippy, rustfmt, and diff checks.
@arthurpaulino

Copy link
Copy Markdown
MemberAuthor

!benchmark fresh

@argument-ci-bot

argument-ci-botBot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

!benchmark — main vs 840932a

backends: aiur=prove · envs: InitStd · baseline: fresh (benchmark products rebuilt, base-SHA run, bencher bypassed)

Warning

  • CPU model mismatch for PR benchmark binaries in this job: built on AMD EPYC 9R45; measured on Intel(R) Xeon(R) 6975P-C. Native Rust code uses -Ctarget-cpu=native.

aiur · InitStd · prove — main from: base run @ e1ca8e2 (fresh — bencher bypassed)

7 constants · 0 with regressions · 0 with improvements (|Δ| > 3.0% on any metric).

IxVM on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append8.874 s💥 CRASHn/a44.291 s💥 CRASHn/a62.650💥 CRASHn/a70.80 GiB💥 CRASHn/a11.33 MiB💥 CRASHn/a77.4 ms💥 CRASHn/a134.35B💥 CRASHn/a
Char.ofOrdinal_le_of_le6.603 s💥 CRASHn/a38.242 s💥 CRASHn/a72.250💥 CRASHn/a63.81 GiB💥 CRASHn/a11.33 MiB💥 CRASHn/a76.9 ms💥 CRASHn/a102.60B💥 CRASHn/a
Array.extract_append6.311 s💥 CRASHn/a33.420 s💥 CRASHn/a48.060💥 CRASHn/a51.95 GiB💥 CRASHn/a11.24 MiB💥 CRASHn/a86.5 ms💥 CRASHn/a97.08B💥 CRASHn/a
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq3.774 s💥 CRASHn/a20.628 s💥 CRASHn/a90.510💥 CRASHn/a33.97 GiB💥 CRASHn/a11.26 MiB💥 CRASHn/a82.6 ms💥 CRASHn/a55.68B💥 CRASHn/a
Std.HashMap4.060 s💥 CRASHn/a22.501 s💥 CRASHn/a90.750💥 CRASHn/a36.36 GiB💥 CRASHn/a11.26 MiB💥 CRASHn/a84.1 ms💥 CRASHn/a61.88B💥 CRASHn/a
String.append708.3 ms💥 CRASHn/a2.863 s💥 CRASHn/a114.220💥 CRASHn/a4.99 GiB💥 CRASHn/a9.94 MiB💥 CRASHn/a70.3 ms💥 CRASHn/a3.37B💥 CRASHn/a
Nat.add_comm496.0 ms💥 CRASHn/a1.330 s💥 CRASHn/a34.580💥 CRASHn/a4.33 GiB💥 CRASHn/a9.09 MiB💥 CRASHn/a58.9 ms💥 CRASHn/a308.40M💥 CRASHn/a
FRI verifier on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append6.880 s💥 CRASHn/a54.265 s💥 CRASHn/a51.140💥 CRASHn/a101.05 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a49.5 ms💥 CRASHn/a210.23B💥 CRASHn/a
Char.ofOrdinal_le_of_le6.813 s💥 CRASHn/a53.449 s💥 CRASHn/a51.690💥 CRASHn/a99.81 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a45.1 ms💥 CRASHn/a207.18B💥 CRASHn/a
Array.extract_append6.371 s💥 CRASHn/a50.695 s💥 CRASHn/a31.680💥 CRASHn/a95.19 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a58.1 ms💥 CRASHn/a200.65B💥 CRASHn/a
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq6.748 s💥 CRASHn/a53.378 s💥 CRASHn/a34.980💥 CRASHn/a98.71 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a46.1 ms💥 CRASHn/a205.59B💥 CRASHn/a
Std.HashMap6.455 s💥 CRASHn/a50.868 s💥 CRASHn/a40.140💥 CRASHn/a95.03 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a45.7 ms💥 CRASHn/a203.35B💥 CRASHn/a
String.append5.307 s💥 CRASHn/a47.595 s💥 CRASHn/a6.870💥 CRASHn/a87.89 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a32.4 ms💥 CRASHn/a168.67B💥 CRASHn/a
Nat.add_comm4.526 s💥 CRASHn/a31.415 s💥 CRASHn/a1.460💥 CRASHn/a58.66 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a28.0 ms💥 CRASHn/a130.84B💥 CRASHn/a
Pipeline total (7 constants)
constanttotal-time (main)total-time (PR)Δ%pipeline-throughput (const/s) (main)pipeline-throughput (const/s) (PR)Δ%pipeline-peak-ram (main)pipeline-peak-ram (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append1m 38.6s💥 CRASHn/a28.160💥 CRASHn/a101.05 GiB💥 CRASHn/a
Char.ofOrdinal_le_of_le1m 31.7s💥 CRASHn/a30.130💥 CRASHn/a99.81 GiB💥 CRASHn/a
Array.extract_append1m 24.1s💥 CRASHn/a19.090💥 CRASHn/a95.19 GiB💥 CRASHn/a
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq1m 14.0s💥 CRASHn/a25.230💥 CRASHn/a98.71 GiB💥 CRASHn/a
Std.HashMap1m 13.4s💥 CRASHn/a27.830💥 CRASHn/a95.03 GiB💥 CRASHn/a
String.append50.458 s💥 CRASHn/a6.480💥 CRASHn/a87.89 GiB💥 CRASHn/a
Nat.add_comm32.745 s💥 CRASHn/a1.400💥 CRASHn/a58.66 GiB💥 CRASHn/a

Workflow logs

samuelburnham added a commit that referenced this pull request Sep 1, 2026
The Warp x64 runner pool mixes Intel Granite Rapids and AMD Zen 5, and a
build job may land on one vendor while the job that runs its binaries
lands on the other. Neither vendor's feature set contains the other's, so
`-Ctarget-cpu=native` does not produce a portable binary: Zen 5 enables
SSE4A, and LLVM emits it. Disassembling the workspace built for znver5
finds 31 SSE4A instructions, all INSERTQ, in `ix-ffi` and in
`aiur_ixvm_witness::add_entries_parallel`. Granite Rapids has no SSE4A,
so the first one executed raises #UD, killing the process with SIGILL
during witness generation. That is what turned every row of #605's
benchmark into a crash.
Pin the measured intersection of the two CPUs instead. x86-64-v4 covers
every AVX-512 subset Plonky3 uses; +avx512vbmi2 preserves its VPSHRDQ
interleave and +gfni preserves LLVM's byte-shift lowering. A workspace
built with these flags contains no instruction absent from either vendor
and has an instruction vocabulary identical to a graniterapids build.
blake3 dispatches on CPUID at runtime and is unaffected either way.
`.cargo/config.toml` keeps `-Ctarget-cpu=native`: a developer builds and
runs on one machine, and x86-64-v4 would exclude every host without
AVX-512. Only CI has the split, so only CI pins the ISA. The new guard
fails the job when a runner lacks a required feature, so the assumption
is enforced rather than assumed, and the shared `warp-x64` cargo cache
key becomes sound now that codegen no longer varies by host.
Pinning also removes a benchmarking hazard that never crashed: LLVM sets
prefer-256-bit for Granite Rapids but not for Zen 5, so the same source
vectorized 3.2x more widely depending on the build host, and main-vs-PR
timings were not comparable across a vendor split.
Drop the CPU-mismatch warning the benchmark comment used to carry. It
detected a real problem, but the flags above prevent that problem, and
computing it in one job to render it in another cost a Markdown file
threaded through cache entries, artifacts, and a `--warning-file` flag on
`ix bench compare`. Warnings belong to the run that finds them.
@arthurpaulino

Copy link
Copy Markdown
MemberAuthor

!benchmark fresh

@argument-ci-bot

argument-ci-botBot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

⚠️ Some benchmark jobs failed — results may be partial.

!benchmark — main vs d7fc2f1

backends: aiur=prove · envs: InitStd · baseline: fresh (benchmark products rebuilt, base-SHA run, bencher bypassed)

No result tables were produced — see the run logs.

Workflow logs

The Warp x64 runner pool mixes Intel Granite Rapids and AMD Zen 5, and a
build job may land on one vendor while the job that runs its binaries
lands on the other. Neither vendor's feature set contains the other's, so
`-Ctarget-cpu=native` does not produce a portable binary: Zen 5 enables
SSE4A, and LLVM emits it. Disassembling the workspace built for znver5
finds 31 SSE4A instructions, all INSERTQ, in `ix-ffi` and in
`aiur_ixvm_witness::add_entries_parallel`. Granite Rapids has no SSE4A,
so the first one executed raises #UD, killing the process with SIGILL
during witness generation. That is what turned every row of #605's
benchmark into a crash.
Pin the measured intersection of the two CPUs instead. x86-64-v4 covers
every AVX-512 subset Plonky3 uses; +avx512vbmi2 preserves its VPSHRDQ
interleave and +gfni preserves LLVM's byte-shift lowering. A workspace
built with these flags contains no instruction absent from either vendor
and has an instruction vocabulary identical to a graniterapids build.
blake3 dispatches on CPUID at runtime and is unaffected either way.
`.cargo/config.toml` keeps `-Ctarget-cpu=native`: a developer builds and
runs on one machine, and x86-64-v4 would exclude every host without
AVX-512. Only CI has the split, so only CI pins the ISA. The new guard
fails the job when a runner lacks a required feature, so the assumption
is enforced rather than assumed, and the shared `warp-x64` cargo cache
key becomes sound now that codegen no longer varies by host. RUSTFLAGS
is hashed into that key, so the flag change rotates it on its own.
Pinning also removes a benchmarking hazard that never crashed: LLVM sets
prefer-256-bit for Granite Rapids but not for Zen 5, so the same source
vectorized 3.2x more widely depending on the build host, and main-vs-PR
timings were not comparable across a vendor split.
@samuelburnham

Copy link
Copy Markdown
Member

!benchmark fresh

@argument-ci-bot

argument-ci-botBot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

!benchmark — main vs feb014d

backends: aiur=prove · envs: InitStd · baseline: fresh (benchmark products rebuilt, base-SHA run, bencher bypassed)

aiur · InitStd · prove — main from: base run @ e1ca8e2 (fresh — bencher bypassed)

7 constants · 3 with regressions · 7 with improvements (|Δ| > 3.0% on any metric).

IxVM on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append9.269 s9.073 s-2.1%31.250 s30.269 s-3.1% 🟢88.80091.680+3.2% 🟢70.72 GiB72.28 GiB+2.2%11.33 MiB11.06 MiB-2.4%68.5 ms62.7 ms-8.5% (1.09× faster) 🟢134.35B134.35B+0.0%
Char.ofOrdinal_le_of_le6.964 s6.879 s-1.2%26.884 s25.793 s-4.1% 🟢102.780107.120+4.2% 🟢63.83 GiB65.31 GiB+2.3%11.33 MiB11.07 MiB-2.3%77.9 ms59.3 ms-23.8% (1.31× faster) 🟢102.60B102.60B+0.0%
Array.extract_append6.676 s6.751 s+1.1%23.925 s23.179 s-3.1% 🟢67.13069.290+3.2% 🟢51.97 GiB52.72 GiB+1.4%11.24 MiB10.99 MiB-2.2%71.1 ms57.2 ms-19.5% (1.24× faster) 🟢97.08B97.08B+0.0%
Std.HashMap4.171 s4.079 s-2.2%16.222 s15.628 s-3.7% 🟢125.880130.670+3.8% 🟢36.29 GiB37.10 GiB+2.2%11.26 MiB11.01 MiB-2.2%74.8 ms66.6 ms-11.0% (1.12× faster) 🟢61.88B61.88B+0.0%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq3.697 s3.648 s-1.3%14.858 s14.287 s-3.8% 🟢125.650130.680+4.0% 🟢34.02 GiB34.75 GiB+2.1%11.26 MiB11.00 MiB-2.2%75.5 ms57.7 ms-23.6% (1.31× faster) 🟢55.68B55.68B+0.0%
String.append435.5 ms434.0 ms-0.4%2.307 s2.129 s-7.7% (1.08× faster) 🟢141.750153.600+8.4% (1.08× faster) 🟢5.74 GiB5.52 GiB-3.9% 🟢9.94 MiB9.74 MiB-2.1%64.4 ms50.2 ms-22.1% (1.28× faster) 🟢3.37B3.37B+0.0%
Nat.add_comm267.1 ms264.3 ms-1.1%1.068 s984.9 ms-7.8% (1.08× faster) 🟢43.06046.710+8.5% (1.08× faster) 🟢4.51 GiB3.99 GiB-11.5% (1.13× smaller) 🟢9.09 MiB8.90 MiB-2.1%53.3 ms47.9 ms-10.2% (1.11× faster) 🟢308.40M308.40M+0.0%
FRI verifier on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append5.680 s4.975 s-12.4% (1.14× faster) 🟢33.109 s30.692 s-7.3% (1.08× faster) 🟢83.82090.410+7.9% (1.08× faster) 🟢101.44 GiB99.74 GiB-1.7%3.97 MiB3.98 MiB+0.2%26.3 ms22.3 ms-15.4% (1.18× faster) 🟢210.23B203.74B-3.1% 🟢
Char.ofOrdinal_le_of_le5.599 s5.014 s-10.5% (1.12× faster) 🟢32.280 s31.379 s-2.8%85.59088.050+2.9%99.81 GiB101.09 GiB+1.3%3.97 MiB3.98 MiB+0.2%27.6 ms28.6 ms+3.6% ⚠️207.18B208.08B+0.4%
Array.extract_append5.315 s4.858 s-8.6% (1.09× faster) 🟢30.962 s29.672 s-4.2% 🟢51.87054.120+4.3% 🟢95.00 GiB95.36 GiB+0.4%3.97 MiB3.99 MiB+0.4%26.0 ms39.4 ms+51.3% (1.51× slower) ⚠️200.65B200.40B-0.1%
Std.HashMap5.445 s4.864 s-10.7% (1.12× faster) 🟢31.021 s29.949 s-3.5% 🟢65.83068.180+3.6% 🟢94.50 GiB95.90 GiB+1.5%3.97 MiB3.98 MiB+0.2%25.4 ms25.6 ms+0.8%203.35B204.11B+0.4%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq5.496 s4.898 s-10.9% (1.12× faster) 🟢32.335 s29.501 s-8.8% (1.10× faster) 🟢57.74063.290+9.6% (1.10× faster) 🟢98.72 GiB95.35 GiB-3.4% 🟢3.97 MiB3.98 MiB+0.3%25.9 ms31.9 ms+23.0% (1.23× slower) ⚠️205.59B199.65B-2.9%
String.append4.487 s3.955 s-11.8% (1.13× faster) 🟢28.710 s27.141 s-5.5% (1.06× faster) 🟢11.39012.050+5.8% (1.06× faster) 🟢87.89 GiB88.65 GiB+0.9%3.97 MiB3.97 MiB+0.1%37.6 ms27.4 ms-27.2% (1.37× faster) 🟢168.67B164.80B-2.3%
Nat.add_comm3.619 s3.077 s-15.0% (1.18× faster) 🟢19.611 s17.954 s-8.5% (1.09× faster) 🟢2.3502.560+8.9% (1.09× faster) 🟢58.66 GiB58.03 GiB-1.1%3.97 MiB3.98 MiB+0.2%27.6 ms22.0 ms-20.4% (1.26× faster) 🟢130.84B124.20B-5.1% (1.05× fewer) 🟢
Pipeline total (7 constants)
constanttotal-time (main)total-time (PR)Δ%pipeline-throughput (const/s) (main)pipeline-throughput (const/s) (PR)Δ%pipeline-peak-ram (main)pipeline-peak-ram (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append1m 4.4s1m 1.0s-5.3% (1.06× faster) 🟢43.12045.520+5.6% (1.06× faster) 🟢101.44 GiB99.74 GiB-1.7%
Char.ofOrdinal_le_of_le59.164 s57.172 s-3.4% 🟢46.70048.330+3.5% 🟢99.81 GiB101.09 GiB+1.3%
Array.extract_append54.887 s52.851 s-3.7% 🟢29.26030.390+3.9% 🟢95.00 GiB95.36 GiB+0.4%
Std.HashMap47.243 s45.577 s-3.5% 🟢43.22044.800+3.7% 🟢94.50 GiB95.90 GiB+1.5%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq47.193 s43.788 s-7.2% (1.08× faster) 🟢39.56042.640+7.8% (1.08× faster) 🟢98.72 GiB95.35 GiB-3.4% 🟢
String.append31.017 s29.270 s-5.6% (1.06× faster) 🟢10.54011.170+6.0% (1.06× faster) 🟢87.89 GiB88.65 GiB+0.9%
Nat.add_comm20.680 s18.939 s-8.4% (1.09× faster) 🟢2.2202.430+9.5% (1.09× faster) 🟢58.66 GiB58.03 GiB-1.1%

Workflow logs

#606)
* Authenticate all frontier group members in the recursive verifier
The in-circuit pruned-multiproof walk (mmcs_verify_multi / frontier_level)
collapses queries that share a parent to a single lead node and hashes
only the lead's rows via inject_maybe(ar, ...). Non-lead members' opened
rows for the shorter (injected) matrices were still consumed in their own
per-query FRI arithmetic (batch_views_at) but never authenticated against
any commitment — the leaf hash covers only the tallest matrices, and
shorter ones are bound solely through injection. A prover could therefore
forge a non-lead member's shorter-matrix opening. Plonky3's reference
verify_batch_pruned guards exactly this with InconsistentGroupOpening
(and InconsistentDuplicateOpenings for equal-index queries); the port had
neither. The prior per-query walk did not have the gap, so it was
introduced with the direct multiproof consumption.
- frontier_level: on a group merge, assert the lead and member agree on
every not-yet-injected matrix (height <= next_lh) via select_rows_le +
pointer equality. Transitive across pairwise merges, so the whole group
is pinned; matches InconsistentGroupOpening.
- frontier_merge: duplicate transcript indices must open the SAME full
rows, not merely the same tallest-matrix leaf digest; matches
InconsistentDuplicateOpenings.
Pointer equality is admissible inside assert_eq! (equal pointers imply
equal content; a spurious mismatch costs only completeness — see
IxVM.Core). select_rows_le selects rows of matrices at height <= target,
mirroring select_rows.
Validated: the group-merge branch is genuinely reached by the factorial
recursion proof (an always-false variant of the new assert fails the
honest test), the honest proof still verifies with the real assert
(completeness preserved), the existing tamper tests still reject, and
the full lake test suite is green (2717 checks). aiur_multi_stark.rs
regenerated; kernel executor unchanged.
* Drop the multi_stark::advice dependency
Companion to multi-stark removing its unused per-query advice module.
ix consumed native pruned multiproofs directly and referenced only
advice::AdviceError, whose two arms (verification failed, serialization
failed) were immediately string-formatted by the FFI. Replace it with a
plain Result<Vec<u8>, String>: AiurSystem::proof_to_advice_bytes maps
both failures to a message, and the FFI passes the string straight to
LeanExcept::error_string. No behavior change; the Lean binding
(Except String ByteArray) is unaffected.
Bump the multi-stark pin to the advice-removed revision. Requires that
multi-stark's ap/bump-p3-drop-advice be pushed first, exactly as with
every other pin in this series.
* Bump multi-stark audit revision
---------
Co-authored-by: Arthur Paulino <arthurleonardo.ap@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@arthurpaulino@samuelburnham
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Update Aiur to Plonky3 0.6 - #605

Draft
arthurpaulino wants to merge 5 commits into
mainfrom
ap/bump-p3
Draft

Update Aiur to Plonky3 0.6#605
arthurpaulino wants to merge 5 commits into
mainfrom
ap/bump-p3

Conversation

@arthurpaulino

Copy link
Copy Markdown
Member

Update the multi-stark dependency and Rust toolchain for Plonky3 0.6, along with the Rust 1.98 lint migrations required to keep the workspace warning-free. Refresh the Rust-compatible Blake3.lean pin in both root and compile-package manifests.

Adapt recursive Aiur verification to Plonky3's pruned FRI multiproofs. Native proofs retain their compact serialized representation and native verification path; the FFI expands authenticated Merkle frontiers into per-query advice only when entering the existing recursive verifier circuit.

Preserve the packed claim-digest convention in the recursion diagnostic and exercise the advice boundary in the end-to-end test and benchmark paths. CPU and CUDA recursive q1 runs produce identical 823,485-byte inner proofs and 331,273-byte outer proofs.

The q50 Vector.extract_append workload retains identical CPU/CUDA proof sizes. Inner plus outer STARK proving measures 65.87s on CPU and 8.81s with CUDA on the RTX PRO 6000, a 7.48x speedup.

Update the multi-stark dependency and Rust toolchain for Plonky3 0.6, along with the Rust 1.98 lint migrations required to keep the workspace warning-free. Refresh the Rust-compatible Blake3.lean pin in both root and compile-package manifests.
Adapt recursive Aiur verification to Plonky3's pruned FRI multiproofs. Native proofs retain their compact serialized representation and native verification path; the FFI expands authenticated Merkle frontiers into per-query advice only when entering the existing recursive verifier circuit.
Preserve the packed claim-digest convention in the recursion diagnostic and exercise the advice boundary in the end-to-end test and benchmark paths. CPU and CUDA recursive q1 runs produce identical 823,485-byte inner proofs and 331,273-byte outer proofs.
The q50 Vector.extract_append workload retains identical CPU/CUDA proof sizes. Inner plus outer STARK proving measures 65.87s on CPU and 8.81s with CUDA on the RTX PRO 6000, a 7.48x speedup.
PR benchmark runs execute trusted workflow YAML from the default branch while loading composite actions from the PR checkout. When Bencher data and binary caches are unavailable, the workflow checks out main under base/ and asks Lake to rebuild it without first installing the Rust channel pinned by that checkout.
Teach the existing CPU provenance action to install the base checkout's validated Rust channel and profile immediately before an uncached base build. The step is a no-op when the toolchain is already available and leaves cached benchmark comparisons unchanged.
@arthurpaulino

Copy link
Copy Markdown
MemberAuthor

!benchmark fresh

@argument-ci-bot

argument-ci-botBot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

!benchmark — main vs eab18e4

backends: aiur=prove · envs: InitStd · baseline: fresh (benchmark products rebuilt, base-SHA run, bencher bypassed)

aiur · InitStd · prove — main from: base run @ e1ca8e2 (fresh — bencher bypassed)

7 constants · 1 with regressions · 7 with improvements (|Δ| > 3.0% on any metric).

IxVM on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append8.816 s8.990 s+2.0%30.095 s29.094 s-3.3% 🟢92.21095.380+3.4% 🟢70.72 GiB72.35 GiB+2.3%11.33 MiB11.07 MiB-2.3%68.7 ms60.4 ms-12.1% (1.14× faster) 🟢134.35B134.35B+0.0%
Char.ofOrdinal_le_of_le6.689 s6.790 s+1.5%25.631 s25.209 s-1.6%107.800109.600+1.7%63.87 GiB65.41 GiB+2.4%11.33 MiB11.07 MiB-2.3%73.7 ms60.3 ms-18.2% (1.22× faster) 🟢102.60B102.60B+0.0%
Array.extract_append6.679 s6.415 s-4.0% 🟢23.076 s22.072 s-4.4% 🟢69.60072.760+4.5% 🟢52.03 GiB52.79 GiB+1.4%11.24 MiB10.99 MiB-2.2%72.3 ms64.6 ms-10.6% (1.12× faster) 🟢97.08B97.08B+0.0%
Std.HashMap3.964 s4.002 s+1.0%15.519 s15.133 s-2.5%131.580134.940+2.6%36.34 GiB37.08 GiB+2.1%11.26 MiB11.00 MiB-2.3%75.1 ms65.5 ms-12.8% (1.15× faster) 🟢61.88B61.88B+0.0%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq3.598 s3.528 s-1.9%14.327 s13.806 s-3.6% 🟢130.310135.230+3.8% 🟢33.96 GiB34.69 GiB+2.1%11.26 MiB11.01 MiB-2.2%69.9 ms59.5 ms-15.0% (1.18× faster) 🟢55.68B55.68B+0.0%
String.append424.3 ms423.0 ms-0.3%2.278 s2.124 s-6.7% (1.07× faster) 🟢143.540153.930+7.2% (1.07× faster) 🟢4.89 GiB5.00 GiB+2.4%9.94 MiB9.74 MiB-2.0%64.2 ms52.0 ms-19.0% (1.24× faster) 🟢3.37B3.37B+0.0%
Nat.add_comm260.1 ms259.8 ms-0.1%1.068 s970.5 ms-9.1% (1.10× faster) 🟢43.07047.400+10.1% (1.10× faster) 🟢3.99 GiB4.64 GiB+16.4% (1.16× larger) ⚠️9.09 MiB8.91 MiB-2.0%53.9 ms41.9 ms-22.3% (1.29× faster) 🟢308.40M308.40M+0.0%
FRI verifier on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append5.488 s5.490 s+0.0%31.664 s30.699 s-3.0% 🟢87.64090.400+3.1% 🟢101.37 GiB101.07 GiB-0.3%3.97 MiB3.72 MiB-6.3% (1.07× smaller) 🟢29.2 ms23.8 ms-18.5% (1.23× faster) 🟢210.23B210.23B+0.0%
Char.ofOrdinal_le_of_le5.386 s5.489 s+1.9%30.846 s30.914 s+0.2%89.57089.380-0.2%100.32 GiB100.32 GiB+0.0%3.97 MiB3.72 MiB-6.3% (1.07× smaller) 🟢24.8 ms23.1 ms-6.6% (1.07× faster) 🟢207.18B207.18B+0.0%
Array.extract_append5.209 s5.235 s+0.5%29.948 s28.844 s-3.7% 🟢53.63055.680+3.8% 🟢94.59 GiB95.37 GiB+0.8%3.97 MiB3.73 MiB-6.1% (1.07× smaller) 🟢25.7 ms21.3 ms-17.1% (1.21× faster) 🟢200.65B200.65B+0.0%
Std.HashMap5.314 s5.306 s-0.1%29.652 s29.584 s-0.2%68.87069.020+0.2%94.56 GiB95.34 GiB+0.8%3.97 MiB3.72 MiB-6.3% (1.07× smaller) 🟢25.2 ms21.2 ms-15.7% (1.19× faster) 🟢203.35B203.35B+0.0%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq5.400 s5.333 s-1.2%31.051 s30.174 s-2.8%60.13061.870+2.9%98.68 GiB99.48 GiB+0.8%3.97 MiB3.73 MiB-6.2% (1.07× smaller) 🟢30.2 ms31.0 ms+2.5%205.59B205.59B+0.0%
String.append4.334 s4.419 s+2.0%27.489 s26.612 s-3.2% 🟢11.90012.290+3.3% 🟢87.87 GiB88.58 GiB+0.8%3.97 MiB3.73 MiB-6.1% (1.07× smaller) 🟢34.5 ms22.7 ms-34.3% (1.52× faster) 🟢168.67B168.67B+0.0%
Nat.add_comm3.518 s3.499 s-0.5%18.993 s17.942 s-5.5% (1.06× faster) 🟢2.4202.560+5.8% (1.06× faster) 🟢58.70 GiB59.42 GiB+1.2%3.97 MiB3.72 MiB-6.2% (1.07× smaller) 🟢24.8 ms20.8 ms-16.1% (1.19× faster) 🟢130.84B130.84B+0.0%
Pipeline total (7 constants)
constanttotal-time (main)total-time (PR)Δ%pipeline-throughput (const/s) (main)pipeline-throughput (const/s) (PR)Δ%pipeline-peak-ram (main)pipeline-peak-ram (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append1m 1.8s59.792 s-3.2% 🟢44.93046.410+3.3% 🟢101.37 GiB101.07 GiB-0.3%
Char.ofOrdinal_le_of_le56.477 s56.124 s-0.6%48.92049.230+0.6%100.32 GiB100.32 GiB+0.0%
Array.extract_append53.024 s50.916 s-4.0% 🟢30.29031.540+4.1% 🟢94.59 GiB95.37 GiB+0.8%
Std.HashMap45.172 s44.717 s-1.0%45.21045.660+1.0%94.56 GiB95.34 GiB+0.8%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq45.377 s43.980 s-3.1% 🟢41.14042.450+3.2% 🟢98.68 GiB99.48 GiB+0.8%
String.append29.767 s28.737 s-3.5% 🟢10.99011.380+3.5% 🟢87.87 GiB88.58 GiB+0.8%
Nat.add_comm20.061 s18.912 s-5.7% (1.06× faster) 🟢2.2902.430+6.1% (1.06× faster) 🟢58.70 GiB59.42 GiB+1.2%

Workflow logs

Consume the Plonky3 0.6 batch-opening layout directly in Aiur instead of expanding every pruned Merkle frontier into one authentication path per FRI query. Sample all query indices from the unchanged transcript, sort and deduplicate them with an O(q log q) merge sort, authenticate each input and commit-phase commitment once, then retain the existing per-query reduced-opening and FRI arithmetic.
Bind every frontier to transcript-derived indices, consume boundary digests in Plonky3's level/parent/child order, reject trailing frontier elements and inconsistent duplicate leaves, and assert all native opening dimensions and sibling counts. Explicitly constrain the digest-bound protocol specialization to cap height 0, binary FRI, and a constant final polynomial. Move memo_u32_less_than into IxVM Core so both substitution and multiproof sorting share its constrained rows.
Strengthen the recursive negative test to mutate a structurally valid stage-1 commitment. Regenerate both checked-in Aiur Rust executors and retain interpreter/codegen query-count parity.
On Vector.extract_append q50, recursive-verifier FFT cost falls from 204.073B to 201.166B. CPU outer proving improves from 50.09s to 45.03s and the full CPU pipeline from 90.64s to 82.90s. GPU outer proving improves from 15.85s to 13.72s and the full GPU pipeline from 28.86s to 26.69s. The outer proof grows from 3.92 MB to 4.17 MB.
Validated with the MultiStark primitive suite, recursive honest/tamper/parity tests, codegen --check, release workspace clippy, release CUDA clippy, rustfmt, and diff checks.
@arthurpaulino

Copy link
Copy Markdown
MemberAuthor

!benchmark fresh

@argument-ci-bot

argument-ci-botBot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

!benchmark — main vs 840932a

backends: aiur=prove · envs: InitStd · baseline: fresh (benchmark products rebuilt, base-SHA run, bencher bypassed)

Warning

  • CPU model mismatch for PR benchmark binaries in this job: built on AMD EPYC 9R45; measured on Intel(R) Xeon(R) 6975P-C. Native Rust code uses -Ctarget-cpu=native.

aiur · InitStd · prove — main from: base run @ e1ca8e2 (fresh — bencher bypassed)

7 constants · 0 with regressions · 0 with improvements (|Δ| > 3.0% on any metric).

IxVM on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append8.874 s💥 CRASHn/a44.291 s💥 CRASHn/a62.650💥 CRASHn/a70.80 GiB💥 CRASHn/a11.33 MiB💥 CRASHn/a77.4 ms💥 CRASHn/a134.35B💥 CRASHn/a
Char.ofOrdinal_le_of_le6.603 s💥 CRASHn/a38.242 s💥 CRASHn/a72.250💥 CRASHn/a63.81 GiB💥 CRASHn/a11.33 MiB💥 CRASHn/a76.9 ms💥 CRASHn/a102.60B💥 CRASHn/a
Array.extract_append6.311 s💥 CRASHn/a33.420 s💥 CRASHn/a48.060💥 CRASHn/a51.95 GiB💥 CRASHn/a11.24 MiB💥 CRASHn/a86.5 ms💥 CRASHn/a97.08B💥 CRASHn/a
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq3.774 s💥 CRASHn/a20.628 s💥 CRASHn/a90.510💥 CRASHn/a33.97 GiB💥 CRASHn/a11.26 MiB💥 CRASHn/a82.6 ms💥 CRASHn/a55.68B💥 CRASHn/a
Std.HashMap4.060 s💥 CRASHn/a22.501 s💥 CRASHn/a90.750💥 CRASHn/a36.36 GiB💥 CRASHn/a11.26 MiB💥 CRASHn/a84.1 ms💥 CRASHn/a61.88B💥 CRASHn/a
String.append708.3 ms💥 CRASHn/a2.863 s💥 CRASHn/a114.220💥 CRASHn/a4.99 GiB💥 CRASHn/a9.94 MiB💥 CRASHn/a70.3 ms💥 CRASHn/a3.37B💥 CRASHn/a
Nat.add_comm496.0 ms💥 CRASHn/a1.330 s💥 CRASHn/a34.580💥 CRASHn/a4.33 GiB💥 CRASHn/a9.09 MiB💥 CRASHn/a58.9 ms💥 CRASHn/a308.40M💥 CRASHn/a
FRI verifier on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append6.880 s💥 CRASHn/a54.265 s💥 CRASHn/a51.140💥 CRASHn/a101.05 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a49.5 ms💥 CRASHn/a210.23B💥 CRASHn/a
Char.ofOrdinal_le_of_le6.813 s💥 CRASHn/a53.449 s💥 CRASHn/a51.690💥 CRASHn/a99.81 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a45.1 ms💥 CRASHn/a207.18B💥 CRASHn/a
Array.extract_append6.371 s💥 CRASHn/a50.695 s💥 CRASHn/a31.680💥 CRASHn/a95.19 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a58.1 ms💥 CRASHn/a200.65B💥 CRASHn/a
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq6.748 s💥 CRASHn/a53.378 s💥 CRASHn/a34.980💥 CRASHn/a98.71 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a46.1 ms💥 CRASHn/a205.59B💥 CRASHn/a
Std.HashMap6.455 s💥 CRASHn/a50.868 s💥 CRASHn/a40.140💥 CRASHn/a95.03 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a45.7 ms💥 CRASHn/a203.35B💥 CRASHn/a
String.append5.307 s💥 CRASHn/a47.595 s💥 CRASHn/a6.870💥 CRASHn/a87.89 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a32.4 ms💥 CRASHn/a168.67B💥 CRASHn/a
Nat.add_comm4.526 s💥 CRASHn/a31.415 s💥 CRASHn/a1.460💥 CRASHn/a58.66 GiB💥 CRASHn/a3.97 MiB💥 CRASHn/a28.0 ms💥 CRASHn/a130.84B💥 CRASHn/a
Pipeline total (7 constants)
constanttotal-time (main)total-time (PR)Δ%pipeline-throughput (const/s) (main)pipeline-throughput (const/s) (PR)Δ%pipeline-peak-ram (main)pipeline-peak-ram (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append1m 38.6s💥 CRASHn/a28.160💥 CRASHn/a101.05 GiB💥 CRASHn/a
Char.ofOrdinal_le_of_le1m 31.7s💥 CRASHn/a30.130💥 CRASHn/a99.81 GiB💥 CRASHn/a
Array.extract_append1m 24.1s💥 CRASHn/a19.090💥 CRASHn/a95.19 GiB💥 CRASHn/a
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq1m 14.0s💥 CRASHn/a25.230💥 CRASHn/a98.71 GiB💥 CRASHn/a
Std.HashMap1m 13.4s💥 CRASHn/a27.830💥 CRASHn/a95.03 GiB💥 CRASHn/a
String.append50.458 s💥 CRASHn/a6.480💥 CRASHn/a87.89 GiB💥 CRASHn/a
Nat.add_comm32.745 s💥 CRASHn/a1.400💥 CRASHn/a58.66 GiB💥 CRASHn/a

Workflow logs

samuelburnham added a commit that referenced this pull request Sep 1, 2026
The Warp x64 runner pool mixes Intel Granite Rapids and AMD Zen 5, and a
build job may land on one vendor while the job that runs its binaries
lands on the other. Neither vendor's feature set contains the other's, so
`-Ctarget-cpu=native` does not produce a portable binary: Zen 5 enables
SSE4A, and LLVM emits it. Disassembling the workspace built for znver5
finds 31 SSE4A instructions, all INSERTQ, in `ix-ffi` and in
`aiur_ixvm_witness::add_entries_parallel`. Granite Rapids has no SSE4A,
so the first one executed raises #UD, killing the process with SIGILL
during witness generation. That is what turned every row of #605's
benchmark into a crash.
Pin the measured intersection of the two CPUs instead. x86-64-v4 covers
every AVX-512 subset Plonky3 uses; +avx512vbmi2 preserves its VPSHRDQ
interleave and +gfni preserves LLVM's byte-shift lowering. A workspace
built with these flags contains no instruction absent from either vendor
and has an instruction vocabulary identical to a graniterapids build.
blake3 dispatches on CPUID at runtime and is unaffected either way.
`.cargo/config.toml` keeps `-Ctarget-cpu=native`: a developer builds and
runs on one machine, and x86-64-v4 would exclude every host without
AVX-512. Only CI has the split, so only CI pins the ISA. The new guard
fails the job when a runner lacks a required feature, so the assumption
is enforced rather than assumed, and the shared `warp-x64` cargo cache
key becomes sound now that codegen no longer varies by host.
Pinning also removes a benchmarking hazard that never crashed: LLVM sets
prefer-256-bit for Granite Rapids but not for Zen 5, so the same source
vectorized 3.2x more widely depending on the build host, and main-vs-PR
timings were not comparable across a vendor split.
Drop the CPU-mismatch warning the benchmark comment used to carry. It
detected a real problem, but the flags above prevent that problem, and
computing it in one job to render it in another cost a Markdown file
threaded through cache entries, artifacts, and a `--warning-file` flag on
`ix bench compare`. Warnings belong to the run that finds them.
@arthurpaulino

Copy link
Copy Markdown
MemberAuthor

!benchmark fresh

@argument-ci-bot

argument-ci-botBot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

⚠️ Some benchmark jobs failed — results may be partial.

!benchmark — main vs d7fc2f1

backends: aiur=prove · envs: InitStd · baseline: fresh (benchmark products rebuilt, base-SHA run, bencher bypassed)

No result tables were produced — see the run logs.

Workflow logs

The Warp x64 runner pool mixes Intel Granite Rapids and AMD Zen 5, and a
build job may land on one vendor while the job that runs its binaries
lands on the other. Neither vendor's feature set contains the other's, so
`-Ctarget-cpu=native` does not produce a portable binary: Zen 5 enables
SSE4A, and LLVM emits it. Disassembling the workspace built for znver5
finds 31 SSE4A instructions, all INSERTQ, in `ix-ffi` and in
`aiur_ixvm_witness::add_entries_parallel`. Granite Rapids has no SSE4A,
so the first one executed raises #UD, killing the process with SIGILL
during witness generation. That is what turned every row of #605's
benchmark into a crash.
Pin the measured intersection of the two CPUs instead. x86-64-v4 covers
every AVX-512 subset Plonky3 uses; +avx512vbmi2 preserves its VPSHRDQ
interleave and +gfni preserves LLVM's byte-shift lowering. A workspace
built with these flags contains no instruction absent from either vendor
and has an instruction vocabulary identical to a graniterapids build.
blake3 dispatches on CPUID at runtime and is unaffected either way.
`.cargo/config.toml` keeps `-Ctarget-cpu=native`: a developer builds and
runs on one machine, and x86-64-v4 would exclude every host without
AVX-512. Only CI has the split, so only CI pins the ISA. The new guard
fails the job when a runner lacks a required feature, so the assumption
is enforced rather than assumed, and the shared `warp-x64` cargo cache
key becomes sound now that codegen no longer varies by host. RUSTFLAGS
is hashed into that key, so the flag change rotates it on its own.
Pinning also removes a benchmarking hazard that never crashed: LLVM sets
prefer-256-bit for Granite Rapids but not for Zen 5, so the same source
vectorized 3.2x more widely depending on the build host, and main-vs-PR
timings were not comparable across a vendor split.
@samuelburnham

Copy link
Copy Markdown
Member

!benchmark fresh

@argument-ci-bot

argument-ci-botBot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

!benchmark — main vs feb014d

backends: aiur=prove · envs: InitStd · baseline: fresh (benchmark products rebuilt, base-SHA run, bencher bypassed)

aiur · InitStd · prove — main from: base run @ e1ca8e2 (fresh — bencher bypassed)

7 constants · 3 with regressions · 7 with improvements (|Δ| > 3.0% on any metric).

IxVM on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append9.269 s9.073 s-2.1%31.250 s30.269 s-3.1% 🟢88.80091.680+3.2% 🟢70.72 GiB72.28 GiB+2.2%11.33 MiB11.06 MiB-2.4%68.5 ms62.7 ms-8.5% (1.09× faster) 🟢134.35B134.35B+0.0%
Char.ofOrdinal_le_of_le6.964 s6.879 s-1.2%26.884 s25.793 s-4.1% 🟢102.780107.120+4.2% 🟢63.83 GiB65.31 GiB+2.3%11.33 MiB11.07 MiB-2.3%77.9 ms59.3 ms-23.8% (1.31× faster) 🟢102.60B102.60B+0.0%
Array.extract_append6.676 s6.751 s+1.1%23.925 s23.179 s-3.1% 🟢67.13069.290+3.2% 🟢51.97 GiB52.72 GiB+1.4%11.24 MiB10.99 MiB-2.2%71.1 ms57.2 ms-19.5% (1.24× faster) 🟢97.08B97.08B+0.0%
Std.HashMap4.171 s4.079 s-2.2%16.222 s15.628 s-3.7% 🟢125.880130.670+3.8% 🟢36.29 GiB37.10 GiB+2.2%11.26 MiB11.01 MiB-2.2%74.8 ms66.6 ms-11.0% (1.12× faster) 🟢61.88B61.88B+0.0%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq3.697 s3.648 s-1.3%14.858 s14.287 s-3.8% 🟢125.650130.680+4.0% 🟢34.02 GiB34.75 GiB+2.1%11.26 MiB11.00 MiB-2.2%75.5 ms57.7 ms-23.6% (1.31× faster) 🟢55.68B55.68B+0.0%
String.append435.5 ms434.0 ms-0.4%2.307 s2.129 s-7.7% (1.08× faster) 🟢141.750153.600+8.4% (1.08× faster) 🟢5.74 GiB5.52 GiB-3.9% 🟢9.94 MiB9.74 MiB-2.1%64.4 ms50.2 ms-22.1% (1.28× faster) 🟢3.37B3.37B+0.0%
Nat.add_comm267.1 ms264.3 ms-1.1%1.068 s984.9 ms-7.8% (1.08× faster) 🟢43.06046.710+8.5% (1.08× faster) 🟢4.51 GiB3.99 GiB-11.5% (1.13× smaller) 🟢9.09 MiB8.90 MiB-2.1%53.3 ms47.9 ms-10.2% (1.11× faster) 🟢308.40M308.40M+0.0%
FRI verifier on FRI (7 constants)
constantexecute-time (main)execute-time (PR)Δ%prove-time (main)prove-time (PR)Δ%throughput (const/s) (main)throughput (const/s) (PR)Δ%peak-ram (main)peak-ram (PR)Δ%proof-size (main)proof-size (PR)Δ%verify-time (main)verify-time (PR)Δ%fft-cost (main)fft-cost (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append5.680 s4.975 s-12.4% (1.14× faster) 🟢33.109 s30.692 s-7.3% (1.08× faster) 🟢83.82090.410+7.9% (1.08× faster) 🟢101.44 GiB99.74 GiB-1.7%3.97 MiB3.98 MiB+0.2%26.3 ms22.3 ms-15.4% (1.18× faster) 🟢210.23B203.74B-3.1% 🟢
Char.ofOrdinal_le_of_le5.599 s5.014 s-10.5% (1.12× faster) 🟢32.280 s31.379 s-2.8%85.59088.050+2.9%99.81 GiB101.09 GiB+1.3%3.97 MiB3.98 MiB+0.2%27.6 ms28.6 ms+3.6% ⚠️207.18B208.08B+0.4%
Array.extract_append5.315 s4.858 s-8.6% (1.09× faster) 🟢30.962 s29.672 s-4.2% 🟢51.87054.120+4.3% 🟢95.00 GiB95.36 GiB+0.4%3.97 MiB3.99 MiB+0.4%26.0 ms39.4 ms+51.3% (1.51× slower) ⚠️200.65B200.40B-0.1%
Std.HashMap5.445 s4.864 s-10.7% (1.12× faster) 🟢31.021 s29.949 s-3.5% 🟢65.83068.180+3.6% 🟢94.50 GiB95.90 GiB+1.5%3.97 MiB3.98 MiB+0.2%25.4 ms25.6 ms+0.8%203.35B204.11B+0.4%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq5.496 s4.898 s-10.9% (1.12× faster) 🟢32.335 s29.501 s-8.8% (1.10× faster) 🟢57.74063.290+9.6% (1.10× faster) 🟢98.72 GiB95.35 GiB-3.4% 🟢3.97 MiB3.98 MiB+0.3%25.9 ms31.9 ms+23.0% (1.23× slower) ⚠️205.59B199.65B-2.9%
String.append4.487 s3.955 s-11.8% (1.13× faster) 🟢28.710 s27.141 s-5.5% (1.06× faster) 🟢11.39012.050+5.8% (1.06× faster) 🟢87.89 GiB88.65 GiB+0.9%3.97 MiB3.97 MiB+0.1%37.6 ms27.4 ms-27.2% (1.37× faster) 🟢168.67B164.80B-2.3%
Nat.add_comm3.619 s3.077 s-15.0% (1.18× faster) 🟢19.611 s17.954 s-8.5% (1.09× faster) 🟢2.3502.560+8.9% (1.09× faster) 🟢58.66 GiB58.03 GiB-1.1%3.97 MiB3.98 MiB+0.2%27.6 ms22.0 ms-20.4% (1.26× faster) 🟢130.84B124.20B-5.1% (1.05× fewer) 🟢
Pipeline total (7 constants)
constanttotal-time (main)total-time (PR)Δ%pipeline-throughput (const/s) (main)pipeline-throughput (const/s) (PR)Δ%pipeline-peak-ram (main)pipeline-peak-ram (PR)Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append1m 4.4s1m 1.0s-5.3% (1.06× faster) 🟢43.12045.520+5.6% (1.06× faster) 🟢101.44 GiB99.74 GiB-1.7%
Char.ofOrdinal_le_of_le59.164 s57.172 s-3.4% 🟢46.70048.330+3.5% 🟢99.81 GiB101.09 GiB+1.3%
Array.extract_append54.887 s52.851 s-3.7% 🟢29.26030.390+3.9% 🟢95.00 GiB95.36 GiB+0.4%
Std.HashMap47.243 s45.577 s-3.5% 🟢43.22044.800+3.7% 🟢94.50 GiB95.90 GiB+1.5%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq47.193 s43.788 s-7.2% (1.08× faster) 🟢39.56042.640+7.8% (1.08× faster) 🟢98.72 GiB95.35 GiB-3.4% 🟢
String.append31.017 s29.270 s-5.6% (1.06× faster) 🟢10.54011.170+6.0% (1.06× faster) 🟢87.89 GiB88.65 GiB+0.9%
Nat.add_comm20.680 s18.939 s-8.4% (1.09× faster) 🟢2.2202.430+9.5% (1.09× faster) 🟢58.66 GiB58.03 GiB-1.1%

Workflow logs

#606)
* Authenticate all frontier group members in the recursive verifier
The in-circuit pruned-multiproof walk (mmcs_verify_multi / frontier_level)
collapses queries that share a parent to a single lead node and hashes
only the lead's rows via inject_maybe(ar, ...). Non-lead members' opened
rows for the shorter (injected) matrices were still consumed in their own
per-query FRI arithmetic (batch_views_at) but never authenticated against
any commitment — the leaf hash covers only the tallest matrices, and
shorter ones are bound solely through injection. A prover could therefore
forge a non-lead member's shorter-matrix opening. Plonky3's reference
verify_batch_pruned guards exactly this with InconsistentGroupOpening
(and InconsistentDuplicateOpenings for equal-index queries); the port had
neither. The prior per-query walk did not have the gap, so it was
introduced with the direct multiproof consumption.
- frontier_level: on a group merge, assert the lead and member agree on
every not-yet-injected matrix (height <= next_lh) via select_rows_le +
pointer equality. Transitive across pairwise merges, so the whole group
is pinned; matches InconsistentGroupOpening.
- frontier_merge: duplicate transcript indices must open the SAME full
rows, not merely the same tallest-matrix leaf digest; matches
InconsistentDuplicateOpenings.
Pointer equality is admissible inside assert_eq! (equal pointers imply
equal content; a spurious mismatch costs only completeness — see
IxVM.Core). select_rows_le selects rows of matrices at height <= target,
mirroring select_rows.
Validated: the group-merge branch is genuinely reached by the factorial
recursion proof (an always-false variant of the new assert fails the
honest test), the honest proof still verifies with the real assert
(completeness preserved), the existing tamper tests still reject, and
the full lake test suite is green (2717 checks). aiur_multi_stark.rs
regenerated; kernel executor unchanged.
* Drop the multi_stark::advice dependency
Companion to multi-stark removing its unused per-query advice module.
ix consumed native pruned multiproofs directly and referenced only
advice::AdviceError, whose two arms (verification failed, serialization
failed) were immediately string-formatted by the FFI. Replace it with a
plain Result<Vec<u8>, String>: AiurSystem::proof_to_advice_bytes maps
both failures to a message, and the FFI passes the string straight to
LeanExcept::error_string. No behavior change; the Lean binding
(Except String ByteArray) is unaffected.
Bump the multi-stark pin to the advice-removed revision. Requires that
multi-stark's ap/bump-p3-drop-advice be pushed first, exactly as with
every other pin in this series.
* Bump multi-stark audit revision
---------
Co-authored-by: Arthur Paulino <arthurleonardo.ap@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@arthurpaulino@samuelburnham