IxVM kernel: unblock UTF-8 decode/encode proof + Nat-layer FFT cuts - #450

Merged
arthurpaulino merged 8 commits into
mainfrom
ap/utf8-tier-1d
Jun 25, 2026
Merged

IxVM kernel: unblock UTF-8 decode/encode proof + Nat-layer FFT cuts#450
arthurpaulino merged 8 commits into
mainfrom
ap/utf8-tier-1d

Conversation

@arthurpaulino

Copy link
Copy Markdown
Member

Headline

Aiur kernel previously OOM'd on
_private.Init.Data.String.Decode.0.ByteArray.utf8DecodeChar?.assemble₄_eq_some_of_toBitVec._proof_1_8.
This constant is a prerequisite of ByteArray.utf8DecodeChar?_utf8EncodeChar_append
(part of the UTF-8 round-trip lemma chain).

Now it typechecks at 37.31B FFT (down from the OOM baseline). The
two-direction lemma …_utf8EncodeChar_append itself remains
out-of-reach (still OOMs further along the same dependency tree under
the user's memory cap), but the prerequisite that was the immediate
blocker is unstuck.

The same kernel changes also deliver large FFT cuts on previously-
expensive targets — Tier 1d's structural short-circuit replaces full
delta-whnf cascades wherever it fires:

ConstantBeforeAfterΔ
Vector.append4.02B3.16B−21.4%
Array.append_assoc3.94B3.08B−21.8%

Commits

1. a6cd34a — Tier 1d def-eq short-circuit (the unblock)

Aiur's k_is_def_eq_core jumped from Tier 1.5 straight to full delta
WHNF (Tier 2). The cascading Nat.rec / Nat.succ iota expansions then
drove whnf_const_head past 1M unique entries before either side
reached a comparable canonical form — OOM. Rust's def-eq settles the
same pair via no-delta whnf + quick structural recursion before any
of that fires.

Three minimum-necessary pieces ported. No KStore. No FVar. No Subst /
KernelTypes change.

  • whnf_nd family (Whnf.lean, mirror Rust
    whnf_no_delta_for_def_eq). Same dispatch tree as whnf, but
    whnf_nd_const_head's Defn arm falls through to a stuck
    apply_spine instead of delta-unfolding. Iota / proj / quot /
    primitives still fire.

  • k_infer_only family (Infer.lean, mirror Rust
    with_infer_only). App drops k_check(a, dom); Lam drops
    k_ensure_sort(ty); Let drops val/ty validation. Used at
    try_proof_irrel, is_prop_type, try_unit_like — the def-eq
    tactics that only need the synthesized type.

  • k_is_def_eq_struct_safe + Tier 1d wiring (DefEq.lean,
    mirror Rust quick_def_eq + post-try_def_eq_app). Sort-Sort via
    level_equal; Lam-Lam / All-All via recursive k_is_def_eq on the
    type and on the body under Cons(ty_a, types) (types-cons, NOT
    FVar opening). Inserted between Tier 1c (string lit) and Tier 2
    (full whnf):

    aw_nd = whnf_nd(a); bw_nd = whnf_nd(b)
    ptr_eq(aw_nd, bw_nd) → 1
    k_is_def_eq_struct_safe(aw_nd, bw_nd) → 1
    try_lazy_delta_app(aw_nd, bw_nd) → 1 (rerun: spine args may have
    reduced past what Tier 1.5's pre-whnf attempt could see)
    

Each piece independently validated necessary. The previously-tried
KStore explicit caches, FVar variant + opens, FVar-based binder
opening — all confirmed NOT necessary for this unblock and left out.

3 files changed, +296/−6 lines.

2. 36d6c3a — drop g_or from u64_sub_with_borrow

u64_sub_with_borrow combined two per-byte borrow bits with g_or.
The two bits are mutually exclusive: u_t = 1 ⇒ intermediate t_i ≥ 1 ⇒ subtracting br_in ∈ {0,1} cannot underflow ⇒ u_r = 0. Field
+ substitutes for g_or directly. Per Aiur cost model g_or adds
+1 aux + 1 lookup per call (≈ 5 width); field + is free. 7 g_ors ×
2.23M rows.

UTF-8 _proof_1_8: 39.12B → 38.14B (−2.6%).

3. 9e787da — drop g_or from klimbs_add_carry / klimbs_sub_borrow

Same mutually-exclusive-carry pattern. Two limb-level borrows /
carries from sequential u64 ops cannot both be 1.

UTF-8 _proof_1_8: 38.14B → 38.07B (−0.18%).

4. 4e379e7 — hot/cold split try_nat_dispatch, extract binop arm

try_nat_dispatch's width 90 was floored by the binop arm (2× whnf

  • 2× try_extract_nat + try_nat_binop_addr + apply_spine), charged on
    every Nat.succ / Nat.pred row. Factor binop dispatch into
    try_nat_binop_dispatch. Main narrows to the max of succ / pred
    arms.

UTF-8 _proof_1_8: 38.07B → 37.80B (−0.7%).

5. 80ce3d2 — hot/cold split expr_lbr, extract Let arm

expr_lbr's width was floored by the Let arm (3 recursive expr_lbr
calls + 2 lbr_max + 1 lbr_dec), charged on every row even though Let
is rare. Factor into expr_lbr_let.

Nat.add_comm: 55.63M → 55.50M (−0.2%). UTF-8 _proof_1_8: 37.80B
→ 37.62B (−0.5%)
.

6. 1f8effd — hot/cold split try_extract_nat, extract App arm

try_extract_nat's width was floored by the App arm (list_lookup +
address_eq + recursive try_extract_nat + klimbs_succ). Factor into
try_extract_nat_app. Main narrows to leaf-arm width.

UTF-8 _proof_1_8: 37.62B → 37.31B (−0.8%).

7. 039e9cf — re-pin IxVM FFT costs

41 pins in Tests/Ix/IxVM.lean::kernelCheckEntries updated. Every
constant got cheaper; none regressed. Largest reductions:

  • Vector.append: 4.02B → 3.16B (−21.4%)
  • Array.append_assoc: 3.94B → 3.08B (−21.8%)

lake test -- --ignored ixvm passes with 0 FFT mismatches.

Cumulative on UTF-8 _proof_1_8

CommitFFT
baseline (main)OOM
a6cd34a Tier 1d39.12B
36d6c3a g_or → + in u64_sub_with_borrow38.14B (−2.6%)
9e787da g_or → + in klimbs_add_carry / klimbs_sub_borrow38.07B (−0.18%)
4e379e7 hot/cold try_nat_dispatch37.80B (−0.7%)
80ce3d2 hot/cold expr_lbr37.62B (−0.5%)
1f8effd hot/cold try_extract_nat37.31B (−0.8%)

Post-unlock optimization: −4.6% (39.12B → 37.31B).

Cost on small targets

  • Main baseline Nat.add_comm: 56.08M FFT.
  • This branch Nat.add_comm: 55.50M FFT.

Tier 1d itself adds no overhead on the common case (whnf_nd +
struct_safe + try_lazy_delta_app are themselves Aiur-memoized); the
follow-up optimizations are net wins.

Test plan

  • lake exe check Nat.add_comm passes (55.50M FFT).
  • lake exe check Vector.extract_append passes.
  • lake exe check "_private.…assemble₄_eq_some_of_toBitVec._proof_1_8"
    passes (37.31B FFT, previously OOM).
  • lake test -- --ignored ixvm — all 41 FFT pins updated, suite
    passes with 0 mismatches.

Comment threadIx/IxVM/Kernel/Infer.lean
…ost-spine-congruence)
UTF-8 `_private.Init.Data.String.Decode.0.ByteArray.utf8DecodeChar?
.assemble₄_eq_some_of_toBitVec._proof_1_8` OOMs on the previous
pipeline because `k_is_def_eq_core` jumps from Tier 1.5 straight to
full delta WHNF (Tier 2); the cascading Nat.rec / Nat.succ iota
expansions then drive `whnf_const_head` past 1M unique entries before
either side reaches a comparable canonical form. Rust's def-eq settles
the same pair via the no-delta whnf + quick structural recursion
before any of that fires.
This patch ports the three pieces of that short-circuit and nothing
else — no FVar variant, no KStore, no Subst changes, no signature
sweep.
* `whnf_nd` family in `Whnf.lean` (mirror Rust `whnf_no_delta_for_def_eq`).
Same dispatch tree as `whnf` (beta / let zeta / iota / proj / quot /
primitives all fire), except `whnf_nd_const_head`'s Defn arm falls
through to a stuck `apply_spine` instead of delta-unfolding.
* `k_infer_only` family in `Infer.lean` (mirror Rust `with_infer_only`).
App drops `k_check(a, dom)`; Lam drops `k_ensure_sort(ty)`; Let drops
the val/ty validations. Distinct Aiur memo from `k_infer`, parity
with Rust's separate `infer_cache` / `infer_only_cache`.
* `k_is_def_eq_struct_safe` in `DefEq.lean` (mirror Rust
`quick_def_eq`). Sort-Sort via `level_equal`; Lam-Lam / All-All
recurse on type and on body under `Cons(ty_a, types)`. Returns 1
only when DEFINITELY def-eq; 0 means fall through. Sound on
partially-whnf'd (no-delta) inputs because the handled shapes
don't depend on further reductions.
* `k_is_def_eq_core` Tier 1d wiring inserted between Tier 1c (string
lit) and Tier 2 (full whnf):
aw_nd = whnf_nd(a); bw_nd = whnf_nd(b)
ptr_eq(aw_nd, bw_nd) → 1
k_is_def_eq_struct_safe(aw_nd, bw_nd) → 1 if 1
try_lazy_delta_app(aw_nd, bw_nd) → 1 if 1 (rerun post-whnf_nd:
spine args may have reduced past what Tier 1.5's pre-whnf attempt
could see, exposing Const-Const congruence that was hidden)
* `try_proof_irrel`, `is_prop_type`, `try_unit_like` switch from
`k_infer` to `k_infer_only` — these helpers only need the synthesized
type, not the full re-validation work that `k_infer` does for each
recursive App/Let/Lam.
Each piece individually validated necessary (removing it puts UTF-8
back into the OOM regime). FVar variant + opens, KStore explicit
caches, infer_only's FVar-based binder opening — all confirmed NOT
necessary for the UTF-8 unblock and left out (see PLAN.md for future
experiments).
Measured (FFT cost):
Nat.add_comm: 56.08M → 55.63M (~stable; new code paths add no
overhead on the common case because Tier 1d's whnf_nd + struct_safe
are themselves Aiur-memoized).
_private.Init.Data.String.Decode.0.ByteArray.utf8DecodeChar?
.assemble₄_eq_some_of_toBitVec._proof_1_8: OOM → 39.12B FFT, passes.
3 files, +296/-6 lines.
`u64_sub_with_borrow` combines two per-byte borrow bits with `g_or`. The
two bits are MUTUALLY EXCLUSIVE: `u_t = borrow(a_i - b_i)` and
`u_r = borrow((a_i + 256 - b_i) - br_in)`. If `u_t = 1` the intermediate
`t_i ≥ 1`, so subtracting `br_in ∈ {0,1}` cannot underflow ⇒ `u_r = 0`.
Field `+` substitutes for `g_or` directly (per the same pattern as
`u64_add` in `ByteStream.lean`).
Per Aiur cost model, `g_or` adds +1 aux + 1 lookup per call; field `+`
is free. 8 g_or call sites in `u64_sub_with_borrow` each charged on every
one of the function's 2.23M rows.
Measured (FFT cost) on UTF-8 `_proof_1_8`:
39.12B → 38.14B (-2.6%)
Nat.add_comm unchanged (55.63M).
See [[reference_aiur_carry_add]].
Same mutually-exclusive-carry pattern as `u64_sub_with_borrow`:
* `klimbs_add_carry`: u64_add of (la, lb) yields carry1; u64_add of
(sum1, carry_in) yields carry2. carry1=1 ⇒ sum1 ≤ 2^64-2 ⇒
carry2=0.
* `klimbs_sub_borrow`: symmetric for borrows.
Replace `g_or(c1, c2)` with `c1 + c2` (field +). Both helpers run on
hot Nat-primitive paths.
Measured on UTF-8 `_proof_1_8`:
38.14B → 38.07B (-0.18%)
Nat.add_comm unchanged.
See [[reference_aiur_carry_add]].
`try_nat_dispatch` ran 1.12M rows in UTF-8 `_proof_1_8` at width 90,
charging 5.16% of total FFT. Width was floored by its widest match arm
(the binop branch with 2× whnf + 2× try_extract_nat + try_nat_binop_addr
+ apply_spine), even on Nat.succ / Nat.pred rows that never touched it.
Factor binop dispatch into its own `try_nat_binop_dispatch` fn. Main
dispatcher narrows to the max of succ / pred arms (single whnf +
try_extract_nat + klimbs_succ/dec + apply_spine). The cold fn's width
only charges the rows that actually dispatch a binop.
Measured on UTF-8 `_proof_1_8`:
38.07B → 37.80B (-0.7%)
Nat.add_comm unchanged.
`expr_lbr` ran 1.47M rows in UTF-8 `_proof_1_8` at width 39, charging
3.01% of total FFT. The Let arm (3 recursive expr_lbr calls + 2 lbr_max
+ 1 lbr_dec) is the widest match arm, charged on every row of expr_lbr
even though Let is rare in most expressions encountered.
Factor the Let arm into `expr_lbr_let(ty, val, body)`. Main expr_lbr
narrows to max of the 2-recursion arms (App / Lam / Forall). Cold fn
only charges Let-arm rows.
Measured:
Nat.add_comm: 55.63M → 55.50M (-0.2%)
UTF-8 `_proof_1_8`: 37.80B → 37.62B (-0.5%)
`try_extract_nat` ran 1.12M rows at width 45, charging 2.68% of UTF-8
`_proof_1_8` total FFT. The App arm (list_lookup + address_eq +
recursive try_extract_nat + klimbs_succ) is the widest match arm; the
Lit / Const / default arms are leaf compares.
Factor App into `try_extract_nat_app(f, a, addrs)`. Main extractor
narrows to leaf-arm width. Cold fn only charges App-arm rows.
Measured on UTF-8 `_proof_1_8`:
37.62B → 37.31B (-0.8%)
Nat.add_comm unchanged.
Updates 41 pinned FFT costs in `Tests/Ix/IxVM.lean::kernelCheckEntries`
to match the new kernel's output. All pins moved DOWN — every constant
got cheaper, none regressed.
Largest reductions (% change):
Vector.append: 4_023_268_168 → 3_160_970_390 (-21.4%)
Array.append_assoc: 3_938_574_533 → 3_079_334_815 (-21.8%)
String.Internal.append: 793_580_333 → 775_968_134 ( -2.2%)
bv_to_nat_lit: 635_780_327 → 619_870_154 ( -2.5%)
nat_gcd_lit: 665_518_356 → 649_859_784 ( -2.4%)
Nat.sub_le_of_le_add: 567_575_653 → 557_867_526 ( -1.7%)
IxVMPrim.nat_mod_lit: 414_695_549 → 407_517_834 ( -1.7%)
IxVMPrim.nat_div_lit: 405_607_545 → 398_641_590 ( -1.7%)
IxVMPrim.nat_shr_lit: 411_128_901 → 404_158_486 ( -1.7%)
Nat.decLe: 209_641_496 → 206_196_563 ( -1.6%)
Nat.add_comm: 56_084_908 → 55_504_714 ( -1.0%)
`lake test -- --ignored ixvm` passes with 0 FFT mismatches.
The `k_infer_only` section header skipped over the fact that the
function is only sound on well-typed inputs (since it drops
`k_check(a, dom)` on `App`, `k_ensure_sort(ty)` on `Lam`, val/ty
checks on `Let`). Spell out:
* The invariant — only call on terms produced by `whnf` of a
well-typed term, never on arbitrary inputs.
* Why the current sites (`try_proof_irrel`, `is_prop_type`,
`try_unit_like`) respect it.
* The planned non-deterministic-hint dispatch (`Hint::{None,
KInfer, KInferOnly}`) that lets us share `k_infer`'s memo where a
hit already exists instead of paying the parallel `infer_only`
memo cost.
@arthurpaulino
arthurpaulino merged commit ad7e383 into mainJun 25, 2026
14 checks passed
@arthurpaulino
arthurpaulino deleted the ap/utf8-tier-1d branch June 25, 2026 13:49
johnchandlerburnham pushed a commit that referenced this pull request Jul 21, 2026
…505)
* IxVM: drop dead KValNode/KVal/KValEnv
From ap/kernel 828fb85 (Arthur Paulino): the NbE value domain is defined
but referenced nowhere — the live kernel runs on de-Bruijn KExpr;
vestigial from an abandoned NbE direction. (The closed-term context
normalization from that commit was measured separately and not taken:
the per-call expr_lbr probe cost +2.9% FFT on recursor loops for a 0.5%
record reduction.)
* IxVM: memoized prim_family dispatch + width-safe offset-stuck placement
Three coordinated changes to Const-head whnf dispatch (cherry-pick of
130f30b, adapted to post-#450/#457 main):
1. prim_family(addr) classifies a head address into the one reducer
family that could fire on it (nat/str/bitvec/native/decidable; the
sets are disjoint). Keyed on the ADDRESS ALONE it memoizes to one row
per distinct constant address per run, and whnf_const_head calls at
most one family reducer. The previous gauntlet ran every reducer in
sequence for guaranteed misses.
2. The symbolic-Nat offset-stuck check moves from a delta-arm probe into
try_nat_dispatch's miss path as a cold function
(try_nat_offset_dispatch, verdict 2 = "already stuck, do not
re-whnf"), with the offset construction shared via
mk_nat_offset_stuck (also used by the linear-rec collapse).
3. nat_lit_to_ctor_or_self exposes ONE constructor layer
(n -> succ(Lit(n-1))) instead of materializing the full succ chain.
Adaptations vs the original patch:
- whnf_nd_const_head (no-delta WHNF, added on main after the patch)
converted to the same family dispatch.
- Kept main's cold-extracted try_nat_binop_dispatch and routed the
symbolic-base case to try_nat_offset_dispatch from its miss arm.
Measured (lake exe ix check Nat.add_comm): total width 34820 -> 34806,
FFT cost 49571210 -> 48860647 (-1.43%). All 53 ixvm-suite FFT pins
decreased (-0.19%..-1.43%); parity and claim smokes pass. Pins updated;
crates/ixvm-codegen/src/aiur_ixvm.rs regenerated via `lake exe ix
codegen`.
* IxVM: port jcb/fixes H-14 — ptr_val skip map + lockstep addr cursor
Two quadratic/constant-factor fixes to check_all_skipping, ported from
jcb/fixes (3763356, John C. Burnham); cherry-pick of 2d83be1 adapted to
post-#457 main:
- The assumption-leaf skip set keys on ptr_val instead of the first 4
address bytes: one tree lookup per constant, no per-lookup address
load, no confirming address_eq. Sound by the build_addr_pos_map
interning invariant (one pointer, one content): a ptr hit implies the
address IS a leaf; a de-interned pointer reads as absent and the
constant just gets checked — fail-closed.
- The iterator walks the addrs list in LOCKSTEP with consts (cur_addrs
suffix) instead of list_lookup(addrs, pos) per constant, which
re-walked the prefix every iteration — a standalone O(closure^2).
Adaptations: kept main's two-arg check_canonical_block_sort call;
addr_key retained (the Inductive.lean block-membership table added on
main after this patch still uses it), comment updated.
Measured (lake exe ix check Nat.add_comm): total width 34806 -> 34781,
FFT cost unchanged (plain checks never take the skip path). Full ixvm
suite green incl. the frontier-assumption claim smoke; all FFT pins
unchanged. crates/ixvm-codegen/src/aiur_ixvm.rs regenerated.
---------
Co-authored-by: samuelburnham <45365069+samuelburnham@users.noreply.github.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@arthurpaulino@gabriel-barrett
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

IxVM kernel: unblock UTF-8 decode/encode proof + Nat-layer FFT cuts - #450

Merged
arthurpaulino merged 8 commits into
mainfrom
ap/utf8-tier-1d
Jun 25, 2026
Merged

IxVM kernel: unblock UTF-8 decode/encode proof + Nat-layer FFT cuts#450
arthurpaulino merged 8 commits into
mainfrom
ap/utf8-tier-1d

Conversation

@arthurpaulino

Copy link
Copy Markdown
Member

Headline

Aiur kernel previously OOM'd on
_private.Init.Data.String.Decode.0.ByteArray.utf8DecodeChar?.assemble₄_eq_some_of_toBitVec._proof_1_8.
This constant is a prerequisite of ByteArray.utf8DecodeChar?_utf8EncodeChar_append
(part of the UTF-8 round-trip lemma chain).

Now it typechecks at 37.31B FFT (down from the OOM baseline). The
two-direction lemma …_utf8EncodeChar_append itself remains
out-of-reach (still OOMs further along the same dependency tree under
the user's memory cap), but the prerequisite that was the immediate
blocker is unstuck.

The same kernel changes also deliver large FFT cuts on previously-
expensive targets — Tier 1d's structural short-circuit replaces full
delta-whnf cascades wherever it fires:

ConstantBeforeAfterΔ
Vector.append4.02B3.16B−21.4%
Array.append_assoc3.94B3.08B−21.8%

Commits

1. a6cd34a — Tier 1d def-eq short-circuit (the unblock)

Aiur's k_is_def_eq_core jumped from Tier 1.5 straight to full delta
WHNF (Tier 2). The cascading Nat.rec / Nat.succ iota expansions then
drove whnf_const_head past 1M unique entries before either side
reached a comparable canonical form — OOM. Rust's def-eq settles the
same pair via no-delta whnf + quick structural recursion before any
of that fires.

Three minimum-necessary pieces ported. No KStore. No FVar. No Subst /
KernelTypes change.

  • whnf_nd family (Whnf.lean, mirror Rust
    whnf_no_delta_for_def_eq). Same dispatch tree as whnf, but
    whnf_nd_const_head's Defn arm falls through to a stuck
    apply_spine instead of delta-unfolding. Iota / proj / quot /
    primitives still fire.

  • k_infer_only family (Infer.lean, mirror Rust
    with_infer_only). App drops k_check(a, dom); Lam drops
    k_ensure_sort(ty); Let drops val/ty validation. Used at
    try_proof_irrel, is_prop_type, try_unit_like — the def-eq
    tactics that only need the synthesized type.

  • k_is_def_eq_struct_safe + Tier 1d wiring (DefEq.lean,
    mirror Rust quick_def_eq + post-try_def_eq_app). Sort-Sort via
    level_equal; Lam-Lam / All-All via recursive k_is_def_eq on the
    type and on the body under Cons(ty_a, types) (types-cons, NOT
    FVar opening). Inserted between Tier 1c (string lit) and Tier 2
    (full whnf):

    aw_nd = whnf_nd(a); bw_nd = whnf_nd(b)
    ptr_eq(aw_nd, bw_nd) → 1
    k_is_def_eq_struct_safe(aw_nd, bw_nd) → 1
    try_lazy_delta_app(aw_nd, bw_nd) → 1 (rerun: spine args may have
    reduced past what Tier 1.5's pre-whnf attempt could see)
    

Each piece independently validated necessary. The previously-tried
KStore explicit caches, FVar variant + opens, FVar-based binder
opening — all confirmed NOT necessary for this unblock and left out.

3 files changed, +296/−6 lines.

2. 36d6c3a — drop g_or from u64_sub_with_borrow

u64_sub_with_borrow combined two per-byte borrow bits with g_or.
The two bits are mutually exclusive: u_t = 1 ⇒ intermediate t_i ≥ 1 ⇒ subtracting br_in ∈ {0,1} cannot underflow ⇒ u_r = 0. Field
+ substitutes for g_or directly. Per Aiur cost model g_or adds
+1 aux + 1 lookup per call (≈ 5 width); field + is free. 7 g_ors ×
2.23M rows.

UTF-8 _proof_1_8: 39.12B → 38.14B (−2.6%).

3. 9e787da — drop g_or from klimbs_add_carry / klimbs_sub_borrow

Same mutually-exclusive-carry pattern. Two limb-level borrows /
carries from sequential u64 ops cannot both be 1.

UTF-8 _proof_1_8: 38.14B → 38.07B (−0.18%).

4. 4e379e7 — hot/cold split try_nat_dispatch, extract binop arm

try_nat_dispatch's width 90 was floored by the binop arm (2× whnf

  • 2× try_extract_nat + try_nat_binop_addr + apply_spine), charged on
    every Nat.succ / Nat.pred row. Factor binop dispatch into
    try_nat_binop_dispatch. Main narrows to the max of succ / pred
    arms.

UTF-8 _proof_1_8: 38.07B → 37.80B (−0.7%).

5. 80ce3d2 — hot/cold split expr_lbr, extract Let arm

expr_lbr's width was floored by the Let arm (3 recursive expr_lbr
calls + 2 lbr_max + 1 lbr_dec), charged on every row even though Let
is rare. Factor into expr_lbr_let.

Nat.add_comm: 55.63M → 55.50M (−0.2%). UTF-8 _proof_1_8: 37.80B
→ 37.62B (−0.5%)
.

6. 1f8effd — hot/cold split try_extract_nat, extract App arm

try_extract_nat's width was floored by the App arm (list_lookup +
address_eq + recursive try_extract_nat + klimbs_succ). Factor into
try_extract_nat_app. Main narrows to leaf-arm width.

UTF-8 _proof_1_8: 37.62B → 37.31B (−0.8%).

7. 039e9cf — re-pin IxVM FFT costs

41 pins in Tests/Ix/IxVM.lean::kernelCheckEntries updated. Every
constant got cheaper; none regressed. Largest reductions:

  • Vector.append: 4.02B → 3.16B (−21.4%)
  • Array.append_assoc: 3.94B → 3.08B (−21.8%)

lake test -- --ignored ixvm passes with 0 FFT mismatches.

Cumulative on UTF-8 _proof_1_8

CommitFFT
baseline (main)OOM
a6cd34a Tier 1d39.12B
36d6c3a g_or → + in u64_sub_with_borrow38.14B (−2.6%)
9e787da g_or → + in klimbs_add_carry / klimbs_sub_borrow38.07B (−0.18%)
4e379e7 hot/cold try_nat_dispatch37.80B (−0.7%)
80ce3d2 hot/cold expr_lbr37.62B (−0.5%)
1f8effd hot/cold try_extract_nat37.31B (−0.8%)

Post-unlock optimization: −4.6% (39.12B → 37.31B).

Cost on small targets

  • Main baseline Nat.add_comm: 56.08M FFT.
  • This branch Nat.add_comm: 55.50M FFT.

Tier 1d itself adds no overhead on the common case (whnf_nd +
struct_safe + try_lazy_delta_app are themselves Aiur-memoized); the
follow-up optimizations are net wins.

Test plan

  • lake exe check Nat.add_comm passes (55.50M FFT).
  • lake exe check Vector.extract_append passes.
  • lake exe check "_private.…assemble₄_eq_some_of_toBitVec._proof_1_8"
    passes (37.31B FFT, previously OOM).
  • lake test -- --ignored ixvm — all 41 FFT pins updated, suite
    passes with 0 mismatches.

Comment threadIx/IxVM/Kernel/Infer.lean
…ost-spine-congruence)
UTF-8 `_private.Init.Data.String.Decode.0.ByteArray.utf8DecodeChar?
.assemble₄_eq_some_of_toBitVec._proof_1_8` OOMs on the previous
pipeline because `k_is_def_eq_core` jumps from Tier 1.5 straight to
full delta WHNF (Tier 2); the cascading Nat.rec / Nat.succ iota
expansions then drive `whnf_const_head` past 1M unique entries before
either side reaches a comparable canonical form. Rust's def-eq settles
the same pair via the no-delta whnf + quick structural recursion
before any of that fires.
This patch ports the three pieces of that short-circuit and nothing
else — no FVar variant, no KStore, no Subst changes, no signature
sweep.
* `whnf_nd` family in `Whnf.lean` (mirror Rust `whnf_no_delta_for_def_eq`).
Same dispatch tree as `whnf` (beta / let zeta / iota / proj / quot /
primitives all fire), except `whnf_nd_const_head`'s Defn arm falls
through to a stuck `apply_spine` instead of delta-unfolding.
* `k_infer_only` family in `Infer.lean` (mirror Rust `with_infer_only`).
App drops `k_check(a, dom)`; Lam drops `k_ensure_sort(ty)`; Let drops
the val/ty validations. Distinct Aiur memo from `k_infer`, parity
with Rust's separate `infer_cache` / `infer_only_cache`.
* `k_is_def_eq_struct_safe` in `DefEq.lean` (mirror Rust
`quick_def_eq`). Sort-Sort via `level_equal`; Lam-Lam / All-All
recurse on type and on body under `Cons(ty_a, types)`. Returns 1
only when DEFINITELY def-eq; 0 means fall through. Sound on
partially-whnf'd (no-delta) inputs because the handled shapes
don't depend on further reductions.
* `k_is_def_eq_core` Tier 1d wiring inserted between Tier 1c (string
lit) and Tier 2 (full whnf):
aw_nd = whnf_nd(a); bw_nd = whnf_nd(b)
ptr_eq(aw_nd, bw_nd) → 1
k_is_def_eq_struct_safe(aw_nd, bw_nd) → 1 if 1
try_lazy_delta_app(aw_nd, bw_nd) → 1 if 1 (rerun post-whnf_nd:
spine args may have reduced past what Tier 1.5's pre-whnf attempt
could see, exposing Const-Const congruence that was hidden)
* `try_proof_irrel`, `is_prop_type`, `try_unit_like` switch from
`k_infer` to `k_infer_only` — these helpers only need the synthesized
type, not the full re-validation work that `k_infer` does for each
recursive App/Let/Lam.
Each piece individually validated necessary (removing it puts UTF-8
back into the OOM regime). FVar variant + opens, KStore explicit
caches, infer_only's FVar-based binder opening — all confirmed NOT
necessary for the UTF-8 unblock and left out (see PLAN.md for future
experiments).
Measured (FFT cost):
Nat.add_comm: 56.08M → 55.63M (~stable; new code paths add no
overhead on the common case because Tier 1d's whnf_nd + struct_safe
are themselves Aiur-memoized).
_private.Init.Data.String.Decode.0.ByteArray.utf8DecodeChar?
.assemble₄_eq_some_of_toBitVec._proof_1_8: OOM → 39.12B FFT, passes.
3 files, +296/-6 lines.
`u64_sub_with_borrow` combines two per-byte borrow bits with `g_or`. The
two bits are MUTUALLY EXCLUSIVE: `u_t = borrow(a_i - b_i)` and
`u_r = borrow((a_i + 256 - b_i) - br_in)`. If `u_t = 1` the intermediate
`t_i ≥ 1`, so subtracting `br_in ∈ {0,1}` cannot underflow ⇒ `u_r = 0`.
Field `+` substitutes for `g_or` directly (per the same pattern as
`u64_add` in `ByteStream.lean`).
Per Aiur cost model, `g_or` adds +1 aux + 1 lookup per call; field `+`
is free. 8 g_or call sites in `u64_sub_with_borrow` each charged on every
one of the function's 2.23M rows.
Measured (FFT cost) on UTF-8 `_proof_1_8`:
39.12B → 38.14B (-2.6%)
Nat.add_comm unchanged (55.63M).
See [[reference_aiur_carry_add]].
Same mutually-exclusive-carry pattern as `u64_sub_with_borrow`:
* `klimbs_add_carry`: u64_add of (la, lb) yields carry1; u64_add of
(sum1, carry_in) yields carry2. carry1=1 ⇒ sum1 ≤ 2^64-2 ⇒
carry2=0.
* `klimbs_sub_borrow`: symmetric for borrows.
Replace `g_or(c1, c2)` with `c1 + c2` (field +). Both helpers run on
hot Nat-primitive paths.
Measured on UTF-8 `_proof_1_8`:
38.14B → 38.07B (-0.18%)
Nat.add_comm unchanged.
See [[reference_aiur_carry_add]].
`try_nat_dispatch` ran 1.12M rows in UTF-8 `_proof_1_8` at width 90,
charging 5.16% of total FFT. Width was floored by its widest match arm
(the binop branch with 2× whnf + 2× try_extract_nat + try_nat_binop_addr
+ apply_spine), even on Nat.succ / Nat.pred rows that never touched it.
Factor binop dispatch into its own `try_nat_binop_dispatch` fn. Main
dispatcher narrows to the max of succ / pred arms (single whnf +
try_extract_nat + klimbs_succ/dec + apply_spine). The cold fn's width
only charges the rows that actually dispatch a binop.
Measured on UTF-8 `_proof_1_8`:
38.07B → 37.80B (-0.7%)
Nat.add_comm unchanged.
`expr_lbr` ran 1.47M rows in UTF-8 `_proof_1_8` at width 39, charging
3.01% of total FFT. The Let arm (3 recursive expr_lbr calls + 2 lbr_max
+ 1 lbr_dec) is the widest match arm, charged on every row of expr_lbr
even though Let is rare in most expressions encountered.
Factor the Let arm into `expr_lbr_let(ty, val, body)`. Main expr_lbr
narrows to max of the 2-recursion arms (App / Lam / Forall). Cold fn
only charges Let-arm rows.
Measured:
Nat.add_comm: 55.63M → 55.50M (-0.2%)
UTF-8 `_proof_1_8`: 37.80B → 37.62B (-0.5%)
`try_extract_nat` ran 1.12M rows at width 45, charging 2.68% of UTF-8
`_proof_1_8` total FFT. The App arm (list_lookup + address_eq +
recursive try_extract_nat + klimbs_succ) is the widest match arm; the
Lit / Const / default arms are leaf compares.
Factor App into `try_extract_nat_app(f, a, addrs)`. Main extractor
narrows to leaf-arm width. Cold fn only charges App-arm rows.
Measured on UTF-8 `_proof_1_8`:
37.62B → 37.31B (-0.8%)
Nat.add_comm unchanged.
Updates 41 pinned FFT costs in `Tests/Ix/IxVM.lean::kernelCheckEntries`
to match the new kernel's output. All pins moved DOWN — every constant
got cheaper, none regressed.
Largest reductions (% change):
Vector.append: 4_023_268_168 → 3_160_970_390 (-21.4%)
Array.append_assoc: 3_938_574_533 → 3_079_334_815 (-21.8%)
String.Internal.append: 793_580_333 → 775_968_134 ( -2.2%)
bv_to_nat_lit: 635_780_327 → 619_870_154 ( -2.5%)
nat_gcd_lit: 665_518_356 → 649_859_784 ( -2.4%)
Nat.sub_le_of_le_add: 567_575_653 → 557_867_526 ( -1.7%)
IxVMPrim.nat_mod_lit: 414_695_549 → 407_517_834 ( -1.7%)
IxVMPrim.nat_div_lit: 405_607_545 → 398_641_590 ( -1.7%)
IxVMPrim.nat_shr_lit: 411_128_901 → 404_158_486 ( -1.7%)
Nat.decLe: 209_641_496 → 206_196_563 ( -1.6%)
Nat.add_comm: 56_084_908 → 55_504_714 ( -1.0%)
`lake test -- --ignored ixvm` passes with 0 FFT mismatches.
The `k_infer_only` section header skipped over the fact that the
function is only sound on well-typed inputs (since it drops
`k_check(a, dom)` on `App`, `k_ensure_sort(ty)` on `Lam`, val/ty
checks on `Let`). Spell out:
* The invariant — only call on terms produced by `whnf` of a
well-typed term, never on arbitrary inputs.
* Why the current sites (`try_proof_irrel`, `is_prop_type`,
`try_unit_like`) respect it.
* The planned non-deterministic-hint dispatch (`Hint::{None,
KInfer, KInferOnly}`) that lets us share `k_infer`'s memo where a
hit already exists instead of paying the parallel `infer_only`
memo cost.
@arthurpaulino
arthurpaulino merged commit ad7e383 into mainJun 25, 2026
14 checks passed
@arthurpaulino
arthurpaulino deleted the ap/utf8-tier-1d branch June 25, 2026 13:49
johnchandlerburnham pushed a commit that referenced this pull request Jul 21, 2026
…505)
* IxVM: drop dead KValNode/KVal/KValEnv
From ap/kernel 828fb85 (Arthur Paulino): the NbE value domain is defined
but referenced nowhere — the live kernel runs on de-Bruijn KExpr;
vestigial from an abandoned NbE direction. (The closed-term context
normalization from that commit was measured separately and not taken:
the per-call expr_lbr probe cost +2.9% FFT on recursor loops for a 0.5%
record reduction.)
* IxVM: memoized prim_family dispatch + width-safe offset-stuck placement
Three coordinated changes to Const-head whnf dispatch (cherry-pick of
130f30b, adapted to post-#450/#457 main):
1. prim_family(addr) classifies a head address into the one reducer
family that could fire on it (nat/str/bitvec/native/decidable; the
sets are disjoint). Keyed on the ADDRESS ALONE it memoizes to one row
per distinct constant address per run, and whnf_const_head calls at
most one family reducer. The previous gauntlet ran every reducer in
sequence for guaranteed misses.
2. The symbolic-Nat offset-stuck check moves from a delta-arm probe into
try_nat_dispatch's miss path as a cold function
(try_nat_offset_dispatch, verdict 2 = "already stuck, do not
re-whnf"), with the offset construction shared via
mk_nat_offset_stuck (also used by the linear-rec collapse).
3. nat_lit_to_ctor_or_self exposes ONE constructor layer
(n -> succ(Lit(n-1))) instead of materializing the full succ chain.
Adaptations vs the original patch:
- whnf_nd_const_head (no-delta WHNF, added on main after the patch)
converted to the same family dispatch.
- Kept main's cold-extracted try_nat_binop_dispatch and routed the
symbolic-base case to try_nat_offset_dispatch from its miss arm.
Measured (lake exe ix check Nat.add_comm): total width 34820 -> 34806,
FFT cost 49571210 -> 48860647 (-1.43%). All 53 ixvm-suite FFT pins
decreased (-0.19%..-1.43%); parity and claim smokes pass. Pins updated;
crates/ixvm-codegen/src/aiur_ixvm.rs regenerated via `lake exe ix
codegen`.
* IxVM: port jcb/fixes H-14 — ptr_val skip map + lockstep addr cursor
Two quadratic/constant-factor fixes to check_all_skipping, ported from
jcb/fixes (3763356, John C. Burnham); cherry-pick of 2d83be1 adapted to
post-#457 main:
- The assumption-leaf skip set keys on ptr_val instead of the first 4
address bytes: one tree lookup per constant, no per-lookup address
load, no confirming address_eq. Sound by the build_addr_pos_map
interning invariant (one pointer, one content): a ptr hit implies the
address IS a leaf; a de-interned pointer reads as absent and the
constant just gets checked — fail-closed.
- The iterator walks the addrs list in LOCKSTEP with consts (cur_addrs
suffix) instead of list_lookup(addrs, pos) per constant, which
re-walked the prefix every iteration — a standalone O(closure^2).
Adaptations: kept main's two-arg check_canonical_block_sort call;
addr_key retained (the Inductive.lean block-membership table added on
main after this patch still uses it), comment updated.
Measured (lake exe ix check Nat.add_comm): total width 34806 -> 34781,
FFT cost unchanged (plain checks never take the skip path). Full ixvm
suite green incl. the frontier-assumption claim smoke; all FFT pins
unchanged. crates/ixvm-codegen/src/aiur_ixvm.rs regenerated.
---------
Co-authored-by: samuelburnham <45365069+samuelburnham@users.noreply.github.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@arthurpaulino@gabriel-barrett
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

IxVM kernel: unblock UTF-8 decode/encode proof + Nat-layer FFT cuts - #450

Merged
arthurpaulino merged 8 commits into
mainfrom
ap/utf8-tier-1d
Jun 25, 2026
Merged

IxVM kernel: unblock UTF-8 decode/encode proof + Nat-layer FFT cuts#450
arthurpaulino merged 8 commits into
mainfrom
ap/utf8-tier-1d

Conversation

@arthurpaulino

Copy link
Copy Markdown
Member

Headline

Aiur kernel previously OOM'd on
_private.Init.Data.String.Decode.0.ByteArray.utf8DecodeChar?.assemble₄_eq_some_of_toBitVec._proof_1_8.
This constant is a prerequisite of ByteArray.utf8DecodeChar?_utf8EncodeChar_append
(part of the UTF-8 round-trip lemma chain).

Now it typechecks at 37.31B FFT (down from the OOM baseline). The
two-direction lemma …_utf8EncodeChar_append itself remains
out-of-reach (still OOMs further along the same dependency tree under
the user's memory cap), but the prerequisite that was the immediate
blocker is unstuck.

The same kernel changes also deliver large FFT cuts on previously-
expensive targets — Tier 1d's structural short-circuit replaces full
delta-whnf cascades wherever it fires:

ConstantBeforeAfterΔ
Vector.append4.02B3.16B−21.4%
Array.append_assoc3.94B3.08B−21.8%

Commits

1. a6cd34a — Tier 1d def-eq short-circuit (the unblock)

Aiur's k_is_def_eq_core jumped from Tier 1.5 straight to full delta
WHNF (Tier 2). The cascading Nat.rec / Nat.succ iota expansions then
drove whnf_const_head past 1M unique entries before either side
reached a comparable canonical form — OOM. Rust's def-eq settles the
same pair via no-delta whnf + quick structural recursion before any
of that fires.

Three minimum-necessary pieces ported. No KStore. No FVar. No Subst /
KernelTypes change.

  • whnf_nd family (Whnf.lean, mirror Rust
    whnf_no_delta_for_def_eq). Same dispatch tree as whnf, but
    whnf_nd_const_head's Defn arm falls through to a stuck
    apply_spine instead of delta-unfolding. Iota / proj / quot /
    primitives still fire.

  • k_infer_only family (Infer.lean, mirror Rust
    with_infer_only). App drops k_check(a, dom); Lam drops
    k_ensure_sort(ty); Let drops val/ty validation. Used at
    try_proof_irrel, is_prop_type, try_unit_like — the def-eq
    tactics that only need the synthesized type.

  • k_is_def_eq_struct_safe + Tier 1d wiring (DefEq.lean,
    mirror Rust quick_def_eq + post-try_def_eq_app). Sort-Sort via
    level_equal; Lam-Lam / All-All via recursive k_is_def_eq on the
    type and on the body under Cons(ty_a, types) (types-cons, NOT
    FVar opening). Inserted between Tier 1c (string lit) and Tier 2
    (full whnf):

    aw_nd = whnf_nd(a); bw_nd = whnf_nd(b)
    ptr_eq(aw_nd, bw_nd) → 1
    k_is_def_eq_struct_safe(aw_nd, bw_nd) → 1
    try_lazy_delta_app(aw_nd, bw_nd) → 1 (rerun: spine args may have
    reduced past what Tier 1.5's pre-whnf attempt could see)
    

Each piece independently validated necessary. The previously-tried
KStore explicit caches, FVar variant + opens, FVar-based binder
opening — all confirmed NOT necessary for this unblock and left out.

3 files changed, +296/−6 lines.

2. 36d6c3a — drop g_or from u64_sub_with_borrow

u64_sub_with_borrow combined two per-byte borrow bits with g_or.
The two bits are mutually exclusive: u_t = 1 ⇒ intermediate t_i ≥ 1 ⇒ subtracting br_in ∈ {0,1} cannot underflow ⇒ u_r = 0. Field
+ substitutes for g_or directly. Per Aiur cost model g_or adds
+1 aux + 1 lookup per call (≈ 5 width); field + is free. 7 g_ors ×
2.23M rows.

UTF-8 _proof_1_8: 39.12B → 38.14B (−2.6%).

3. 9e787da — drop g_or from klimbs_add_carry / klimbs_sub_borrow

Same mutually-exclusive-carry pattern. Two limb-level borrows /
carries from sequential u64 ops cannot both be 1.

UTF-8 _proof_1_8: 38.14B → 38.07B (−0.18%).

4. 4e379e7 — hot/cold split try_nat_dispatch, extract binop arm

try_nat_dispatch's width 90 was floored by the binop arm (2× whnf

  • 2× try_extract_nat + try_nat_binop_addr + apply_spine), charged on
    every Nat.succ / Nat.pred row. Factor binop dispatch into
    try_nat_binop_dispatch. Main narrows to the max of succ / pred
    arms.

UTF-8 _proof_1_8: 38.07B → 37.80B (−0.7%).

5. 80ce3d2 — hot/cold split expr_lbr, extract Let arm

expr_lbr's width was floored by the Let arm (3 recursive expr_lbr
calls + 2 lbr_max + 1 lbr_dec), charged on every row even though Let
is rare. Factor into expr_lbr_let.

Nat.add_comm: 55.63M → 55.50M (−0.2%). UTF-8 _proof_1_8: 37.80B
→ 37.62B (−0.5%)
.

6. 1f8effd — hot/cold split try_extract_nat, extract App arm

try_extract_nat's width was floored by the App arm (list_lookup +
address_eq + recursive try_extract_nat + klimbs_succ). Factor into
try_extract_nat_app. Main narrows to leaf-arm width.

UTF-8 _proof_1_8: 37.62B → 37.31B (−0.8%).

7. 039e9cf — re-pin IxVM FFT costs

41 pins in Tests/Ix/IxVM.lean::kernelCheckEntries updated. Every
constant got cheaper; none regressed. Largest reductions:

  • Vector.append: 4.02B → 3.16B (−21.4%)
  • Array.append_assoc: 3.94B → 3.08B (−21.8%)

lake test -- --ignored ixvm passes with 0 FFT mismatches.

Cumulative on UTF-8 _proof_1_8

CommitFFT
baseline (main)OOM
a6cd34a Tier 1d39.12B
36d6c3a g_or → + in u64_sub_with_borrow38.14B (−2.6%)
9e787da g_or → + in klimbs_add_carry / klimbs_sub_borrow38.07B (−0.18%)
4e379e7 hot/cold try_nat_dispatch37.80B (−0.7%)
80ce3d2 hot/cold expr_lbr37.62B (−0.5%)
1f8effd hot/cold try_extract_nat37.31B (−0.8%)

Post-unlock optimization: −4.6% (39.12B → 37.31B).

Cost on small targets

  • Main baseline Nat.add_comm: 56.08M FFT.
  • This branch Nat.add_comm: 55.50M FFT.

Tier 1d itself adds no overhead on the common case (whnf_nd +
struct_safe + try_lazy_delta_app are themselves Aiur-memoized); the
follow-up optimizations are net wins.

Test plan

  • lake exe check Nat.add_comm passes (55.50M FFT).
  • lake exe check Vector.extract_append passes.
  • lake exe check "_private.…assemble₄_eq_some_of_toBitVec._proof_1_8"
    passes (37.31B FFT, previously OOM).
  • lake test -- --ignored ixvm — all 41 FFT pins updated, suite
    passes with 0 mismatches.

Comment threadIx/IxVM/Kernel/Infer.lean
…ost-spine-congruence)
UTF-8 `_private.Init.Data.String.Decode.0.ByteArray.utf8DecodeChar?
.assemble₄_eq_some_of_toBitVec._proof_1_8` OOMs on the previous
pipeline because `k_is_def_eq_core` jumps from Tier 1.5 straight to
full delta WHNF (Tier 2); the cascading Nat.rec / Nat.succ iota
expansions then drive `whnf_const_head` past 1M unique entries before
either side reaches a comparable canonical form. Rust's def-eq settles
the same pair via the no-delta whnf + quick structural recursion
before any of that fires.
This patch ports the three pieces of that short-circuit and nothing
else — no FVar variant, no KStore, no Subst changes, no signature
sweep.
* `whnf_nd` family in `Whnf.lean` (mirror Rust `whnf_no_delta_for_def_eq`).
Same dispatch tree as `whnf` (beta / let zeta / iota / proj / quot /
primitives all fire), except `whnf_nd_const_head`'s Defn arm falls
through to a stuck `apply_spine` instead of delta-unfolding.
* `k_infer_only` family in `Infer.lean` (mirror Rust `with_infer_only`).
App drops `k_check(a, dom)`; Lam drops `k_ensure_sort(ty)`; Let drops
the val/ty validations. Distinct Aiur memo from `k_infer`, parity
with Rust's separate `infer_cache` / `infer_only_cache`.
* `k_is_def_eq_struct_safe` in `DefEq.lean` (mirror Rust
`quick_def_eq`). Sort-Sort via `level_equal`; Lam-Lam / All-All
recurse on type and on body under `Cons(ty_a, types)`. Returns 1
only when DEFINITELY def-eq; 0 means fall through. Sound on
partially-whnf'd (no-delta) inputs because the handled shapes
don't depend on further reductions.
* `k_is_def_eq_core` Tier 1d wiring inserted between Tier 1c (string
lit) and Tier 2 (full whnf):
aw_nd = whnf_nd(a); bw_nd = whnf_nd(b)
ptr_eq(aw_nd, bw_nd) → 1
k_is_def_eq_struct_safe(aw_nd, bw_nd) → 1 if 1
try_lazy_delta_app(aw_nd, bw_nd) → 1 if 1 (rerun post-whnf_nd:
spine args may have reduced past what Tier 1.5's pre-whnf attempt
could see, exposing Const-Const congruence that was hidden)
* `try_proof_irrel`, `is_prop_type`, `try_unit_like` switch from
`k_infer` to `k_infer_only` — these helpers only need the synthesized
type, not the full re-validation work that `k_infer` does for each
recursive App/Let/Lam.
Each piece individually validated necessary (removing it puts UTF-8
back into the OOM regime). FVar variant + opens, KStore explicit
caches, infer_only's FVar-based binder opening — all confirmed NOT
necessary for the UTF-8 unblock and left out (see PLAN.md for future
experiments).
Measured (FFT cost):
Nat.add_comm: 56.08M → 55.63M (~stable; new code paths add no
overhead on the common case because Tier 1d's whnf_nd + struct_safe
are themselves Aiur-memoized).
_private.Init.Data.String.Decode.0.ByteArray.utf8DecodeChar?
.assemble₄_eq_some_of_toBitVec._proof_1_8: OOM → 39.12B FFT, passes.
3 files, +296/-6 lines.
`u64_sub_with_borrow` combines two per-byte borrow bits with `g_or`. The
two bits are MUTUALLY EXCLUSIVE: `u_t = borrow(a_i - b_i)` and
`u_r = borrow((a_i + 256 - b_i) - br_in)`. If `u_t = 1` the intermediate
`t_i ≥ 1`, so subtracting `br_in ∈ {0,1}` cannot underflow ⇒ `u_r = 0`.
Field `+` substitutes for `g_or` directly (per the same pattern as
`u64_add` in `ByteStream.lean`).
Per Aiur cost model, `g_or` adds +1 aux + 1 lookup per call; field `+`
is free. 8 g_or call sites in `u64_sub_with_borrow` each charged on every
one of the function's 2.23M rows.
Measured (FFT cost) on UTF-8 `_proof_1_8`:
39.12B → 38.14B (-2.6%)
Nat.add_comm unchanged (55.63M).
See [[reference_aiur_carry_add]].
Same mutually-exclusive-carry pattern as `u64_sub_with_borrow`:
* `klimbs_add_carry`: u64_add of (la, lb) yields carry1; u64_add of
(sum1, carry_in) yields carry2. carry1=1 ⇒ sum1 ≤ 2^64-2 ⇒
carry2=0.
* `klimbs_sub_borrow`: symmetric for borrows.
Replace `g_or(c1, c2)` with `c1 + c2` (field +). Both helpers run on
hot Nat-primitive paths.
Measured on UTF-8 `_proof_1_8`:
38.14B → 38.07B (-0.18%)
Nat.add_comm unchanged.
See [[reference_aiur_carry_add]].
`try_nat_dispatch` ran 1.12M rows in UTF-8 `_proof_1_8` at width 90,
charging 5.16% of total FFT. Width was floored by its widest match arm
(the binop branch with 2× whnf + 2× try_extract_nat + try_nat_binop_addr
+ apply_spine), even on Nat.succ / Nat.pred rows that never touched it.
Factor binop dispatch into its own `try_nat_binop_dispatch` fn. Main
dispatcher narrows to the max of succ / pred arms (single whnf +
try_extract_nat + klimbs_succ/dec + apply_spine). The cold fn's width
only charges the rows that actually dispatch a binop.
Measured on UTF-8 `_proof_1_8`:
38.07B → 37.80B (-0.7%)
Nat.add_comm unchanged.
`expr_lbr` ran 1.47M rows in UTF-8 `_proof_1_8` at width 39, charging
3.01% of total FFT. The Let arm (3 recursive expr_lbr calls + 2 lbr_max
+ 1 lbr_dec) is the widest match arm, charged on every row of expr_lbr
even though Let is rare in most expressions encountered.
Factor the Let arm into `expr_lbr_let(ty, val, body)`. Main expr_lbr
narrows to max of the 2-recursion arms (App / Lam / Forall). Cold fn
only charges Let-arm rows.
Measured:
Nat.add_comm: 55.63M → 55.50M (-0.2%)
UTF-8 `_proof_1_8`: 37.80B → 37.62B (-0.5%)
`try_extract_nat` ran 1.12M rows at width 45, charging 2.68% of UTF-8
`_proof_1_8` total FFT. The App arm (list_lookup + address_eq +
recursive try_extract_nat + klimbs_succ) is the widest match arm; the
Lit / Const / default arms are leaf compares.
Factor App into `try_extract_nat_app(f, a, addrs)`. Main extractor
narrows to leaf-arm width. Cold fn only charges App-arm rows.
Measured on UTF-8 `_proof_1_8`:
37.62B → 37.31B (-0.8%)
Nat.add_comm unchanged.
Updates 41 pinned FFT costs in `Tests/Ix/IxVM.lean::kernelCheckEntries`
to match the new kernel's output. All pins moved DOWN — every constant
got cheaper, none regressed.
Largest reductions (% change):
Vector.append: 4_023_268_168 → 3_160_970_390 (-21.4%)
Array.append_assoc: 3_938_574_533 → 3_079_334_815 (-21.8%)
String.Internal.append: 793_580_333 → 775_968_134 ( -2.2%)
bv_to_nat_lit: 635_780_327 → 619_870_154 ( -2.5%)
nat_gcd_lit: 665_518_356 → 649_859_784 ( -2.4%)
Nat.sub_le_of_le_add: 567_575_653 → 557_867_526 ( -1.7%)
IxVMPrim.nat_mod_lit: 414_695_549 → 407_517_834 ( -1.7%)
IxVMPrim.nat_div_lit: 405_607_545 → 398_641_590 ( -1.7%)
IxVMPrim.nat_shr_lit: 411_128_901 → 404_158_486 ( -1.7%)
Nat.decLe: 209_641_496 → 206_196_563 ( -1.6%)
Nat.add_comm: 56_084_908 → 55_504_714 ( -1.0%)
`lake test -- --ignored ixvm` passes with 0 FFT mismatches.
The `k_infer_only` section header skipped over the fact that the
function is only sound on well-typed inputs (since it drops
`k_check(a, dom)` on `App`, `k_ensure_sort(ty)` on `Lam`, val/ty
checks on `Let`). Spell out:
* The invariant — only call on terms produced by `whnf` of a
well-typed term, never on arbitrary inputs.
* Why the current sites (`try_proof_irrel`, `is_prop_type`,
`try_unit_like`) respect it.
* The planned non-deterministic-hint dispatch (`Hint::{None,
KInfer, KInferOnly}`) that lets us share `k_infer`'s memo where a
hit already exists instead of paying the parallel `infer_only`
memo cost.
@arthurpaulino
arthurpaulino merged commit ad7e383 into mainJun 25, 2026
14 checks passed
@arthurpaulino
arthurpaulino deleted the ap/utf8-tier-1d branch June 25, 2026 13:49
johnchandlerburnham pushed a commit that referenced this pull request Jul 21, 2026
…505)
* IxVM: drop dead KValNode/KVal/KValEnv
From ap/kernel 828fb85 (Arthur Paulino): the NbE value domain is defined
but referenced nowhere — the live kernel runs on de-Bruijn KExpr;
vestigial from an abandoned NbE direction. (The closed-term context
normalization from that commit was measured separately and not taken:
the per-call expr_lbr probe cost +2.9% FFT on recursor loops for a 0.5%
record reduction.)
* IxVM: memoized prim_family dispatch + width-safe offset-stuck placement
Three coordinated changes to Const-head whnf dispatch (cherry-pick of
130f30b, adapted to post-#450/#457 main):
1. prim_family(addr) classifies a head address into the one reducer
family that could fire on it (nat/str/bitvec/native/decidable; the
sets are disjoint). Keyed on the ADDRESS ALONE it memoizes to one row
per distinct constant address per run, and whnf_const_head calls at
most one family reducer. The previous gauntlet ran every reducer in
sequence for guaranteed misses.
2. The symbolic-Nat offset-stuck check moves from a delta-arm probe into
try_nat_dispatch's miss path as a cold function
(try_nat_offset_dispatch, verdict 2 = "already stuck, do not
re-whnf"), with the offset construction shared via
mk_nat_offset_stuck (also used by the linear-rec collapse).
3. nat_lit_to_ctor_or_self exposes ONE constructor layer
(n -> succ(Lit(n-1))) instead of materializing the full succ chain.
Adaptations vs the original patch:
- whnf_nd_const_head (no-delta WHNF, added on main after the patch)
converted to the same family dispatch.
- Kept main's cold-extracted try_nat_binop_dispatch and routed the
symbolic-base case to try_nat_offset_dispatch from its miss arm.
Measured (lake exe ix check Nat.add_comm): total width 34820 -> 34806,
FFT cost 49571210 -> 48860647 (-1.43%). All 53 ixvm-suite FFT pins
decreased (-0.19%..-1.43%); parity and claim smokes pass. Pins updated;
crates/ixvm-codegen/src/aiur_ixvm.rs regenerated via `lake exe ix
codegen`.
* IxVM: port jcb/fixes H-14 — ptr_val skip map + lockstep addr cursor
Two quadratic/constant-factor fixes to check_all_skipping, ported from
jcb/fixes (3763356, John C. Burnham); cherry-pick of 2d83be1 adapted to
post-#457 main:
- The assumption-leaf skip set keys on ptr_val instead of the first 4
address bytes: one tree lookup per constant, no per-lookup address
load, no confirming address_eq. Sound by the build_addr_pos_map
interning invariant (one pointer, one content): a ptr hit implies the
address IS a leaf; a de-interned pointer reads as absent and the
constant just gets checked — fail-closed.
- The iterator walks the addrs list in LOCKSTEP with consts (cur_addrs
suffix) instead of list_lookup(addrs, pos) per constant, which
re-walked the prefix every iteration — a standalone O(closure^2).
Adaptations: kept main's two-arg check_canonical_block_sort call;
addr_key retained (the Inductive.lean block-membership table added on
main after this patch still uses it), comment updated.
Measured (lake exe ix check Nat.add_comm): total width 34806 -> 34781,
FFT cost unchanged (plain checks never take the skip path). Full ixvm
suite green incl. the frontier-assumption claim smoke; all FFT pins
unchanged. crates/ixvm-codegen/src/aiur_ixvm.rs regenerated.
---------
Co-authored-by: samuelburnham <45365069+samuelburnham@users.noreply.github.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@arthurpaulino@gabriel-barrett
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

IxVM kernel: unblock UTF-8 decode/encode proof + Nat-layer FFT cuts - #450

Merged
arthurpaulino merged 8 commits into
mainfrom
ap/utf8-tier-1d
Jun 25, 2026
Merged

IxVM kernel: unblock UTF-8 decode/encode proof + Nat-layer FFT cuts#450
arthurpaulino merged 8 commits into
mainfrom
ap/utf8-tier-1d

Conversation

@arthurpaulino

Copy link
Copy Markdown
Member

Headline

Aiur kernel previously OOM'd on
_private.Init.Data.String.Decode.0.ByteArray.utf8DecodeChar?.assemble₄_eq_some_of_toBitVec._proof_1_8.
This constant is a prerequisite of ByteArray.utf8DecodeChar?_utf8EncodeChar_append
(part of the UTF-8 round-trip lemma chain).

Now it typechecks at 37.31B FFT (down from the OOM baseline). The
two-direction lemma …_utf8EncodeChar_append itself remains
out-of-reach (still OOMs further along the same dependency tree under
the user's memory cap), but the prerequisite that was the immediate
blocker is unstuck.

The same kernel changes also deliver large FFT cuts on previously-
expensive targets — Tier 1d's structural short-circuit replaces full
delta-whnf cascades wherever it fires:

ConstantBeforeAfterΔ
Vector.append4.02B3.16B−21.4%
Array.append_assoc3.94B3.08B−21.8%

Commits

1. a6cd34a — Tier 1d def-eq short-circuit (the unblock)

Aiur's k_is_def_eq_core jumped from Tier 1.5 straight to full delta
WHNF (Tier 2). The cascading Nat.rec / Nat.succ iota expansions then
drove whnf_const_head past 1M unique entries before either side
reached a comparable canonical form — OOM. Rust's def-eq settles the
same pair via no-delta whnf + quick structural recursion before any
of that fires.

Three minimum-necessary pieces ported. No KStore. No FVar. No Subst /
KernelTypes change.

  • whnf_nd family (Whnf.lean, mirror Rust
    whnf_no_delta_for_def_eq). Same dispatch tree as whnf, but
    whnf_nd_const_head's Defn arm falls through to a stuck
    apply_spine instead of delta-unfolding. Iota / proj / quot /
    primitives still fire.

  • k_infer_only family (Infer.lean, mirror Rust
    with_infer_only). App drops k_check(a, dom); Lam drops
    k_ensure_sort(ty); Let drops val/ty validation. Used at
    try_proof_irrel, is_prop_type, try_unit_like — the def-eq
    tactics that only need the synthesized type.

  • k_is_def_eq_struct_safe + Tier 1d wiring (DefEq.lean,
    mirror Rust quick_def_eq + post-try_def_eq_app). Sort-Sort via
    level_equal; Lam-Lam / All-All via recursive k_is_def_eq on the
    type and on the body under Cons(ty_a, types) (types-cons, NOT
    FVar opening). Inserted between Tier 1c (string lit) and Tier 2
    (full whnf):

    aw_nd = whnf_nd(a); bw_nd = whnf_nd(b)
    ptr_eq(aw_nd, bw_nd) → 1
    k_is_def_eq_struct_safe(aw_nd, bw_nd) → 1
    try_lazy_delta_app(aw_nd, bw_nd) → 1 (rerun: spine args may have
    reduced past what Tier 1.5's pre-whnf attempt could see)
    

Each piece independently validated necessary. The previously-tried
KStore explicit caches, FVar variant + opens, FVar-based binder
opening — all confirmed NOT necessary for this unblock and left out.

3 files changed, +296/−6 lines.

2. 36d6c3a — drop g_or from u64_sub_with_borrow

u64_sub_with_borrow combined two per-byte borrow bits with g_or.
The two bits are mutually exclusive: u_t = 1 ⇒ intermediate t_i ≥ 1 ⇒ subtracting br_in ∈ {0,1} cannot underflow ⇒ u_r = 0. Field
+ substitutes for g_or directly. Per Aiur cost model g_or adds
+1 aux + 1 lookup per call (≈ 5 width); field + is free. 7 g_ors ×
2.23M rows.

UTF-8 _proof_1_8: 39.12B → 38.14B (−2.6%).

3. 9e787da — drop g_or from klimbs_add_carry / klimbs_sub_borrow

Same mutually-exclusive-carry pattern. Two limb-level borrows /
carries from sequential u64 ops cannot both be 1.

UTF-8 _proof_1_8: 38.14B → 38.07B (−0.18%).

4. 4e379e7 — hot/cold split try_nat_dispatch, extract binop arm

try_nat_dispatch's width 90 was floored by the binop arm (2× whnf

  • 2× try_extract_nat + try_nat_binop_addr + apply_spine), charged on
    every Nat.succ / Nat.pred row. Factor binop dispatch into
    try_nat_binop_dispatch. Main narrows to the max of succ / pred
    arms.

UTF-8 _proof_1_8: 38.07B → 37.80B (−0.7%).

5. 80ce3d2 — hot/cold split expr_lbr, extract Let arm

expr_lbr's width was floored by the Let arm (3 recursive expr_lbr
calls + 2 lbr_max + 1 lbr_dec), charged on every row even though Let
is rare. Factor into expr_lbr_let.

Nat.add_comm: 55.63M → 55.50M (−0.2%). UTF-8 _proof_1_8: 37.80B
→ 37.62B (−0.5%)
.

6. 1f8effd — hot/cold split try_extract_nat, extract App arm

try_extract_nat's width was floored by the App arm (list_lookup +
address_eq + recursive try_extract_nat + klimbs_succ). Factor into
try_extract_nat_app. Main narrows to leaf-arm width.

UTF-8 _proof_1_8: 37.62B → 37.31B (−0.8%).

7. 039e9cf — re-pin IxVM FFT costs

41 pins in Tests/Ix/IxVM.lean::kernelCheckEntries updated. Every
constant got cheaper; none regressed. Largest reductions:

  • Vector.append: 4.02B → 3.16B (−21.4%)
  • Array.append_assoc: 3.94B → 3.08B (−21.8%)

lake test -- --ignored ixvm passes with 0 FFT mismatches.

Cumulative on UTF-8 _proof_1_8

CommitFFT
baseline (main)OOM
a6cd34a Tier 1d39.12B
36d6c3a g_or → + in u64_sub_with_borrow38.14B (−2.6%)
9e787da g_or → + in klimbs_add_carry / klimbs_sub_borrow38.07B (−0.18%)
4e379e7 hot/cold try_nat_dispatch37.80B (−0.7%)
80ce3d2 hot/cold expr_lbr37.62B (−0.5%)
1f8effd hot/cold try_extract_nat37.31B (−0.8%)

Post-unlock optimization: −4.6% (39.12B → 37.31B).

Cost on small targets

  • Main baseline Nat.add_comm: 56.08M FFT.
  • This branch Nat.add_comm: 55.50M FFT.

Tier 1d itself adds no overhead on the common case (whnf_nd +
struct_safe + try_lazy_delta_app are themselves Aiur-memoized); the
follow-up optimizations are net wins.

Test plan

  • lake exe check Nat.add_comm passes (55.50M FFT).
  • lake exe check Vector.extract_append passes.
  • lake exe check "_private.…assemble₄_eq_some_of_toBitVec._proof_1_8"
    passes (37.31B FFT, previously OOM).
  • lake test -- --ignored ixvm — all 41 FFT pins updated, suite
    passes with 0 mismatches.

Comment threadIx/IxVM/Kernel/Infer.lean
…ost-spine-congruence)
UTF-8 `_private.Init.Data.String.Decode.0.ByteArray.utf8DecodeChar?
.assemble₄_eq_some_of_toBitVec._proof_1_8` OOMs on the previous
pipeline because `k_is_def_eq_core` jumps from Tier 1.5 straight to
full delta WHNF (Tier 2); the cascading Nat.rec / Nat.succ iota
expansions then drive `whnf_const_head` past 1M unique entries before
either side reaches a comparable canonical form. Rust's def-eq settles
the same pair via the no-delta whnf + quick structural recursion
before any of that fires.
This patch ports the three pieces of that short-circuit and nothing
else — no FVar variant, no KStore, no Subst changes, no signature
sweep.
* `whnf_nd` family in `Whnf.lean` (mirror Rust `whnf_no_delta_for_def_eq`).
Same dispatch tree as `whnf` (beta / let zeta / iota / proj / quot /
primitives all fire), except `whnf_nd_const_head`'s Defn arm falls
through to a stuck `apply_spine` instead of delta-unfolding.
* `k_infer_only` family in `Infer.lean` (mirror Rust `with_infer_only`).
App drops `k_check(a, dom)`; Lam drops `k_ensure_sort(ty)`; Let drops
the val/ty validations. Distinct Aiur memo from `k_infer`, parity
with Rust's separate `infer_cache` / `infer_only_cache`.
* `k_is_def_eq_struct_safe` in `DefEq.lean` (mirror Rust
`quick_def_eq`). Sort-Sort via `level_equal`; Lam-Lam / All-All
recurse on type and on body under `Cons(ty_a, types)`. Returns 1
only when DEFINITELY def-eq; 0 means fall through. Sound on
partially-whnf'd (no-delta) inputs because the handled shapes
don't depend on further reductions.
* `k_is_def_eq_core` Tier 1d wiring inserted between Tier 1c (string
lit) and Tier 2 (full whnf):
aw_nd = whnf_nd(a); bw_nd = whnf_nd(b)
ptr_eq(aw_nd, bw_nd) → 1
k_is_def_eq_struct_safe(aw_nd, bw_nd) → 1 if 1
try_lazy_delta_app(aw_nd, bw_nd) → 1 if 1 (rerun post-whnf_nd:
spine args may have reduced past what Tier 1.5's pre-whnf attempt
could see, exposing Const-Const congruence that was hidden)
* `try_proof_irrel`, `is_prop_type`, `try_unit_like` switch from
`k_infer` to `k_infer_only` — these helpers only need the synthesized
type, not the full re-validation work that `k_infer` does for each
recursive App/Let/Lam.
Each piece individually validated necessary (removing it puts UTF-8
back into the OOM regime). FVar variant + opens, KStore explicit
caches, infer_only's FVar-based binder opening — all confirmed NOT
necessary for the UTF-8 unblock and left out (see PLAN.md for future
experiments).
Measured (FFT cost):
Nat.add_comm: 56.08M → 55.63M (~stable; new code paths add no
overhead on the common case because Tier 1d's whnf_nd + struct_safe
are themselves Aiur-memoized).
_private.Init.Data.String.Decode.0.ByteArray.utf8DecodeChar?
.assemble₄_eq_some_of_toBitVec._proof_1_8: OOM → 39.12B FFT, passes.
3 files, +296/-6 lines.
`u64_sub_with_borrow` combines two per-byte borrow bits with `g_or`. The
two bits are MUTUALLY EXCLUSIVE: `u_t = borrow(a_i - b_i)` and
`u_r = borrow((a_i + 256 - b_i) - br_in)`. If `u_t = 1` the intermediate
`t_i ≥ 1`, so subtracting `br_in ∈ {0,1}` cannot underflow ⇒ `u_r = 0`.
Field `+` substitutes for `g_or` directly (per the same pattern as
`u64_add` in `ByteStream.lean`).
Per Aiur cost model, `g_or` adds +1 aux + 1 lookup per call; field `+`
is free. 8 g_or call sites in `u64_sub_with_borrow` each charged on every
one of the function's 2.23M rows.
Measured (FFT cost) on UTF-8 `_proof_1_8`:
39.12B → 38.14B (-2.6%)
Nat.add_comm unchanged (55.63M).
See [[reference_aiur_carry_add]].
Same mutually-exclusive-carry pattern as `u64_sub_with_borrow`:
* `klimbs_add_carry`: u64_add of (la, lb) yields carry1; u64_add of
(sum1, carry_in) yields carry2. carry1=1 ⇒ sum1 ≤ 2^64-2 ⇒
carry2=0.
* `klimbs_sub_borrow`: symmetric for borrows.
Replace `g_or(c1, c2)` with `c1 + c2` (field +). Both helpers run on
hot Nat-primitive paths.
Measured on UTF-8 `_proof_1_8`:
38.14B → 38.07B (-0.18%)
Nat.add_comm unchanged.
See [[reference_aiur_carry_add]].
`try_nat_dispatch` ran 1.12M rows in UTF-8 `_proof_1_8` at width 90,
charging 5.16% of total FFT. Width was floored by its widest match arm
(the binop branch with 2× whnf + 2× try_extract_nat + try_nat_binop_addr
+ apply_spine), even on Nat.succ / Nat.pred rows that never touched it.
Factor binop dispatch into its own `try_nat_binop_dispatch` fn. Main
dispatcher narrows to the max of succ / pred arms (single whnf +
try_extract_nat + klimbs_succ/dec + apply_spine). The cold fn's width
only charges the rows that actually dispatch a binop.
Measured on UTF-8 `_proof_1_8`:
38.07B → 37.80B (-0.7%)
Nat.add_comm unchanged.
`expr_lbr` ran 1.47M rows in UTF-8 `_proof_1_8` at width 39, charging
3.01% of total FFT. The Let arm (3 recursive expr_lbr calls + 2 lbr_max
+ 1 lbr_dec) is the widest match arm, charged on every row of expr_lbr
even though Let is rare in most expressions encountered.
Factor the Let arm into `expr_lbr_let(ty, val, body)`. Main expr_lbr
narrows to max of the 2-recursion arms (App / Lam / Forall). Cold fn
only charges Let-arm rows.
Measured:
Nat.add_comm: 55.63M → 55.50M (-0.2%)
UTF-8 `_proof_1_8`: 37.80B → 37.62B (-0.5%)
`try_extract_nat` ran 1.12M rows at width 45, charging 2.68% of UTF-8
`_proof_1_8` total FFT. The App arm (list_lookup + address_eq +
recursive try_extract_nat + klimbs_succ) is the widest match arm; the
Lit / Const / default arms are leaf compares.
Factor App into `try_extract_nat_app(f, a, addrs)`. Main extractor
narrows to leaf-arm width. Cold fn only charges App-arm rows.
Measured on UTF-8 `_proof_1_8`:
37.62B → 37.31B (-0.8%)
Nat.add_comm unchanged.
Updates 41 pinned FFT costs in `Tests/Ix/IxVM.lean::kernelCheckEntries`
to match the new kernel's output. All pins moved DOWN — every constant
got cheaper, none regressed.
Largest reductions (% change):
Vector.append: 4_023_268_168 → 3_160_970_390 (-21.4%)
Array.append_assoc: 3_938_574_533 → 3_079_334_815 (-21.8%)
String.Internal.append: 793_580_333 → 775_968_134 ( -2.2%)
bv_to_nat_lit: 635_780_327 → 619_870_154 ( -2.5%)
nat_gcd_lit: 665_518_356 → 649_859_784 ( -2.4%)
Nat.sub_le_of_le_add: 567_575_653 → 557_867_526 ( -1.7%)
IxVMPrim.nat_mod_lit: 414_695_549 → 407_517_834 ( -1.7%)
IxVMPrim.nat_div_lit: 405_607_545 → 398_641_590 ( -1.7%)
IxVMPrim.nat_shr_lit: 411_128_901 → 404_158_486 ( -1.7%)
Nat.decLe: 209_641_496 → 206_196_563 ( -1.6%)
Nat.add_comm: 56_084_908 → 55_504_714 ( -1.0%)
`lake test -- --ignored ixvm` passes with 0 FFT mismatches.
The `k_infer_only` section header skipped over the fact that the
function is only sound on well-typed inputs (since it drops
`k_check(a, dom)` on `App`, `k_ensure_sort(ty)` on `Lam`, val/ty
checks on `Let`). Spell out:
* The invariant — only call on terms produced by `whnf` of a
well-typed term, never on arbitrary inputs.
* Why the current sites (`try_proof_irrel`, `is_prop_type`,
`try_unit_like`) respect it.
* The planned non-deterministic-hint dispatch (`Hint::{None,
KInfer, KInferOnly}`) that lets us share `k_infer`'s memo where a
hit already exists instead of paying the parallel `infer_only`
memo cost.
@arthurpaulino
arthurpaulino merged commit ad7e383 into mainJun 25, 2026
14 checks passed
@arthurpaulino
arthurpaulino deleted the ap/utf8-tier-1d branch June 25, 2026 13:49
johnchandlerburnham pushed a commit that referenced this pull request Jul 21, 2026
…505)
* IxVM: drop dead KValNode/KVal/KValEnv
From ap/kernel 828fb85 (Arthur Paulino): the NbE value domain is defined
but referenced nowhere — the live kernel runs on de-Bruijn KExpr;
vestigial from an abandoned NbE direction. (The closed-term context
normalization from that commit was measured separately and not taken:
the per-call expr_lbr probe cost +2.9% FFT on recursor loops for a 0.5%
record reduction.)
* IxVM: memoized prim_family dispatch + width-safe offset-stuck placement
Three coordinated changes to Const-head whnf dispatch (cherry-pick of
130f30b, adapted to post-#450/#457 main):
1. prim_family(addr) classifies a head address into the one reducer
family that could fire on it (nat/str/bitvec/native/decidable; the
sets are disjoint). Keyed on the ADDRESS ALONE it memoizes to one row
per distinct constant address per run, and whnf_const_head calls at
most one family reducer. The previous gauntlet ran every reducer in
sequence for guaranteed misses.
2. The symbolic-Nat offset-stuck check moves from a delta-arm probe into
try_nat_dispatch's miss path as a cold function
(try_nat_offset_dispatch, verdict 2 = "already stuck, do not
re-whnf"), with the offset construction shared via
mk_nat_offset_stuck (also used by the linear-rec collapse).
3. nat_lit_to_ctor_or_self exposes ONE constructor layer
(n -> succ(Lit(n-1))) instead of materializing the full succ chain.
Adaptations vs the original patch:
- whnf_nd_const_head (no-delta WHNF, added on main after the patch)
converted to the same family dispatch.
- Kept main's cold-extracted try_nat_binop_dispatch and routed the
symbolic-base case to try_nat_offset_dispatch from its miss arm.
Measured (lake exe ix check Nat.add_comm): total width 34820 -> 34806,
FFT cost 49571210 -> 48860647 (-1.43%). All 53 ixvm-suite FFT pins
decreased (-0.19%..-1.43%); parity and claim smokes pass. Pins updated;
crates/ixvm-codegen/src/aiur_ixvm.rs regenerated via `lake exe ix
codegen`.
* IxVM: port jcb/fixes H-14 — ptr_val skip map + lockstep addr cursor
Two quadratic/constant-factor fixes to check_all_skipping, ported from
jcb/fixes (3763356, John C. Burnham); cherry-pick of 2d83be1 adapted to
post-#457 main:
- The assumption-leaf skip set keys on ptr_val instead of the first 4
address bytes: one tree lookup per constant, no per-lookup address
load, no confirming address_eq. Sound by the build_addr_pos_map
interning invariant (one pointer, one content): a ptr hit implies the
address IS a leaf; a de-interned pointer reads as absent and the
constant just gets checked — fail-closed.
- The iterator walks the addrs list in LOCKSTEP with consts (cur_addrs
suffix) instead of list_lookup(addrs, pos) per constant, which
re-walked the prefix every iteration — a standalone O(closure^2).
Adaptations: kept main's two-arg check_canonical_block_sort call;
addr_key retained (the Inductive.lean block-membership table added on
main after this patch still uses it), comment updated.
Measured (lake exe ix check Nat.add_comm): total width 34806 -> 34781,
FFT cost unchanged (plain checks never take the skip path). Full ixvm
suite green incl. the frontier-assumption claim smoke; all FFT pins
unchanged. crates/ixvm-codegen/src/aiur_ixvm.rs regenerated.
---------
Co-authored-by: samuelburnham <45365069+samuelburnham@users.noreply.github.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@arthurpaulino@gabriel-barrett
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

IxVM kernel: unblock UTF-8 decode/encode proof + Nat-layer FFT cuts - #450

Merged
arthurpaulino merged 8 commits into
mainfrom
ap/utf8-tier-1d
Jun 25, 2026
Merged

IxVM kernel: unblock UTF-8 decode/encode proof + Nat-layer FFT cuts#450
arthurpaulino merged 8 commits into
mainfrom
ap/utf8-tier-1d

Conversation

@arthurpaulino

Copy link
Copy Markdown
Member

Headline

Aiur kernel previously OOM'd on
_private.Init.Data.String.Decode.0.ByteArray.utf8DecodeChar?.assemble₄_eq_some_of_toBitVec._proof_1_8.
This constant is a prerequisite of ByteArray.utf8DecodeChar?_utf8EncodeChar_append
(part of the UTF-8 round-trip lemma chain).

Now it typechecks at 37.31B FFT (down from the OOM baseline). The
two-direction lemma …_utf8EncodeChar_append itself remains
out-of-reach (still OOMs further along the same dependency tree under
the user's memory cap), but the prerequisite that was the immediate
blocker is unstuck.

The same kernel changes also deliver large FFT cuts on previously-
expensive targets — Tier 1d's structural short-circuit replaces full
delta-whnf cascades wherever it fires:

ConstantBeforeAfterΔ
Vector.append4.02B3.16B−21.4%
Array.append_assoc3.94B3.08B−21.8%

Commits

1. a6cd34a — Tier 1d def-eq short-circuit (the unblock)

Aiur's k_is_def_eq_core jumped from Tier 1.5 straight to full delta
WHNF (Tier 2). The cascading Nat.rec / Nat.succ iota expansions then
drove whnf_const_head past 1M unique entries before either side
reached a comparable canonical form — OOM. Rust's def-eq settles the
same pair via no-delta whnf + quick structural recursion before any
of that fires.

Three minimum-necessary pieces ported. No KStore. No FVar. No Subst /
KernelTypes change.

  • whnf_nd family (Whnf.lean, mirror Rust
    whnf_no_delta_for_def_eq). Same dispatch tree as whnf, but
    whnf_nd_const_head's Defn arm falls through to a stuck
    apply_spine instead of delta-unfolding. Iota / proj / quot /
    primitives still fire.

  • k_infer_only family (Infer.lean, mirror Rust
    with_infer_only). App drops k_check(a, dom); Lam drops
    k_ensure_sort(ty); Let drops val/ty validation. Used at
    try_proof_irrel, is_prop_type, try_unit_like — the def-eq
    tactics that only need the synthesized type.

  • k_is_def_eq_struct_safe + Tier 1d wiring (DefEq.lean,
    mirror Rust quick_def_eq + post-try_def_eq_app). Sort-Sort via
    level_equal; Lam-Lam / All-All via recursive k_is_def_eq on the
    type and on the body under Cons(ty_a, types) (types-cons, NOT
    FVar opening). Inserted between Tier 1c (string lit) and Tier 2
    (full whnf):

    aw_nd = whnf_nd(a); bw_nd = whnf_nd(b)
    ptr_eq(aw_nd, bw_nd) → 1
    k_is_def_eq_struct_safe(aw_nd, bw_nd) → 1
    try_lazy_delta_app(aw_nd, bw_nd) → 1 (rerun: spine args may have
    reduced past what Tier 1.5's pre-whnf attempt could see)
    

Each piece independently validated necessary. The previously-tried
KStore explicit caches, FVar variant + opens, FVar-based binder
opening — all confirmed NOT necessary for this unblock and left out.

3 files changed, +296/−6 lines.

2. 36d6c3a — drop g_or from u64_sub_with_borrow

u64_sub_with_borrow combined two per-byte borrow bits with g_or.
The two bits are mutually exclusive: u_t = 1 ⇒ intermediate t_i ≥ 1 ⇒ subtracting br_in ∈ {0,1} cannot underflow ⇒ u_r = 0. Field
+ substitutes for g_or directly. Per Aiur cost model g_or adds
+1 aux + 1 lookup per call (≈ 5 width); field + is free. 7 g_ors ×
2.23M rows.

UTF-8 _proof_1_8: 39.12B → 38.14B (−2.6%).

3. 9e787da — drop g_or from klimbs_add_carry / klimbs_sub_borrow

Same mutually-exclusive-carry pattern. Two limb-level borrows /
carries from sequential u64 ops cannot both be 1.

UTF-8 _proof_1_8: 38.14B → 38.07B (−0.18%).

4. 4e379e7 — hot/cold split try_nat_dispatch, extract binop arm

try_nat_dispatch's width 90 was floored by the binop arm (2× whnf

  • 2× try_extract_nat + try_nat_binop_addr + apply_spine), charged on
    every Nat.succ / Nat.pred row. Factor binop dispatch into
    try_nat_binop_dispatch. Main narrows to the max of succ / pred
    arms.

UTF-8 _proof_1_8: 38.07B → 37.80B (−0.7%).

5. 80ce3d2 — hot/cold split expr_lbr, extract Let arm

expr_lbr's width was floored by the Let arm (3 recursive expr_lbr
calls + 2 lbr_max + 1 lbr_dec), charged on every row even though Let
is rare. Factor into expr_lbr_let.

Nat.add_comm: 55.63M → 55.50M (−0.2%). UTF-8 _proof_1_8: 37.80B
→ 37.62B (−0.5%)
.

6. 1f8effd — hot/cold split try_extract_nat, extract App arm

try_extract_nat's width was floored by the App arm (list_lookup +
address_eq + recursive try_extract_nat + klimbs_succ). Factor into
try_extract_nat_app. Main narrows to leaf-arm width.

UTF-8 _proof_1_8: 37.62B → 37.31B (−0.8%).

7. 039e9cf — re-pin IxVM FFT costs

41 pins in Tests/Ix/IxVM.lean::kernelCheckEntries updated. Every
constant got cheaper; none regressed. Largest reductions:

  • Vector.append: 4.02B → 3.16B (−21.4%)
  • Array.append_assoc: 3.94B → 3.08B (−21.8%)

lake test -- --ignored ixvm passes with 0 FFT mismatches.

Cumulative on UTF-8 _proof_1_8

CommitFFT
baseline (main)OOM
a6cd34a Tier 1d39.12B
36d6c3a g_or → + in u64_sub_with_borrow38.14B (−2.6%)
9e787da g_or → + in klimbs_add_carry / klimbs_sub_borrow38.07B (−0.18%)
4e379e7 hot/cold try_nat_dispatch37.80B (−0.7%)
80ce3d2 hot/cold expr_lbr37.62B (−0.5%)
1f8effd hot/cold try_extract_nat37.31B (−0.8%)

Post-unlock optimization: −4.6% (39.12B → 37.31B).

Cost on small targets

  • Main baseline Nat.add_comm: 56.08M FFT.
  • This branch Nat.add_comm: 55.50M FFT.

Tier 1d itself adds no overhead on the common case (whnf_nd +
struct_safe + try_lazy_delta_app are themselves Aiur-memoized); the
follow-up optimizations are net wins.

Test plan

  • lake exe check Nat.add_comm passes (55.50M FFT).
  • lake exe check Vector.extract_append passes.
  • lake exe check "_private.…assemble₄_eq_some_of_toBitVec._proof_1_8"
    passes (37.31B FFT, previously OOM).
  • lake test -- --ignored ixvm — all 41 FFT pins updated, suite
    passes with 0 mismatches.

Comment threadIx/IxVM/Kernel/Infer.lean
…ost-spine-congruence)
UTF-8 `_private.Init.Data.String.Decode.0.ByteArray.utf8DecodeChar?
.assemble₄_eq_some_of_toBitVec._proof_1_8` OOMs on the previous
pipeline because `k_is_def_eq_core` jumps from Tier 1.5 straight to
full delta WHNF (Tier 2); the cascading Nat.rec / Nat.succ iota
expansions then drive `whnf_const_head` past 1M unique entries before
either side reaches a comparable canonical form. Rust's def-eq settles
the same pair via the no-delta whnf + quick structural recursion
before any of that fires.
This patch ports the three pieces of that short-circuit and nothing
else — no FVar variant, no KStore, no Subst changes, no signature
sweep.
* `whnf_nd` family in `Whnf.lean` (mirror Rust `whnf_no_delta_for_def_eq`).
Same dispatch tree as `whnf` (beta / let zeta / iota / proj / quot /
primitives all fire), except `whnf_nd_const_head`'s Defn arm falls
through to a stuck `apply_spine` instead of delta-unfolding.
* `k_infer_only` family in `Infer.lean` (mirror Rust `with_infer_only`).
App drops `k_check(a, dom)`; Lam drops `k_ensure_sort(ty)`; Let drops
the val/ty validations. Distinct Aiur memo from `k_infer`, parity
with Rust's separate `infer_cache` / `infer_only_cache`.
* `k_is_def_eq_struct_safe` in `DefEq.lean` (mirror Rust
`quick_def_eq`). Sort-Sort via `level_equal`; Lam-Lam / All-All
recurse on type and on body under `Cons(ty_a, types)`. Returns 1
only when DEFINITELY def-eq; 0 means fall through. Sound on
partially-whnf'd (no-delta) inputs because the handled shapes
don't depend on further reductions.
* `k_is_def_eq_core` Tier 1d wiring inserted between Tier 1c (string
lit) and Tier 2 (full whnf):
aw_nd = whnf_nd(a); bw_nd = whnf_nd(b)
ptr_eq(aw_nd, bw_nd) → 1
k_is_def_eq_struct_safe(aw_nd, bw_nd) → 1 if 1
try_lazy_delta_app(aw_nd, bw_nd) → 1 if 1 (rerun post-whnf_nd:
spine args may have reduced past what Tier 1.5's pre-whnf attempt
could see, exposing Const-Const congruence that was hidden)
* `try_proof_irrel`, `is_prop_type`, `try_unit_like` switch from
`k_infer` to `k_infer_only` — these helpers only need the synthesized
type, not the full re-validation work that `k_infer` does for each
recursive App/Let/Lam.
Each piece individually validated necessary (removing it puts UTF-8
back into the OOM regime). FVar variant + opens, KStore explicit
caches, infer_only's FVar-based binder opening — all confirmed NOT
necessary for the UTF-8 unblock and left out (see PLAN.md for future
experiments).
Measured (FFT cost):
Nat.add_comm: 56.08M → 55.63M (~stable; new code paths add no
overhead on the common case because Tier 1d's whnf_nd + struct_safe
are themselves Aiur-memoized).
_private.Init.Data.String.Decode.0.ByteArray.utf8DecodeChar?
.assemble₄_eq_some_of_toBitVec._proof_1_8: OOM → 39.12B FFT, passes.
3 files, +296/-6 lines.
`u64_sub_with_borrow` combines two per-byte borrow bits with `g_or`. The
two bits are MUTUALLY EXCLUSIVE: `u_t = borrow(a_i - b_i)` and
`u_r = borrow((a_i + 256 - b_i) - br_in)`. If `u_t = 1` the intermediate
`t_i ≥ 1`, so subtracting `br_in ∈ {0,1}` cannot underflow ⇒ `u_r = 0`.
Field `+` substitutes for `g_or` directly (per the same pattern as
`u64_add` in `ByteStream.lean`).
Per Aiur cost model, `g_or` adds +1 aux + 1 lookup per call; field `+`
is free. 8 g_or call sites in `u64_sub_with_borrow` each charged on every
one of the function's 2.23M rows.
Measured (FFT cost) on UTF-8 `_proof_1_8`:
39.12B → 38.14B (-2.6%)
Nat.add_comm unchanged (55.63M).
See [[reference_aiur_carry_add]].
Same mutually-exclusive-carry pattern as `u64_sub_with_borrow`:
* `klimbs_add_carry`: u64_add of (la, lb) yields carry1; u64_add of
(sum1, carry_in) yields carry2. carry1=1 ⇒ sum1 ≤ 2^64-2 ⇒
carry2=0.
* `klimbs_sub_borrow`: symmetric for borrows.
Replace `g_or(c1, c2)` with `c1 + c2` (field +). Both helpers run on
hot Nat-primitive paths.
Measured on UTF-8 `_proof_1_8`:
38.14B → 38.07B (-0.18%)
Nat.add_comm unchanged.
See [[reference_aiur_carry_add]].
`try_nat_dispatch` ran 1.12M rows in UTF-8 `_proof_1_8` at width 90,
charging 5.16% of total FFT. Width was floored by its widest match arm
(the binop branch with 2× whnf + 2× try_extract_nat + try_nat_binop_addr
+ apply_spine), even on Nat.succ / Nat.pred rows that never touched it.
Factor binop dispatch into its own `try_nat_binop_dispatch` fn. Main
dispatcher narrows to the max of succ / pred arms (single whnf +
try_extract_nat + klimbs_succ/dec + apply_spine). The cold fn's width
only charges the rows that actually dispatch a binop.
Measured on UTF-8 `_proof_1_8`:
38.07B → 37.80B (-0.7%)
Nat.add_comm unchanged.
`expr_lbr` ran 1.47M rows in UTF-8 `_proof_1_8` at width 39, charging
3.01% of total FFT. The Let arm (3 recursive expr_lbr calls + 2 lbr_max
+ 1 lbr_dec) is the widest match arm, charged on every row of expr_lbr
even though Let is rare in most expressions encountered.
Factor the Let arm into `expr_lbr_let(ty, val, body)`. Main expr_lbr
narrows to max of the 2-recursion arms (App / Lam / Forall). Cold fn
only charges Let-arm rows.
Measured:
Nat.add_comm: 55.63M → 55.50M (-0.2%)
UTF-8 `_proof_1_8`: 37.80B → 37.62B (-0.5%)
`try_extract_nat` ran 1.12M rows at width 45, charging 2.68% of UTF-8
`_proof_1_8` total FFT. The App arm (list_lookup + address_eq +
recursive try_extract_nat + klimbs_succ) is the widest match arm; the
Lit / Const / default arms are leaf compares.
Factor App into `try_extract_nat_app(f, a, addrs)`. Main extractor
narrows to leaf-arm width. Cold fn only charges App-arm rows.
Measured on UTF-8 `_proof_1_8`:
37.62B → 37.31B (-0.8%)
Nat.add_comm unchanged.
Updates 41 pinned FFT costs in `Tests/Ix/IxVM.lean::kernelCheckEntries`
to match the new kernel's output. All pins moved DOWN — every constant
got cheaper, none regressed.
Largest reductions (% change):
Vector.append: 4_023_268_168 → 3_160_970_390 (-21.4%)
Array.append_assoc: 3_938_574_533 → 3_079_334_815 (-21.8%)
String.Internal.append: 793_580_333 → 775_968_134 ( -2.2%)
bv_to_nat_lit: 635_780_327 → 619_870_154 ( -2.5%)
nat_gcd_lit: 665_518_356 → 649_859_784 ( -2.4%)
Nat.sub_le_of_le_add: 567_575_653 → 557_867_526 ( -1.7%)
IxVMPrim.nat_mod_lit: 414_695_549 → 407_517_834 ( -1.7%)
IxVMPrim.nat_div_lit: 405_607_545 → 398_641_590 ( -1.7%)
IxVMPrim.nat_shr_lit: 411_128_901 → 404_158_486 ( -1.7%)
Nat.decLe: 209_641_496 → 206_196_563 ( -1.6%)
Nat.add_comm: 56_084_908 → 55_504_714 ( -1.0%)
`lake test -- --ignored ixvm` passes with 0 FFT mismatches.
The `k_infer_only` section header skipped over the fact that the
function is only sound on well-typed inputs (since it drops
`k_check(a, dom)` on `App`, `k_ensure_sort(ty)` on `Lam`, val/ty
checks on `Let`). Spell out:
* The invariant — only call on terms produced by `whnf` of a
well-typed term, never on arbitrary inputs.
* Why the current sites (`try_proof_irrel`, `is_prop_type`,
`try_unit_like`) respect it.
* The planned non-deterministic-hint dispatch (`Hint::{None,
KInfer, KInferOnly}`) that lets us share `k_infer`'s memo where a
hit already exists instead of paying the parallel `infer_only`
memo cost.
@arthurpaulino
arthurpaulino merged commit ad7e383 into mainJun 25, 2026
14 checks passed
@arthurpaulino
arthurpaulino deleted the ap/utf8-tier-1d branch June 25, 2026 13:49
johnchandlerburnham pushed a commit that referenced this pull request Jul 21, 2026
…505)
* IxVM: drop dead KValNode/KVal/KValEnv
From ap/kernel 828fb85 (Arthur Paulino): the NbE value domain is defined
but referenced nowhere — the live kernel runs on de-Bruijn KExpr;
vestigial from an abandoned NbE direction. (The closed-term context
normalization from that commit was measured separately and not taken:
the per-call expr_lbr probe cost +2.9% FFT on recursor loops for a 0.5%
record reduction.)
* IxVM: memoized prim_family dispatch + width-safe offset-stuck placement
Three coordinated changes to Const-head whnf dispatch (cherry-pick of
130f30b, adapted to post-#450/#457 main):
1. prim_family(addr) classifies a head address into the one reducer
family that could fire on it (nat/str/bitvec/native/decidable; the
sets are disjoint). Keyed on the ADDRESS ALONE it memoizes to one row
per distinct constant address per run, and whnf_const_head calls at
most one family reducer. The previous gauntlet ran every reducer in
sequence for guaranteed misses.
2. The symbolic-Nat offset-stuck check moves from a delta-arm probe into
try_nat_dispatch's miss path as a cold function
(try_nat_offset_dispatch, verdict 2 = "already stuck, do not
re-whnf"), with the offset construction shared via
mk_nat_offset_stuck (also used by the linear-rec collapse).
3. nat_lit_to_ctor_or_self exposes ONE constructor layer
(n -> succ(Lit(n-1))) instead of materializing the full succ chain.
Adaptations vs the original patch:
- whnf_nd_const_head (no-delta WHNF, added on main after the patch)
converted to the same family dispatch.
- Kept main's cold-extracted try_nat_binop_dispatch and routed the
symbolic-base case to try_nat_offset_dispatch from its miss arm.
Measured (lake exe ix check Nat.add_comm): total width 34820 -> 34806,
FFT cost 49571210 -> 48860647 (-1.43%). All 53 ixvm-suite FFT pins
decreased (-0.19%..-1.43%); parity and claim smokes pass. Pins updated;
crates/ixvm-codegen/src/aiur_ixvm.rs regenerated via `lake exe ix
codegen`.
* IxVM: port jcb/fixes H-14 — ptr_val skip map + lockstep addr cursor
Two quadratic/constant-factor fixes to check_all_skipping, ported from
jcb/fixes (3763356, John C. Burnham); cherry-pick of 2d83be1 adapted to
post-#457 main:
- The assumption-leaf skip set keys on ptr_val instead of the first 4
address bytes: one tree lookup per constant, no per-lookup address
load, no confirming address_eq. Sound by the build_addr_pos_map
interning invariant (one pointer, one content): a ptr hit implies the
address IS a leaf; a de-interned pointer reads as absent and the
constant just gets checked — fail-closed.
- The iterator walks the addrs list in LOCKSTEP with consts (cur_addrs
suffix) instead of list_lookup(addrs, pos) per constant, which
re-walked the prefix every iteration — a standalone O(closure^2).
Adaptations: kept main's two-arg check_canonical_block_sort call;
addr_key retained (the Inductive.lean block-membership table added on
main after this patch still uses it), comment updated.
Measured (lake exe ix check Nat.add_comm): total width 34806 -> 34781,
FFT cost unchanged (plain checks never take the skip path). Full ixvm
suite green incl. the frontier-assumption claim smoke; all FFT pins
unchanged. crates/ixvm-codegen/src/aiur_ixvm.rs regenerated.
---------
Co-authored-by: samuelburnham <45365069+samuelburnham@users.noreply.github.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@arthurpaulino@gabriel-barrett
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

IxVM kernel: unblock UTF-8 decode/encode proof + Nat-layer FFT cuts - #450

Merged
arthurpaulino merged 8 commits into
mainfrom
ap/utf8-tier-1d
Jun 25, 2026
Merged

IxVM kernel: unblock UTF-8 decode/encode proof + Nat-layer FFT cuts#450
arthurpaulino merged 8 commits into
mainfrom
ap/utf8-tier-1d

Conversation

@arthurpaulino

Copy link
Copy Markdown
Member

Headline

Aiur kernel previously OOM'd on
_private.Init.Data.String.Decode.0.ByteArray.utf8DecodeChar?.assemble₄_eq_some_of_toBitVec._proof_1_8.
This constant is a prerequisite of ByteArray.utf8DecodeChar?_utf8EncodeChar_append
(part of the UTF-8 round-trip lemma chain).

Now it typechecks at 37.31B FFT (down from the OOM baseline). The
two-direction lemma …_utf8EncodeChar_append itself remains
out-of-reach (still OOMs further along the same dependency tree under
the user's memory cap), but the prerequisite that was the immediate
blocker is unstuck.

The same kernel changes also deliver large FFT cuts on previously-
expensive targets — Tier 1d's structural short-circuit replaces full
delta-whnf cascades wherever it fires:

ConstantBeforeAfterΔ
Vector.append4.02B3.16B−21.4%
Array.append_assoc3.94B3.08B−21.8%

Commits

1. a6cd34a — Tier 1d def-eq short-circuit (the unblock)

Aiur's k_is_def_eq_core jumped from Tier 1.5 straight to full delta
WHNF (Tier 2). The cascading Nat.rec / Nat.succ iota expansions then
drove whnf_const_head past 1M unique entries before either side
reached a comparable canonical form — OOM. Rust's def-eq settles the
same pair via no-delta whnf + quick structural recursion before any
of that fires.

Three minimum-necessary pieces ported. No KStore. No FVar. No Subst /
KernelTypes change.

  • whnf_nd family (Whnf.lean, mirror Rust
    whnf_no_delta_for_def_eq). Same dispatch tree as whnf, but
    whnf_nd_const_head's Defn arm falls through to a stuck
    apply_spine instead of delta-unfolding. Iota / proj / quot /
    primitives still fire.

  • k_infer_only family (Infer.lean, mirror Rust
    with_infer_only). App drops k_check(a, dom); Lam drops
    k_ensure_sort(ty); Let drops val/ty validation. Used at
    try_proof_irrel, is_prop_type, try_unit_like — the def-eq
    tactics that only need the synthesized type.

  • k_is_def_eq_struct_safe + Tier 1d wiring (DefEq.lean,
    mirror Rust quick_def_eq + post-try_def_eq_app). Sort-Sort via
    level_equal; Lam-Lam / All-All via recursive k_is_def_eq on the
    type and on the body under Cons(ty_a, types) (types-cons, NOT
    FVar opening). Inserted between Tier 1c (string lit) and Tier 2
    (full whnf):

    aw_nd = whnf_nd(a); bw_nd = whnf_nd(b)
    ptr_eq(aw_nd, bw_nd) → 1
    k_is_def_eq_struct_safe(aw_nd, bw_nd) → 1
    try_lazy_delta_app(aw_nd, bw_nd) → 1 (rerun: spine args may have
    reduced past what Tier 1.5's pre-whnf attempt could see)
    

Each piece independently validated necessary. The previously-tried
KStore explicit caches, FVar variant + opens, FVar-based binder
opening — all confirmed NOT necessary for this unblock and left out.

3 files changed, +296/−6 lines.

2. 36d6c3a — drop g_or from u64_sub_with_borrow

u64_sub_with_borrow combined two per-byte borrow bits with g_or.
The two bits are mutually exclusive: u_t = 1 ⇒ intermediate t_i ≥ 1 ⇒ subtracting br_in ∈ {0,1} cannot underflow ⇒ u_r = 0. Field
+ substitutes for g_or directly. Per Aiur cost model g_or adds
+1 aux + 1 lookup per call (≈ 5 width); field + is free. 7 g_ors ×
2.23M rows.

UTF-8 _proof_1_8: 39.12B → 38.14B (−2.6%).

3. 9e787da — drop g_or from klimbs_add_carry / klimbs_sub_borrow

Same mutually-exclusive-carry pattern. Two limb-level borrows /
carries from sequential u64 ops cannot both be 1.

UTF-8 _proof_1_8: 38.14B → 38.07B (−0.18%).

4. 4e379e7 — hot/cold split try_nat_dispatch, extract binop arm

try_nat_dispatch's width 90 was floored by the binop arm (2× whnf

  • 2× try_extract_nat + try_nat_binop_addr + apply_spine), charged on
    every Nat.succ / Nat.pred row. Factor binop dispatch into
    try_nat_binop_dispatch. Main narrows to the max of succ / pred
    arms.

UTF-8 _proof_1_8: 38.07B → 37.80B (−0.7%).

5. 80ce3d2 — hot/cold split expr_lbr, extract Let arm

expr_lbr's width was floored by the Let arm (3 recursive expr_lbr
calls + 2 lbr_max + 1 lbr_dec), charged on every row even though Let
is rare. Factor into expr_lbr_let.

Nat.add_comm: 55.63M → 55.50M (−0.2%). UTF-8 _proof_1_8: 37.80B
→ 37.62B (−0.5%)
.

6. 1f8effd — hot/cold split try_extract_nat, extract App arm

try_extract_nat's width was floored by the App arm (list_lookup +
address_eq + recursive try_extract_nat + klimbs_succ). Factor into
try_extract_nat_app. Main narrows to leaf-arm width.

UTF-8 _proof_1_8: 37.62B → 37.31B (−0.8%).

7. 039e9cf — re-pin IxVM FFT costs

41 pins in Tests/Ix/IxVM.lean::kernelCheckEntries updated. Every
constant got cheaper; none regressed. Largest reductions:

  • Vector.append: 4.02B → 3.16B (−21.4%)
  • Array.append_assoc: 3.94B → 3.08B (−21.8%)

lake test -- --ignored ixvm passes with 0 FFT mismatches.

Cumulative on UTF-8 _proof_1_8

CommitFFT
baseline (main)OOM
a6cd34a Tier 1d39.12B
36d6c3a g_or → + in u64_sub_with_borrow38.14B (−2.6%)
9e787da g_or → + in klimbs_add_carry / klimbs_sub_borrow38.07B (−0.18%)
4e379e7 hot/cold try_nat_dispatch37.80B (−0.7%)
80ce3d2 hot/cold expr_lbr37.62B (−0.5%)
1f8effd hot/cold try_extract_nat37.31B (−0.8%)

Post-unlock optimization: −4.6% (39.12B → 37.31B).

Cost on small targets

  • Main baseline Nat.add_comm: 56.08M FFT.
  • This branch Nat.add_comm: 55.50M FFT.

Tier 1d itself adds no overhead on the common case (whnf_nd +
struct_safe + try_lazy_delta_app are themselves Aiur-memoized); the
follow-up optimizations are net wins.

Test plan

  • lake exe check Nat.add_comm passes (55.50M FFT).
  • lake exe check Vector.extract_append passes.
  • lake exe check "_private.…assemble₄_eq_some_of_toBitVec._proof_1_8"
    passes (37.31B FFT, previously OOM).
  • lake test -- --ignored ixvm — all 41 FFT pins updated, suite
    passes with 0 mismatches.

Comment threadIx/IxVM/Kernel/Infer.lean
…ost-spine-congruence)
UTF-8 `_private.Init.Data.String.Decode.0.ByteArray.utf8DecodeChar?
.assemble₄_eq_some_of_toBitVec._proof_1_8` OOMs on the previous
pipeline because `k_is_def_eq_core` jumps from Tier 1.5 straight to
full delta WHNF (Tier 2); the cascading Nat.rec / Nat.succ iota
expansions then drive `whnf_const_head` past 1M unique entries before
either side reaches a comparable canonical form. Rust's def-eq settles
the same pair via the no-delta whnf + quick structural recursion
before any of that fires.
This patch ports the three pieces of that short-circuit and nothing
else — no FVar variant, no KStore, no Subst changes, no signature
sweep.
* `whnf_nd` family in `Whnf.lean` (mirror Rust `whnf_no_delta_for_def_eq`).
Same dispatch tree as `whnf` (beta / let zeta / iota / proj / quot /
primitives all fire), except `whnf_nd_const_head`'s Defn arm falls
through to a stuck `apply_spine` instead of delta-unfolding.
* `k_infer_only` family in `Infer.lean` (mirror Rust `with_infer_only`).
App drops `k_check(a, dom)`; Lam drops `k_ensure_sort(ty)`; Let drops
the val/ty validations. Distinct Aiur memo from `k_infer`, parity
with Rust's separate `infer_cache` / `infer_only_cache`.
* `k_is_def_eq_struct_safe` in `DefEq.lean` (mirror Rust
`quick_def_eq`). Sort-Sort via `level_equal`; Lam-Lam / All-All
recurse on type and on body under `Cons(ty_a, types)`. Returns 1
only when DEFINITELY def-eq; 0 means fall through. Sound on
partially-whnf'd (no-delta) inputs because the handled shapes
don't depend on further reductions.
* `k_is_def_eq_core` Tier 1d wiring inserted between Tier 1c (string
lit) and Tier 2 (full whnf):
aw_nd = whnf_nd(a); bw_nd = whnf_nd(b)
ptr_eq(aw_nd, bw_nd) → 1
k_is_def_eq_struct_safe(aw_nd, bw_nd) → 1 if 1
try_lazy_delta_app(aw_nd, bw_nd) → 1 if 1 (rerun post-whnf_nd:
spine args may have reduced past what Tier 1.5's pre-whnf attempt
could see, exposing Const-Const congruence that was hidden)
* `try_proof_irrel`, `is_prop_type`, `try_unit_like` switch from
`k_infer` to `k_infer_only` — these helpers only need the synthesized
type, not the full re-validation work that `k_infer` does for each
recursive App/Let/Lam.
Each piece individually validated necessary (removing it puts UTF-8
back into the OOM regime). FVar variant + opens, KStore explicit
caches, infer_only's FVar-based binder opening — all confirmed NOT
necessary for the UTF-8 unblock and left out (see PLAN.md for future
experiments).
Measured (FFT cost):
Nat.add_comm: 56.08M → 55.63M (~stable; new code paths add no
overhead on the common case because Tier 1d's whnf_nd + struct_safe
are themselves Aiur-memoized).
_private.Init.Data.String.Decode.0.ByteArray.utf8DecodeChar?
.assemble₄_eq_some_of_toBitVec._proof_1_8: OOM → 39.12B FFT, passes.
3 files, +296/-6 lines.
`u64_sub_with_borrow` combines two per-byte borrow bits with `g_or`. The
two bits are MUTUALLY EXCLUSIVE: `u_t = borrow(a_i - b_i)` and
`u_r = borrow((a_i + 256 - b_i) - br_in)`. If `u_t = 1` the intermediate
`t_i ≥ 1`, so subtracting `br_in ∈ {0,1}` cannot underflow ⇒ `u_r = 0`.
Field `+` substitutes for `g_or` directly (per the same pattern as
`u64_add` in `ByteStream.lean`).
Per Aiur cost model, `g_or` adds +1 aux + 1 lookup per call; field `+`
is free. 8 g_or call sites in `u64_sub_with_borrow` each charged on every
one of the function's 2.23M rows.
Measured (FFT cost) on UTF-8 `_proof_1_8`:
39.12B → 38.14B (-2.6%)
Nat.add_comm unchanged (55.63M).
See [[reference_aiur_carry_add]].
Same mutually-exclusive-carry pattern as `u64_sub_with_borrow`:
* `klimbs_add_carry`: u64_add of (la, lb) yields carry1; u64_add of
(sum1, carry_in) yields carry2. carry1=1 ⇒ sum1 ≤ 2^64-2 ⇒
carry2=0.
* `klimbs_sub_borrow`: symmetric for borrows.
Replace `g_or(c1, c2)` with `c1 + c2` (field +). Both helpers run on
hot Nat-primitive paths.
Measured on UTF-8 `_proof_1_8`:
38.14B → 38.07B (-0.18%)
Nat.add_comm unchanged.
See [[reference_aiur_carry_add]].
`try_nat_dispatch` ran 1.12M rows in UTF-8 `_proof_1_8` at width 90,
charging 5.16% of total FFT. Width was floored by its widest match arm
(the binop branch with 2× whnf + 2× try_extract_nat + try_nat_binop_addr
+ apply_spine), even on Nat.succ / Nat.pred rows that never touched it.
Factor binop dispatch into its own `try_nat_binop_dispatch` fn. Main
dispatcher narrows to the max of succ / pred arms (single whnf +
try_extract_nat + klimbs_succ/dec + apply_spine). The cold fn's width
only charges the rows that actually dispatch a binop.
Measured on UTF-8 `_proof_1_8`:
38.07B → 37.80B (-0.7%)
Nat.add_comm unchanged.
`expr_lbr` ran 1.47M rows in UTF-8 `_proof_1_8` at width 39, charging
3.01% of total FFT. The Let arm (3 recursive expr_lbr calls + 2 lbr_max
+ 1 lbr_dec) is the widest match arm, charged on every row of expr_lbr
even though Let is rare in most expressions encountered.
Factor the Let arm into `expr_lbr_let(ty, val, body)`. Main expr_lbr
narrows to max of the 2-recursion arms (App / Lam / Forall). Cold fn
only charges Let-arm rows.
Measured:
Nat.add_comm: 55.63M → 55.50M (-0.2%)
UTF-8 `_proof_1_8`: 37.80B → 37.62B (-0.5%)
`try_extract_nat` ran 1.12M rows at width 45, charging 2.68% of UTF-8
`_proof_1_8` total FFT. The App arm (list_lookup + address_eq +
recursive try_extract_nat + klimbs_succ) is the widest match arm; the
Lit / Const / default arms are leaf compares.
Factor App into `try_extract_nat_app(f, a, addrs)`. Main extractor
narrows to leaf-arm width. Cold fn only charges App-arm rows.
Measured on UTF-8 `_proof_1_8`:
37.62B → 37.31B (-0.8%)
Nat.add_comm unchanged.
Updates 41 pinned FFT costs in `Tests/Ix/IxVM.lean::kernelCheckEntries`
to match the new kernel's output. All pins moved DOWN — every constant
got cheaper, none regressed.
Largest reductions (% change):
Vector.append: 4_023_268_168 → 3_160_970_390 (-21.4%)
Array.append_assoc: 3_938_574_533 → 3_079_334_815 (-21.8%)
String.Internal.append: 793_580_333 → 775_968_134 ( -2.2%)
bv_to_nat_lit: 635_780_327 → 619_870_154 ( -2.5%)
nat_gcd_lit: 665_518_356 → 649_859_784 ( -2.4%)
Nat.sub_le_of_le_add: 567_575_653 → 557_867_526 ( -1.7%)
IxVMPrim.nat_mod_lit: 414_695_549 → 407_517_834 ( -1.7%)
IxVMPrim.nat_div_lit: 405_607_545 → 398_641_590 ( -1.7%)
IxVMPrim.nat_shr_lit: 411_128_901 → 404_158_486 ( -1.7%)
Nat.decLe: 209_641_496 → 206_196_563 ( -1.6%)
Nat.add_comm: 56_084_908 → 55_504_714 ( -1.0%)
`lake test -- --ignored ixvm` passes with 0 FFT mismatches.
The `k_infer_only` section header skipped over the fact that the
function is only sound on well-typed inputs (since it drops
`k_check(a, dom)` on `App`, `k_ensure_sort(ty)` on `Lam`, val/ty
checks on `Let`). Spell out:
* The invariant — only call on terms produced by `whnf` of a
well-typed term, never on arbitrary inputs.
* Why the current sites (`try_proof_irrel`, `is_prop_type`,
`try_unit_like`) respect it.
* The planned non-deterministic-hint dispatch (`Hint::{None,
KInfer, KInferOnly}`) that lets us share `k_infer`'s memo where a
hit already exists instead of paying the parallel `infer_only`
memo cost.
@arthurpaulino
arthurpaulino merged commit ad7e383 into mainJun 25, 2026
14 checks passed
@arthurpaulino
arthurpaulino deleted the ap/utf8-tier-1d branch June 25, 2026 13:49
johnchandlerburnham pushed a commit that referenced this pull request Jul 21, 2026
…505)
* IxVM: drop dead KValNode/KVal/KValEnv
From ap/kernel 828fb85 (Arthur Paulino): the NbE value domain is defined
but referenced nowhere — the live kernel runs on de-Bruijn KExpr;
vestigial from an abandoned NbE direction. (The closed-term context
normalization from that commit was measured separately and not taken:
the per-call expr_lbr probe cost +2.9% FFT on recursor loops for a 0.5%
record reduction.)
* IxVM: memoized prim_family dispatch + width-safe offset-stuck placement
Three coordinated changes to Const-head whnf dispatch (cherry-pick of
130f30b, adapted to post-#450/#457 main):
1. prim_family(addr) classifies a head address into the one reducer
family that could fire on it (nat/str/bitvec/native/decidable; the
sets are disjoint). Keyed on the ADDRESS ALONE it memoizes to one row
per distinct constant address per run, and whnf_const_head calls at
most one family reducer. The previous gauntlet ran every reducer in
sequence for guaranteed misses.
2. The symbolic-Nat offset-stuck check moves from a delta-arm probe into
try_nat_dispatch's miss path as a cold function
(try_nat_offset_dispatch, verdict 2 = "already stuck, do not
re-whnf"), with the offset construction shared via
mk_nat_offset_stuck (also used by the linear-rec collapse).
3. nat_lit_to_ctor_or_self exposes ONE constructor layer
(n -> succ(Lit(n-1))) instead of materializing the full succ chain.
Adaptations vs the original patch:
- whnf_nd_const_head (no-delta WHNF, added on main after the patch)
converted to the same family dispatch.
- Kept main's cold-extracted try_nat_binop_dispatch and routed the
symbolic-base case to try_nat_offset_dispatch from its miss arm.
Measured (lake exe ix check Nat.add_comm): total width 34820 -> 34806,
FFT cost 49571210 -> 48860647 (-1.43%). All 53 ixvm-suite FFT pins
decreased (-0.19%..-1.43%); parity and claim smokes pass. Pins updated;
crates/ixvm-codegen/src/aiur_ixvm.rs regenerated via `lake exe ix
codegen`.
* IxVM: port jcb/fixes H-14 — ptr_val skip map + lockstep addr cursor
Two quadratic/constant-factor fixes to check_all_skipping, ported from
jcb/fixes (3763356, John C. Burnham); cherry-pick of 2d83be1 adapted to
post-#457 main:
- The assumption-leaf skip set keys on ptr_val instead of the first 4
address bytes: one tree lookup per constant, no per-lookup address
load, no confirming address_eq. Sound by the build_addr_pos_map
interning invariant (one pointer, one content): a ptr hit implies the
address IS a leaf; a de-interned pointer reads as absent and the
constant just gets checked — fail-closed.
- The iterator walks the addrs list in LOCKSTEP with consts (cur_addrs
suffix) instead of list_lookup(addrs, pos) per constant, which
re-walked the prefix every iteration — a standalone O(closure^2).
Adaptations: kept main's two-arg check_canonical_block_sort call;
addr_key retained (the Inductive.lean block-membership table added on
main after this patch still uses it), comment updated.
Measured (lake exe ix check Nat.add_comm): total width 34806 -> 34781,
FFT cost unchanged (plain checks never take the skip path). Full ixvm
suite green incl. the frontier-assumption claim smoke; all FFT pins
unchanged. crates/ixvm-codegen/src/aiur_ixvm.rs regenerated.
---------
Co-authored-by: samuelburnham <45365069+samuelburnham@users.noreply.github.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@arthurpaulino@gabriel-barrett
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

IxVM kernel: unblock UTF-8 decode/encode proof + Nat-layer FFT cuts - #450

Merged
arthurpaulino merged 8 commits into
mainfrom
ap/utf8-tier-1d
Jun 25, 2026
Merged

IxVM kernel: unblock UTF-8 decode/encode proof + Nat-layer FFT cuts#450
arthurpaulino merged 8 commits into
mainfrom
ap/utf8-tier-1d

Conversation

@arthurpaulino

Copy link
Copy Markdown
Member

Headline

Aiur kernel previously OOM'd on
_private.Init.Data.String.Decode.0.ByteArray.utf8DecodeChar?.assemble₄_eq_some_of_toBitVec._proof_1_8.
This constant is a prerequisite of ByteArray.utf8DecodeChar?_utf8EncodeChar_append
(part of the UTF-8 round-trip lemma chain).

Now it typechecks at 37.31B FFT (down from the OOM baseline). The
two-direction lemma …_utf8EncodeChar_append itself remains
out-of-reach (still OOMs further along the same dependency tree under
the user's memory cap), but the prerequisite that was the immediate
blocker is unstuck.

The same kernel changes also deliver large FFT cuts on previously-
expensive targets — Tier 1d's structural short-circuit replaces full
delta-whnf cascades wherever it fires:

ConstantBeforeAfterΔ
Vector.append4.02B3.16B−21.4%
Array.append_assoc3.94B3.08B−21.8%

Commits

1. a6cd34a — Tier 1d def-eq short-circuit (the unblock)

Aiur's k_is_def_eq_core jumped from Tier 1.5 straight to full delta
WHNF (Tier 2). The cascading Nat.rec / Nat.succ iota expansions then
drove whnf_const_head past 1M unique entries before either side
reached a comparable canonical form — OOM. Rust's def-eq settles the
same pair via no-delta whnf + quick structural recursion before any
of that fires.

Three minimum-necessary pieces ported. No KStore. No FVar. No Subst /
KernelTypes change.

  • whnf_nd family (Whnf.lean, mirror Rust
    whnf_no_delta_for_def_eq). Same dispatch tree as whnf, but
    whnf_nd_const_head's Defn arm falls through to a stuck
    apply_spine instead of delta-unfolding. Iota / proj / quot /
    primitives still fire.

  • k_infer_only family (Infer.lean, mirror Rust
    with_infer_only). App drops k_check(a, dom); Lam drops
    k_ensure_sort(ty); Let drops val/ty validation. Used at
    try_proof_irrel, is_prop_type, try_unit_like — the def-eq
    tactics that only need the synthesized type.

  • k_is_def_eq_struct_safe + Tier 1d wiring (DefEq.lean,
    mirror Rust quick_def_eq + post-try_def_eq_app). Sort-Sort via
    level_equal; Lam-Lam / All-All via recursive k_is_def_eq on the
    type and on the body under Cons(ty_a, types) (types-cons, NOT
    FVar opening). Inserted between Tier 1c (string lit) and Tier 2
    (full whnf):

    aw_nd = whnf_nd(a); bw_nd = whnf_nd(b)
    ptr_eq(aw_nd, bw_nd) → 1
    k_is_def_eq_struct_safe(aw_nd, bw_nd) → 1
    try_lazy_delta_app(aw_nd, bw_nd) → 1 (rerun: spine args may have
    reduced past what Tier 1.5's pre-whnf attempt could see)
    

Each piece independently validated necessary. The previously-tried
KStore explicit caches, FVar variant + opens, FVar-based binder
opening — all confirmed NOT necessary for this unblock and left out.

3 files changed, +296/−6 lines.

2. 36d6c3a — drop g_or from u64_sub_with_borrow

u64_sub_with_borrow combined two per-byte borrow bits with g_or.
The two bits are mutually exclusive: u_t = 1 ⇒ intermediate t_i ≥ 1 ⇒ subtracting br_in ∈ {0,1} cannot underflow ⇒ u_r = 0. Field
+ substitutes for g_or directly. Per Aiur cost model g_or adds
+1 aux + 1 lookup per call (≈ 5 width); field + is free. 7 g_ors ×
2.23M rows.

UTF-8 _proof_1_8: 39.12B → 38.14B (−2.6%).

3. 9e787da — drop g_or from klimbs_add_carry / klimbs_sub_borrow

Same mutually-exclusive-carry pattern. Two limb-level borrows /
carries from sequential u64 ops cannot both be 1.

UTF-8 _proof_1_8: 38.14B → 38.07B (−0.18%).

4. 4e379e7 — hot/cold split try_nat_dispatch, extract binop arm

try_nat_dispatch's width 90 was floored by the binop arm (2× whnf

  • 2× try_extract_nat + try_nat_binop_addr + apply_spine), charged on
    every Nat.succ / Nat.pred row. Factor binop dispatch into
    try_nat_binop_dispatch. Main narrows to the max of succ / pred
    arms.

UTF-8 _proof_1_8: 38.07B → 37.80B (−0.7%).

5. 80ce3d2 — hot/cold split expr_lbr, extract Let arm

expr_lbr's width was floored by the Let arm (3 recursive expr_lbr
calls + 2 lbr_max + 1 lbr_dec), charged on every row even though Let
is rare. Factor into expr_lbr_let.

Nat.add_comm: 55.63M → 55.50M (−0.2%). UTF-8 _proof_1_8: 37.80B
→ 37.62B (−0.5%)
.

6. 1f8effd — hot/cold split try_extract_nat, extract App arm

try_extract_nat's width was floored by the App arm (list_lookup +
address_eq + recursive try_extract_nat + klimbs_succ). Factor into
try_extract_nat_app. Main narrows to leaf-arm width.

UTF-8 _proof_1_8: 37.62B → 37.31B (−0.8%).

7. 039e9cf — re-pin IxVM FFT costs

41 pins in Tests/Ix/IxVM.lean::kernelCheckEntries updated. Every
constant got cheaper; none regressed. Largest reductions:

  • Vector.append: 4.02B → 3.16B (−21.4%)
  • Array.append_assoc: 3.94B → 3.08B (−21.8%)

lake test -- --ignored ixvm passes with 0 FFT mismatches.

Cumulative on UTF-8 _proof_1_8

CommitFFT
baseline (main)OOM
a6cd34a Tier 1d39.12B
36d6c3a g_or → + in u64_sub_with_borrow38.14B (−2.6%)
9e787da g_or → + in klimbs_add_carry / klimbs_sub_borrow38.07B (−0.18%)
4e379e7 hot/cold try_nat_dispatch37.80B (−0.7%)
80ce3d2 hot/cold expr_lbr37.62B (−0.5%)
1f8effd hot/cold try_extract_nat37.31B (−0.8%)

Post-unlock optimization: −4.6% (39.12B → 37.31B).

Cost on small targets

  • Main baseline Nat.add_comm: 56.08M FFT.
  • This branch Nat.add_comm: 55.50M FFT.

Tier 1d itself adds no overhead on the common case (whnf_nd +
struct_safe + try_lazy_delta_app are themselves Aiur-memoized); the
follow-up optimizations are net wins.

Test plan

  • lake exe check Nat.add_comm passes (55.50M FFT).
  • lake exe check Vector.extract_append passes.
  • lake exe check "_private.…assemble₄_eq_some_of_toBitVec._proof_1_8"
    passes (37.31B FFT, previously OOM).
  • lake test -- --ignored ixvm — all 41 FFT pins updated, suite
    passes with 0 mismatches.

Comment threadIx/IxVM/Kernel/Infer.lean
…ost-spine-congruence)
UTF-8 `_private.Init.Data.String.Decode.0.ByteArray.utf8DecodeChar?
.assemble₄_eq_some_of_toBitVec._proof_1_8` OOMs on the previous
pipeline because `k_is_def_eq_core` jumps from Tier 1.5 straight to
full delta WHNF (Tier 2); the cascading Nat.rec / Nat.succ iota
expansions then drive `whnf_const_head` past 1M unique entries before
either side reaches a comparable canonical form. Rust's def-eq settles
the same pair via the no-delta whnf + quick structural recursion
before any of that fires.
This patch ports the three pieces of that short-circuit and nothing
else — no FVar variant, no KStore, no Subst changes, no signature
sweep.
* `whnf_nd` family in `Whnf.lean` (mirror Rust `whnf_no_delta_for_def_eq`).
Same dispatch tree as `whnf` (beta / let zeta / iota / proj / quot /
primitives all fire), except `whnf_nd_const_head`'s Defn arm falls
through to a stuck `apply_spine` instead of delta-unfolding.
* `k_infer_only` family in `Infer.lean` (mirror Rust `with_infer_only`).
App drops `k_check(a, dom)`; Lam drops `k_ensure_sort(ty)`; Let drops
the val/ty validations. Distinct Aiur memo from `k_infer`, parity
with Rust's separate `infer_cache` / `infer_only_cache`.
* `k_is_def_eq_struct_safe` in `DefEq.lean` (mirror Rust
`quick_def_eq`). Sort-Sort via `level_equal`; Lam-Lam / All-All
recurse on type and on body under `Cons(ty_a, types)`. Returns 1
only when DEFINITELY def-eq; 0 means fall through. Sound on
partially-whnf'd (no-delta) inputs because the handled shapes
don't depend on further reductions.
* `k_is_def_eq_core` Tier 1d wiring inserted between Tier 1c (string
lit) and Tier 2 (full whnf):
aw_nd = whnf_nd(a); bw_nd = whnf_nd(b)
ptr_eq(aw_nd, bw_nd) → 1
k_is_def_eq_struct_safe(aw_nd, bw_nd) → 1 if 1
try_lazy_delta_app(aw_nd, bw_nd) → 1 if 1 (rerun post-whnf_nd:
spine args may have reduced past what Tier 1.5's pre-whnf attempt
could see, exposing Const-Const congruence that was hidden)
* `try_proof_irrel`, `is_prop_type`, `try_unit_like` switch from
`k_infer` to `k_infer_only` — these helpers only need the synthesized
type, not the full re-validation work that `k_infer` does for each
recursive App/Let/Lam.
Each piece individually validated necessary (removing it puts UTF-8
back into the OOM regime). FVar variant + opens, KStore explicit
caches, infer_only's FVar-based binder opening — all confirmed NOT
necessary for the UTF-8 unblock and left out (see PLAN.md for future
experiments).
Measured (FFT cost):
Nat.add_comm: 56.08M → 55.63M (~stable; new code paths add no
overhead on the common case because Tier 1d's whnf_nd + struct_safe
are themselves Aiur-memoized).
_private.Init.Data.String.Decode.0.ByteArray.utf8DecodeChar?
.assemble₄_eq_some_of_toBitVec._proof_1_8: OOM → 39.12B FFT, passes.
3 files, +296/-6 lines.
`u64_sub_with_borrow` combines two per-byte borrow bits with `g_or`. The
two bits are MUTUALLY EXCLUSIVE: `u_t = borrow(a_i - b_i)` and
`u_r = borrow((a_i + 256 - b_i) - br_in)`. If `u_t = 1` the intermediate
`t_i ≥ 1`, so subtracting `br_in ∈ {0,1}` cannot underflow ⇒ `u_r = 0`.
Field `+` substitutes for `g_or` directly (per the same pattern as
`u64_add` in `ByteStream.lean`).
Per Aiur cost model, `g_or` adds +1 aux + 1 lookup per call; field `+`
is free. 8 g_or call sites in `u64_sub_with_borrow` each charged on every
one of the function's 2.23M rows.
Measured (FFT cost) on UTF-8 `_proof_1_8`:
39.12B → 38.14B (-2.6%)
Nat.add_comm unchanged (55.63M).
See [[reference_aiur_carry_add]].
Same mutually-exclusive-carry pattern as `u64_sub_with_borrow`:
* `klimbs_add_carry`: u64_add of (la, lb) yields carry1; u64_add of
(sum1, carry_in) yields carry2. carry1=1 ⇒ sum1 ≤ 2^64-2 ⇒
carry2=0.
* `klimbs_sub_borrow`: symmetric for borrows.
Replace `g_or(c1, c2)` with `c1 + c2` (field +). Both helpers run on
hot Nat-primitive paths.
Measured on UTF-8 `_proof_1_8`:
38.14B → 38.07B (-0.18%)
Nat.add_comm unchanged.
See [[reference_aiur_carry_add]].
`try_nat_dispatch` ran 1.12M rows in UTF-8 `_proof_1_8` at width 90,
charging 5.16% of total FFT. Width was floored by its widest match arm
(the binop branch with 2× whnf + 2× try_extract_nat + try_nat_binop_addr
+ apply_spine), even on Nat.succ / Nat.pred rows that never touched it.
Factor binop dispatch into its own `try_nat_binop_dispatch` fn. Main
dispatcher narrows to the max of succ / pred arms (single whnf +
try_extract_nat + klimbs_succ/dec + apply_spine). The cold fn's width
only charges the rows that actually dispatch a binop.
Measured on UTF-8 `_proof_1_8`:
38.07B → 37.80B (-0.7%)
Nat.add_comm unchanged.
`expr_lbr` ran 1.47M rows in UTF-8 `_proof_1_8` at width 39, charging
3.01% of total FFT. The Let arm (3 recursive expr_lbr calls + 2 lbr_max
+ 1 lbr_dec) is the widest match arm, charged on every row of expr_lbr
even though Let is rare in most expressions encountered.
Factor the Let arm into `expr_lbr_let(ty, val, body)`. Main expr_lbr
narrows to max of the 2-recursion arms (App / Lam / Forall). Cold fn
only charges Let-arm rows.
Measured:
Nat.add_comm: 55.63M → 55.50M (-0.2%)
UTF-8 `_proof_1_8`: 37.80B → 37.62B (-0.5%)
`try_extract_nat` ran 1.12M rows at width 45, charging 2.68% of UTF-8
`_proof_1_8` total FFT. The App arm (list_lookup + address_eq +
recursive try_extract_nat + klimbs_succ) is the widest match arm; the
Lit / Const / default arms are leaf compares.
Factor App into `try_extract_nat_app(f, a, addrs)`. Main extractor
narrows to leaf-arm width. Cold fn only charges App-arm rows.
Measured on UTF-8 `_proof_1_8`:
37.62B → 37.31B (-0.8%)
Nat.add_comm unchanged.
Updates 41 pinned FFT costs in `Tests/Ix/IxVM.lean::kernelCheckEntries`
to match the new kernel's output. All pins moved DOWN — every constant
got cheaper, none regressed.
Largest reductions (% change):
Vector.append: 4_023_268_168 → 3_160_970_390 (-21.4%)
Array.append_assoc: 3_938_574_533 → 3_079_334_815 (-21.8%)
String.Internal.append: 793_580_333 → 775_968_134 ( -2.2%)
bv_to_nat_lit: 635_780_327 → 619_870_154 ( -2.5%)
nat_gcd_lit: 665_518_356 → 649_859_784 ( -2.4%)
Nat.sub_le_of_le_add: 567_575_653 → 557_867_526 ( -1.7%)
IxVMPrim.nat_mod_lit: 414_695_549 → 407_517_834 ( -1.7%)
IxVMPrim.nat_div_lit: 405_607_545 → 398_641_590 ( -1.7%)
IxVMPrim.nat_shr_lit: 411_128_901 → 404_158_486 ( -1.7%)
Nat.decLe: 209_641_496 → 206_196_563 ( -1.6%)
Nat.add_comm: 56_084_908 → 55_504_714 ( -1.0%)
`lake test -- --ignored ixvm` passes with 0 FFT mismatches.
The `k_infer_only` section header skipped over the fact that the
function is only sound on well-typed inputs (since it drops
`k_check(a, dom)` on `App`, `k_ensure_sort(ty)` on `Lam`, val/ty
checks on `Let`). Spell out:
* The invariant — only call on terms produced by `whnf` of a
well-typed term, never on arbitrary inputs.
* Why the current sites (`try_proof_irrel`, `is_prop_type`,
`try_unit_like`) respect it.
* The planned non-deterministic-hint dispatch (`Hint::{None,
KInfer, KInferOnly}`) that lets us share `k_infer`'s memo where a
hit already exists instead of paying the parallel `infer_only`
memo cost.
@arthurpaulino
arthurpaulino merged commit ad7e383 into mainJun 25, 2026
14 checks passed
@arthurpaulino
arthurpaulino deleted the ap/utf8-tier-1d branch June 25, 2026 13:49
johnchandlerburnham pushed a commit that referenced this pull request Jul 21, 2026
…505)
* IxVM: drop dead KValNode/KVal/KValEnv
From ap/kernel 828fb85 (Arthur Paulino): the NbE value domain is defined
but referenced nowhere — the live kernel runs on de-Bruijn KExpr;
vestigial from an abandoned NbE direction. (The closed-term context
normalization from that commit was measured separately and not taken:
the per-call expr_lbr probe cost +2.9% FFT on recursor loops for a 0.5%
record reduction.)
* IxVM: memoized prim_family dispatch + width-safe offset-stuck placement
Three coordinated changes to Const-head whnf dispatch (cherry-pick of
130f30b, adapted to post-#450/#457 main):
1. prim_family(addr) classifies a head address into the one reducer
family that could fire on it (nat/str/bitvec/native/decidable; the
sets are disjoint). Keyed on the ADDRESS ALONE it memoizes to one row
per distinct constant address per run, and whnf_const_head calls at
most one family reducer. The previous gauntlet ran every reducer in
sequence for guaranteed misses.
2. The symbolic-Nat offset-stuck check moves from a delta-arm probe into
try_nat_dispatch's miss path as a cold function
(try_nat_offset_dispatch, verdict 2 = "already stuck, do not
re-whnf"), with the offset construction shared via
mk_nat_offset_stuck (also used by the linear-rec collapse).
3. nat_lit_to_ctor_or_self exposes ONE constructor layer
(n -> succ(Lit(n-1))) instead of materializing the full succ chain.
Adaptations vs the original patch:
- whnf_nd_const_head (no-delta WHNF, added on main after the patch)
converted to the same family dispatch.
- Kept main's cold-extracted try_nat_binop_dispatch and routed the
symbolic-base case to try_nat_offset_dispatch from its miss arm.
Measured (lake exe ix check Nat.add_comm): total width 34820 -> 34806,
FFT cost 49571210 -> 48860647 (-1.43%). All 53 ixvm-suite FFT pins
decreased (-0.19%..-1.43%); parity and claim smokes pass. Pins updated;
crates/ixvm-codegen/src/aiur_ixvm.rs regenerated via `lake exe ix
codegen`.
* IxVM: port jcb/fixes H-14 — ptr_val skip map + lockstep addr cursor
Two quadratic/constant-factor fixes to check_all_skipping, ported from
jcb/fixes (3763356, John C. Burnham); cherry-pick of 2d83be1 adapted to
post-#457 main:
- The assumption-leaf skip set keys on ptr_val instead of the first 4
address bytes: one tree lookup per constant, no per-lookup address
load, no confirming address_eq. Sound by the build_addr_pos_map
interning invariant (one pointer, one content): a ptr hit implies the
address IS a leaf; a de-interned pointer reads as absent and the
constant just gets checked — fail-closed.
- The iterator walks the addrs list in LOCKSTEP with consts (cur_addrs
suffix) instead of list_lookup(addrs, pos) per constant, which
re-walked the prefix every iteration — a standalone O(closure^2).
Adaptations: kept main's two-arg check_canonical_block_sort call;
addr_key retained (the Inductive.lean block-membership table added on
main after this patch still uses it), comment updated.
Measured (lake exe ix check Nat.add_comm): total width 34806 -> 34781,
FFT cost unchanged (plain checks never take the skip path). Full ixvm
suite green incl. the frontier-assumption claim smoke; all FFT pins
unchanged. crates/ixvm-codegen/src/aiur_ixvm.rs regenerated.
---------
Co-authored-by: samuelburnham <45365069+samuelburnham@users.noreply.github.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@arthurpaulino@gabriel-barrett
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

IxVM kernel: unblock UTF-8 decode/encode proof + Nat-layer FFT cuts - #450

Merged
arthurpaulino merged 8 commits into
mainfrom
ap/utf8-tier-1d
Jun 25, 2026
Merged

IxVM kernel: unblock UTF-8 decode/encode proof + Nat-layer FFT cuts#450
arthurpaulino merged 8 commits into
mainfrom
ap/utf8-tier-1d

Conversation

@arthurpaulino

Copy link
Copy Markdown
Member

Headline

Aiur kernel previously OOM'd on
_private.Init.Data.String.Decode.0.ByteArray.utf8DecodeChar?.assemble₄_eq_some_of_toBitVec._proof_1_8.
This constant is a prerequisite of ByteArray.utf8DecodeChar?_utf8EncodeChar_append
(part of the UTF-8 round-trip lemma chain).

Now it typechecks at 37.31B FFT (down from the OOM baseline). The
two-direction lemma …_utf8EncodeChar_append itself remains
out-of-reach (still OOMs further along the same dependency tree under
the user's memory cap), but the prerequisite that was the immediate
blocker is unstuck.

The same kernel changes also deliver large FFT cuts on previously-
expensive targets — Tier 1d's structural short-circuit replaces full
delta-whnf cascades wherever it fires:

ConstantBeforeAfterΔ
Vector.append4.02B3.16B−21.4%
Array.append_assoc3.94B3.08B−21.8%

Commits

1. a6cd34a — Tier 1d def-eq short-circuit (the unblock)

Aiur's k_is_def_eq_core jumped from Tier 1.5 straight to full delta
WHNF (Tier 2). The cascading Nat.rec / Nat.succ iota expansions then
drove whnf_const_head past 1M unique entries before either side
reached a comparable canonical form — OOM. Rust's def-eq settles the
same pair via no-delta whnf + quick structural recursion before any
of that fires.

Three minimum-necessary pieces ported. No KStore. No FVar. No Subst /
KernelTypes change.

  • whnf_nd family (Whnf.lean, mirror Rust
    whnf_no_delta_for_def_eq). Same dispatch tree as whnf, but
    whnf_nd_const_head's Defn arm falls through to a stuck
    apply_spine instead of delta-unfolding. Iota / proj / quot /
    primitives still fire.

  • k_infer_only family (Infer.lean, mirror Rust
    with_infer_only). App drops k_check(a, dom); Lam drops
    k_ensure_sort(ty); Let drops val/ty validation. Used at
    try_proof_irrel, is_prop_type, try_unit_like — the def-eq
    tactics that only need the synthesized type.

  • k_is_def_eq_struct_safe + Tier 1d wiring (DefEq.lean,
    mirror Rust quick_def_eq + post-try_def_eq_app). Sort-Sort via
    level_equal; Lam-Lam / All-All via recursive k_is_def_eq on the
    type and on the body under Cons(ty_a, types) (types-cons, NOT
    FVar opening). Inserted between Tier 1c (string lit) and Tier 2
    (full whnf):

    aw_nd = whnf_nd(a); bw_nd = whnf_nd(b)
    ptr_eq(aw_nd, bw_nd) → 1
    k_is_def_eq_struct_safe(aw_nd, bw_nd) → 1
    try_lazy_delta_app(aw_nd, bw_nd) → 1 (rerun: spine args may have
    reduced past what Tier 1.5's pre-whnf attempt could see)
    

Each piece independently validated necessary. The previously-tried
KStore explicit caches, FVar variant + opens, FVar-based binder
opening — all confirmed NOT necessary for this unblock and left out.

3 files changed, +296/−6 lines.

2. 36d6c3a — drop g_or from u64_sub_with_borrow

u64_sub_with_borrow combined two per-byte borrow bits with g_or.
The two bits are mutually exclusive: u_t = 1 ⇒ intermediate t_i ≥ 1 ⇒ subtracting br_in ∈ {0,1} cannot underflow ⇒ u_r = 0. Field
+ substitutes for g_or directly. Per Aiur cost model g_or adds
+1 aux + 1 lookup per call (≈ 5 width); field + is free. 7 g_ors ×
2.23M rows.

UTF-8 _proof_1_8: 39.12B → 38.14B (−2.6%).

3. 9e787da — drop g_or from klimbs_add_carry / klimbs_sub_borrow

Same mutually-exclusive-carry pattern. Two limb-level borrows /
carries from sequential u64 ops cannot both be 1.

UTF-8 _proof_1_8: 38.14B → 38.07B (−0.18%).

4. 4e379e7 — hot/cold split try_nat_dispatch, extract binop arm

try_nat_dispatch's width 90 was floored by the binop arm (2× whnf

  • 2× try_extract_nat + try_nat_binop_addr + apply_spine), charged on
    every Nat.succ / Nat.pred row. Factor binop dispatch into
    try_nat_binop_dispatch. Main narrows to the max of succ / pred
    arms.

UTF-8 _proof_1_8: 38.07B → 37.80B (−0.7%).

5. 80ce3d2 — hot/cold split expr_lbr, extract Let arm

expr_lbr's width was floored by the Let arm (3 recursive expr_lbr
calls + 2 lbr_max + 1 lbr_dec), charged on every row even though Let
is rare. Factor into expr_lbr_let.

Nat.add_comm: 55.63M → 55.50M (−0.2%). UTF-8 _proof_1_8: 37.80B
→ 37.62B (−0.5%)
.

6. 1f8effd — hot/cold split try_extract_nat, extract App arm

try_extract_nat's width was floored by the App arm (list_lookup +
address_eq + recursive try_extract_nat + klimbs_succ). Factor into
try_extract_nat_app. Main narrows to leaf-arm width.

UTF-8 _proof_1_8: 37.62B → 37.31B (−0.8%).

7. 039e9cf — re-pin IxVM FFT costs

41 pins in Tests/Ix/IxVM.lean::kernelCheckEntries updated. Every
constant got cheaper; none regressed. Largest reductions:

  • Vector.append: 4.02B → 3.16B (−21.4%)
  • Array.append_assoc: 3.94B → 3.08B (−21.8%)

lake test -- --ignored ixvm passes with 0 FFT mismatches.

Cumulative on UTF-8 _proof_1_8

CommitFFT
baseline (main)OOM
a6cd34a Tier 1d39.12B
36d6c3a g_or → + in u64_sub_with_borrow38.14B (−2.6%)
9e787da g_or → + in klimbs_add_carry / klimbs_sub_borrow38.07B (−0.18%)
4e379e7 hot/cold try_nat_dispatch37.80B (−0.7%)
80ce3d2 hot/cold expr_lbr37.62B (−0.5%)
1f8effd hot/cold try_extract_nat37.31B (−0.8%)

Post-unlock optimization: −4.6% (39.12B → 37.31B).

Cost on small targets

  • Main baseline Nat.add_comm: 56.08M FFT.
  • This branch Nat.add_comm: 55.50M FFT.

Tier 1d itself adds no overhead on the common case (whnf_nd +
struct_safe + try_lazy_delta_app are themselves Aiur-memoized); the
follow-up optimizations are net wins.

Test plan

  • lake exe check Nat.add_comm passes (55.50M FFT).
  • lake exe check Vector.extract_append passes.
  • lake exe check "_private.…assemble₄_eq_some_of_toBitVec._proof_1_8"
    passes (37.31B FFT, previously OOM).
  • lake test -- --ignored ixvm — all 41 FFT pins updated, suite
    passes with 0 mismatches.

Comment threadIx/IxVM/Kernel/Infer.lean
…ost-spine-congruence)
UTF-8 `_private.Init.Data.String.Decode.0.ByteArray.utf8DecodeChar?
.assemble₄_eq_some_of_toBitVec._proof_1_8` OOMs on the previous
pipeline because `k_is_def_eq_core` jumps from Tier 1.5 straight to
full delta WHNF (Tier 2); the cascading Nat.rec / Nat.succ iota
expansions then drive `whnf_const_head` past 1M unique entries before
either side reaches a comparable canonical form. Rust's def-eq settles
the same pair via the no-delta whnf + quick structural recursion
before any of that fires.
This patch ports the three pieces of that short-circuit and nothing
else — no FVar variant, no KStore, no Subst changes, no signature
sweep.
* `whnf_nd` family in `Whnf.lean` (mirror Rust `whnf_no_delta_for_def_eq`).
Same dispatch tree as `whnf` (beta / let zeta / iota / proj / quot /
primitives all fire), except `whnf_nd_const_head`'s Defn arm falls
through to a stuck `apply_spine` instead of delta-unfolding.
* `k_infer_only` family in `Infer.lean` (mirror Rust `with_infer_only`).
App drops `k_check(a, dom)`; Lam drops `k_ensure_sort(ty)`; Let drops
the val/ty validations. Distinct Aiur memo from `k_infer`, parity
with Rust's separate `infer_cache` / `infer_only_cache`.
* `k_is_def_eq_struct_safe` in `DefEq.lean` (mirror Rust
`quick_def_eq`). Sort-Sort via `level_equal`; Lam-Lam / All-All
recurse on type and on body under `Cons(ty_a, types)`. Returns 1
only when DEFINITELY def-eq; 0 means fall through. Sound on
partially-whnf'd (no-delta) inputs because the handled shapes
don't depend on further reductions.
* `k_is_def_eq_core` Tier 1d wiring inserted between Tier 1c (string
lit) and Tier 2 (full whnf):
aw_nd = whnf_nd(a); bw_nd = whnf_nd(b)
ptr_eq(aw_nd, bw_nd) → 1
k_is_def_eq_struct_safe(aw_nd, bw_nd) → 1 if 1
try_lazy_delta_app(aw_nd, bw_nd) → 1 if 1 (rerun post-whnf_nd:
spine args may have reduced past what Tier 1.5's pre-whnf attempt
could see, exposing Const-Const congruence that was hidden)
* `try_proof_irrel`, `is_prop_type`, `try_unit_like` switch from
`k_infer` to `k_infer_only` — these helpers only need the synthesized
type, not the full re-validation work that `k_infer` does for each
recursive App/Let/Lam.
Each piece individually validated necessary (removing it puts UTF-8
back into the OOM regime). FVar variant + opens, KStore explicit
caches, infer_only's FVar-based binder opening — all confirmed NOT
necessary for the UTF-8 unblock and left out (see PLAN.md for future
experiments).
Measured (FFT cost):
Nat.add_comm: 56.08M → 55.63M (~stable; new code paths add no
overhead on the common case because Tier 1d's whnf_nd + struct_safe
are themselves Aiur-memoized).
_private.Init.Data.String.Decode.0.ByteArray.utf8DecodeChar?
.assemble₄_eq_some_of_toBitVec._proof_1_8: OOM → 39.12B FFT, passes.
3 files, +296/-6 lines.
`u64_sub_with_borrow` combines two per-byte borrow bits with `g_or`. The
two bits are MUTUALLY EXCLUSIVE: `u_t = borrow(a_i - b_i)` and
`u_r = borrow((a_i + 256 - b_i) - br_in)`. If `u_t = 1` the intermediate
`t_i ≥ 1`, so subtracting `br_in ∈ {0,1}` cannot underflow ⇒ `u_r = 0`.
Field `+` substitutes for `g_or` directly (per the same pattern as
`u64_add` in `ByteStream.lean`).
Per Aiur cost model, `g_or` adds +1 aux + 1 lookup per call; field `+`
is free. 8 g_or call sites in `u64_sub_with_borrow` each charged on every
one of the function's 2.23M rows.
Measured (FFT cost) on UTF-8 `_proof_1_8`:
39.12B → 38.14B (-2.6%)
Nat.add_comm unchanged (55.63M).
See [[reference_aiur_carry_add]].
Same mutually-exclusive-carry pattern as `u64_sub_with_borrow`:
* `klimbs_add_carry`: u64_add of (la, lb) yields carry1; u64_add of
(sum1, carry_in) yields carry2. carry1=1 ⇒ sum1 ≤ 2^64-2 ⇒
carry2=0.
* `klimbs_sub_borrow`: symmetric for borrows.
Replace `g_or(c1, c2)` with `c1 + c2` (field +). Both helpers run on
hot Nat-primitive paths.
Measured on UTF-8 `_proof_1_8`:
38.14B → 38.07B (-0.18%)
Nat.add_comm unchanged.
See [[reference_aiur_carry_add]].
`try_nat_dispatch` ran 1.12M rows in UTF-8 `_proof_1_8` at width 90,
charging 5.16% of total FFT. Width was floored by its widest match arm
(the binop branch with 2× whnf + 2× try_extract_nat + try_nat_binop_addr
+ apply_spine), even on Nat.succ / Nat.pred rows that never touched it.
Factor binop dispatch into its own `try_nat_binop_dispatch` fn. Main
dispatcher narrows to the max of succ / pred arms (single whnf +
try_extract_nat + klimbs_succ/dec + apply_spine). The cold fn's width
only charges the rows that actually dispatch a binop.
Measured on UTF-8 `_proof_1_8`:
38.07B → 37.80B (-0.7%)
Nat.add_comm unchanged.
`expr_lbr` ran 1.47M rows in UTF-8 `_proof_1_8` at width 39, charging
3.01% of total FFT. The Let arm (3 recursive expr_lbr calls + 2 lbr_max
+ 1 lbr_dec) is the widest match arm, charged on every row of expr_lbr
even though Let is rare in most expressions encountered.
Factor the Let arm into `expr_lbr_let(ty, val, body)`. Main expr_lbr
narrows to max of the 2-recursion arms (App / Lam / Forall). Cold fn
only charges Let-arm rows.
Measured:
Nat.add_comm: 55.63M → 55.50M (-0.2%)
UTF-8 `_proof_1_8`: 37.80B → 37.62B (-0.5%)
`try_extract_nat` ran 1.12M rows at width 45, charging 2.68% of UTF-8
`_proof_1_8` total FFT. The App arm (list_lookup + address_eq +
recursive try_extract_nat + klimbs_succ) is the widest match arm; the
Lit / Const / default arms are leaf compares.
Factor App into `try_extract_nat_app(f, a, addrs)`. Main extractor
narrows to leaf-arm width. Cold fn only charges App-arm rows.
Measured on UTF-8 `_proof_1_8`:
37.62B → 37.31B (-0.8%)
Nat.add_comm unchanged.
Updates 41 pinned FFT costs in `Tests/Ix/IxVM.lean::kernelCheckEntries`
to match the new kernel's output. All pins moved DOWN — every constant
got cheaper, none regressed.
Largest reductions (% change):
Vector.append: 4_023_268_168 → 3_160_970_390 (-21.4%)
Array.append_assoc: 3_938_574_533 → 3_079_334_815 (-21.8%)
String.Internal.append: 793_580_333 → 775_968_134 ( -2.2%)
bv_to_nat_lit: 635_780_327 → 619_870_154 ( -2.5%)
nat_gcd_lit: 665_518_356 → 649_859_784 ( -2.4%)
Nat.sub_le_of_le_add: 567_575_653 → 557_867_526 ( -1.7%)
IxVMPrim.nat_mod_lit: 414_695_549 → 407_517_834 ( -1.7%)
IxVMPrim.nat_div_lit: 405_607_545 → 398_641_590 ( -1.7%)
IxVMPrim.nat_shr_lit: 411_128_901 → 404_158_486 ( -1.7%)
Nat.decLe: 209_641_496 → 206_196_563 ( -1.6%)
Nat.add_comm: 56_084_908 → 55_504_714 ( -1.0%)
`lake test -- --ignored ixvm` passes with 0 FFT mismatches.
The `k_infer_only` section header skipped over the fact that the
function is only sound on well-typed inputs (since it drops
`k_check(a, dom)` on `App`, `k_ensure_sort(ty)` on `Lam`, val/ty
checks on `Let`). Spell out:
* The invariant — only call on terms produced by `whnf` of a
well-typed term, never on arbitrary inputs.
* Why the current sites (`try_proof_irrel`, `is_prop_type`,
`try_unit_like`) respect it.
* The planned non-deterministic-hint dispatch (`Hint::{None,
KInfer, KInferOnly}`) that lets us share `k_infer`'s memo where a
hit already exists instead of paying the parallel `infer_only`
memo cost.
@arthurpaulino
arthurpaulino merged commit ad7e383 into mainJun 25, 2026
14 checks passed
@arthurpaulino
arthurpaulino deleted the ap/utf8-tier-1d branch June 25, 2026 13:49
johnchandlerburnham pushed a commit that referenced this pull request Jul 21, 2026
…505)
* IxVM: drop dead KValNode/KVal/KValEnv
From ap/kernel 828fb85 (Arthur Paulino): the NbE value domain is defined
but referenced nowhere — the live kernel runs on de-Bruijn KExpr;
vestigial from an abandoned NbE direction. (The closed-term context
normalization from that commit was measured separately and not taken:
the per-call expr_lbr probe cost +2.9% FFT on recursor loops for a 0.5%
record reduction.)
* IxVM: memoized prim_family dispatch + width-safe offset-stuck placement
Three coordinated changes to Const-head whnf dispatch (cherry-pick of
130f30b, adapted to post-#450/#457 main):
1. prim_family(addr) classifies a head address into the one reducer
family that could fire on it (nat/str/bitvec/native/decidable; the
sets are disjoint). Keyed on the ADDRESS ALONE it memoizes to one row
per distinct constant address per run, and whnf_const_head calls at
most one family reducer. The previous gauntlet ran every reducer in
sequence for guaranteed misses.
2. The symbolic-Nat offset-stuck check moves from a delta-arm probe into
try_nat_dispatch's miss path as a cold function
(try_nat_offset_dispatch, verdict 2 = "already stuck, do not
re-whnf"), with the offset construction shared via
mk_nat_offset_stuck (also used by the linear-rec collapse).
3. nat_lit_to_ctor_or_self exposes ONE constructor layer
(n -> succ(Lit(n-1))) instead of materializing the full succ chain.
Adaptations vs the original patch:
- whnf_nd_const_head (no-delta WHNF, added on main after the patch)
converted to the same family dispatch.
- Kept main's cold-extracted try_nat_binop_dispatch and routed the
symbolic-base case to try_nat_offset_dispatch from its miss arm.
Measured (lake exe ix check Nat.add_comm): total width 34820 -> 34806,
FFT cost 49571210 -> 48860647 (-1.43%). All 53 ixvm-suite FFT pins
decreased (-0.19%..-1.43%); parity and claim smokes pass. Pins updated;
crates/ixvm-codegen/src/aiur_ixvm.rs regenerated via `lake exe ix
codegen`.
* IxVM: port jcb/fixes H-14 — ptr_val skip map + lockstep addr cursor
Two quadratic/constant-factor fixes to check_all_skipping, ported from
jcb/fixes (3763356, John C. Burnham); cherry-pick of 2d83be1 adapted to
post-#457 main:
- The assumption-leaf skip set keys on ptr_val instead of the first 4
address bytes: one tree lookup per constant, no per-lookup address
load, no confirming address_eq. Sound by the build_addr_pos_map
interning invariant (one pointer, one content): a ptr hit implies the
address IS a leaf; a de-interned pointer reads as absent and the
constant just gets checked — fail-closed.
- The iterator walks the addrs list in LOCKSTEP with consts (cur_addrs
suffix) instead of list_lookup(addrs, pos) per constant, which
re-walked the prefix every iteration — a standalone O(closure^2).
Adaptations: kept main's two-arg check_canonical_block_sort call;
addr_key retained (the Inductive.lean block-membership table added on
main after this patch still uses it), comment updated.
Measured (lake exe ix check Nat.add_comm): total width 34806 -> 34781,
FFT cost unchanged (plain checks never take the skip path). Full ixvm
suite green incl. the frontier-assumption claim smoke; all FFT pins
unchanged. crates/ixvm-codegen/src/aiur_ixvm.rs regenerated.
---------
Co-authored-by: samuelburnham <45365069+samuelburnham@users.noreply.github.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@arthurpaulino@gabriel-barrett