Hyperoptimize the blake3_compress_chunks circuit - #413

Merged
arthurpaulino merged 5 commits into
mainfrom
ap/blake3
May 18, 2026
Merged

Hyperoptimize the blake3_compress_chunks circuit#413
arthurpaulino merged 5 commits into
mainfrom
ap/blake3

Conversation

@arthurpaulino

@arthurpaulinoarthurpaulino commented May 18, 2026

Copy link
Copy Markdown
Member

Summary

The blake3 circuits dominated the lake exe kernel String.Internal.append
proving-cost profile. This PR cuts total FFT cost by 19.1% (3109M →
2516M) through five incremental, individually-measured changes, and
removes two now-dead helpers.

Background

Aiur compiles each function to a zk circuit. A circuit's width is
inputSize + selectors + auxiliaries + 4·(1 + lookups), where
auxiliaries and lookups are the maximum over the match arms — so
width is set by the widest arm; a wide cold arm wastes columns on every
row. Cost is width × height × log2(height). Op costs that matter here:
call +outputSize aux +1 lookup; store/load +1 lookup; u8_* ops +1
lookup each; pure field add/sub/mul-by-const cost zero columns.

Changes

1. Factor block-completion out of the hot loop

blake3_compress_chunks width 341 was set by its coldest arms
(Cons,63,1023) / (Cons,63,_) — the only ones carrying both
assign_block_value (+64 aux) and blake3_compress (+32 aux) plus flag
helpers and store, yet firing on ~1.8% of rows. Extracted that tail
into a cold blake3_compress_block circuit; the two arms merged into
one. 341 → 281.

2. Thread Layer by pointer

enum Layer { Push(&Layer, [[G;4];8]), Nil } is a 34-wide value type.
blake3_compress_chunks took and returned it by value — 34 columns of
inputSize every row plus 34 aux on every recursive call. Switched to
layer: &Layer -> &Layer; blake3 adds a store/load at the
boundary so blake3_compress_layer keeps its by-value signature.
281 → 215.

3. Accumulate block bytes in a list, not a 64-wide buffer

blake3_compress_chunks threaded block_buffer: [[G;4];16] (64 columns
of inputSize) and filled it via assign_block_value, whose
[[G;4];16] return cost +64 aux every hot row. Circuit memory is
content-addressed (no in-place mutation), so a threaded 64-element
accumulator costs 64 columns somewhere on every row — a linked list does
not. Replaced it with a reverse-ordered byte accumulator list:

  • Hot loop: store(ListNode.Cons(head, byte_acc)), one store/row.
  • bytes_to_block materializes a full 64-byte list into [[G;4];16] via
    64 unrolled loads — one cold row per completed block.
  • pad_block zero-pads the trailing partial block.
  • blake3_finish absorbs the input-exhausted arms.
  • assign_block_value is now dead and removed.

215 → 75.

4. Combine carry bits with field add, not u8 xor

The byte adders (u32_add, u32_be_add, u64_add,
relaxed_u64_be_add_2_bytes) propagated carries with
u8_xor(overflow_i, carry_ia) — a circuit op (+1 aux, +1 lookup). The
two bits are mutually exclusive (a carry out of a_i + b_i forces the
partial sum <= 254, so + carry_in cannot carry), so their xor is a
plain field add costing zero columns. Replaced every carry-combining
u8_xor with +. u32_add 70 → 60.

5. Inline u8_recompose and delete it

u8_recompose([b0..b7]) = b0 + 2*b1 + ... + 128*b7 is pure arithmetic,
but each call cost +1 aux +1 lookup. Inlined all uses as explicit
weighted sums — 8 in blake3_g_function, ~36 across sha256_compress
and define_W_i — then deleted the now-unused definition.
blake3_g_function 250 → 210.

Results

Measured with lake exe kernel String.Internal.append:

metricbaselinethis PR
blake3_compress_chunks width34175 (−78%)
blake3_compress_chunks FFT426M94M
u32_add width7060
u32_add FFT602M516M
blake3_g_function width250210
blake3_g_function FFT310M261M
assign_block_value FFT157Meliminated
total FFT cost3109M2516M (−19.1%)

New cold circuits introduced: bytes_to_block, pad_block,
blake3_finish, blake3_compress_block — all small.

Validation

lake test -- --ignored aiur-hashes passes (145 assertions, exit 0),
covering blake3 and sha256 across sizes 0–3168 including block and chunk
boundaries. lake exe kernel String.Internal.append runs clean.

The `blake3_compress_chunks` circuit width was set by its coldest match
arms `(Cons,63,1023)` / `(Cons,63,_)` — the only arms carrying both the
`assign_block_value` call (+64 aux) and the `blake3_compress` call
(+32 aux) plus flag helpers and `store`. Those arms fire on ~1.8% of
chunk-loop rows; the hot buffer-fill arm needs only `assign_block_value`,
so ~60 columns were trash-zero on 98% of rows.
Extract the post-assign tail into a cold `blake3_compress_block` circuit.
The two `(Cons,63,...)` arms were identical once the `blake3_compress`
machinery moved out, so they merge into one `(Cons,63,_)` arm.
Measured with `lake exe kernel String.Internal.append`:
blake3_compress_chunks width 341 -> 281; total FFT cost -72M (-2.3%).
`enum Layer { Push(&Layer, [[G;4];8]), Nil }` is a 34-wide value type
(pointer 1 + digest 32 + tag 1). `blake3_compress_chunks` and
`blake3_compress_block` took and returned it by value, paying 34 columns
of `inputSize` on every row plus 34 aux on every recursive call output.
Switch both to `layer: &Layer -> &Layer`. The push sites change from
`Layer.Push(store(layer), d)` to `store(Layer.Push(layer, d))` — same
single 34-wide store, but the threaded value is now a 1-wide pointer.
`blake3` wraps the initial `Layer.Nil` in `store` and `load`s the result
before `blake3_compress_layer`, which keeps its by-value signature.
Measured with `lake exe kernel String.Internal.append`:
blake3_compress_chunks width 281 -> 215; total FFT cost -83M.
`blake3_compress_chunks` threaded `block_buffer: [[G;4];16]` — 64 columns
of `inputSize` on every row — and filled it via `assign_block_value`,
whose `[[G;4];16]` return cost +64 aux on every hot row. Circuit memory
is content-addressed (no in-place mutation), so a threaded 64-element
accumulator costs 64 columns somewhere on every row. A linked list does
not: appending a byte is a single O(1)-width `store`.
Replace the buffer with a reverse-ordered byte accumulator list:
- Hot chunk loop: `store(ListNode.Cons(head, byte_acc))`, one store/row.
- `bytes_to_block` materializes a full 64-byte list into `[[G;4];16]`
via 64 unrolled `load`s — one cold circuit row per completed block.
- `pad_block` zero-pads the trailing partial block before materializing.
- `blake3_finish` absorbs the three input-exhausted match arms, keeping
their `blake3_compress` call out of the hot loop's width.
- `assign_block_value` is now dead and removed.
Measured with `lake exe kernel String.Internal.append`:
blake3_compress_chunks width 215 -> 75; assign_block_value (157M FFT)
eliminated; total FFT cost -302M (-10.2%).
The byte adders (`u32_add`, `u32_be_add`, `u64_add`,
`relaxed_u64_be_add_2_bytes`) propagate carries as
`u8_xor(overflow_i, carry_ia)`. `u8_xor` is a circuit op costing +1 aux
and +1 lookup. But these two bits are mutually exclusive: a carry out of
`a_i + b_i` forces the partial sum `<= 254`, so the subsequent
`+ carry_in` cannot itself carry. Their xor is therefore a plain field
add, which costs zero columns.
Replace every carry-combining `u8_xor(x, y)` with `x + y`. The data
xors in `u32_xor` are untouched.
Measured with `lake exe kernel String.Internal.append`:
u32_add width 70 -> 60; total FFT cost -86M (-3.2%).
Verified with `lake test -- --ignored aiur-hashes` (145 assertions, exit 0).
`u8_recompose([b0..b7])` is `b0 + 2*b1 + ... + 128*b7` — pure field
arithmetic. As a function call each of its uses cost +1 aux and +1
lookup for nothing; inlined it costs zero columns (no setup op is
spread, unlike a constant-building wrapper).
Inline all uses as explicit weighted sums: 8 in `blake3_g_function`,
~36 across `sha256_compress` and `define_W_i`. Statically-zero terms
collapse (`u8_recompose([0;8])` becomes `0`; trailing `0`s drop). With
no callers left, delete the `u8_recompose` definition.
Measured with `lake exe kernel String.Internal.append`:
blake3_g_function width 250 -> 210; total FFT cost -50M. The sha256
circuits are not exercised by this kernel entry point, so their share
of the gain does not show in this metric.
Verified with `lake test -- --ignored aiur-hashes` (145 assertions, exit 0).
@arthurpaulino
arthurpaulino enabled auto-merge (squash) May 18, 2026 22:18
@arthurpaulino
arthurpaulino merged commit eb4509a into mainMay 18, 2026
14 checks passed
@arthurpaulino
arthurpaulino deleted the ap/blake3 branch May 18, 2026 22:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@arthurpaulino@gabriel-barrett
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Hyperoptimize the blake3_compress_chunks circuit - #413

Merged
arthurpaulino merged 5 commits into
mainfrom
ap/blake3
May 18, 2026
Merged

Hyperoptimize the blake3_compress_chunks circuit#413
arthurpaulino merged 5 commits into
mainfrom
ap/blake3

Conversation

@arthurpaulino

@arthurpaulinoarthurpaulino commented May 18, 2026

Copy link
Copy Markdown
Member

Summary

The blake3 circuits dominated the lake exe kernel String.Internal.append
proving-cost profile. This PR cuts total FFT cost by 19.1% (3109M →
2516M) through five incremental, individually-measured changes, and
removes two now-dead helpers.

Background

Aiur compiles each function to a zk circuit. A circuit's width is
inputSize + selectors + auxiliaries + 4·(1 + lookups), where
auxiliaries and lookups are the maximum over the match arms — so
width is set by the widest arm; a wide cold arm wastes columns on every
row. Cost is width × height × log2(height). Op costs that matter here:
call +outputSize aux +1 lookup; store/load +1 lookup; u8_* ops +1
lookup each; pure field add/sub/mul-by-const cost zero columns.

Changes

1. Factor block-completion out of the hot loop

blake3_compress_chunks width 341 was set by its coldest arms
(Cons,63,1023) / (Cons,63,_) — the only ones carrying both
assign_block_value (+64 aux) and blake3_compress (+32 aux) plus flag
helpers and store, yet firing on ~1.8% of rows. Extracted that tail
into a cold blake3_compress_block circuit; the two arms merged into
one. 341 → 281.

2. Thread Layer by pointer

enum Layer { Push(&Layer, [[G;4];8]), Nil } is a 34-wide value type.
blake3_compress_chunks took and returned it by value — 34 columns of
inputSize every row plus 34 aux on every recursive call. Switched to
layer: &Layer -> &Layer; blake3 adds a store/load at the
boundary so blake3_compress_layer keeps its by-value signature.
281 → 215.

3. Accumulate block bytes in a list, not a 64-wide buffer

blake3_compress_chunks threaded block_buffer: [[G;4];16] (64 columns
of inputSize) and filled it via assign_block_value, whose
[[G;4];16] return cost +64 aux every hot row. Circuit memory is
content-addressed (no in-place mutation), so a threaded 64-element
accumulator costs 64 columns somewhere on every row — a linked list does
not. Replaced it with a reverse-ordered byte accumulator list:

  • Hot loop: store(ListNode.Cons(head, byte_acc)), one store/row.
  • bytes_to_block materializes a full 64-byte list into [[G;4];16] via
    64 unrolled loads — one cold row per completed block.
  • pad_block zero-pads the trailing partial block.
  • blake3_finish absorbs the input-exhausted arms.
  • assign_block_value is now dead and removed.

215 → 75.

4. Combine carry bits with field add, not u8 xor

The byte adders (u32_add, u32_be_add, u64_add,
relaxed_u64_be_add_2_bytes) propagated carries with
u8_xor(overflow_i, carry_ia) — a circuit op (+1 aux, +1 lookup). The
two bits are mutually exclusive (a carry out of a_i + b_i forces the
partial sum <= 254, so + carry_in cannot carry), so their xor is a
plain field add costing zero columns. Replaced every carry-combining
u8_xor with +. u32_add 70 → 60.

5. Inline u8_recompose and delete it

u8_recompose([b0..b7]) = b0 + 2*b1 + ... + 128*b7 is pure arithmetic,
but each call cost +1 aux +1 lookup. Inlined all uses as explicit
weighted sums — 8 in blake3_g_function, ~36 across sha256_compress
and define_W_i — then deleted the now-unused definition.
blake3_g_function 250 → 210.

Results

Measured with lake exe kernel String.Internal.append:

metricbaselinethis PR
blake3_compress_chunks width34175 (−78%)
blake3_compress_chunks FFT426M94M
u32_add width7060
u32_add FFT602M516M
blake3_g_function width250210
blake3_g_function FFT310M261M
assign_block_value FFT157Meliminated
total FFT cost3109M2516M (−19.1%)

New cold circuits introduced: bytes_to_block, pad_block,
blake3_finish, blake3_compress_block — all small.

Validation

lake test -- --ignored aiur-hashes passes (145 assertions, exit 0),
covering blake3 and sha256 across sizes 0–3168 including block and chunk
boundaries. lake exe kernel String.Internal.append runs clean.

The `blake3_compress_chunks` circuit width was set by its coldest match
arms `(Cons,63,1023)` / `(Cons,63,_)` — the only arms carrying both the
`assign_block_value` call (+64 aux) and the `blake3_compress` call
(+32 aux) plus flag helpers and `store`. Those arms fire on ~1.8% of
chunk-loop rows; the hot buffer-fill arm needs only `assign_block_value`,
so ~60 columns were trash-zero on 98% of rows.
Extract the post-assign tail into a cold `blake3_compress_block` circuit.
The two `(Cons,63,...)` arms were identical once the `blake3_compress`
machinery moved out, so they merge into one `(Cons,63,_)` arm.
Measured with `lake exe kernel String.Internal.append`:
blake3_compress_chunks width 341 -> 281; total FFT cost -72M (-2.3%).
`enum Layer { Push(&Layer, [[G;4];8]), Nil }` is a 34-wide value type
(pointer 1 + digest 32 + tag 1). `blake3_compress_chunks` and
`blake3_compress_block` took and returned it by value, paying 34 columns
of `inputSize` on every row plus 34 aux on every recursive call output.
Switch both to `layer: &Layer -> &Layer`. The push sites change from
`Layer.Push(store(layer), d)` to `store(Layer.Push(layer, d))` — same
single 34-wide store, but the threaded value is now a 1-wide pointer.
`blake3` wraps the initial `Layer.Nil` in `store` and `load`s the result
before `blake3_compress_layer`, which keeps its by-value signature.
Measured with `lake exe kernel String.Internal.append`:
blake3_compress_chunks width 281 -> 215; total FFT cost -83M.
`blake3_compress_chunks` threaded `block_buffer: [[G;4];16]` — 64 columns
of `inputSize` on every row — and filled it via `assign_block_value`,
whose `[[G;4];16]` return cost +64 aux on every hot row. Circuit memory
is content-addressed (no in-place mutation), so a threaded 64-element
accumulator costs 64 columns somewhere on every row. A linked list does
not: appending a byte is a single O(1)-width `store`.
Replace the buffer with a reverse-ordered byte accumulator list:
- Hot chunk loop: `store(ListNode.Cons(head, byte_acc))`, one store/row.
- `bytes_to_block` materializes a full 64-byte list into `[[G;4];16]`
via 64 unrolled `load`s — one cold circuit row per completed block.
- `pad_block` zero-pads the trailing partial block before materializing.
- `blake3_finish` absorbs the three input-exhausted match arms, keeping
their `blake3_compress` call out of the hot loop's width.
- `assign_block_value` is now dead and removed.
Measured with `lake exe kernel String.Internal.append`:
blake3_compress_chunks width 215 -> 75; assign_block_value (157M FFT)
eliminated; total FFT cost -302M (-10.2%).
The byte adders (`u32_add`, `u32_be_add`, `u64_add`,
`relaxed_u64_be_add_2_bytes`) propagate carries as
`u8_xor(overflow_i, carry_ia)`. `u8_xor` is a circuit op costing +1 aux
and +1 lookup. But these two bits are mutually exclusive: a carry out of
`a_i + b_i` forces the partial sum `<= 254`, so the subsequent
`+ carry_in` cannot itself carry. Their xor is therefore a plain field
add, which costs zero columns.
Replace every carry-combining `u8_xor(x, y)` with `x + y`. The data
xors in `u32_xor` are untouched.
Measured with `lake exe kernel String.Internal.append`:
u32_add width 70 -> 60; total FFT cost -86M (-3.2%).
Verified with `lake test -- --ignored aiur-hashes` (145 assertions, exit 0).
`u8_recompose([b0..b7])` is `b0 + 2*b1 + ... + 128*b7` — pure field
arithmetic. As a function call each of its uses cost +1 aux and +1
lookup for nothing; inlined it costs zero columns (no setup op is
spread, unlike a constant-building wrapper).
Inline all uses as explicit weighted sums: 8 in `blake3_g_function`,
~36 across `sha256_compress` and `define_W_i`. Statically-zero terms
collapse (`u8_recompose([0;8])` becomes `0`; trailing `0`s drop). With
no callers left, delete the `u8_recompose` definition.
Measured with `lake exe kernel String.Internal.append`:
blake3_g_function width 250 -> 210; total FFT cost -50M. The sha256
circuits are not exercised by this kernel entry point, so their share
of the gain does not show in this metric.
Verified with `lake test -- --ignored aiur-hashes` (145 assertions, exit 0).
@arthurpaulino
arthurpaulino enabled auto-merge (squash) May 18, 2026 22:18
@arthurpaulino
arthurpaulino merged commit eb4509a into mainMay 18, 2026
14 checks passed
@arthurpaulino
arthurpaulino deleted the ap/blake3 branch May 18, 2026 22:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@arthurpaulino@gabriel-barrett
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Hyperoptimize the blake3_compress_chunks circuit - #413

Merged
arthurpaulino merged 5 commits into
mainfrom
ap/blake3
May 18, 2026
Merged

Hyperoptimize the blake3_compress_chunks circuit#413
arthurpaulino merged 5 commits into
mainfrom
ap/blake3

Conversation

@arthurpaulino

@arthurpaulinoarthurpaulino commented May 18, 2026

Copy link
Copy Markdown
Member

Summary

The blake3 circuits dominated the lake exe kernel String.Internal.append
proving-cost profile. This PR cuts total FFT cost by 19.1% (3109M →
2516M) through five incremental, individually-measured changes, and
removes two now-dead helpers.

Background

Aiur compiles each function to a zk circuit. A circuit's width is
inputSize + selectors + auxiliaries + 4·(1 + lookups), where
auxiliaries and lookups are the maximum over the match arms — so
width is set by the widest arm; a wide cold arm wastes columns on every
row. Cost is width × height × log2(height). Op costs that matter here:
call +outputSize aux +1 lookup; store/load +1 lookup; u8_* ops +1
lookup each; pure field add/sub/mul-by-const cost zero columns.

Changes

1. Factor block-completion out of the hot loop

blake3_compress_chunks width 341 was set by its coldest arms
(Cons,63,1023) / (Cons,63,_) — the only ones carrying both
assign_block_value (+64 aux) and blake3_compress (+32 aux) plus flag
helpers and store, yet firing on ~1.8% of rows. Extracted that tail
into a cold blake3_compress_block circuit; the two arms merged into
one. 341 → 281.

2. Thread Layer by pointer

enum Layer { Push(&Layer, [[G;4];8]), Nil } is a 34-wide value type.
blake3_compress_chunks took and returned it by value — 34 columns of
inputSize every row plus 34 aux on every recursive call. Switched to
layer: &Layer -> &Layer; blake3 adds a store/load at the
boundary so blake3_compress_layer keeps its by-value signature.
281 → 215.

3. Accumulate block bytes in a list, not a 64-wide buffer

blake3_compress_chunks threaded block_buffer: [[G;4];16] (64 columns
of inputSize) and filled it via assign_block_value, whose
[[G;4];16] return cost +64 aux every hot row. Circuit memory is
content-addressed (no in-place mutation), so a threaded 64-element
accumulator costs 64 columns somewhere on every row — a linked list does
not. Replaced it with a reverse-ordered byte accumulator list:

  • Hot loop: store(ListNode.Cons(head, byte_acc)), one store/row.
  • bytes_to_block materializes a full 64-byte list into [[G;4];16] via
    64 unrolled loads — one cold row per completed block.
  • pad_block zero-pads the trailing partial block.
  • blake3_finish absorbs the input-exhausted arms.
  • assign_block_value is now dead and removed.

215 → 75.

4. Combine carry bits with field add, not u8 xor

The byte adders (u32_add, u32_be_add, u64_add,
relaxed_u64_be_add_2_bytes) propagated carries with
u8_xor(overflow_i, carry_ia) — a circuit op (+1 aux, +1 lookup). The
two bits are mutually exclusive (a carry out of a_i + b_i forces the
partial sum <= 254, so + carry_in cannot carry), so their xor is a
plain field add costing zero columns. Replaced every carry-combining
u8_xor with +. u32_add 70 → 60.

5. Inline u8_recompose and delete it

u8_recompose([b0..b7]) = b0 + 2*b1 + ... + 128*b7 is pure arithmetic,
but each call cost +1 aux +1 lookup. Inlined all uses as explicit
weighted sums — 8 in blake3_g_function, ~36 across sha256_compress
and define_W_i — then deleted the now-unused definition.
blake3_g_function 250 → 210.

Results

Measured with lake exe kernel String.Internal.append:

metricbaselinethis PR
blake3_compress_chunks width34175 (−78%)
blake3_compress_chunks FFT426M94M
u32_add width7060
u32_add FFT602M516M
blake3_g_function width250210
blake3_g_function FFT310M261M
assign_block_value FFT157Meliminated
total FFT cost3109M2516M (−19.1%)

New cold circuits introduced: bytes_to_block, pad_block,
blake3_finish, blake3_compress_block — all small.

Validation

lake test -- --ignored aiur-hashes passes (145 assertions, exit 0),
covering blake3 and sha256 across sizes 0–3168 including block and chunk
boundaries. lake exe kernel String.Internal.append runs clean.

The `blake3_compress_chunks` circuit width was set by its coldest match
arms `(Cons,63,1023)` / `(Cons,63,_)` — the only arms carrying both the
`assign_block_value` call (+64 aux) and the `blake3_compress` call
(+32 aux) plus flag helpers and `store`. Those arms fire on ~1.8% of
chunk-loop rows; the hot buffer-fill arm needs only `assign_block_value`,
so ~60 columns were trash-zero on 98% of rows.
Extract the post-assign tail into a cold `blake3_compress_block` circuit.
The two `(Cons,63,...)` arms were identical once the `blake3_compress`
machinery moved out, so they merge into one `(Cons,63,_)` arm.
Measured with `lake exe kernel String.Internal.append`:
blake3_compress_chunks width 341 -> 281; total FFT cost -72M (-2.3%).
`enum Layer { Push(&Layer, [[G;4];8]), Nil }` is a 34-wide value type
(pointer 1 + digest 32 + tag 1). `blake3_compress_chunks` and
`blake3_compress_block` took and returned it by value, paying 34 columns
of `inputSize` on every row plus 34 aux on every recursive call output.
Switch both to `layer: &Layer -> &Layer`. The push sites change from
`Layer.Push(store(layer), d)` to `store(Layer.Push(layer, d))` — same
single 34-wide store, but the threaded value is now a 1-wide pointer.
`blake3` wraps the initial `Layer.Nil` in `store` and `load`s the result
before `blake3_compress_layer`, which keeps its by-value signature.
Measured with `lake exe kernel String.Internal.append`:
blake3_compress_chunks width 281 -> 215; total FFT cost -83M.
`blake3_compress_chunks` threaded `block_buffer: [[G;4];16]` — 64 columns
of `inputSize` on every row — and filled it via `assign_block_value`,
whose `[[G;4];16]` return cost +64 aux on every hot row. Circuit memory
is content-addressed (no in-place mutation), so a threaded 64-element
accumulator costs 64 columns somewhere on every row. A linked list does
not: appending a byte is a single O(1)-width `store`.
Replace the buffer with a reverse-ordered byte accumulator list:
- Hot chunk loop: `store(ListNode.Cons(head, byte_acc))`, one store/row.
- `bytes_to_block` materializes a full 64-byte list into `[[G;4];16]`
via 64 unrolled `load`s — one cold circuit row per completed block.
- `pad_block` zero-pads the trailing partial block before materializing.
- `blake3_finish` absorbs the three input-exhausted match arms, keeping
their `blake3_compress` call out of the hot loop's width.
- `assign_block_value` is now dead and removed.
Measured with `lake exe kernel String.Internal.append`:
blake3_compress_chunks width 215 -> 75; assign_block_value (157M FFT)
eliminated; total FFT cost -302M (-10.2%).
The byte adders (`u32_add`, `u32_be_add`, `u64_add`,
`relaxed_u64_be_add_2_bytes`) propagate carries as
`u8_xor(overflow_i, carry_ia)`. `u8_xor` is a circuit op costing +1 aux
and +1 lookup. But these two bits are mutually exclusive: a carry out of
`a_i + b_i` forces the partial sum `<= 254`, so the subsequent
`+ carry_in` cannot itself carry. Their xor is therefore a plain field
add, which costs zero columns.
Replace every carry-combining `u8_xor(x, y)` with `x + y`. The data
xors in `u32_xor` are untouched.
Measured with `lake exe kernel String.Internal.append`:
u32_add width 70 -> 60; total FFT cost -86M (-3.2%).
Verified with `lake test -- --ignored aiur-hashes` (145 assertions, exit 0).
`u8_recompose([b0..b7])` is `b0 + 2*b1 + ... + 128*b7` — pure field
arithmetic. As a function call each of its uses cost +1 aux and +1
lookup for nothing; inlined it costs zero columns (no setup op is
spread, unlike a constant-building wrapper).
Inline all uses as explicit weighted sums: 8 in `blake3_g_function`,
~36 across `sha256_compress` and `define_W_i`. Statically-zero terms
collapse (`u8_recompose([0;8])` becomes `0`; trailing `0`s drop). With
no callers left, delete the `u8_recompose` definition.
Measured with `lake exe kernel String.Internal.append`:
blake3_g_function width 250 -> 210; total FFT cost -50M. The sha256
circuits are not exercised by this kernel entry point, so their share
of the gain does not show in this metric.
Verified with `lake test -- --ignored aiur-hashes` (145 assertions, exit 0).
@arthurpaulino
arthurpaulino enabled auto-merge (squash) May 18, 2026 22:18
@arthurpaulino
arthurpaulino merged commit eb4509a into mainMay 18, 2026
14 checks passed
@arthurpaulino
arthurpaulino deleted the ap/blake3 branch May 18, 2026 22:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@arthurpaulino@gabriel-barrett
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Hyperoptimize the blake3_compress_chunks circuit - #413

Merged
arthurpaulino merged 5 commits into
mainfrom
ap/blake3
May 18, 2026
Merged

Hyperoptimize the blake3_compress_chunks circuit#413
arthurpaulino merged 5 commits into
mainfrom
ap/blake3

Conversation

@arthurpaulino

@arthurpaulinoarthurpaulino commented May 18, 2026

Copy link
Copy Markdown
Member

Summary

The blake3 circuits dominated the lake exe kernel String.Internal.append
proving-cost profile. This PR cuts total FFT cost by 19.1% (3109M →
2516M) through five incremental, individually-measured changes, and
removes two now-dead helpers.

Background

Aiur compiles each function to a zk circuit. A circuit's width is
inputSize + selectors + auxiliaries + 4·(1 + lookups), where
auxiliaries and lookups are the maximum over the match arms — so
width is set by the widest arm; a wide cold arm wastes columns on every
row. Cost is width × height × log2(height). Op costs that matter here:
call +outputSize aux +1 lookup; store/load +1 lookup; u8_* ops +1
lookup each; pure field add/sub/mul-by-const cost zero columns.

Changes

1. Factor block-completion out of the hot loop

blake3_compress_chunks width 341 was set by its coldest arms
(Cons,63,1023) / (Cons,63,_) — the only ones carrying both
assign_block_value (+64 aux) and blake3_compress (+32 aux) plus flag
helpers and store, yet firing on ~1.8% of rows. Extracted that tail
into a cold blake3_compress_block circuit; the two arms merged into
one. 341 → 281.

2. Thread Layer by pointer

enum Layer { Push(&Layer, [[G;4];8]), Nil } is a 34-wide value type.
blake3_compress_chunks took and returned it by value — 34 columns of
inputSize every row plus 34 aux on every recursive call. Switched to
layer: &Layer -> &Layer; blake3 adds a store/load at the
boundary so blake3_compress_layer keeps its by-value signature.
281 → 215.

3. Accumulate block bytes in a list, not a 64-wide buffer

blake3_compress_chunks threaded block_buffer: [[G;4];16] (64 columns
of inputSize) and filled it via assign_block_value, whose
[[G;4];16] return cost +64 aux every hot row. Circuit memory is
content-addressed (no in-place mutation), so a threaded 64-element
accumulator costs 64 columns somewhere on every row — a linked list does
not. Replaced it with a reverse-ordered byte accumulator list:

  • Hot loop: store(ListNode.Cons(head, byte_acc)), one store/row.
  • bytes_to_block materializes a full 64-byte list into [[G;4];16] via
    64 unrolled loads — one cold row per completed block.
  • pad_block zero-pads the trailing partial block.
  • blake3_finish absorbs the input-exhausted arms.
  • assign_block_value is now dead and removed.

215 → 75.

4. Combine carry bits with field add, not u8 xor

The byte adders (u32_add, u32_be_add, u64_add,
relaxed_u64_be_add_2_bytes) propagated carries with
u8_xor(overflow_i, carry_ia) — a circuit op (+1 aux, +1 lookup). The
two bits are mutually exclusive (a carry out of a_i + b_i forces the
partial sum <= 254, so + carry_in cannot carry), so their xor is a
plain field add costing zero columns. Replaced every carry-combining
u8_xor with +. u32_add 70 → 60.

5. Inline u8_recompose and delete it

u8_recompose([b0..b7]) = b0 + 2*b1 + ... + 128*b7 is pure arithmetic,
but each call cost +1 aux +1 lookup. Inlined all uses as explicit
weighted sums — 8 in blake3_g_function, ~36 across sha256_compress
and define_W_i — then deleted the now-unused definition.
blake3_g_function 250 → 210.

Results

Measured with lake exe kernel String.Internal.append:

metricbaselinethis PR
blake3_compress_chunks width34175 (−78%)
blake3_compress_chunks FFT426M94M
u32_add width7060
u32_add FFT602M516M
blake3_g_function width250210
blake3_g_function FFT310M261M
assign_block_value FFT157Meliminated
total FFT cost3109M2516M (−19.1%)

New cold circuits introduced: bytes_to_block, pad_block,
blake3_finish, blake3_compress_block — all small.

Validation

lake test -- --ignored aiur-hashes passes (145 assertions, exit 0),
covering blake3 and sha256 across sizes 0–3168 including block and chunk
boundaries. lake exe kernel String.Internal.append runs clean.

The `blake3_compress_chunks` circuit width was set by its coldest match
arms `(Cons,63,1023)` / `(Cons,63,_)` — the only arms carrying both the
`assign_block_value` call (+64 aux) and the `blake3_compress` call
(+32 aux) plus flag helpers and `store`. Those arms fire on ~1.8% of
chunk-loop rows; the hot buffer-fill arm needs only `assign_block_value`,
so ~60 columns were trash-zero on 98% of rows.
Extract the post-assign tail into a cold `blake3_compress_block` circuit.
The two `(Cons,63,...)` arms were identical once the `blake3_compress`
machinery moved out, so they merge into one `(Cons,63,_)` arm.
Measured with `lake exe kernel String.Internal.append`:
blake3_compress_chunks width 341 -> 281; total FFT cost -72M (-2.3%).
`enum Layer { Push(&Layer, [[G;4];8]), Nil }` is a 34-wide value type
(pointer 1 + digest 32 + tag 1). `blake3_compress_chunks` and
`blake3_compress_block` took and returned it by value, paying 34 columns
of `inputSize` on every row plus 34 aux on every recursive call output.
Switch both to `layer: &Layer -> &Layer`. The push sites change from
`Layer.Push(store(layer), d)` to `store(Layer.Push(layer, d))` — same
single 34-wide store, but the threaded value is now a 1-wide pointer.
`blake3` wraps the initial `Layer.Nil` in `store` and `load`s the result
before `blake3_compress_layer`, which keeps its by-value signature.
Measured with `lake exe kernel String.Internal.append`:
blake3_compress_chunks width 281 -> 215; total FFT cost -83M.
`blake3_compress_chunks` threaded `block_buffer: [[G;4];16]` — 64 columns
of `inputSize` on every row — and filled it via `assign_block_value`,
whose `[[G;4];16]` return cost +64 aux on every hot row. Circuit memory
is content-addressed (no in-place mutation), so a threaded 64-element
accumulator costs 64 columns somewhere on every row. A linked list does
not: appending a byte is a single O(1)-width `store`.
Replace the buffer with a reverse-ordered byte accumulator list:
- Hot chunk loop: `store(ListNode.Cons(head, byte_acc))`, one store/row.
- `bytes_to_block` materializes a full 64-byte list into `[[G;4];16]`
via 64 unrolled `load`s — one cold circuit row per completed block.
- `pad_block` zero-pads the trailing partial block before materializing.
- `blake3_finish` absorbs the three input-exhausted match arms, keeping
their `blake3_compress` call out of the hot loop's width.
- `assign_block_value` is now dead and removed.
Measured with `lake exe kernel String.Internal.append`:
blake3_compress_chunks width 215 -> 75; assign_block_value (157M FFT)
eliminated; total FFT cost -302M (-10.2%).
The byte adders (`u32_add`, `u32_be_add`, `u64_add`,
`relaxed_u64_be_add_2_bytes`) propagate carries as
`u8_xor(overflow_i, carry_ia)`. `u8_xor` is a circuit op costing +1 aux
and +1 lookup. But these two bits are mutually exclusive: a carry out of
`a_i + b_i` forces the partial sum `<= 254`, so the subsequent
`+ carry_in` cannot itself carry. Their xor is therefore a plain field
add, which costs zero columns.
Replace every carry-combining `u8_xor(x, y)` with `x + y`. The data
xors in `u32_xor` are untouched.
Measured with `lake exe kernel String.Internal.append`:
u32_add width 70 -> 60; total FFT cost -86M (-3.2%).
Verified with `lake test -- --ignored aiur-hashes` (145 assertions, exit 0).
`u8_recompose([b0..b7])` is `b0 + 2*b1 + ... + 128*b7` — pure field
arithmetic. As a function call each of its uses cost +1 aux and +1
lookup for nothing; inlined it costs zero columns (no setup op is
spread, unlike a constant-building wrapper).
Inline all uses as explicit weighted sums: 8 in `blake3_g_function`,
~36 across `sha256_compress` and `define_W_i`. Statically-zero terms
collapse (`u8_recompose([0;8])` becomes `0`; trailing `0`s drop). With
no callers left, delete the `u8_recompose` definition.
Measured with `lake exe kernel String.Internal.append`:
blake3_g_function width 250 -> 210; total FFT cost -50M. The sha256
circuits are not exercised by this kernel entry point, so their share
of the gain does not show in this metric.
Verified with `lake test -- --ignored aiur-hashes` (145 assertions, exit 0).
@arthurpaulino
arthurpaulino enabled auto-merge (squash) May 18, 2026 22:18
@arthurpaulino
arthurpaulino merged commit eb4509a into mainMay 18, 2026
14 checks passed
@arthurpaulino
arthurpaulino deleted the ap/blake3 branch May 18, 2026 22:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@arthurpaulino@gabriel-barrett
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Hyperoptimize the blake3_compress_chunks circuit - #413

Merged
arthurpaulino merged 5 commits into
mainfrom
ap/blake3
May 18, 2026
Merged

Hyperoptimize the blake3_compress_chunks circuit#413
arthurpaulino merged 5 commits into
mainfrom
ap/blake3

Conversation

@arthurpaulino

@arthurpaulinoarthurpaulino commented May 18, 2026

Copy link
Copy Markdown
Member

Summary

The blake3 circuits dominated the lake exe kernel String.Internal.append
proving-cost profile. This PR cuts total FFT cost by 19.1% (3109M →
2516M) through five incremental, individually-measured changes, and
removes two now-dead helpers.

Background

Aiur compiles each function to a zk circuit. A circuit's width is
inputSize + selectors + auxiliaries + 4·(1 + lookups), where
auxiliaries and lookups are the maximum over the match arms — so
width is set by the widest arm; a wide cold arm wastes columns on every
row. Cost is width × height × log2(height). Op costs that matter here:
call +outputSize aux +1 lookup; store/load +1 lookup; u8_* ops +1
lookup each; pure field add/sub/mul-by-const cost zero columns.

Changes

1. Factor block-completion out of the hot loop

blake3_compress_chunks width 341 was set by its coldest arms
(Cons,63,1023) / (Cons,63,_) — the only ones carrying both
assign_block_value (+64 aux) and blake3_compress (+32 aux) plus flag
helpers and store, yet firing on ~1.8% of rows. Extracted that tail
into a cold blake3_compress_block circuit; the two arms merged into
one. 341 → 281.

2. Thread Layer by pointer

enum Layer { Push(&Layer, [[G;4];8]), Nil } is a 34-wide value type.
blake3_compress_chunks took and returned it by value — 34 columns of
inputSize every row plus 34 aux on every recursive call. Switched to
layer: &Layer -> &Layer; blake3 adds a store/load at the
boundary so blake3_compress_layer keeps its by-value signature.
281 → 215.

3. Accumulate block bytes in a list, not a 64-wide buffer

blake3_compress_chunks threaded block_buffer: [[G;4];16] (64 columns
of inputSize) and filled it via assign_block_value, whose
[[G;4];16] return cost +64 aux every hot row. Circuit memory is
content-addressed (no in-place mutation), so a threaded 64-element
accumulator costs 64 columns somewhere on every row — a linked list does
not. Replaced it with a reverse-ordered byte accumulator list:

  • Hot loop: store(ListNode.Cons(head, byte_acc)), one store/row.
  • bytes_to_block materializes a full 64-byte list into [[G;4];16] via
    64 unrolled loads — one cold row per completed block.
  • pad_block zero-pads the trailing partial block.
  • blake3_finish absorbs the input-exhausted arms.
  • assign_block_value is now dead and removed.

215 → 75.

4. Combine carry bits with field add, not u8 xor

The byte adders (u32_add, u32_be_add, u64_add,
relaxed_u64_be_add_2_bytes) propagated carries with
u8_xor(overflow_i, carry_ia) — a circuit op (+1 aux, +1 lookup). The
two bits are mutually exclusive (a carry out of a_i + b_i forces the
partial sum <= 254, so + carry_in cannot carry), so their xor is a
plain field add costing zero columns. Replaced every carry-combining
u8_xor with +. u32_add 70 → 60.

5. Inline u8_recompose and delete it

u8_recompose([b0..b7]) = b0 + 2*b1 + ... + 128*b7 is pure arithmetic,
but each call cost +1 aux +1 lookup. Inlined all uses as explicit
weighted sums — 8 in blake3_g_function, ~36 across sha256_compress
and define_W_i — then deleted the now-unused definition.
blake3_g_function 250 → 210.

Results

Measured with lake exe kernel String.Internal.append:

metricbaselinethis PR
blake3_compress_chunks width34175 (−78%)
blake3_compress_chunks FFT426M94M
u32_add width7060
u32_add FFT602M516M
blake3_g_function width250210
blake3_g_function FFT310M261M
assign_block_value FFT157Meliminated
total FFT cost3109M2516M (−19.1%)

New cold circuits introduced: bytes_to_block, pad_block,
blake3_finish, blake3_compress_block — all small.

Validation

lake test -- --ignored aiur-hashes passes (145 assertions, exit 0),
covering blake3 and sha256 across sizes 0–3168 including block and chunk
boundaries. lake exe kernel String.Internal.append runs clean.

The `blake3_compress_chunks` circuit width was set by its coldest match
arms `(Cons,63,1023)` / `(Cons,63,_)` — the only arms carrying both the
`assign_block_value` call (+64 aux) and the `blake3_compress` call
(+32 aux) plus flag helpers and `store`. Those arms fire on ~1.8% of
chunk-loop rows; the hot buffer-fill arm needs only `assign_block_value`,
so ~60 columns were trash-zero on 98% of rows.
Extract the post-assign tail into a cold `blake3_compress_block` circuit.
The two `(Cons,63,...)` arms were identical once the `blake3_compress`
machinery moved out, so they merge into one `(Cons,63,_)` arm.
Measured with `lake exe kernel String.Internal.append`:
blake3_compress_chunks width 341 -> 281; total FFT cost -72M (-2.3%).
`enum Layer { Push(&Layer, [[G;4];8]), Nil }` is a 34-wide value type
(pointer 1 + digest 32 + tag 1). `blake3_compress_chunks` and
`blake3_compress_block` took and returned it by value, paying 34 columns
of `inputSize` on every row plus 34 aux on every recursive call output.
Switch both to `layer: &Layer -> &Layer`. The push sites change from
`Layer.Push(store(layer), d)` to `store(Layer.Push(layer, d))` — same
single 34-wide store, but the threaded value is now a 1-wide pointer.
`blake3` wraps the initial `Layer.Nil` in `store` and `load`s the result
before `blake3_compress_layer`, which keeps its by-value signature.
Measured with `lake exe kernel String.Internal.append`:
blake3_compress_chunks width 281 -> 215; total FFT cost -83M.
`blake3_compress_chunks` threaded `block_buffer: [[G;4];16]` — 64 columns
of `inputSize` on every row — and filled it via `assign_block_value`,
whose `[[G;4];16]` return cost +64 aux on every hot row. Circuit memory
is content-addressed (no in-place mutation), so a threaded 64-element
accumulator costs 64 columns somewhere on every row. A linked list does
not: appending a byte is a single O(1)-width `store`.
Replace the buffer with a reverse-ordered byte accumulator list:
- Hot chunk loop: `store(ListNode.Cons(head, byte_acc))`, one store/row.
- `bytes_to_block` materializes a full 64-byte list into `[[G;4];16]`
via 64 unrolled `load`s — one cold circuit row per completed block.
- `pad_block` zero-pads the trailing partial block before materializing.
- `blake3_finish` absorbs the three input-exhausted match arms, keeping
their `blake3_compress` call out of the hot loop's width.
- `assign_block_value` is now dead and removed.
Measured with `lake exe kernel String.Internal.append`:
blake3_compress_chunks width 215 -> 75; assign_block_value (157M FFT)
eliminated; total FFT cost -302M (-10.2%).
The byte adders (`u32_add`, `u32_be_add`, `u64_add`,
`relaxed_u64_be_add_2_bytes`) propagate carries as
`u8_xor(overflow_i, carry_ia)`. `u8_xor` is a circuit op costing +1 aux
and +1 lookup. But these two bits are mutually exclusive: a carry out of
`a_i + b_i` forces the partial sum `<= 254`, so the subsequent
`+ carry_in` cannot itself carry. Their xor is therefore a plain field
add, which costs zero columns.
Replace every carry-combining `u8_xor(x, y)` with `x + y`. The data
xors in `u32_xor` are untouched.
Measured with `lake exe kernel String.Internal.append`:
u32_add width 70 -> 60; total FFT cost -86M (-3.2%).
Verified with `lake test -- --ignored aiur-hashes` (145 assertions, exit 0).
`u8_recompose([b0..b7])` is `b0 + 2*b1 + ... + 128*b7` — pure field
arithmetic. As a function call each of its uses cost +1 aux and +1
lookup for nothing; inlined it costs zero columns (no setup op is
spread, unlike a constant-building wrapper).
Inline all uses as explicit weighted sums: 8 in `blake3_g_function`,
~36 across `sha256_compress` and `define_W_i`. Statically-zero terms
collapse (`u8_recompose([0;8])` becomes `0`; trailing `0`s drop). With
no callers left, delete the `u8_recompose` definition.
Measured with `lake exe kernel String.Internal.append`:
blake3_g_function width 250 -> 210; total FFT cost -50M. The sha256
circuits are not exercised by this kernel entry point, so their share
of the gain does not show in this metric.
Verified with `lake test -- --ignored aiur-hashes` (145 assertions, exit 0).
@arthurpaulino
arthurpaulino enabled auto-merge (squash) May 18, 2026 22:18
@arthurpaulino
arthurpaulino merged commit eb4509a into mainMay 18, 2026
14 checks passed
@arthurpaulino
arthurpaulino deleted the ap/blake3 branch May 18, 2026 22:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@arthurpaulino@gabriel-barrett
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Hyperoptimize the blake3_compress_chunks circuit - #413

Merged
arthurpaulino merged 5 commits into
mainfrom
ap/blake3
May 18, 2026
Merged

Hyperoptimize the blake3_compress_chunks circuit#413
arthurpaulino merged 5 commits into
mainfrom
ap/blake3

Conversation

@arthurpaulino

@arthurpaulinoarthurpaulino commented May 18, 2026

Copy link
Copy Markdown
Member

Summary

The blake3 circuits dominated the lake exe kernel String.Internal.append
proving-cost profile. This PR cuts total FFT cost by 19.1% (3109M →
2516M) through five incremental, individually-measured changes, and
removes two now-dead helpers.

Background

Aiur compiles each function to a zk circuit. A circuit's width is
inputSize + selectors + auxiliaries + 4·(1 + lookups), where
auxiliaries and lookups are the maximum over the match arms — so
width is set by the widest arm; a wide cold arm wastes columns on every
row. Cost is width × height × log2(height). Op costs that matter here:
call +outputSize aux +1 lookup; store/load +1 lookup; u8_* ops +1
lookup each; pure field add/sub/mul-by-const cost zero columns.

Changes

1. Factor block-completion out of the hot loop

blake3_compress_chunks width 341 was set by its coldest arms
(Cons,63,1023) / (Cons,63,_) — the only ones carrying both
assign_block_value (+64 aux) and blake3_compress (+32 aux) plus flag
helpers and store, yet firing on ~1.8% of rows. Extracted that tail
into a cold blake3_compress_block circuit; the two arms merged into
one. 341 → 281.

2. Thread Layer by pointer

enum Layer { Push(&Layer, [[G;4];8]), Nil } is a 34-wide value type.
blake3_compress_chunks took and returned it by value — 34 columns of
inputSize every row plus 34 aux on every recursive call. Switched to
layer: &Layer -> &Layer; blake3 adds a store/load at the
boundary so blake3_compress_layer keeps its by-value signature.
281 → 215.

3. Accumulate block bytes in a list, not a 64-wide buffer

blake3_compress_chunks threaded block_buffer: [[G;4];16] (64 columns
of inputSize) and filled it via assign_block_value, whose
[[G;4];16] return cost +64 aux every hot row. Circuit memory is
content-addressed (no in-place mutation), so a threaded 64-element
accumulator costs 64 columns somewhere on every row — a linked list does
not. Replaced it with a reverse-ordered byte accumulator list:

  • Hot loop: store(ListNode.Cons(head, byte_acc)), one store/row.
  • bytes_to_block materializes a full 64-byte list into [[G;4];16] via
    64 unrolled loads — one cold row per completed block.
  • pad_block zero-pads the trailing partial block.
  • blake3_finish absorbs the input-exhausted arms.
  • assign_block_value is now dead and removed.

215 → 75.

4. Combine carry bits with field add, not u8 xor

The byte adders (u32_add, u32_be_add, u64_add,
relaxed_u64_be_add_2_bytes) propagated carries with
u8_xor(overflow_i, carry_ia) — a circuit op (+1 aux, +1 lookup). The
two bits are mutually exclusive (a carry out of a_i + b_i forces the
partial sum <= 254, so + carry_in cannot carry), so their xor is a
plain field add costing zero columns. Replaced every carry-combining
u8_xor with +. u32_add 70 → 60.

5. Inline u8_recompose and delete it

u8_recompose([b0..b7]) = b0 + 2*b1 + ... + 128*b7 is pure arithmetic,
but each call cost +1 aux +1 lookup. Inlined all uses as explicit
weighted sums — 8 in blake3_g_function, ~36 across sha256_compress
and define_W_i — then deleted the now-unused definition.
blake3_g_function 250 → 210.

Results

Measured with lake exe kernel String.Internal.append:

metricbaselinethis PR
blake3_compress_chunks width34175 (−78%)
blake3_compress_chunks FFT426M94M
u32_add width7060
u32_add FFT602M516M
blake3_g_function width250210
blake3_g_function FFT310M261M
assign_block_value FFT157Meliminated
total FFT cost3109M2516M (−19.1%)

New cold circuits introduced: bytes_to_block, pad_block,
blake3_finish, blake3_compress_block — all small.

Validation

lake test -- --ignored aiur-hashes passes (145 assertions, exit 0),
covering blake3 and sha256 across sizes 0–3168 including block and chunk
boundaries. lake exe kernel String.Internal.append runs clean.

The `blake3_compress_chunks` circuit width was set by its coldest match
arms `(Cons,63,1023)` / `(Cons,63,_)` — the only arms carrying both the
`assign_block_value` call (+64 aux) and the `blake3_compress` call
(+32 aux) plus flag helpers and `store`. Those arms fire on ~1.8% of
chunk-loop rows; the hot buffer-fill arm needs only `assign_block_value`,
so ~60 columns were trash-zero on 98% of rows.
Extract the post-assign tail into a cold `blake3_compress_block` circuit.
The two `(Cons,63,...)` arms were identical once the `blake3_compress`
machinery moved out, so they merge into one `(Cons,63,_)` arm.
Measured with `lake exe kernel String.Internal.append`:
blake3_compress_chunks width 341 -> 281; total FFT cost -72M (-2.3%).
`enum Layer { Push(&Layer, [[G;4];8]), Nil }` is a 34-wide value type
(pointer 1 + digest 32 + tag 1). `blake3_compress_chunks` and
`blake3_compress_block` took and returned it by value, paying 34 columns
of `inputSize` on every row plus 34 aux on every recursive call output.
Switch both to `layer: &Layer -> &Layer`. The push sites change from
`Layer.Push(store(layer), d)` to `store(Layer.Push(layer, d))` — same
single 34-wide store, but the threaded value is now a 1-wide pointer.
`blake3` wraps the initial `Layer.Nil` in `store` and `load`s the result
before `blake3_compress_layer`, which keeps its by-value signature.
Measured with `lake exe kernel String.Internal.append`:
blake3_compress_chunks width 281 -> 215; total FFT cost -83M.
`blake3_compress_chunks` threaded `block_buffer: [[G;4];16]` — 64 columns
of `inputSize` on every row — and filled it via `assign_block_value`,
whose `[[G;4];16]` return cost +64 aux on every hot row. Circuit memory
is content-addressed (no in-place mutation), so a threaded 64-element
accumulator costs 64 columns somewhere on every row. A linked list does
not: appending a byte is a single O(1)-width `store`.
Replace the buffer with a reverse-ordered byte accumulator list:
- Hot chunk loop: `store(ListNode.Cons(head, byte_acc))`, one store/row.
- `bytes_to_block` materializes a full 64-byte list into `[[G;4];16]`
via 64 unrolled `load`s — one cold circuit row per completed block.
- `pad_block` zero-pads the trailing partial block before materializing.
- `blake3_finish` absorbs the three input-exhausted match arms, keeping
their `blake3_compress` call out of the hot loop's width.
- `assign_block_value` is now dead and removed.
Measured with `lake exe kernel String.Internal.append`:
blake3_compress_chunks width 215 -> 75; assign_block_value (157M FFT)
eliminated; total FFT cost -302M (-10.2%).
The byte adders (`u32_add`, `u32_be_add`, `u64_add`,
`relaxed_u64_be_add_2_bytes`) propagate carries as
`u8_xor(overflow_i, carry_ia)`. `u8_xor` is a circuit op costing +1 aux
and +1 lookup. But these two bits are mutually exclusive: a carry out of
`a_i + b_i` forces the partial sum `<= 254`, so the subsequent
`+ carry_in` cannot itself carry. Their xor is therefore a plain field
add, which costs zero columns.
Replace every carry-combining `u8_xor(x, y)` with `x + y`. The data
xors in `u32_xor` are untouched.
Measured with `lake exe kernel String.Internal.append`:
u32_add width 70 -> 60; total FFT cost -86M (-3.2%).
Verified with `lake test -- --ignored aiur-hashes` (145 assertions, exit 0).
`u8_recompose([b0..b7])` is `b0 + 2*b1 + ... + 128*b7` — pure field
arithmetic. As a function call each of its uses cost +1 aux and +1
lookup for nothing; inlined it costs zero columns (no setup op is
spread, unlike a constant-building wrapper).
Inline all uses as explicit weighted sums: 8 in `blake3_g_function`,
~36 across `sha256_compress` and `define_W_i`. Statically-zero terms
collapse (`u8_recompose([0;8])` becomes `0`; trailing `0`s drop). With
no callers left, delete the `u8_recompose` definition.
Measured with `lake exe kernel String.Internal.append`:
blake3_g_function width 250 -> 210; total FFT cost -50M. The sha256
circuits are not exercised by this kernel entry point, so their share
of the gain does not show in this metric.
Verified with `lake test -- --ignored aiur-hashes` (145 assertions, exit 0).
@arthurpaulino
arthurpaulino enabled auto-merge (squash) May 18, 2026 22:18
@arthurpaulino
arthurpaulino merged commit eb4509a into mainMay 18, 2026
14 checks passed
@arthurpaulino
arthurpaulino deleted the ap/blake3 branch May 18, 2026 22:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@arthurpaulino@gabriel-barrett
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Hyperoptimize the blake3_compress_chunks circuit - #413

Merged
arthurpaulino merged 5 commits into
mainfrom
ap/blake3
May 18, 2026
Merged

Hyperoptimize the blake3_compress_chunks circuit#413
arthurpaulino merged 5 commits into
mainfrom
ap/blake3

Conversation

@arthurpaulino

@arthurpaulinoarthurpaulino commented May 18, 2026

Copy link
Copy Markdown
Member

Summary

The blake3 circuits dominated the lake exe kernel String.Internal.append
proving-cost profile. This PR cuts total FFT cost by 19.1% (3109M →
2516M) through five incremental, individually-measured changes, and
removes two now-dead helpers.

Background

Aiur compiles each function to a zk circuit. A circuit's width is
inputSize + selectors + auxiliaries + 4·(1 + lookups), where
auxiliaries and lookups are the maximum over the match arms — so
width is set by the widest arm; a wide cold arm wastes columns on every
row. Cost is width × height × log2(height). Op costs that matter here:
call +outputSize aux +1 lookup; store/load +1 lookup; u8_* ops +1
lookup each; pure field add/sub/mul-by-const cost zero columns.

Changes

1. Factor block-completion out of the hot loop

blake3_compress_chunks width 341 was set by its coldest arms
(Cons,63,1023) / (Cons,63,_) — the only ones carrying both
assign_block_value (+64 aux) and blake3_compress (+32 aux) plus flag
helpers and store, yet firing on ~1.8% of rows. Extracted that tail
into a cold blake3_compress_block circuit; the two arms merged into
one. 341 → 281.

2. Thread Layer by pointer

enum Layer { Push(&Layer, [[G;4];8]), Nil } is a 34-wide value type.
blake3_compress_chunks took and returned it by value — 34 columns of
inputSize every row plus 34 aux on every recursive call. Switched to
layer: &Layer -> &Layer; blake3 adds a store/load at the
boundary so blake3_compress_layer keeps its by-value signature.
281 → 215.

3. Accumulate block bytes in a list, not a 64-wide buffer

blake3_compress_chunks threaded block_buffer: [[G;4];16] (64 columns
of inputSize) and filled it via assign_block_value, whose
[[G;4];16] return cost +64 aux every hot row. Circuit memory is
content-addressed (no in-place mutation), so a threaded 64-element
accumulator costs 64 columns somewhere on every row — a linked list does
not. Replaced it with a reverse-ordered byte accumulator list:

  • Hot loop: store(ListNode.Cons(head, byte_acc)), one store/row.
  • bytes_to_block materializes a full 64-byte list into [[G;4];16] via
    64 unrolled loads — one cold row per completed block.
  • pad_block zero-pads the trailing partial block.
  • blake3_finish absorbs the input-exhausted arms.
  • assign_block_value is now dead and removed.

215 → 75.

4. Combine carry bits with field add, not u8 xor

The byte adders (u32_add, u32_be_add, u64_add,
relaxed_u64_be_add_2_bytes) propagated carries with
u8_xor(overflow_i, carry_ia) — a circuit op (+1 aux, +1 lookup). The
two bits are mutually exclusive (a carry out of a_i + b_i forces the
partial sum <= 254, so + carry_in cannot carry), so their xor is a
plain field add costing zero columns. Replaced every carry-combining
u8_xor with +. u32_add 70 → 60.

5. Inline u8_recompose and delete it

u8_recompose([b0..b7]) = b0 + 2*b1 + ... + 128*b7 is pure arithmetic,
but each call cost +1 aux +1 lookup. Inlined all uses as explicit
weighted sums — 8 in blake3_g_function, ~36 across sha256_compress
and define_W_i — then deleted the now-unused definition.
blake3_g_function 250 → 210.

Results

Measured with lake exe kernel String.Internal.append:

metricbaselinethis PR
blake3_compress_chunks width34175 (−78%)
blake3_compress_chunks FFT426M94M
u32_add width7060
u32_add FFT602M516M
blake3_g_function width250210
blake3_g_function FFT310M261M
assign_block_value FFT157Meliminated
total FFT cost3109M2516M (−19.1%)

New cold circuits introduced: bytes_to_block, pad_block,
blake3_finish, blake3_compress_block — all small.

Validation

lake test -- --ignored aiur-hashes passes (145 assertions, exit 0),
covering blake3 and sha256 across sizes 0–3168 including block and chunk
boundaries. lake exe kernel String.Internal.append runs clean.

The `blake3_compress_chunks` circuit width was set by its coldest match
arms `(Cons,63,1023)` / `(Cons,63,_)` — the only arms carrying both the
`assign_block_value` call (+64 aux) and the `blake3_compress` call
(+32 aux) plus flag helpers and `store`. Those arms fire on ~1.8% of
chunk-loop rows; the hot buffer-fill arm needs only `assign_block_value`,
so ~60 columns were trash-zero on 98% of rows.
Extract the post-assign tail into a cold `blake3_compress_block` circuit.
The two `(Cons,63,...)` arms were identical once the `blake3_compress`
machinery moved out, so they merge into one `(Cons,63,_)` arm.
Measured with `lake exe kernel String.Internal.append`:
blake3_compress_chunks width 341 -> 281; total FFT cost -72M (-2.3%).
`enum Layer { Push(&Layer, [[G;4];8]), Nil }` is a 34-wide value type
(pointer 1 + digest 32 + tag 1). `blake3_compress_chunks` and
`blake3_compress_block` took and returned it by value, paying 34 columns
of `inputSize` on every row plus 34 aux on every recursive call output.
Switch both to `layer: &Layer -> &Layer`. The push sites change from
`Layer.Push(store(layer), d)` to `store(Layer.Push(layer, d))` — same
single 34-wide store, but the threaded value is now a 1-wide pointer.
`blake3` wraps the initial `Layer.Nil` in `store` and `load`s the result
before `blake3_compress_layer`, which keeps its by-value signature.
Measured with `lake exe kernel String.Internal.append`:
blake3_compress_chunks width 281 -> 215; total FFT cost -83M.
`blake3_compress_chunks` threaded `block_buffer: [[G;4];16]` — 64 columns
of `inputSize` on every row — and filled it via `assign_block_value`,
whose `[[G;4];16]` return cost +64 aux on every hot row. Circuit memory
is content-addressed (no in-place mutation), so a threaded 64-element
accumulator costs 64 columns somewhere on every row. A linked list does
not: appending a byte is a single O(1)-width `store`.
Replace the buffer with a reverse-ordered byte accumulator list:
- Hot chunk loop: `store(ListNode.Cons(head, byte_acc))`, one store/row.
- `bytes_to_block` materializes a full 64-byte list into `[[G;4];16]`
via 64 unrolled `load`s — one cold circuit row per completed block.
- `pad_block` zero-pads the trailing partial block before materializing.
- `blake3_finish` absorbs the three input-exhausted match arms, keeping
their `blake3_compress` call out of the hot loop's width.
- `assign_block_value` is now dead and removed.
Measured with `lake exe kernel String.Internal.append`:
blake3_compress_chunks width 215 -> 75; assign_block_value (157M FFT)
eliminated; total FFT cost -302M (-10.2%).
The byte adders (`u32_add`, `u32_be_add`, `u64_add`,
`relaxed_u64_be_add_2_bytes`) propagate carries as
`u8_xor(overflow_i, carry_ia)`. `u8_xor` is a circuit op costing +1 aux
and +1 lookup. But these two bits are mutually exclusive: a carry out of
`a_i + b_i` forces the partial sum `<= 254`, so the subsequent
`+ carry_in` cannot itself carry. Their xor is therefore a plain field
add, which costs zero columns.
Replace every carry-combining `u8_xor(x, y)` with `x + y`. The data
xors in `u32_xor` are untouched.
Measured with `lake exe kernel String.Internal.append`:
u32_add width 70 -> 60; total FFT cost -86M (-3.2%).
Verified with `lake test -- --ignored aiur-hashes` (145 assertions, exit 0).
`u8_recompose([b0..b7])` is `b0 + 2*b1 + ... + 128*b7` — pure field
arithmetic. As a function call each of its uses cost +1 aux and +1
lookup for nothing; inlined it costs zero columns (no setup op is
spread, unlike a constant-building wrapper).
Inline all uses as explicit weighted sums: 8 in `blake3_g_function`,
~36 across `sha256_compress` and `define_W_i`. Statically-zero terms
collapse (`u8_recompose([0;8])` becomes `0`; trailing `0`s drop). With
no callers left, delete the `u8_recompose` definition.
Measured with `lake exe kernel String.Internal.append`:
blake3_g_function width 250 -> 210; total FFT cost -50M. The sha256
circuits are not exercised by this kernel entry point, so their share
of the gain does not show in this metric.
Verified with `lake test -- --ignored aiur-hashes` (145 assertions, exit 0).
@arthurpaulino
arthurpaulino enabled auto-merge (squash) May 18, 2026 22:18
@arthurpaulino
arthurpaulino merged commit eb4509a into mainMay 18, 2026
14 checks passed
@arthurpaulino
arthurpaulino deleted the ap/blake3 branch May 18, 2026 22:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@arthurpaulino@gabriel-barrett
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Hyperoptimize the blake3_compress_chunks circuit - #413

Merged
arthurpaulino merged 5 commits into
mainfrom
ap/blake3
May 18, 2026
Merged

Hyperoptimize the blake3_compress_chunks circuit#413
arthurpaulino merged 5 commits into
mainfrom
ap/blake3

Conversation

@arthurpaulino

@arthurpaulinoarthurpaulino commented May 18, 2026

Copy link
Copy Markdown
Member

Summary

The blake3 circuits dominated the lake exe kernel String.Internal.append
proving-cost profile. This PR cuts total FFT cost by 19.1% (3109M →
2516M) through five incremental, individually-measured changes, and
removes two now-dead helpers.

Background

Aiur compiles each function to a zk circuit. A circuit's width is
inputSize + selectors + auxiliaries + 4·(1 + lookups), where
auxiliaries and lookups are the maximum over the match arms — so
width is set by the widest arm; a wide cold arm wastes columns on every
row. Cost is width × height × log2(height). Op costs that matter here:
call +outputSize aux +1 lookup; store/load +1 lookup; u8_* ops +1
lookup each; pure field add/sub/mul-by-const cost zero columns.

Changes

1. Factor block-completion out of the hot loop

blake3_compress_chunks width 341 was set by its coldest arms
(Cons,63,1023) / (Cons,63,_) — the only ones carrying both
assign_block_value (+64 aux) and blake3_compress (+32 aux) plus flag
helpers and store, yet firing on ~1.8% of rows. Extracted that tail
into a cold blake3_compress_block circuit; the two arms merged into
one. 341 → 281.

2. Thread Layer by pointer

enum Layer { Push(&Layer, [[G;4];8]), Nil } is a 34-wide value type.
blake3_compress_chunks took and returned it by value — 34 columns of
inputSize every row plus 34 aux on every recursive call. Switched to
layer: &Layer -> &Layer; blake3 adds a store/load at the
boundary so blake3_compress_layer keeps its by-value signature.
281 → 215.

3. Accumulate block bytes in a list, not a 64-wide buffer

blake3_compress_chunks threaded block_buffer: [[G;4];16] (64 columns
of inputSize) and filled it via assign_block_value, whose
[[G;4];16] return cost +64 aux every hot row. Circuit memory is
content-addressed (no in-place mutation), so a threaded 64-element
accumulator costs 64 columns somewhere on every row — a linked list does
not. Replaced it with a reverse-ordered byte accumulator list:

  • Hot loop: store(ListNode.Cons(head, byte_acc)), one store/row.
  • bytes_to_block materializes a full 64-byte list into [[G;4];16] via
    64 unrolled loads — one cold row per completed block.
  • pad_block zero-pads the trailing partial block.
  • blake3_finish absorbs the input-exhausted arms.
  • assign_block_value is now dead and removed.

215 → 75.

4. Combine carry bits with field add, not u8 xor

The byte adders (u32_add, u32_be_add, u64_add,
relaxed_u64_be_add_2_bytes) propagated carries with
u8_xor(overflow_i, carry_ia) — a circuit op (+1 aux, +1 lookup). The
two bits are mutually exclusive (a carry out of a_i + b_i forces the
partial sum <= 254, so + carry_in cannot carry), so their xor is a
plain field add costing zero columns. Replaced every carry-combining
u8_xor with +. u32_add 70 → 60.

5. Inline u8_recompose and delete it

u8_recompose([b0..b7]) = b0 + 2*b1 + ... + 128*b7 is pure arithmetic,
but each call cost +1 aux +1 lookup. Inlined all uses as explicit
weighted sums — 8 in blake3_g_function, ~36 across sha256_compress
and define_W_i — then deleted the now-unused definition.
blake3_g_function 250 → 210.

Results

Measured with lake exe kernel String.Internal.append:

metricbaselinethis PR
blake3_compress_chunks width34175 (−78%)
blake3_compress_chunks FFT426M94M
u32_add width7060
u32_add FFT602M516M
blake3_g_function width250210
blake3_g_function FFT310M261M
assign_block_value FFT157Meliminated
total FFT cost3109M2516M (−19.1%)

New cold circuits introduced: bytes_to_block, pad_block,
blake3_finish, blake3_compress_block — all small.

Validation

lake test -- --ignored aiur-hashes passes (145 assertions, exit 0),
covering blake3 and sha256 across sizes 0–3168 including block and chunk
boundaries. lake exe kernel String.Internal.append runs clean.

The `blake3_compress_chunks` circuit width was set by its coldest match
arms `(Cons,63,1023)` / `(Cons,63,_)` — the only arms carrying both the
`assign_block_value` call (+64 aux) and the `blake3_compress` call
(+32 aux) plus flag helpers and `store`. Those arms fire on ~1.8% of
chunk-loop rows; the hot buffer-fill arm needs only `assign_block_value`,
so ~60 columns were trash-zero on 98% of rows.
Extract the post-assign tail into a cold `blake3_compress_block` circuit.
The two `(Cons,63,...)` arms were identical once the `blake3_compress`
machinery moved out, so they merge into one `(Cons,63,_)` arm.
Measured with `lake exe kernel String.Internal.append`:
blake3_compress_chunks width 341 -> 281; total FFT cost -72M (-2.3%).
`enum Layer { Push(&Layer, [[G;4];8]), Nil }` is a 34-wide value type
(pointer 1 + digest 32 + tag 1). `blake3_compress_chunks` and
`blake3_compress_block` took and returned it by value, paying 34 columns
of `inputSize` on every row plus 34 aux on every recursive call output.
Switch both to `layer: &Layer -> &Layer`. The push sites change from
`Layer.Push(store(layer), d)` to `store(Layer.Push(layer, d))` — same
single 34-wide store, but the threaded value is now a 1-wide pointer.
`blake3` wraps the initial `Layer.Nil` in `store` and `load`s the result
before `blake3_compress_layer`, which keeps its by-value signature.
Measured with `lake exe kernel String.Internal.append`:
blake3_compress_chunks width 281 -> 215; total FFT cost -83M.
`blake3_compress_chunks` threaded `block_buffer: [[G;4];16]` — 64 columns
of `inputSize` on every row — and filled it via `assign_block_value`,
whose `[[G;4];16]` return cost +64 aux on every hot row. Circuit memory
is content-addressed (no in-place mutation), so a threaded 64-element
accumulator costs 64 columns somewhere on every row. A linked list does
not: appending a byte is a single O(1)-width `store`.
Replace the buffer with a reverse-ordered byte accumulator list:
- Hot chunk loop: `store(ListNode.Cons(head, byte_acc))`, one store/row.
- `bytes_to_block` materializes a full 64-byte list into `[[G;4];16]`
via 64 unrolled `load`s — one cold circuit row per completed block.
- `pad_block` zero-pads the trailing partial block before materializing.
- `blake3_finish` absorbs the three input-exhausted match arms, keeping
their `blake3_compress` call out of the hot loop's width.
- `assign_block_value` is now dead and removed.
Measured with `lake exe kernel String.Internal.append`:
blake3_compress_chunks width 215 -> 75; assign_block_value (157M FFT)
eliminated; total FFT cost -302M (-10.2%).
The byte adders (`u32_add`, `u32_be_add`, `u64_add`,
`relaxed_u64_be_add_2_bytes`) propagate carries as
`u8_xor(overflow_i, carry_ia)`. `u8_xor` is a circuit op costing +1 aux
and +1 lookup. But these two bits are mutually exclusive: a carry out of
`a_i + b_i` forces the partial sum `<= 254`, so the subsequent
`+ carry_in` cannot itself carry. Their xor is therefore a plain field
add, which costs zero columns.
Replace every carry-combining `u8_xor(x, y)` with `x + y`. The data
xors in `u32_xor` are untouched.
Measured with `lake exe kernel String.Internal.append`:
u32_add width 70 -> 60; total FFT cost -86M (-3.2%).
Verified with `lake test -- --ignored aiur-hashes` (145 assertions, exit 0).
`u8_recompose([b0..b7])` is `b0 + 2*b1 + ... + 128*b7` — pure field
arithmetic. As a function call each of its uses cost +1 aux and +1
lookup for nothing; inlined it costs zero columns (no setup op is
spread, unlike a constant-building wrapper).
Inline all uses as explicit weighted sums: 8 in `blake3_g_function`,
~36 across `sha256_compress` and `define_W_i`. Statically-zero terms
collapse (`u8_recompose([0;8])` becomes `0`; trailing `0`s drop). With
no callers left, delete the `u8_recompose` definition.
Measured with `lake exe kernel String.Internal.append`:
blake3_g_function width 250 -> 210; total FFT cost -50M. The sha256
circuits are not exercised by this kernel entry point, so their share
of the gain does not show in this metric.
Verified with `lake test -- --ignored aiur-hashes` (145 assertions, exit 0).
@arthurpaulino
arthurpaulino enabled auto-merge (squash) May 18, 2026 22:18
@arthurpaulino
arthurpaulino merged commit eb4509a into mainMay 18, 2026
14 checks passed
@arthurpaulino
arthurpaulino deleted the ap/blake3 branch May 18, 2026 22:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@arthurpaulino@gabriel-barrett