Replace Espresso with a hardened ANE-LM backend - #86

Merged
IchenDEV merged 17 commits into
mainfrom
lava-topaz
Aug 31, 2026
Merged

Replace Espresso with a hardened ANE-LM backend#86
IchenDEV merged 17 commits into
mainfrom
lava-topaz

Conversation

@IchenDEV

@IchenDEVIchenDEV commented Aug 27, 2026

Copy link
Copy Markdown
Owner

Summary

  • replace the non-working Espresso dependency with the exact-pinned IchenDEV/ANE-LM@033472ec12ea796fc7ea4f8cefd7ed456f69900b SwiftPM runtime
  • add selectable local Qwen3 Hugging Face directory loading, native ANE generation, cancellation, reset, unload, and user-controlled MLX fallback
  • validate safetensors structure, required Qwen3 shapes and BF16 tensors, tokenizer byte-level coverage, contiguous IDs, merges, ChatML boundary tokens, and EOS mapping before saving or loading a model
  • release mapped native allocations deterministically and clean partial ANE state after failed loads
  • disclose private API, Mac App Store, compatibility, and delayed RSS-reclaim limitations in both localizations
  • use the macOS semantic under-page color for the shared Settings background

M5 evidence

  • official Qwen3-0.6B produced non-empty text through Utter on M5 Max / macOS 27 without MLX fallback
  • 20 requests passed across three complete load/generate/unload lifecycles, with one cancellation in the final lifecycle
  • within-lifecycle RSS growth was 80 KB, 32 KB, and 0 KB; every unload released more than 512 MiB from its peak
  • post-unload RSS remained nondeterministic, but final cross-lifecycle growth was 26,080 KB and below the enforced 128 MiB bound

This evidence is limited to Qwen3-0.6B on the tested M5 Max / macOS 27 host. Other model sizes, Apple Silicon devices, and OS versions remain unverified. The backend uses private Apple APIs and is not suitable for Mac App Store distribution; MLX remains the stable fallback.

Verification

  • DEVELOPER_DIR=/Applications/Xcode-beta.app/Contents/Developer swift test — 592 XCTest tests passed, 10 environment-gated tests skipped, and the Swift Testing suite passed after merging current main
  • real ANE lifecycle test with UTTER_ANE_TEST_ITERATIONS=20 — passed in 84.367 seconds
  • bash scripts/sdlc-checks.sh — all 12 change bundles passed the strict shell gate
  • bash scripts/ci-basic-checks.sh — passed after conflict resolution and all approval-state updates
  • DEVELOPER_DIR=/Applications/Xcode-beta.app/Contents/Developer bash scripts/build-app.sh — release App and DMG built; app and mounted-DMG signatures and DMG checksum passed after merging current main
  • independent P1/P2 review — CLEAN

The two active intents, specs, plans, and verification artifacts were approved by the user on 2026-08-31. Merge and release remain separate human gates.

Evidence

  • docs/research/2026-08-31-apple-neural-engine-m5-compatibility.md
  • docs/sdlc/changes/2026-08-31-ane-lm-runtime/
  • docs/sdlc/changes/2026-08-31-settings-semantic-background/

@IchenDEVIchenDEV changed the title Add Espresso ANE inference backendReplace Espresso with a hardened ANE-LM backendAug 30, 2026

@IchenDEVIchenDEV left a comment

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I found three correctness/resource issues in the new ANE path and one defect in the environment-gated fallback test. The current smoke coverage can pass on partial Qwen reasoning output, long prompts silently exceed the native KV window, and a successful generation fallback leaves both local backends resident. These should be addressed before merge.

if modelName.lowercased().contains("qwen") {
return "<|im_start|>system\n\(system)<|im_end|>\n"
+ "<|im_start|>user\n\(user)<|im_end|>\n"
+ "<|im_start|>assistant\n"

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Disable Qwen3 thinking for direct-output requests

Qwen3 treats this bare assistant\n suffix as its default thinking mode; its chat template only inserts the empty <think>\n\n</think>\n\n prefix when thinking is disabled. The formatting and edit-command paths expect direct text/JSON, use budgets as low as 256 tokens, and often pass temperature 0, which the pinned sampler implements as greedy decoding. The model can therefore spend the whole budget inside <think> before emitting the final answer, while the current real-model test still passes on any non-empty reasoning fragment. Please use the non-thinking suffix (or an equivalent hard /no_think switch) and assert final user-facing text/valid JSON in the real-model test.

ane_lm_generate(
model.runtime,
tokens.baseAddress,
tokens.count,

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Reject prompts that overflow the 2,048-token ANE cache

The pinned Qwen3 runtime has a fixed 2,048-entry KV cache and overwrites cache.start once it fills, while ane_lm_generate accepts every prompt length. Passing tokens.count here therefore makes prompts over 2,048 tokens silently forget their beginning before the first generated token. Screen context, memory, or longer transcripts can cross this even though the selected model advertises a much larger context. Guard the prompt length (ideally reserving output headroom) and fall back to MLX or truncate explicitly; add a >2,048-token regression test.

)
}
)
if result.usedMLX {

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Unload ANE after a successful MLX fallback

If ANE loading succeeded but espressoLLM.generate throws, the actor still owns its native LoadedModel. This branch then returns a valid result after loading MLX, and the fallback policy only changes the saved backend to .mlx; it never runs the model lifecycle unload path. The process therefore keeps both full models resident until a later manual unload or exit, which can turn a recoverable ANE failure into memory pressure or an OOM. Unload espressoLLM before returning the successful MLX fallback (after preserving the failure message/outcome), and cover the generation-failure path with an isLoaded == false assertion.

)
if index == 0 {
baselineFootprint = currentMemoryFootprint()
options.localLLMBackend = .mlx

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Preserve the fallback observation before switching iterations to MLX

With OPENTYPE_FALLBACK_ITERATIONS > 1, iteration 1 runs the .mlx branch, which calls clearEspressoOutcome(). Because the test consumes the tracker only after the loop, outcome is nil and XCTAssertEqual(outcome, .fallback) fails before the repeated-request memory check can validate anything. Consume/capture the fallback immediately after iteration 0, or keep a separate sawFallback flag.

@IchenDEV
IchenDEV merged commit bfa6216 into mainAug 31, 2026
3 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@IchenDEV
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Replace Espresso with a hardened ANE-LM backend - #86

Merged
IchenDEV merged 17 commits into
mainfrom
lava-topaz
Aug 31, 2026
Merged

Replace Espresso with a hardened ANE-LM backend#86
IchenDEV merged 17 commits into
mainfrom
lava-topaz

Conversation

@IchenDEV

@IchenDEVIchenDEV commented Aug 27, 2026

Copy link
Copy Markdown
Owner

Summary

  • replace the non-working Espresso dependency with the exact-pinned IchenDEV/ANE-LM@033472ec12ea796fc7ea4f8cefd7ed456f69900b SwiftPM runtime
  • add selectable local Qwen3 Hugging Face directory loading, native ANE generation, cancellation, reset, unload, and user-controlled MLX fallback
  • validate safetensors structure, required Qwen3 shapes and BF16 tensors, tokenizer byte-level coverage, contiguous IDs, merges, ChatML boundary tokens, and EOS mapping before saving or loading a model
  • release mapped native allocations deterministically and clean partial ANE state after failed loads
  • disclose private API, Mac App Store, compatibility, and delayed RSS-reclaim limitations in both localizations
  • use the macOS semantic under-page color for the shared Settings background

M5 evidence

  • official Qwen3-0.6B produced non-empty text through Utter on M5 Max / macOS 27 without MLX fallback
  • 20 requests passed across three complete load/generate/unload lifecycles, with one cancellation in the final lifecycle
  • within-lifecycle RSS growth was 80 KB, 32 KB, and 0 KB; every unload released more than 512 MiB from its peak
  • post-unload RSS remained nondeterministic, but final cross-lifecycle growth was 26,080 KB and below the enforced 128 MiB bound

This evidence is limited to Qwen3-0.6B on the tested M5 Max / macOS 27 host. Other model sizes, Apple Silicon devices, and OS versions remain unverified. The backend uses private Apple APIs and is not suitable for Mac App Store distribution; MLX remains the stable fallback.

Verification

  • DEVELOPER_DIR=/Applications/Xcode-beta.app/Contents/Developer swift test — 592 XCTest tests passed, 10 environment-gated tests skipped, and the Swift Testing suite passed after merging current main
  • real ANE lifecycle test with UTTER_ANE_TEST_ITERATIONS=20 — passed in 84.367 seconds
  • bash scripts/sdlc-checks.sh — all 12 change bundles passed the strict shell gate
  • bash scripts/ci-basic-checks.sh — passed after conflict resolution and all approval-state updates
  • DEVELOPER_DIR=/Applications/Xcode-beta.app/Contents/Developer bash scripts/build-app.sh — release App and DMG built; app and mounted-DMG signatures and DMG checksum passed after merging current main
  • independent P1/P2 review — CLEAN

The two active intents, specs, plans, and verification artifacts were approved by the user on 2026-08-31. Merge and release remain separate human gates.

Evidence

  • docs/research/2026-08-31-apple-neural-engine-m5-compatibility.md
  • docs/sdlc/changes/2026-08-31-ane-lm-runtime/
  • docs/sdlc/changes/2026-08-31-settings-semantic-background/

@IchenDEVIchenDEV changed the title Add Espresso ANE inference backendReplace Espresso with a hardened ANE-LM backendAug 30, 2026

@IchenDEVIchenDEV left a comment

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I found three correctness/resource issues in the new ANE path and one defect in the environment-gated fallback test. The current smoke coverage can pass on partial Qwen reasoning output, long prompts silently exceed the native KV window, and a successful generation fallback leaves both local backends resident. These should be addressed before merge.

if modelName.lowercased().contains("qwen") {
return "<|im_start|>system\n\(system)<|im_end|>\n"
+ "<|im_start|>user\n\(user)<|im_end|>\n"
+ "<|im_start|>assistant\n"

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Disable Qwen3 thinking for direct-output requests

Qwen3 treats this bare assistant\n suffix as its default thinking mode; its chat template only inserts the empty <think>\n\n</think>\n\n prefix when thinking is disabled. The formatting and edit-command paths expect direct text/JSON, use budgets as low as 256 tokens, and often pass temperature 0, which the pinned sampler implements as greedy decoding. The model can therefore spend the whole budget inside <think> before emitting the final answer, while the current real-model test still passes on any non-empty reasoning fragment. Please use the non-thinking suffix (or an equivalent hard /no_think switch) and assert final user-facing text/valid JSON in the real-model test.

ane_lm_generate(
model.runtime,
tokens.baseAddress,
tokens.count,

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Reject prompts that overflow the 2,048-token ANE cache

The pinned Qwen3 runtime has a fixed 2,048-entry KV cache and overwrites cache.start once it fills, while ane_lm_generate accepts every prompt length. Passing tokens.count here therefore makes prompts over 2,048 tokens silently forget their beginning before the first generated token. Screen context, memory, or longer transcripts can cross this even though the selected model advertises a much larger context. Guard the prompt length (ideally reserving output headroom) and fall back to MLX or truncate explicitly; add a >2,048-token regression test.

)
}
)
if result.usedMLX {

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Unload ANE after a successful MLX fallback

If ANE loading succeeded but espressoLLM.generate throws, the actor still owns its native LoadedModel. This branch then returns a valid result after loading MLX, and the fallback policy only changes the saved backend to .mlx; it never runs the model lifecycle unload path. The process therefore keeps both full models resident until a later manual unload or exit, which can turn a recoverable ANE failure into memory pressure or an OOM. Unload espressoLLM before returning the successful MLX fallback (after preserving the failure message/outcome), and cover the generation-failure path with an isLoaded == false assertion.

)
if index == 0 {
baselineFootprint = currentMemoryFootprint()
options.localLLMBackend = .mlx

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Preserve the fallback observation before switching iterations to MLX

With OPENTYPE_FALLBACK_ITERATIONS > 1, iteration 1 runs the .mlx branch, which calls clearEspressoOutcome(). Because the test consumes the tracker only after the loop, outcome is nil and XCTAssertEqual(outcome, .fallback) fails before the repeated-request memory check can validate anything. Consume/capture the fallback immediately after iteration 0, or keep a separate sawFallback flag.

@IchenDEV
IchenDEV merged commit bfa6216 into mainAug 31, 2026
3 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@IchenDEV
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Replace Espresso with a hardened ANE-LM backend - #86

Merged
IchenDEV merged 17 commits into
mainfrom
lava-topaz
Aug 31, 2026
Merged

Replace Espresso with a hardened ANE-LM backend#86
IchenDEV merged 17 commits into
mainfrom
lava-topaz

Conversation

@IchenDEV

@IchenDEVIchenDEV commented Aug 27, 2026

Copy link
Copy Markdown
Owner

Summary

  • replace the non-working Espresso dependency with the exact-pinned IchenDEV/ANE-LM@033472ec12ea796fc7ea4f8cefd7ed456f69900b SwiftPM runtime
  • add selectable local Qwen3 Hugging Face directory loading, native ANE generation, cancellation, reset, unload, and user-controlled MLX fallback
  • validate safetensors structure, required Qwen3 shapes and BF16 tensors, tokenizer byte-level coverage, contiguous IDs, merges, ChatML boundary tokens, and EOS mapping before saving or loading a model
  • release mapped native allocations deterministically and clean partial ANE state after failed loads
  • disclose private API, Mac App Store, compatibility, and delayed RSS-reclaim limitations in both localizations
  • use the macOS semantic under-page color for the shared Settings background

M5 evidence

  • official Qwen3-0.6B produced non-empty text through Utter on M5 Max / macOS 27 without MLX fallback
  • 20 requests passed across three complete load/generate/unload lifecycles, with one cancellation in the final lifecycle
  • within-lifecycle RSS growth was 80 KB, 32 KB, and 0 KB; every unload released more than 512 MiB from its peak
  • post-unload RSS remained nondeterministic, but final cross-lifecycle growth was 26,080 KB and below the enforced 128 MiB bound

This evidence is limited to Qwen3-0.6B on the tested M5 Max / macOS 27 host. Other model sizes, Apple Silicon devices, and OS versions remain unverified. The backend uses private Apple APIs and is not suitable for Mac App Store distribution; MLX remains the stable fallback.

Verification

  • DEVELOPER_DIR=/Applications/Xcode-beta.app/Contents/Developer swift test — 592 XCTest tests passed, 10 environment-gated tests skipped, and the Swift Testing suite passed after merging current main
  • real ANE lifecycle test with UTTER_ANE_TEST_ITERATIONS=20 — passed in 84.367 seconds
  • bash scripts/sdlc-checks.sh — all 12 change bundles passed the strict shell gate
  • bash scripts/ci-basic-checks.sh — passed after conflict resolution and all approval-state updates
  • DEVELOPER_DIR=/Applications/Xcode-beta.app/Contents/Developer bash scripts/build-app.sh — release App and DMG built; app and mounted-DMG signatures and DMG checksum passed after merging current main
  • independent P1/P2 review — CLEAN

The two active intents, specs, plans, and verification artifacts were approved by the user on 2026-08-31. Merge and release remain separate human gates.

Evidence

  • docs/research/2026-08-31-apple-neural-engine-m5-compatibility.md
  • docs/sdlc/changes/2026-08-31-ane-lm-runtime/
  • docs/sdlc/changes/2026-08-31-settings-semantic-background/

@IchenDEVIchenDEV changed the title Add Espresso ANE inference backendReplace Espresso with a hardened ANE-LM backendAug 30, 2026

@IchenDEVIchenDEV left a comment

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I found three correctness/resource issues in the new ANE path and one defect in the environment-gated fallback test. The current smoke coverage can pass on partial Qwen reasoning output, long prompts silently exceed the native KV window, and a successful generation fallback leaves both local backends resident. These should be addressed before merge.

if modelName.lowercased().contains("qwen") {
return "<|im_start|>system\n\(system)<|im_end|>\n"
+ "<|im_start|>user\n\(user)<|im_end|>\n"
+ "<|im_start|>assistant\n"

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Disable Qwen3 thinking for direct-output requests

Qwen3 treats this bare assistant\n suffix as its default thinking mode; its chat template only inserts the empty <think>\n\n</think>\n\n prefix when thinking is disabled. The formatting and edit-command paths expect direct text/JSON, use budgets as low as 256 tokens, and often pass temperature 0, which the pinned sampler implements as greedy decoding. The model can therefore spend the whole budget inside <think> before emitting the final answer, while the current real-model test still passes on any non-empty reasoning fragment. Please use the non-thinking suffix (or an equivalent hard /no_think switch) and assert final user-facing text/valid JSON in the real-model test.

ane_lm_generate(
model.runtime,
tokens.baseAddress,
tokens.count,

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Reject prompts that overflow the 2,048-token ANE cache

The pinned Qwen3 runtime has a fixed 2,048-entry KV cache and overwrites cache.start once it fills, while ane_lm_generate accepts every prompt length. Passing tokens.count here therefore makes prompts over 2,048 tokens silently forget their beginning before the first generated token. Screen context, memory, or longer transcripts can cross this even though the selected model advertises a much larger context. Guard the prompt length (ideally reserving output headroom) and fall back to MLX or truncate explicitly; add a >2,048-token regression test.

)
}
)
if result.usedMLX {

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Unload ANE after a successful MLX fallback

If ANE loading succeeded but espressoLLM.generate throws, the actor still owns its native LoadedModel. This branch then returns a valid result after loading MLX, and the fallback policy only changes the saved backend to .mlx; it never runs the model lifecycle unload path. The process therefore keeps both full models resident until a later manual unload or exit, which can turn a recoverable ANE failure into memory pressure or an OOM. Unload espressoLLM before returning the successful MLX fallback (after preserving the failure message/outcome), and cover the generation-failure path with an isLoaded == false assertion.

)
if index == 0 {
baselineFootprint = currentMemoryFootprint()
options.localLLMBackend = .mlx

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Preserve the fallback observation before switching iterations to MLX

With OPENTYPE_FALLBACK_ITERATIONS > 1, iteration 1 runs the .mlx branch, which calls clearEspressoOutcome(). Because the test consumes the tracker only after the loop, outcome is nil and XCTAssertEqual(outcome, .fallback) fails before the repeated-request memory check can validate anything. Consume/capture the fallback immediately after iteration 0, or keep a separate sawFallback flag.

@IchenDEV
IchenDEV merged commit bfa6216 into mainAug 31, 2026
3 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@IchenDEV
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Replace Espresso with a hardened ANE-LM backend - #86

Merged
IchenDEV merged 17 commits into
mainfrom
lava-topaz
Aug 31, 2026
Merged

Replace Espresso with a hardened ANE-LM backend#86
IchenDEV merged 17 commits into
mainfrom
lava-topaz

Conversation

@IchenDEV

@IchenDEVIchenDEV commented Aug 27, 2026

Copy link
Copy Markdown
Owner

Summary

  • replace the non-working Espresso dependency with the exact-pinned IchenDEV/ANE-LM@033472ec12ea796fc7ea4f8cefd7ed456f69900b SwiftPM runtime
  • add selectable local Qwen3 Hugging Face directory loading, native ANE generation, cancellation, reset, unload, and user-controlled MLX fallback
  • validate safetensors structure, required Qwen3 shapes and BF16 tensors, tokenizer byte-level coverage, contiguous IDs, merges, ChatML boundary tokens, and EOS mapping before saving or loading a model
  • release mapped native allocations deterministically and clean partial ANE state after failed loads
  • disclose private API, Mac App Store, compatibility, and delayed RSS-reclaim limitations in both localizations
  • use the macOS semantic under-page color for the shared Settings background

M5 evidence

  • official Qwen3-0.6B produced non-empty text through Utter on M5 Max / macOS 27 without MLX fallback
  • 20 requests passed across three complete load/generate/unload lifecycles, with one cancellation in the final lifecycle
  • within-lifecycle RSS growth was 80 KB, 32 KB, and 0 KB; every unload released more than 512 MiB from its peak
  • post-unload RSS remained nondeterministic, but final cross-lifecycle growth was 26,080 KB and below the enforced 128 MiB bound

This evidence is limited to Qwen3-0.6B on the tested M5 Max / macOS 27 host. Other model sizes, Apple Silicon devices, and OS versions remain unverified. The backend uses private Apple APIs and is not suitable for Mac App Store distribution; MLX remains the stable fallback.

Verification

  • DEVELOPER_DIR=/Applications/Xcode-beta.app/Contents/Developer swift test — 592 XCTest tests passed, 10 environment-gated tests skipped, and the Swift Testing suite passed after merging current main
  • real ANE lifecycle test with UTTER_ANE_TEST_ITERATIONS=20 — passed in 84.367 seconds
  • bash scripts/sdlc-checks.sh — all 12 change bundles passed the strict shell gate
  • bash scripts/ci-basic-checks.sh — passed after conflict resolution and all approval-state updates
  • DEVELOPER_DIR=/Applications/Xcode-beta.app/Contents/Developer bash scripts/build-app.sh — release App and DMG built; app and mounted-DMG signatures and DMG checksum passed after merging current main
  • independent P1/P2 review — CLEAN

The two active intents, specs, plans, and verification artifacts were approved by the user on 2026-08-31. Merge and release remain separate human gates.

Evidence

  • docs/research/2026-08-31-apple-neural-engine-m5-compatibility.md
  • docs/sdlc/changes/2026-08-31-ane-lm-runtime/
  • docs/sdlc/changes/2026-08-31-settings-semantic-background/

@IchenDEVIchenDEV changed the title Add Espresso ANE inference backendReplace Espresso with a hardened ANE-LM backendAug 30, 2026

@IchenDEVIchenDEV left a comment

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I found three correctness/resource issues in the new ANE path and one defect in the environment-gated fallback test. The current smoke coverage can pass on partial Qwen reasoning output, long prompts silently exceed the native KV window, and a successful generation fallback leaves both local backends resident. These should be addressed before merge.

if modelName.lowercased().contains("qwen") {
return "<|im_start|>system\n\(system)<|im_end|>\n"
+ "<|im_start|>user\n\(user)<|im_end|>\n"
+ "<|im_start|>assistant\n"

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Disable Qwen3 thinking for direct-output requests

Qwen3 treats this bare assistant\n suffix as its default thinking mode; its chat template only inserts the empty <think>\n\n</think>\n\n prefix when thinking is disabled. The formatting and edit-command paths expect direct text/JSON, use budgets as low as 256 tokens, and often pass temperature 0, which the pinned sampler implements as greedy decoding. The model can therefore spend the whole budget inside <think> before emitting the final answer, while the current real-model test still passes on any non-empty reasoning fragment. Please use the non-thinking suffix (or an equivalent hard /no_think switch) and assert final user-facing text/valid JSON in the real-model test.

ane_lm_generate(
model.runtime,
tokens.baseAddress,
tokens.count,

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Reject prompts that overflow the 2,048-token ANE cache

The pinned Qwen3 runtime has a fixed 2,048-entry KV cache and overwrites cache.start once it fills, while ane_lm_generate accepts every prompt length. Passing tokens.count here therefore makes prompts over 2,048 tokens silently forget their beginning before the first generated token. Screen context, memory, or longer transcripts can cross this even though the selected model advertises a much larger context. Guard the prompt length (ideally reserving output headroom) and fall back to MLX or truncate explicitly; add a >2,048-token regression test.

)
}
)
if result.usedMLX {

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Unload ANE after a successful MLX fallback

If ANE loading succeeded but espressoLLM.generate throws, the actor still owns its native LoadedModel. This branch then returns a valid result after loading MLX, and the fallback policy only changes the saved backend to .mlx; it never runs the model lifecycle unload path. The process therefore keeps both full models resident until a later manual unload or exit, which can turn a recoverable ANE failure into memory pressure or an OOM. Unload espressoLLM before returning the successful MLX fallback (after preserving the failure message/outcome), and cover the generation-failure path with an isLoaded == false assertion.

)
if index == 0 {
baselineFootprint = currentMemoryFootprint()
options.localLLMBackend = .mlx

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Preserve the fallback observation before switching iterations to MLX

With OPENTYPE_FALLBACK_ITERATIONS > 1, iteration 1 runs the .mlx branch, which calls clearEspressoOutcome(). Because the test consumes the tracker only after the loop, outcome is nil and XCTAssertEqual(outcome, .fallback) fails before the repeated-request memory check can validate anything. Consume/capture the fallback immediately after iteration 0, or keep a separate sawFallback flag.

@IchenDEV
IchenDEV merged commit bfa6216 into mainAug 31, 2026
3 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@IchenDEV
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Replace Espresso with a hardened ANE-LM backend - #86

Merged
IchenDEV merged 17 commits into
mainfrom
lava-topaz
Aug 31, 2026
Merged

Replace Espresso with a hardened ANE-LM backend#86
IchenDEV merged 17 commits into
mainfrom
lava-topaz

Conversation

@IchenDEV

@IchenDEVIchenDEV commented Aug 27, 2026

Copy link
Copy Markdown
Owner

Summary

  • replace the non-working Espresso dependency with the exact-pinned IchenDEV/ANE-LM@033472ec12ea796fc7ea4f8cefd7ed456f69900b SwiftPM runtime
  • add selectable local Qwen3 Hugging Face directory loading, native ANE generation, cancellation, reset, unload, and user-controlled MLX fallback
  • validate safetensors structure, required Qwen3 shapes and BF16 tensors, tokenizer byte-level coverage, contiguous IDs, merges, ChatML boundary tokens, and EOS mapping before saving or loading a model
  • release mapped native allocations deterministically and clean partial ANE state after failed loads
  • disclose private API, Mac App Store, compatibility, and delayed RSS-reclaim limitations in both localizations
  • use the macOS semantic under-page color for the shared Settings background

M5 evidence

  • official Qwen3-0.6B produced non-empty text through Utter on M5 Max / macOS 27 without MLX fallback
  • 20 requests passed across three complete load/generate/unload lifecycles, with one cancellation in the final lifecycle
  • within-lifecycle RSS growth was 80 KB, 32 KB, and 0 KB; every unload released more than 512 MiB from its peak
  • post-unload RSS remained nondeterministic, but final cross-lifecycle growth was 26,080 KB and below the enforced 128 MiB bound

This evidence is limited to Qwen3-0.6B on the tested M5 Max / macOS 27 host. Other model sizes, Apple Silicon devices, and OS versions remain unverified. The backend uses private Apple APIs and is not suitable for Mac App Store distribution; MLX remains the stable fallback.

Verification

  • DEVELOPER_DIR=/Applications/Xcode-beta.app/Contents/Developer swift test — 592 XCTest tests passed, 10 environment-gated tests skipped, and the Swift Testing suite passed after merging current main
  • real ANE lifecycle test with UTTER_ANE_TEST_ITERATIONS=20 — passed in 84.367 seconds
  • bash scripts/sdlc-checks.sh — all 12 change bundles passed the strict shell gate
  • bash scripts/ci-basic-checks.sh — passed after conflict resolution and all approval-state updates
  • DEVELOPER_DIR=/Applications/Xcode-beta.app/Contents/Developer bash scripts/build-app.sh — release App and DMG built; app and mounted-DMG signatures and DMG checksum passed after merging current main
  • independent P1/P2 review — CLEAN

The two active intents, specs, plans, and verification artifacts were approved by the user on 2026-08-31. Merge and release remain separate human gates.

Evidence

  • docs/research/2026-08-31-apple-neural-engine-m5-compatibility.md
  • docs/sdlc/changes/2026-08-31-ane-lm-runtime/
  • docs/sdlc/changes/2026-08-31-settings-semantic-background/

@IchenDEVIchenDEV changed the title Add Espresso ANE inference backendReplace Espresso with a hardened ANE-LM backendAug 30, 2026

@IchenDEVIchenDEV left a comment

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I found three correctness/resource issues in the new ANE path and one defect in the environment-gated fallback test. The current smoke coverage can pass on partial Qwen reasoning output, long prompts silently exceed the native KV window, and a successful generation fallback leaves both local backends resident. These should be addressed before merge.

if modelName.lowercased().contains("qwen") {
return "<|im_start|>system\n\(system)<|im_end|>\n"
+ "<|im_start|>user\n\(user)<|im_end|>\n"
+ "<|im_start|>assistant\n"

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Disable Qwen3 thinking for direct-output requests

Qwen3 treats this bare assistant\n suffix as its default thinking mode; its chat template only inserts the empty <think>\n\n</think>\n\n prefix when thinking is disabled. The formatting and edit-command paths expect direct text/JSON, use budgets as low as 256 tokens, and often pass temperature 0, which the pinned sampler implements as greedy decoding. The model can therefore spend the whole budget inside <think> before emitting the final answer, while the current real-model test still passes on any non-empty reasoning fragment. Please use the non-thinking suffix (or an equivalent hard /no_think switch) and assert final user-facing text/valid JSON in the real-model test.

ane_lm_generate(
model.runtime,
tokens.baseAddress,
tokens.count,

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Reject prompts that overflow the 2,048-token ANE cache

The pinned Qwen3 runtime has a fixed 2,048-entry KV cache and overwrites cache.start once it fills, while ane_lm_generate accepts every prompt length. Passing tokens.count here therefore makes prompts over 2,048 tokens silently forget their beginning before the first generated token. Screen context, memory, or longer transcripts can cross this even though the selected model advertises a much larger context. Guard the prompt length (ideally reserving output headroom) and fall back to MLX or truncate explicitly; add a >2,048-token regression test.

)
}
)
if result.usedMLX {

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Unload ANE after a successful MLX fallback

If ANE loading succeeded but espressoLLM.generate throws, the actor still owns its native LoadedModel. This branch then returns a valid result after loading MLX, and the fallback policy only changes the saved backend to .mlx; it never runs the model lifecycle unload path. The process therefore keeps both full models resident until a later manual unload or exit, which can turn a recoverable ANE failure into memory pressure or an OOM. Unload espressoLLM before returning the successful MLX fallback (after preserving the failure message/outcome), and cover the generation-failure path with an isLoaded == false assertion.

)
if index == 0 {
baselineFootprint = currentMemoryFootprint()
options.localLLMBackend = .mlx

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Preserve the fallback observation before switching iterations to MLX

With OPENTYPE_FALLBACK_ITERATIONS > 1, iteration 1 runs the .mlx branch, which calls clearEspressoOutcome(). Because the test consumes the tracker only after the loop, outcome is nil and XCTAssertEqual(outcome, .fallback) fails before the repeated-request memory check can validate anything. Consume/capture the fallback immediately after iteration 0, or keep a separate sawFallback flag.

@IchenDEV
IchenDEV merged commit bfa6216 into mainAug 31, 2026
3 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@IchenDEV
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Replace Espresso with a hardened ANE-LM backend - #86

Merged
IchenDEV merged 17 commits into
mainfrom
lava-topaz
Aug 31, 2026
Merged

Replace Espresso with a hardened ANE-LM backend#86
IchenDEV merged 17 commits into
mainfrom
lava-topaz

Conversation

@IchenDEV

@IchenDEVIchenDEV commented Aug 27, 2026

Copy link
Copy Markdown
Owner

Summary

  • replace the non-working Espresso dependency with the exact-pinned IchenDEV/ANE-LM@033472ec12ea796fc7ea4f8cefd7ed456f69900b SwiftPM runtime
  • add selectable local Qwen3 Hugging Face directory loading, native ANE generation, cancellation, reset, unload, and user-controlled MLX fallback
  • validate safetensors structure, required Qwen3 shapes and BF16 tensors, tokenizer byte-level coverage, contiguous IDs, merges, ChatML boundary tokens, and EOS mapping before saving or loading a model
  • release mapped native allocations deterministically and clean partial ANE state after failed loads
  • disclose private API, Mac App Store, compatibility, and delayed RSS-reclaim limitations in both localizations
  • use the macOS semantic under-page color for the shared Settings background

M5 evidence

  • official Qwen3-0.6B produced non-empty text through Utter on M5 Max / macOS 27 without MLX fallback
  • 20 requests passed across three complete load/generate/unload lifecycles, with one cancellation in the final lifecycle
  • within-lifecycle RSS growth was 80 KB, 32 KB, and 0 KB; every unload released more than 512 MiB from its peak
  • post-unload RSS remained nondeterministic, but final cross-lifecycle growth was 26,080 KB and below the enforced 128 MiB bound

This evidence is limited to Qwen3-0.6B on the tested M5 Max / macOS 27 host. Other model sizes, Apple Silicon devices, and OS versions remain unverified. The backend uses private Apple APIs and is not suitable for Mac App Store distribution; MLX remains the stable fallback.

Verification

  • DEVELOPER_DIR=/Applications/Xcode-beta.app/Contents/Developer swift test — 592 XCTest tests passed, 10 environment-gated tests skipped, and the Swift Testing suite passed after merging current main
  • real ANE lifecycle test with UTTER_ANE_TEST_ITERATIONS=20 — passed in 84.367 seconds
  • bash scripts/sdlc-checks.sh — all 12 change bundles passed the strict shell gate
  • bash scripts/ci-basic-checks.sh — passed after conflict resolution and all approval-state updates
  • DEVELOPER_DIR=/Applications/Xcode-beta.app/Contents/Developer bash scripts/build-app.sh — release App and DMG built; app and mounted-DMG signatures and DMG checksum passed after merging current main
  • independent P1/P2 review — CLEAN

The two active intents, specs, plans, and verification artifacts were approved by the user on 2026-08-31. Merge and release remain separate human gates.

Evidence

  • docs/research/2026-08-31-apple-neural-engine-m5-compatibility.md
  • docs/sdlc/changes/2026-08-31-ane-lm-runtime/
  • docs/sdlc/changes/2026-08-31-settings-semantic-background/

@IchenDEVIchenDEV changed the title Add Espresso ANE inference backendReplace Espresso with a hardened ANE-LM backendAug 30, 2026

@IchenDEVIchenDEV left a comment

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I found three correctness/resource issues in the new ANE path and one defect in the environment-gated fallback test. The current smoke coverage can pass on partial Qwen reasoning output, long prompts silently exceed the native KV window, and a successful generation fallback leaves both local backends resident. These should be addressed before merge.

if modelName.lowercased().contains("qwen") {
return "<|im_start|>system\n\(system)<|im_end|>\n"
+ "<|im_start|>user\n\(user)<|im_end|>\n"
+ "<|im_start|>assistant\n"

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Disable Qwen3 thinking for direct-output requests

Qwen3 treats this bare assistant\n suffix as its default thinking mode; its chat template only inserts the empty <think>\n\n</think>\n\n prefix when thinking is disabled. The formatting and edit-command paths expect direct text/JSON, use budgets as low as 256 tokens, and often pass temperature 0, which the pinned sampler implements as greedy decoding. The model can therefore spend the whole budget inside <think> before emitting the final answer, while the current real-model test still passes on any non-empty reasoning fragment. Please use the non-thinking suffix (or an equivalent hard /no_think switch) and assert final user-facing text/valid JSON in the real-model test.

ane_lm_generate(
model.runtime,
tokens.baseAddress,
tokens.count,

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Reject prompts that overflow the 2,048-token ANE cache

The pinned Qwen3 runtime has a fixed 2,048-entry KV cache and overwrites cache.start once it fills, while ane_lm_generate accepts every prompt length. Passing tokens.count here therefore makes prompts over 2,048 tokens silently forget their beginning before the first generated token. Screen context, memory, or longer transcripts can cross this even though the selected model advertises a much larger context. Guard the prompt length (ideally reserving output headroom) and fall back to MLX or truncate explicitly; add a >2,048-token regression test.

)
}
)
if result.usedMLX {

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Unload ANE after a successful MLX fallback

If ANE loading succeeded but espressoLLM.generate throws, the actor still owns its native LoadedModel. This branch then returns a valid result after loading MLX, and the fallback policy only changes the saved backend to .mlx; it never runs the model lifecycle unload path. The process therefore keeps both full models resident until a later manual unload or exit, which can turn a recoverable ANE failure into memory pressure or an OOM. Unload espressoLLM before returning the successful MLX fallback (after preserving the failure message/outcome), and cover the generation-failure path with an isLoaded == false assertion.

)
if index == 0 {
baselineFootprint = currentMemoryFootprint()
options.localLLMBackend = .mlx

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Preserve the fallback observation before switching iterations to MLX

With OPENTYPE_FALLBACK_ITERATIONS > 1, iteration 1 runs the .mlx branch, which calls clearEspressoOutcome(). Because the test consumes the tracker only after the loop, outcome is nil and XCTAssertEqual(outcome, .fallback) fails before the repeated-request memory check can validate anything. Consume/capture the fallback immediately after iteration 0, or keep a separate sawFallback flag.

@IchenDEV
IchenDEV merged commit bfa6216 into mainAug 31, 2026
3 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@IchenDEV
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Replace Espresso with a hardened ANE-LM backend - #86

Merged
IchenDEV merged 17 commits into
mainfrom
lava-topaz
Aug 31, 2026
Merged

Replace Espresso with a hardened ANE-LM backend#86
IchenDEV merged 17 commits into
mainfrom
lava-topaz

Conversation

@IchenDEV

@IchenDEVIchenDEV commented Aug 27, 2026

Copy link
Copy Markdown
Owner

Summary

  • replace the non-working Espresso dependency with the exact-pinned IchenDEV/ANE-LM@033472ec12ea796fc7ea4f8cefd7ed456f69900b SwiftPM runtime
  • add selectable local Qwen3 Hugging Face directory loading, native ANE generation, cancellation, reset, unload, and user-controlled MLX fallback
  • validate safetensors structure, required Qwen3 shapes and BF16 tensors, tokenizer byte-level coverage, contiguous IDs, merges, ChatML boundary tokens, and EOS mapping before saving or loading a model
  • release mapped native allocations deterministically and clean partial ANE state after failed loads
  • disclose private API, Mac App Store, compatibility, and delayed RSS-reclaim limitations in both localizations
  • use the macOS semantic under-page color for the shared Settings background

M5 evidence

  • official Qwen3-0.6B produced non-empty text through Utter on M5 Max / macOS 27 without MLX fallback
  • 20 requests passed across three complete load/generate/unload lifecycles, with one cancellation in the final lifecycle
  • within-lifecycle RSS growth was 80 KB, 32 KB, and 0 KB; every unload released more than 512 MiB from its peak
  • post-unload RSS remained nondeterministic, but final cross-lifecycle growth was 26,080 KB and below the enforced 128 MiB bound

This evidence is limited to Qwen3-0.6B on the tested M5 Max / macOS 27 host. Other model sizes, Apple Silicon devices, and OS versions remain unverified. The backend uses private Apple APIs and is not suitable for Mac App Store distribution; MLX remains the stable fallback.

Verification

  • DEVELOPER_DIR=/Applications/Xcode-beta.app/Contents/Developer swift test — 592 XCTest tests passed, 10 environment-gated tests skipped, and the Swift Testing suite passed after merging current main
  • real ANE lifecycle test with UTTER_ANE_TEST_ITERATIONS=20 — passed in 84.367 seconds
  • bash scripts/sdlc-checks.sh — all 12 change bundles passed the strict shell gate
  • bash scripts/ci-basic-checks.sh — passed after conflict resolution and all approval-state updates
  • DEVELOPER_DIR=/Applications/Xcode-beta.app/Contents/Developer bash scripts/build-app.sh — release App and DMG built; app and mounted-DMG signatures and DMG checksum passed after merging current main
  • independent P1/P2 review — CLEAN

The two active intents, specs, plans, and verification artifacts were approved by the user on 2026-08-31. Merge and release remain separate human gates.

Evidence

  • docs/research/2026-08-31-apple-neural-engine-m5-compatibility.md
  • docs/sdlc/changes/2026-08-31-ane-lm-runtime/
  • docs/sdlc/changes/2026-08-31-settings-semantic-background/

@IchenDEVIchenDEV changed the title Add Espresso ANE inference backendReplace Espresso with a hardened ANE-LM backendAug 30, 2026

@IchenDEVIchenDEV left a comment

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I found three correctness/resource issues in the new ANE path and one defect in the environment-gated fallback test. The current smoke coverage can pass on partial Qwen reasoning output, long prompts silently exceed the native KV window, and a successful generation fallback leaves both local backends resident. These should be addressed before merge.

if modelName.lowercased().contains("qwen") {
return "<|im_start|>system\n\(system)<|im_end|>\n"
+ "<|im_start|>user\n\(user)<|im_end|>\n"
+ "<|im_start|>assistant\n"

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Disable Qwen3 thinking for direct-output requests

Qwen3 treats this bare assistant\n suffix as its default thinking mode; its chat template only inserts the empty <think>\n\n</think>\n\n prefix when thinking is disabled. The formatting and edit-command paths expect direct text/JSON, use budgets as low as 256 tokens, and often pass temperature 0, which the pinned sampler implements as greedy decoding. The model can therefore spend the whole budget inside <think> before emitting the final answer, while the current real-model test still passes on any non-empty reasoning fragment. Please use the non-thinking suffix (or an equivalent hard /no_think switch) and assert final user-facing text/valid JSON in the real-model test.

ane_lm_generate(
model.runtime,
tokens.baseAddress,
tokens.count,

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Reject prompts that overflow the 2,048-token ANE cache

The pinned Qwen3 runtime has a fixed 2,048-entry KV cache and overwrites cache.start once it fills, while ane_lm_generate accepts every prompt length. Passing tokens.count here therefore makes prompts over 2,048 tokens silently forget their beginning before the first generated token. Screen context, memory, or longer transcripts can cross this even though the selected model advertises a much larger context. Guard the prompt length (ideally reserving output headroom) and fall back to MLX or truncate explicitly; add a >2,048-token regression test.

)
}
)
if result.usedMLX {

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Unload ANE after a successful MLX fallback

If ANE loading succeeded but espressoLLM.generate throws, the actor still owns its native LoadedModel. This branch then returns a valid result after loading MLX, and the fallback policy only changes the saved backend to .mlx; it never runs the model lifecycle unload path. The process therefore keeps both full models resident until a later manual unload or exit, which can turn a recoverable ANE failure into memory pressure or an OOM. Unload espressoLLM before returning the successful MLX fallback (after preserving the failure message/outcome), and cover the generation-failure path with an isLoaded == false assertion.

)
if index == 0 {
baselineFootprint = currentMemoryFootprint()
options.localLLMBackend = .mlx

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Preserve the fallback observation before switching iterations to MLX

With OPENTYPE_FALLBACK_ITERATIONS > 1, iteration 1 runs the .mlx branch, which calls clearEspressoOutcome(). Because the test consumes the tracker only after the loop, outcome is nil and XCTAssertEqual(outcome, .fallback) fails before the repeated-request memory check can validate anything. Consume/capture the fallback immediately after iteration 0, or keep a separate sawFallback flag.

@IchenDEV
IchenDEV merged commit bfa6216 into mainAug 31, 2026
3 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@IchenDEV
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Replace Espresso with a hardened ANE-LM backend - #86

Merged
IchenDEV merged 17 commits into
mainfrom
lava-topaz
Aug 31, 2026
Merged

Replace Espresso with a hardened ANE-LM backend#86
IchenDEV merged 17 commits into
mainfrom
lava-topaz

Conversation

@IchenDEV

@IchenDEVIchenDEV commented Aug 27, 2026

Copy link
Copy Markdown
Owner

Summary

  • replace the non-working Espresso dependency with the exact-pinned IchenDEV/ANE-LM@033472ec12ea796fc7ea4f8cefd7ed456f69900b SwiftPM runtime
  • add selectable local Qwen3 Hugging Face directory loading, native ANE generation, cancellation, reset, unload, and user-controlled MLX fallback
  • validate safetensors structure, required Qwen3 shapes and BF16 tensors, tokenizer byte-level coverage, contiguous IDs, merges, ChatML boundary tokens, and EOS mapping before saving or loading a model
  • release mapped native allocations deterministically and clean partial ANE state after failed loads
  • disclose private API, Mac App Store, compatibility, and delayed RSS-reclaim limitations in both localizations
  • use the macOS semantic under-page color for the shared Settings background

M5 evidence

  • official Qwen3-0.6B produced non-empty text through Utter on M5 Max / macOS 27 without MLX fallback
  • 20 requests passed across three complete load/generate/unload lifecycles, with one cancellation in the final lifecycle
  • within-lifecycle RSS growth was 80 KB, 32 KB, and 0 KB; every unload released more than 512 MiB from its peak
  • post-unload RSS remained nondeterministic, but final cross-lifecycle growth was 26,080 KB and below the enforced 128 MiB bound

This evidence is limited to Qwen3-0.6B on the tested M5 Max / macOS 27 host. Other model sizes, Apple Silicon devices, and OS versions remain unverified. The backend uses private Apple APIs and is not suitable for Mac App Store distribution; MLX remains the stable fallback.

Verification

  • DEVELOPER_DIR=/Applications/Xcode-beta.app/Contents/Developer swift test — 592 XCTest tests passed, 10 environment-gated tests skipped, and the Swift Testing suite passed after merging current main
  • real ANE lifecycle test with UTTER_ANE_TEST_ITERATIONS=20 — passed in 84.367 seconds
  • bash scripts/sdlc-checks.sh — all 12 change bundles passed the strict shell gate
  • bash scripts/ci-basic-checks.sh — passed after conflict resolution and all approval-state updates
  • DEVELOPER_DIR=/Applications/Xcode-beta.app/Contents/Developer bash scripts/build-app.sh — release App and DMG built; app and mounted-DMG signatures and DMG checksum passed after merging current main
  • independent P1/P2 review — CLEAN

The two active intents, specs, plans, and verification artifacts were approved by the user on 2026-08-31. Merge and release remain separate human gates.

Evidence

  • docs/research/2026-08-31-apple-neural-engine-m5-compatibility.md
  • docs/sdlc/changes/2026-08-31-ane-lm-runtime/
  • docs/sdlc/changes/2026-08-31-settings-semantic-background/

@IchenDEVIchenDEV changed the title Add Espresso ANE inference backendReplace Espresso with a hardened ANE-LM backendAug 30, 2026

@IchenDEVIchenDEV left a comment

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I found three correctness/resource issues in the new ANE path and one defect in the environment-gated fallback test. The current smoke coverage can pass on partial Qwen reasoning output, long prompts silently exceed the native KV window, and a successful generation fallback leaves both local backends resident. These should be addressed before merge.

if modelName.lowercased().contains("qwen") {
return "<|im_start|>system\n\(system)<|im_end|>\n"
+ "<|im_start|>user\n\(user)<|im_end|>\n"
+ "<|im_start|>assistant\n"

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Disable Qwen3 thinking for direct-output requests

Qwen3 treats this bare assistant\n suffix as its default thinking mode; its chat template only inserts the empty <think>\n\n</think>\n\n prefix when thinking is disabled. The formatting and edit-command paths expect direct text/JSON, use budgets as low as 256 tokens, and often pass temperature 0, which the pinned sampler implements as greedy decoding. The model can therefore spend the whole budget inside <think> before emitting the final answer, while the current real-model test still passes on any non-empty reasoning fragment. Please use the non-thinking suffix (or an equivalent hard /no_think switch) and assert final user-facing text/valid JSON in the real-model test.

ane_lm_generate(
model.runtime,
tokens.baseAddress,
tokens.count,

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Reject prompts that overflow the 2,048-token ANE cache

The pinned Qwen3 runtime has a fixed 2,048-entry KV cache and overwrites cache.start once it fills, while ane_lm_generate accepts every prompt length. Passing tokens.count here therefore makes prompts over 2,048 tokens silently forget their beginning before the first generated token. Screen context, memory, or longer transcripts can cross this even though the selected model advertises a much larger context. Guard the prompt length (ideally reserving output headroom) and fall back to MLX or truncate explicitly; add a >2,048-token regression test.

)
}
)
if result.usedMLX {

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Unload ANE after a successful MLX fallback

If ANE loading succeeded but espressoLLM.generate throws, the actor still owns its native LoadedModel. This branch then returns a valid result after loading MLX, and the fallback policy only changes the saved backend to .mlx; it never runs the model lifecycle unload path. The process therefore keeps both full models resident until a later manual unload or exit, which can turn a recoverable ANE failure into memory pressure or an OOM. Unload espressoLLM before returning the successful MLX fallback (after preserving the failure message/outcome), and cover the generation-failure path with an isLoaded == false assertion.

)
if index == 0 {
baselineFootprint = currentMemoryFootprint()
options.localLLMBackend = .mlx

Copy link
Copy Markdown
OwnerAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Preserve the fallback observation before switching iterations to MLX

With OPENTYPE_FALLBACK_ITERATIONS > 1, iteration 1 runs the .mlx branch, which calls clearEspressoOutcome(). Because the test consumes the tracker only after the loop, outcome is nil and XCTAssertEqual(outcome, .fallback) fails before the repeated-request memory check can validate anything. Consume/capture the fallback immediately after iteration 0, or keep a separate sawFallback flag.

@IchenDEV
IchenDEV merged commit bfa6216 into mainAug 31, 2026
3 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@IchenDEV