feat(runtime): add attention-first semantic compaction - #986

Merged
likun666661 merged 6 commits into
apache:mainfrom
likun666661:feat/attention-first-semantic-compaction
Jul 16, 2026
Merged

feat(runtime): add attention-first semantic compaction#986
likun666661 merged 6 commits into
apache:mainfrom
likun666661:feat/attention-first-semantic-compaction

Conversation

@likun666661

@likun666661likun666661 commented Jul 14, 2026

Copy link
Copy Markdown
Member

Summary

  • preserve prior replay and the exact current-turn user message as an immutable head anchor
  • replace only the structurally complete middle span before the open provider-protocol tail
  • use the 8K total-context / 4K completed-span attention trigger and separate generation / accepted-projection budgets
  • add bounded structured and complete-sentence/line fallback validation, V2 rolling lineage, and shared attention/emergency scheduling
  • keep the provider-visible continuation strictly exact head anchor + LLM-authored projection + exact protocol tail
  • keep Runtime coverage, source refs, archive refs, preserved-tail indexes, and legacy state cards in the diagnostic side channel only

Safety properties

  • the head anchor is signature-checked on every prepareStep and is never semantic coverage
  • consecutive assistant reasoning/text/tool-call messages form one provider episode; incomplete multi-call episodes remain verbatim in the tail
  • empty, truncated, source-invalid, private-verifier-invalid, protocol-invalid, or over-budget projections do not replace source messages
  • token savings, compact-call cost, and cache loss remain diagnostics in attention mode
  • one prepareStep accepts at most one attention or emergency replacement
  • emergency compaction reports explicit head_anchor_exceeds_capacity / no_safe_completed_span outcomes instead of rewriting the user instruction
  • semanticCompact no longer generates, prompts with, records on new blocks, or renders deterministic state cards
  • rolling successors re-render the predecessor's LLM summary and ignore legacy state cards instead of replaying the old provider message verbatim

MIPS smoke finding

A GLM-5.2 make-mips-interpreter smoke on the earlier version was stopped after it exceeded 135 steps without creating /app/vm.js.

The retained trace showed:

  • 24 accepted semantic compactions and 228 tool calls
  • 19 later summaries whose next action was to write /app/vm.js
  • provider-visible state cards that instead repeated assistant planning such as “re-read” or “re-verify”
  • zero creation of /app/vm.js

The deterministic VM matcher treated any assistant text containing vm as operational state. Those inferred cards were rendered after every LLM summary and then copied into rolling successors, creating a positive feedback loop. The fix in 4eebe953 removes that semantic path rather than attempting to repair it with more regexes.

This smoke is mechanism evidence, not an A/B quality result. The old main-branch arm separately failed to activate semantic compaction because of tool_pair_split and eventually timed out near provider capacity, so it is not a usable success baseline.

Validation

After merging current upstream/main:

  • focused semantic + backend + mid-turn capacity integration: 187 passed, 0 failed
  • npm --workspace @maka/headless test: 964 passed, 0 failed, 1 skipped
  • npm run typecheck: passes across all workspaces after rebuilding changed workspace dependencies
  • Runtime full suite: 1943 passed, 7 skipped, with one unrelated local-only failure in main's macOS sandbox executable-root index assertion because this machine runs Node from /usr/local/bin/node
  • git diff --check: passes

No attention benefit is claimed yet, and this PR does not expand any surface default. A fresh same-commit A/B is still required after the mechanism correction.

Refs #981

Astro-Han added a commit that referenced this pull request Jul 15, 2026
* feat(runtime): default mid-turn capacity compaction on, runtime-owned
The mid_turn capacity invariant becomes a runtime-owned default: whenever
history compaction is enabled the runtime derives midTurn from the selected
model window (contextWindow - reserve), so Desktop, TUI, and headless inherit
it without copying config. MAKA_CONTEXT_HISTORY_COMPACT_MID_TURN=off stays as
the explicit escape hatch.
The backend still gates activation on the checkpoint seams, the persisted head
anchor, and a KNOWN context window, so a session without model metadata or a
child session with no anchor seam never misfires even though the default is on
(covered by new backend safety guards).
Refs #882
* feat: default semantic compaction off, runtime-owned across surfaces
Semantic compaction is the #981/#986 attention-first experiment, not the
standard historyCompact mechanism. Make the runtime the owner of its default:
it stays off unless a surface explicitly opts in (MAKA_CONTEXT_SEMANTIC_COMPACT
truthy, or a non-off MAKA_CONTEXT_SEMANTIC_COMPACT_MODE).
This lets the CLI drop its local semanticCompact strip and eat the runtime
default directly, so the surfaces stop duplicating the same off decision.
Desktop and headless inherit the off default without any per-surface wiring.
Refs #882
* test(runtime): lock the shipped default's proactive long-turn journey
Wire buildDefaultContextBudgetPolicy — the exact default every surface now
inherits — into the streaming backend and prove a long turn near the window
engages mid-turn compaction and continues to a clean end rather than being
truncated or surfacing a raw provider error.
The fake-backend Playwright suite cannot exercise this path: it is a separate
canned-event backend that never runs AiSdkBackend's prepareStep machinery, so a
desktop E2E could only fake the very compaction under test. This deterministic
runtime-layer test is the honest representative journey for the proactive path.
Refs #882
* fix(runtime): bound the derived compaction reserve by the model window
The flat 16384 reserve assumed large-window models: on gpt-4 (8192 window)
the default policy derived maxHistoryEstimatedTokens = max(1, 8192 - 16384) = 1
— a pre-existing turn-boundary pathology on main — and with mid_turn now on by
default the 1-token high water made every multi-step turn run the summarizer
for a checkpoint the replay gate could never admit.
Make defaultCompactReserveTokens the single owner of the reserve default,
bounded by the KNOWN window: min(16384, floor(window / 4)), so both the
turn-boundary budget and the mid_turn high water derive from one value (8K
window -> 2048 reserve / 6144 budget; >= 64K windows byte-identical to before;
unknown window keeps the classic constant). An explicit
MAKA_CONTEXT_HISTORY_COMPACT_RESERVE_TOKENS is respected verbatim. Peers bound
the same way: opencode caps its compaction buffer by the model's output limit,
gemini-cli triggers at a window fraction; pi shares the flat-16384 blind spot.
Strengthen the shipped-defaults journey to fully-default derivation with
durable-persistence assertions (checkpoint recorded, compact block present,
replaced raw span gone), and add the unrescuable path: under the derived
default a turn with no safe span ends with the explicit
context_budget_exhausted outcome. A small-window guard locks the P2 behavior:
on an 8K model the default is completely inert when under budget — never a
pointless summarizer call.
Refs #882
* docs(runtime): align midTurn policy comments with the shipped runtime default
@likun666661
likun666661 merged commit 01fc111 into apache:mainJul 16, 2026
2 of 3 checks passed
Astro-Han added a commit that referenced this pull request Jul 16, 2026
…c CLI color tests) (#1077)
* fix(runtime): let active prune cover the newest step under mid-turn capacity
#986 unconditionally excluded the newest completed step from active
tool-result pruning, breaking three reviewed capacity-rescue invariants
(mid-turn finding C, capacity-replacement re-convergence, overflow
round-3 P1) and CI on main. Include the newest step exactly when
mid-turn capacity compaction is active, keeping the keep-newest-raw
default for every other configuration.
* test(cli): pin color level so ANSI assertions are hermetic
tui-ansi detects color capability from TERM/COLORTERM at module load, so
the pi-transcript and pi-tui-runner ANSI assertions inherited ambient
terminal capability: truecolor locally, level 0 on CI runners with no
TERM/COLORTERM, where every color function is a no-op and the #1064/#1066
color assertions fail deterministically. Pin level 3 via the existing
_setColorLevelForTesting seam, matching tui-ansi.test.ts.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@likun666661
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

feat(runtime): add attention-first semantic compaction - #986

Merged
likun666661 merged 6 commits into
apache:mainfrom
likun666661:feat/attention-first-semantic-compaction
Jul 16, 2026
Merged

feat(runtime): add attention-first semantic compaction#986
likun666661 merged 6 commits into
apache:mainfrom
likun666661:feat/attention-first-semantic-compaction

Conversation

@likun666661

@likun666661likun666661 commented Jul 14, 2026

Copy link
Copy Markdown
Member

Summary

  • preserve prior replay and the exact current-turn user message as an immutable head anchor
  • replace only the structurally complete middle span before the open provider-protocol tail
  • use the 8K total-context / 4K completed-span attention trigger and separate generation / accepted-projection budgets
  • add bounded structured and complete-sentence/line fallback validation, V2 rolling lineage, and shared attention/emergency scheduling
  • keep the provider-visible continuation strictly exact head anchor + LLM-authored projection + exact protocol tail
  • keep Runtime coverage, source refs, archive refs, preserved-tail indexes, and legacy state cards in the diagnostic side channel only

Safety properties

  • the head anchor is signature-checked on every prepareStep and is never semantic coverage
  • consecutive assistant reasoning/text/tool-call messages form one provider episode; incomplete multi-call episodes remain verbatim in the tail
  • empty, truncated, source-invalid, private-verifier-invalid, protocol-invalid, or over-budget projections do not replace source messages
  • token savings, compact-call cost, and cache loss remain diagnostics in attention mode
  • one prepareStep accepts at most one attention or emergency replacement
  • emergency compaction reports explicit head_anchor_exceeds_capacity / no_safe_completed_span outcomes instead of rewriting the user instruction
  • semanticCompact no longer generates, prompts with, records on new blocks, or renders deterministic state cards
  • rolling successors re-render the predecessor's LLM summary and ignore legacy state cards instead of replaying the old provider message verbatim

MIPS smoke finding

A GLM-5.2 make-mips-interpreter smoke on the earlier version was stopped after it exceeded 135 steps without creating /app/vm.js.

The retained trace showed:

  • 24 accepted semantic compactions and 228 tool calls
  • 19 later summaries whose next action was to write /app/vm.js
  • provider-visible state cards that instead repeated assistant planning such as “re-read” or “re-verify”
  • zero creation of /app/vm.js

The deterministic VM matcher treated any assistant text containing vm as operational state. Those inferred cards were rendered after every LLM summary and then copied into rolling successors, creating a positive feedback loop. The fix in 4eebe953 removes that semantic path rather than attempting to repair it with more regexes.

This smoke is mechanism evidence, not an A/B quality result. The old main-branch arm separately failed to activate semantic compaction because of tool_pair_split and eventually timed out near provider capacity, so it is not a usable success baseline.

Validation

After merging current upstream/main:

  • focused semantic + backend + mid-turn capacity integration: 187 passed, 0 failed
  • npm --workspace @maka/headless test: 964 passed, 0 failed, 1 skipped
  • npm run typecheck: passes across all workspaces after rebuilding changed workspace dependencies
  • Runtime full suite: 1943 passed, 7 skipped, with one unrelated local-only failure in main's macOS sandbox executable-root index assertion because this machine runs Node from /usr/local/bin/node
  • git diff --check: passes

No attention benefit is claimed yet, and this PR does not expand any surface default. A fresh same-commit A/B is still required after the mechanism correction.

Refs #981

Astro-Han added a commit that referenced this pull request Jul 15, 2026
* feat(runtime): default mid-turn capacity compaction on, runtime-owned
The mid_turn capacity invariant becomes a runtime-owned default: whenever
history compaction is enabled the runtime derives midTurn from the selected
model window (contextWindow - reserve), so Desktop, TUI, and headless inherit
it without copying config. MAKA_CONTEXT_HISTORY_COMPACT_MID_TURN=off stays as
the explicit escape hatch.
The backend still gates activation on the checkpoint seams, the persisted head
anchor, and a KNOWN context window, so a session without model metadata or a
child session with no anchor seam never misfires even though the default is on
(covered by new backend safety guards).
Refs #882
* feat: default semantic compaction off, runtime-owned across surfaces
Semantic compaction is the #981/#986 attention-first experiment, not the
standard historyCompact mechanism. Make the runtime the owner of its default:
it stays off unless a surface explicitly opts in (MAKA_CONTEXT_SEMANTIC_COMPACT
truthy, or a non-off MAKA_CONTEXT_SEMANTIC_COMPACT_MODE).
This lets the CLI drop its local semanticCompact strip and eat the runtime
default directly, so the surfaces stop duplicating the same off decision.
Desktop and headless inherit the off default without any per-surface wiring.
Refs #882
* test(runtime): lock the shipped default's proactive long-turn journey
Wire buildDefaultContextBudgetPolicy — the exact default every surface now
inherits — into the streaming backend and prove a long turn near the window
engages mid-turn compaction and continues to a clean end rather than being
truncated or surfacing a raw provider error.
The fake-backend Playwright suite cannot exercise this path: it is a separate
canned-event backend that never runs AiSdkBackend's prepareStep machinery, so a
desktop E2E could only fake the very compaction under test. This deterministic
runtime-layer test is the honest representative journey for the proactive path.
Refs #882
* fix(runtime): bound the derived compaction reserve by the model window
The flat 16384 reserve assumed large-window models: on gpt-4 (8192 window)
the default policy derived maxHistoryEstimatedTokens = max(1, 8192 - 16384) = 1
— a pre-existing turn-boundary pathology on main — and with mid_turn now on by
default the 1-token high water made every multi-step turn run the summarizer
for a checkpoint the replay gate could never admit.
Make defaultCompactReserveTokens the single owner of the reserve default,
bounded by the KNOWN window: min(16384, floor(window / 4)), so both the
turn-boundary budget and the mid_turn high water derive from one value (8K
window -> 2048 reserve / 6144 budget; >= 64K windows byte-identical to before;
unknown window keeps the classic constant). An explicit
MAKA_CONTEXT_HISTORY_COMPACT_RESERVE_TOKENS is respected verbatim. Peers bound
the same way: opencode caps its compaction buffer by the model's output limit,
gemini-cli triggers at a window fraction; pi shares the flat-16384 blind spot.
Strengthen the shipped-defaults journey to fully-default derivation with
durable-persistence assertions (checkpoint recorded, compact block present,
replaced raw span gone), and add the unrescuable path: under the derived
default a turn with no safe span ends with the explicit
context_budget_exhausted outcome. A small-window guard locks the P2 behavior:
on an 8K model the default is completely inert when under budget — never a
pointless summarizer call.
Refs #882
* docs(runtime): align midTurn policy comments with the shipped runtime default
@likun666661
likun666661 merged commit 01fc111 into apache:mainJul 16, 2026
2 of 3 checks passed
Astro-Han added a commit that referenced this pull request Jul 16, 2026
…c CLI color tests) (#1077)
* fix(runtime): let active prune cover the newest step under mid-turn capacity
#986 unconditionally excluded the newest completed step from active
tool-result pruning, breaking three reviewed capacity-rescue invariants
(mid-turn finding C, capacity-replacement re-convergence, overflow
round-3 P1) and CI on main. Include the newest step exactly when
mid-turn capacity compaction is active, keeping the keep-newest-raw
default for every other configuration.
* test(cli): pin color level so ANSI assertions are hermetic
tui-ansi detects color capability from TERM/COLORTERM at module load, so
the pi-transcript and pi-tui-runner ANSI assertions inherited ambient
terminal capability: truecolor locally, level 0 on CI runners with no
TERM/COLORTERM, where every color function is a no-op and the #1064/#1066
color assertions fail deterministically. Pin level 3 via the existing
_setColorLevelForTesting seam, matching tui-ansi.test.ts.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@likun666661
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(runtime): add attention-first semantic compaction - #986

Merged
likun666661 merged 6 commits into
apache:mainfrom
likun666661:feat/attention-first-semantic-compaction
Jul 16, 2026
Merged

feat(runtime): add attention-first semantic compaction#986
likun666661 merged 6 commits into
apache:mainfrom
likun666661:feat/attention-first-semantic-compaction

Conversation

@likun666661

@likun666661likun666661 commented Jul 14, 2026

Copy link
Copy Markdown
Member

Summary

  • preserve prior replay and the exact current-turn user message as an immutable head anchor
  • replace only the structurally complete middle span before the open provider-protocol tail
  • use the 8K total-context / 4K completed-span attention trigger and separate generation / accepted-projection budgets
  • add bounded structured and complete-sentence/line fallback validation, V2 rolling lineage, and shared attention/emergency scheduling
  • keep the provider-visible continuation strictly exact head anchor + LLM-authored projection + exact protocol tail
  • keep Runtime coverage, source refs, archive refs, preserved-tail indexes, and legacy state cards in the diagnostic side channel only

Safety properties

  • the head anchor is signature-checked on every prepareStep and is never semantic coverage
  • consecutive assistant reasoning/text/tool-call messages form one provider episode; incomplete multi-call episodes remain verbatim in the tail
  • empty, truncated, source-invalid, private-verifier-invalid, protocol-invalid, or over-budget projections do not replace source messages
  • token savings, compact-call cost, and cache loss remain diagnostics in attention mode
  • one prepareStep accepts at most one attention or emergency replacement
  • emergency compaction reports explicit head_anchor_exceeds_capacity / no_safe_completed_span outcomes instead of rewriting the user instruction
  • semanticCompact no longer generates, prompts with, records on new blocks, or renders deterministic state cards
  • rolling successors re-render the predecessor's LLM summary and ignore legacy state cards instead of replaying the old provider message verbatim

MIPS smoke finding

A GLM-5.2 make-mips-interpreter smoke on the earlier version was stopped after it exceeded 135 steps without creating /app/vm.js.

The retained trace showed:

  • 24 accepted semantic compactions and 228 tool calls
  • 19 later summaries whose next action was to write /app/vm.js
  • provider-visible state cards that instead repeated assistant planning such as “re-read” or “re-verify”
  • zero creation of /app/vm.js

The deterministic VM matcher treated any assistant text containing vm as operational state. Those inferred cards were rendered after every LLM summary and then copied into rolling successors, creating a positive feedback loop. The fix in 4eebe953 removes that semantic path rather than attempting to repair it with more regexes.

This smoke is mechanism evidence, not an A/B quality result. The old main-branch arm separately failed to activate semantic compaction because of tool_pair_split and eventually timed out near provider capacity, so it is not a usable success baseline.

Validation

After merging current upstream/main:

  • focused semantic + backend + mid-turn capacity integration: 187 passed, 0 failed
  • npm --workspace @maka/headless test: 964 passed, 0 failed, 1 skipped
  • npm run typecheck: passes across all workspaces after rebuilding changed workspace dependencies
  • Runtime full suite: 1943 passed, 7 skipped, with one unrelated local-only failure in main's macOS sandbox executable-root index assertion because this machine runs Node from /usr/local/bin/node
  • git diff --check: passes

No attention benefit is claimed yet, and this PR does not expand any surface default. A fresh same-commit A/B is still required after the mechanism correction.

Refs #981

Astro-Han added a commit that referenced this pull request Jul 15, 2026
* feat(runtime): default mid-turn capacity compaction on, runtime-owned
The mid_turn capacity invariant becomes a runtime-owned default: whenever
history compaction is enabled the runtime derives midTurn from the selected
model window (contextWindow - reserve), so Desktop, TUI, and headless inherit
it without copying config. MAKA_CONTEXT_HISTORY_COMPACT_MID_TURN=off stays as
the explicit escape hatch.
The backend still gates activation on the checkpoint seams, the persisted head
anchor, and a KNOWN context window, so a session without model metadata or a
child session with no anchor seam never misfires even though the default is on
(covered by new backend safety guards).
Refs #882
* feat: default semantic compaction off, runtime-owned across surfaces
Semantic compaction is the #981/#986 attention-first experiment, not the
standard historyCompact mechanism. Make the runtime the owner of its default:
it stays off unless a surface explicitly opts in (MAKA_CONTEXT_SEMANTIC_COMPACT
truthy, or a non-off MAKA_CONTEXT_SEMANTIC_COMPACT_MODE).
This lets the CLI drop its local semanticCompact strip and eat the runtime
default directly, so the surfaces stop duplicating the same off decision.
Desktop and headless inherit the off default without any per-surface wiring.
Refs #882
* test(runtime): lock the shipped default's proactive long-turn journey
Wire buildDefaultContextBudgetPolicy — the exact default every surface now
inherits — into the streaming backend and prove a long turn near the window
engages mid-turn compaction and continues to a clean end rather than being
truncated or surfacing a raw provider error.
The fake-backend Playwright suite cannot exercise this path: it is a separate
canned-event backend that never runs AiSdkBackend's prepareStep machinery, so a
desktop E2E could only fake the very compaction under test. This deterministic
runtime-layer test is the honest representative journey for the proactive path.
Refs #882
* fix(runtime): bound the derived compaction reserve by the model window
The flat 16384 reserve assumed large-window models: on gpt-4 (8192 window)
the default policy derived maxHistoryEstimatedTokens = max(1, 8192 - 16384) = 1
— a pre-existing turn-boundary pathology on main — and with mid_turn now on by
default the 1-token high water made every multi-step turn run the summarizer
for a checkpoint the replay gate could never admit.
Make defaultCompactReserveTokens the single owner of the reserve default,
bounded by the KNOWN window: min(16384, floor(window / 4)), so both the
turn-boundary budget and the mid_turn high water derive from one value (8K
window -> 2048 reserve / 6144 budget; >= 64K windows byte-identical to before;
unknown window keeps the classic constant). An explicit
MAKA_CONTEXT_HISTORY_COMPACT_RESERVE_TOKENS is respected verbatim. Peers bound
the same way: opencode caps its compaction buffer by the model's output limit,
gemini-cli triggers at a window fraction; pi shares the flat-16384 blind spot.
Strengthen the shipped-defaults journey to fully-default derivation with
durable-persistence assertions (checkpoint recorded, compact block present,
replaced raw span gone), and add the unrescuable path: under the derived
default a turn with no safe span ends with the explicit
context_budget_exhausted outcome. A small-window guard locks the P2 behavior:
on an 8K model the default is completely inert when under budget — never a
pointless summarizer call.
Refs #882
* docs(runtime): align midTurn policy comments with the shipped runtime default
@likun666661
likun666661 merged commit 01fc111 into apache:mainJul 16, 2026
2 of 3 checks passed
Astro-Han added a commit that referenced this pull request Jul 16, 2026
…c CLI color tests) (#1077)
* fix(runtime): let active prune cover the newest step under mid-turn capacity
#986 unconditionally excluded the newest completed step from active
tool-result pruning, breaking three reviewed capacity-rescue invariants
(mid-turn finding C, capacity-replacement re-convergence, overflow
round-3 P1) and CI on main. Include the newest step exactly when
mid-turn capacity compaction is active, keeping the keep-newest-raw
default for every other configuration.
* test(cli): pin color level so ANSI assertions are hermetic
tui-ansi detects color capability from TERM/COLORTERM at module load, so
the pi-transcript and pi-tui-runner ANSI assertions inherited ambient
terminal capability: truecolor locally, level 0 on CI runners with no
TERM/COLORTERM, where every color function is a no-op and the #1064/#1066
color assertions fail deterministically. Pin level 3 via the existing
_setColorLevelForTesting seam, matching tui-ansi.test.ts.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@likun666661
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(runtime): add attention-first semantic compaction - #986

Merged
likun666661 merged 6 commits into
apache:mainfrom
likun666661:feat/attention-first-semantic-compaction
Jul 16, 2026
Merged

feat(runtime): add attention-first semantic compaction#986
likun666661 merged 6 commits into
apache:mainfrom
likun666661:feat/attention-first-semantic-compaction

Conversation

@likun666661

@likun666661likun666661 commented Jul 14, 2026

Copy link
Copy Markdown
Member

Summary

  • preserve prior replay and the exact current-turn user message as an immutable head anchor
  • replace only the structurally complete middle span before the open provider-protocol tail
  • use the 8K total-context / 4K completed-span attention trigger and separate generation / accepted-projection budgets
  • add bounded structured and complete-sentence/line fallback validation, V2 rolling lineage, and shared attention/emergency scheduling
  • keep the provider-visible continuation strictly exact head anchor + LLM-authored projection + exact protocol tail
  • keep Runtime coverage, source refs, archive refs, preserved-tail indexes, and legacy state cards in the diagnostic side channel only

Safety properties

  • the head anchor is signature-checked on every prepareStep and is never semantic coverage
  • consecutive assistant reasoning/text/tool-call messages form one provider episode; incomplete multi-call episodes remain verbatim in the tail
  • empty, truncated, source-invalid, private-verifier-invalid, protocol-invalid, or over-budget projections do not replace source messages
  • token savings, compact-call cost, and cache loss remain diagnostics in attention mode
  • one prepareStep accepts at most one attention or emergency replacement
  • emergency compaction reports explicit head_anchor_exceeds_capacity / no_safe_completed_span outcomes instead of rewriting the user instruction
  • semanticCompact no longer generates, prompts with, records on new blocks, or renders deterministic state cards
  • rolling successors re-render the predecessor's LLM summary and ignore legacy state cards instead of replaying the old provider message verbatim

MIPS smoke finding

A GLM-5.2 make-mips-interpreter smoke on the earlier version was stopped after it exceeded 135 steps without creating /app/vm.js.

The retained trace showed:

  • 24 accepted semantic compactions and 228 tool calls
  • 19 later summaries whose next action was to write /app/vm.js
  • provider-visible state cards that instead repeated assistant planning such as “re-read” or “re-verify”
  • zero creation of /app/vm.js

The deterministic VM matcher treated any assistant text containing vm as operational state. Those inferred cards were rendered after every LLM summary and then copied into rolling successors, creating a positive feedback loop. The fix in 4eebe953 removes that semantic path rather than attempting to repair it with more regexes.

This smoke is mechanism evidence, not an A/B quality result. The old main-branch arm separately failed to activate semantic compaction because of tool_pair_split and eventually timed out near provider capacity, so it is not a usable success baseline.

Validation

After merging current upstream/main:

  • focused semantic + backend + mid-turn capacity integration: 187 passed, 0 failed
  • npm --workspace @maka/headless test: 964 passed, 0 failed, 1 skipped
  • npm run typecheck: passes across all workspaces after rebuilding changed workspace dependencies
  • Runtime full suite: 1943 passed, 7 skipped, with one unrelated local-only failure in main's macOS sandbox executable-root index assertion because this machine runs Node from /usr/local/bin/node
  • git diff --check: passes

No attention benefit is claimed yet, and this PR does not expand any surface default. A fresh same-commit A/B is still required after the mechanism correction.

Refs #981

Astro-Han added a commit that referenced this pull request Jul 15, 2026
* feat(runtime): default mid-turn capacity compaction on, runtime-owned
The mid_turn capacity invariant becomes a runtime-owned default: whenever
history compaction is enabled the runtime derives midTurn from the selected
model window (contextWindow - reserve), so Desktop, TUI, and headless inherit
it without copying config. MAKA_CONTEXT_HISTORY_COMPACT_MID_TURN=off stays as
the explicit escape hatch.
The backend still gates activation on the checkpoint seams, the persisted head
anchor, and a KNOWN context window, so a session without model metadata or a
child session with no anchor seam never misfires even though the default is on
(covered by new backend safety guards).
Refs #882
* feat: default semantic compaction off, runtime-owned across surfaces
Semantic compaction is the #981/#986 attention-first experiment, not the
standard historyCompact mechanism. Make the runtime the owner of its default:
it stays off unless a surface explicitly opts in (MAKA_CONTEXT_SEMANTIC_COMPACT
truthy, or a non-off MAKA_CONTEXT_SEMANTIC_COMPACT_MODE).
This lets the CLI drop its local semanticCompact strip and eat the runtime
default directly, so the surfaces stop duplicating the same off decision.
Desktop and headless inherit the off default without any per-surface wiring.
Refs #882
* test(runtime): lock the shipped default's proactive long-turn journey
Wire buildDefaultContextBudgetPolicy — the exact default every surface now
inherits — into the streaming backend and prove a long turn near the window
engages mid-turn compaction and continues to a clean end rather than being
truncated or surfacing a raw provider error.
The fake-backend Playwright suite cannot exercise this path: it is a separate
canned-event backend that never runs AiSdkBackend's prepareStep machinery, so a
desktop E2E could only fake the very compaction under test. This deterministic
runtime-layer test is the honest representative journey for the proactive path.
Refs #882
* fix(runtime): bound the derived compaction reserve by the model window
The flat 16384 reserve assumed large-window models: on gpt-4 (8192 window)
the default policy derived maxHistoryEstimatedTokens = max(1, 8192 - 16384) = 1
— a pre-existing turn-boundary pathology on main — and with mid_turn now on by
default the 1-token high water made every multi-step turn run the summarizer
for a checkpoint the replay gate could never admit.
Make defaultCompactReserveTokens the single owner of the reserve default,
bounded by the KNOWN window: min(16384, floor(window / 4)), so both the
turn-boundary budget and the mid_turn high water derive from one value (8K
window -> 2048 reserve / 6144 budget; >= 64K windows byte-identical to before;
unknown window keeps the classic constant). An explicit
MAKA_CONTEXT_HISTORY_COMPACT_RESERVE_TOKENS is respected verbatim. Peers bound
the same way: opencode caps its compaction buffer by the model's output limit,
gemini-cli triggers at a window fraction; pi shares the flat-16384 blind spot.
Strengthen the shipped-defaults journey to fully-default derivation with
durable-persistence assertions (checkpoint recorded, compact block present,
replaced raw span gone), and add the unrescuable path: under the derived
default a turn with no safe span ends with the explicit
context_budget_exhausted outcome. A small-window guard locks the P2 behavior:
on an 8K model the default is completely inert when under budget — never a
pointless summarizer call.
Refs #882
* docs(runtime): align midTurn policy comments with the shipped runtime default
@likun666661
likun666661 merged commit 01fc111 into apache:mainJul 16, 2026
2 of 3 checks passed
Astro-Han added a commit that referenced this pull request Jul 16, 2026
…c CLI color tests) (#1077)
* fix(runtime): let active prune cover the newest step under mid-turn capacity
#986 unconditionally excluded the newest completed step from active
tool-result pruning, breaking three reviewed capacity-rescue invariants
(mid-turn finding C, capacity-replacement re-convergence, overflow
round-3 P1) and CI on main. Include the newest step exactly when
mid-turn capacity compaction is active, keeping the keep-newest-raw
default for every other configuration.
* test(cli): pin color level so ANSI assertions are hermetic
tui-ansi detects color capability from TERM/COLORTERM at module load, so
the pi-transcript and pi-tui-runner ANSI assertions inherited ambient
terminal capability: truecolor locally, level 0 on CI runners with no
TERM/COLORTERM, where every color function is a no-op and the #1064/#1066
color assertions fail deterministically. Pin level 3 via the existing
_setColorLevelForTesting seam, matching tui-ansi.test.ts.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@likun666661
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

feat(runtime): add attention-first semantic compaction - #986

Merged
likun666661 merged 6 commits into
apache:mainfrom
likun666661:feat/attention-first-semantic-compaction
Jul 16, 2026
Merged

feat(runtime): add attention-first semantic compaction#986
likun666661 merged 6 commits into
apache:mainfrom
likun666661:feat/attention-first-semantic-compaction

Conversation

@likun666661

@likun666661likun666661 commented Jul 14, 2026

Copy link
Copy Markdown
Member

Summary

  • preserve prior replay and the exact current-turn user message as an immutable head anchor
  • replace only the structurally complete middle span before the open provider-protocol tail
  • use the 8K total-context / 4K completed-span attention trigger and separate generation / accepted-projection budgets
  • add bounded structured and complete-sentence/line fallback validation, V2 rolling lineage, and shared attention/emergency scheduling
  • keep the provider-visible continuation strictly exact head anchor + LLM-authored projection + exact protocol tail
  • keep Runtime coverage, source refs, archive refs, preserved-tail indexes, and legacy state cards in the diagnostic side channel only

Safety properties

  • the head anchor is signature-checked on every prepareStep and is never semantic coverage
  • consecutive assistant reasoning/text/tool-call messages form one provider episode; incomplete multi-call episodes remain verbatim in the tail
  • empty, truncated, source-invalid, private-verifier-invalid, protocol-invalid, or over-budget projections do not replace source messages
  • token savings, compact-call cost, and cache loss remain diagnostics in attention mode
  • one prepareStep accepts at most one attention or emergency replacement
  • emergency compaction reports explicit head_anchor_exceeds_capacity / no_safe_completed_span outcomes instead of rewriting the user instruction
  • semanticCompact no longer generates, prompts with, records on new blocks, or renders deterministic state cards
  • rolling successors re-render the predecessor's LLM summary and ignore legacy state cards instead of replaying the old provider message verbatim

MIPS smoke finding

A GLM-5.2 make-mips-interpreter smoke on the earlier version was stopped after it exceeded 135 steps without creating /app/vm.js.

The retained trace showed:

  • 24 accepted semantic compactions and 228 tool calls
  • 19 later summaries whose next action was to write /app/vm.js
  • provider-visible state cards that instead repeated assistant planning such as “re-read” or “re-verify”
  • zero creation of /app/vm.js

The deterministic VM matcher treated any assistant text containing vm as operational state. Those inferred cards were rendered after every LLM summary and then copied into rolling successors, creating a positive feedback loop. The fix in 4eebe953 removes that semantic path rather than attempting to repair it with more regexes.

This smoke is mechanism evidence, not an A/B quality result. The old main-branch arm separately failed to activate semantic compaction because of tool_pair_split and eventually timed out near provider capacity, so it is not a usable success baseline.

Validation

After merging current upstream/main:

  • focused semantic + backend + mid-turn capacity integration: 187 passed, 0 failed
  • npm --workspace @maka/headless test: 964 passed, 0 failed, 1 skipped
  • npm run typecheck: passes across all workspaces after rebuilding changed workspace dependencies
  • Runtime full suite: 1943 passed, 7 skipped, with one unrelated local-only failure in main's macOS sandbox executable-root index assertion because this machine runs Node from /usr/local/bin/node
  • git diff --check: passes

No attention benefit is claimed yet, and this PR does not expand any surface default. A fresh same-commit A/B is still required after the mechanism correction.

Refs #981

Astro-Han added a commit that referenced this pull request Jul 15, 2026
* feat(runtime): default mid-turn capacity compaction on, runtime-owned
The mid_turn capacity invariant becomes a runtime-owned default: whenever
history compaction is enabled the runtime derives midTurn from the selected
model window (contextWindow - reserve), so Desktop, TUI, and headless inherit
it without copying config. MAKA_CONTEXT_HISTORY_COMPACT_MID_TURN=off stays as
the explicit escape hatch.
The backend still gates activation on the checkpoint seams, the persisted head
anchor, and a KNOWN context window, so a session without model metadata or a
child session with no anchor seam never misfires even though the default is on
(covered by new backend safety guards).
Refs #882
* feat: default semantic compaction off, runtime-owned across surfaces
Semantic compaction is the #981/#986 attention-first experiment, not the
standard historyCompact mechanism. Make the runtime the owner of its default:
it stays off unless a surface explicitly opts in (MAKA_CONTEXT_SEMANTIC_COMPACT
truthy, or a non-off MAKA_CONTEXT_SEMANTIC_COMPACT_MODE).
This lets the CLI drop its local semanticCompact strip and eat the runtime
default directly, so the surfaces stop duplicating the same off decision.
Desktop and headless inherit the off default without any per-surface wiring.
Refs #882
* test(runtime): lock the shipped default's proactive long-turn journey
Wire buildDefaultContextBudgetPolicy — the exact default every surface now
inherits — into the streaming backend and prove a long turn near the window
engages mid-turn compaction and continues to a clean end rather than being
truncated or surfacing a raw provider error.
The fake-backend Playwright suite cannot exercise this path: it is a separate
canned-event backend that never runs AiSdkBackend's prepareStep machinery, so a
desktop E2E could only fake the very compaction under test. This deterministic
runtime-layer test is the honest representative journey for the proactive path.
Refs #882
* fix(runtime): bound the derived compaction reserve by the model window
The flat 16384 reserve assumed large-window models: on gpt-4 (8192 window)
the default policy derived maxHistoryEstimatedTokens = max(1, 8192 - 16384) = 1
— a pre-existing turn-boundary pathology on main — and with mid_turn now on by
default the 1-token high water made every multi-step turn run the summarizer
for a checkpoint the replay gate could never admit.
Make defaultCompactReserveTokens the single owner of the reserve default,
bounded by the KNOWN window: min(16384, floor(window / 4)), so both the
turn-boundary budget and the mid_turn high water derive from one value (8K
window -> 2048 reserve / 6144 budget; >= 64K windows byte-identical to before;
unknown window keeps the classic constant). An explicit
MAKA_CONTEXT_HISTORY_COMPACT_RESERVE_TOKENS is respected verbatim. Peers bound
the same way: opencode caps its compaction buffer by the model's output limit,
gemini-cli triggers at a window fraction; pi shares the flat-16384 blind spot.
Strengthen the shipped-defaults journey to fully-default derivation with
durable-persistence assertions (checkpoint recorded, compact block present,
replaced raw span gone), and add the unrescuable path: under the derived
default a turn with no safe span ends with the explicit
context_budget_exhausted outcome. A small-window guard locks the P2 behavior:
on an 8K model the default is completely inert when under budget — never a
pointless summarizer call.
Refs #882
* docs(runtime): align midTurn policy comments with the shipped runtime default
@likun666661
likun666661 merged commit 01fc111 into apache:mainJul 16, 2026
2 of 3 checks passed
Astro-Han added a commit that referenced this pull request Jul 16, 2026
…c CLI color tests) (#1077)
* fix(runtime): let active prune cover the newest step under mid-turn capacity
#986 unconditionally excluded the newest completed step from active
tool-result pruning, breaking three reviewed capacity-rescue invariants
(mid-turn finding C, capacity-replacement re-convergence, overflow
round-3 P1) and CI on main. Include the newest step exactly when
mid-turn capacity compaction is active, keeping the keep-newest-raw
default for every other configuration.
* test(cli): pin color level so ANSI assertions are hermetic
tui-ansi detects color capability from TERM/COLORTERM at module load, so
the pi-transcript and pi-tui-runner ANSI assertions inherited ambient
terminal capability: truecolor locally, level 0 on CI runners with no
TERM/COLORTERM, where every color function is a no-op and the #1064/#1066
color assertions fail deterministically. Pin level 3 via the existing
_setColorLevelForTesting seam, matching tui-ansi.test.ts.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@likun666661
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(runtime): add attention-first semantic compaction - #986

Merged
likun666661 merged 6 commits into
apache:mainfrom
likun666661:feat/attention-first-semantic-compaction
Jul 16, 2026
Merged

feat(runtime): add attention-first semantic compaction#986
likun666661 merged 6 commits into
apache:mainfrom
likun666661:feat/attention-first-semantic-compaction

Conversation

@likun666661

@likun666661likun666661 commented Jul 14, 2026

Copy link
Copy Markdown
Member

Summary

  • preserve prior replay and the exact current-turn user message as an immutable head anchor
  • replace only the structurally complete middle span before the open provider-protocol tail
  • use the 8K total-context / 4K completed-span attention trigger and separate generation / accepted-projection budgets
  • add bounded structured and complete-sentence/line fallback validation, V2 rolling lineage, and shared attention/emergency scheduling
  • keep the provider-visible continuation strictly exact head anchor + LLM-authored projection + exact protocol tail
  • keep Runtime coverage, source refs, archive refs, preserved-tail indexes, and legacy state cards in the diagnostic side channel only

Safety properties

  • the head anchor is signature-checked on every prepareStep and is never semantic coverage
  • consecutive assistant reasoning/text/tool-call messages form one provider episode; incomplete multi-call episodes remain verbatim in the tail
  • empty, truncated, source-invalid, private-verifier-invalid, protocol-invalid, or over-budget projections do not replace source messages
  • token savings, compact-call cost, and cache loss remain diagnostics in attention mode
  • one prepareStep accepts at most one attention or emergency replacement
  • emergency compaction reports explicit head_anchor_exceeds_capacity / no_safe_completed_span outcomes instead of rewriting the user instruction
  • semanticCompact no longer generates, prompts with, records on new blocks, or renders deterministic state cards
  • rolling successors re-render the predecessor's LLM summary and ignore legacy state cards instead of replaying the old provider message verbatim

MIPS smoke finding

A GLM-5.2 make-mips-interpreter smoke on the earlier version was stopped after it exceeded 135 steps without creating /app/vm.js.

The retained trace showed:

  • 24 accepted semantic compactions and 228 tool calls
  • 19 later summaries whose next action was to write /app/vm.js
  • provider-visible state cards that instead repeated assistant planning such as “re-read” or “re-verify”
  • zero creation of /app/vm.js

The deterministic VM matcher treated any assistant text containing vm as operational state. Those inferred cards were rendered after every LLM summary and then copied into rolling successors, creating a positive feedback loop. The fix in 4eebe953 removes that semantic path rather than attempting to repair it with more regexes.

This smoke is mechanism evidence, not an A/B quality result. The old main-branch arm separately failed to activate semantic compaction because of tool_pair_split and eventually timed out near provider capacity, so it is not a usable success baseline.

Validation

After merging current upstream/main:

  • focused semantic + backend + mid-turn capacity integration: 187 passed, 0 failed
  • npm --workspace @maka/headless test: 964 passed, 0 failed, 1 skipped
  • npm run typecheck: passes across all workspaces after rebuilding changed workspace dependencies
  • Runtime full suite: 1943 passed, 7 skipped, with one unrelated local-only failure in main's macOS sandbox executable-root index assertion because this machine runs Node from /usr/local/bin/node
  • git diff --check: passes

No attention benefit is claimed yet, and this PR does not expand any surface default. A fresh same-commit A/B is still required after the mechanism correction.

Refs #981

Astro-Han added a commit that referenced this pull request Jul 15, 2026
* feat(runtime): default mid-turn capacity compaction on, runtime-owned
The mid_turn capacity invariant becomes a runtime-owned default: whenever
history compaction is enabled the runtime derives midTurn from the selected
model window (contextWindow - reserve), so Desktop, TUI, and headless inherit
it without copying config. MAKA_CONTEXT_HISTORY_COMPACT_MID_TURN=off stays as
the explicit escape hatch.
The backend still gates activation on the checkpoint seams, the persisted head
anchor, and a KNOWN context window, so a session without model metadata or a
child session with no anchor seam never misfires even though the default is on
(covered by new backend safety guards).
Refs #882
* feat: default semantic compaction off, runtime-owned across surfaces
Semantic compaction is the #981/#986 attention-first experiment, not the
standard historyCompact mechanism. Make the runtime the owner of its default:
it stays off unless a surface explicitly opts in (MAKA_CONTEXT_SEMANTIC_COMPACT
truthy, or a non-off MAKA_CONTEXT_SEMANTIC_COMPACT_MODE).
This lets the CLI drop its local semanticCompact strip and eat the runtime
default directly, so the surfaces stop duplicating the same off decision.
Desktop and headless inherit the off default without any per-surface wiring.
Refs #882
* test(runtime): lock the shipped default's proactive long-turn journey
Wire buildDefaultContextBudgetPolicy — the exact default every surface now
inherits — into the streaming backend and prove a long turn near the window
engages mid-turn compaction and continues to a clean end rather than being
truncated or surfacing a raw provider error.
The fake-backend Playwright suite cannot exercise this path: it is a separate
canned-event backend that never runs AiSdkBackend's prepareStep machinery, so a
desktop E2E could only fake the very compaction under test. This deterministic
runtime-layer test is the honest representative journey for the proactive path.
Refs #882
* fix(runtime): bound the derived compaction reserve by the model window
The flat 16384 reserve assumed large-window models: on gpt-4 (8192 window)
the default policy derived maxHistoryEstimatedTokens = max(1, 8192 - 16384) = 1
— a pre-existing turn-boundary pathology on main — and with mid_turn now on by
default the 1-token high water made every multi-step turn run the summarizer
for a checkpoint the replay gate could never admit.
Make defaultCompactReserveTokens the single owner of the reserve default,
bounded by the KNOWN window: min(16384, floor(window / 4)), so both the
turn-boundary budget and the mid_turn high water derive from one value (8K
window -> 2048 reserve / 6144 budget; >= 64K windows byte-identical to before;
unknown window keeps the classic constant). An explicit
MAKA_CONTEXT_HISTORY_COMPACT_RESERVE_TOKENS is respected verbatim. Peers bound
the same way: opencode caps its compaction buffer by the model's output limit,
gemini-cli triggers at a window fraction; pi shares the flat-16384 blind spot.
Strengthen the shipped-defaults journey to fully-default derivation with
durable-persistence assertions (checkpoint recorded, compact block present,
replaced raw span gone), and add the unrescuable path: under the derived
default a turn with no safe span ends with the explicit
context_budget_exhausted outcome. A small-window guard locks the P2 behavior:
on an 8K model the default is completely inert when under budget — never a
pointless summarizer call.
Refs #882
* docs(runtime): align midTurn policy comments with the shipped runtime default
@likun666661
likun666661 merged commit 01fc111 into apache:mainJul 16, 2026
2 of 3 checks passed
Astro-Han added a commit that referenced this pull request Jul 16, 2026
…c CLI color tests) (#1077)
* fix(runtime): let active prune cover the newest step under mid-turn capacity
#986 unconditionally excluded the newest completed step from active
tool-result pruning, breaking three reviewed capacity-rescue invariants
(mid-turn finding C, capacity-replacement re-convergence, overflow
round-3 P1) and CI on main. Include the newest step exactly when
mid-turn capacity compaction is active, keeping the keep-newest-raw
default for every other configuration.
* test(cli): pin color level so ANSI assertions are hermetic
tui-ansi detects color capability from TERM/COLORTERM at module load, so
the pi-transcript and pi-tui-runner ANSI assertions inherited ambient
terminal capability: truecolor locally, level 0 on CI runners with no
TERM/COLORTERM, where every color function is a no-op and the #1064/#1066
color assertions fail deterministically. Pin level 3 via the existing
_setColorLevelForTesting seam, matching tui-ansi.test.ts.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@likun666661
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(runtime): add attention-first semantic compaction - #986

Merged
likun666661 merged 6 commits into
apache:mainfrom
likun666661:feat/attention-first-semantic-compaction
Jul 16, 2026
Merged

feat(runtime): add attention-first semantic compaction#986
likun666661 merged 6 commits into
apache:mainfrom
likun666661:feat/attention-first-semantic-compaction

Conversation

@likun666661

@likun666661likun666661 commented Jul 14, 2026

Copy link
Copy Markdown
Member

Summary

  • preserve prior replay and the exact current-turn user message as an immutable head anchor
  • replace only the structurally complete middle span before the open provider-protocol tail
  • use the 8K total-context / 4K completed-span attention trigger and separate generation / accepted-projection budgets
  • add bounded structured and complete-sentence/line fallback validation, V2 rolling lineage, and shared attention/emergency scheduling
  • keep the provider-visible continuation strictly exact head anchor + LLM-authored projection + exact protocol tail
  • keep Runtime coverage, source refs, archive refs, preserved-tail indexes, and legacy state cards in the diagnostic side channel only

Safety properties

  • the head anchor is signature-checked on every prepareStep and is never semantic coverage
  • consecutive assistant reasoning/text/tool-call messages form one provider episode; incomplete multi-call episodes remain verbatim in the tail
  • empty, truncated, source-invalid, private-verifier-invalid, protocol-invalid, or over-budget projections do not replace source messages
  • token savings, compact-call cost, and cache loss remain diagnostics in attention mode
  • one prepareStep accepts at most one attention or emergency replacement
  • emergency compaction reports explicit head_anchor_exceeds_capacity / no_safe_completed_span outcomes instead of rewriting the user instruction
  • semanticCompact no longer generates, prompts with, records on new blocks, or renders deterministic state cards
  • rolling successors re-render the predecessor's LLM summary and ignore legacy state cards instead of replaying the old provider message verbatim

MIPS smoke finding

A GLM-5.2 make-mips-interpreter smoke on the earlier version was stopped after it exceeded 135 steps without creating /app/vm.js.

The retained trace showed:

  • 24 accepted semantic compactions and 228 tool calls
  • 19 later summaries whose next action was to write /app/vm.js
  • provider-visible state cards that instead repeated assistant planning such as “re-read” or “re-verify”
  • zero creation of /app/vm.js

The deterministic VM matcher treated any assistant text containing vm as operational state. Those inferred cards were rendered after every LLM summary and then copied into rolling successors, creating a positive feedback loop. The fix in 4eebe953 removes that semantic path rather than attempting to repair it with more regexes.

This smoke is mechanism evidence, not an A/B quality result. The old main-branch arm separately failed to activate semantic compaction because of tool_pair_split and eventually timed out near provider capacity, so it is not a usable success baseline.

Validation

After merging current upstream/main:

  • focused semantic + backend + mid-turn capacity integration: 187 passed, 0 failed
  • npm --workspace @maka/headless test: 964 passed, 0 failed, 1 skipped
  • npm run typecheck: passes across all workspaces after rebuilding changed workspace dependencies
  • Runtime full suite: 1943 passed, 7 skipped, with one unrelated local-only failure in main's macOS sandbox executable-root index assertion because this machine runs Node from /usr/local/bin/node
  • git diff --check: passes

No attention benefit is claimed yet, and this PR does not expand any surface default. A fresh same-commit A/B is still required after the mechanism correction.

Refs #981

Astro-Han added a commit that referenced this pull request Jul 15, 2026
* feat(runtime): default mid-turn capacity compaction on, runtime-owned
The mid_turn capacity invariant becomes a runtime-owned default: whenever
history compaction is enabled the runtime derives midTurn from the selected
model window (contextWindow - reserve), so Desktop, TUI, and headless inherit
it without copying config. MAKA_CONTEXT_HISTORY_COMPACT_MID_TURN=off stays as
the explicit escape hatch.
The backend still gates activation on the checkpoint seams, the persisted head
anchor, and a KNOWN context window, so a session without model metadata or a
child session with no anchor seam never misfires even though the default is on
(covered by new backend safety guards).
Refs #882
* feat: default semantic compaction off, runtime-owned across surfaces
Semantic compaction is the #981/#986 attention-first experiment, not the
standard historyCompact mechanism. Make the runtime the owner of its default:
it stays off unless a surface explicitly opts in (MAKA_CONTEXT_SEMANTIC_COMPACT
truthy, or a non-off MAKA_CONTEXT_SEMANTIC_COMPACT_MODE).
This lets the CLI drop its local semanticCompact strip and eat the runtime
default directly, so the surfaces stop duplicating the same off decision.
Desktop and headless inherit the off default without any per-surface wiring.
Refs #882
* test(runtime): lock the shipped default's proactive long-turn journey
Wire buildDefaultContextBudgetPolicy — the exact default every surface now
inherits — into the streaming backend and prove a long turn near the window
engages mid-turn compaction and continues to a clean end rather than being
truncated or surfacing a raw provider error.
The fake-backend Playwright suite cannot exercise this path: it is a separate
canned-event backend that never runs AiSdkBackend's prepareStep machinery, so a
desktop E2E could only fake the very compaction under test. This deterministic
runtime-layer test is the honest representative journey for the proactive path.
Refs #882
* fix(runtime): bound the derived compaction reserve by the model window
The flat 16384 reserve assumed large-window models: on gpt-4 (8192 window)
the default policy derived maxHistoryEstimatedTokens = max(1, 8192 - 16384) = 1
— a pre-existing turn-boundary pathology on main — and with mid_turn now on by
default the 1-token high water made every multi-step turn run the summarizer
for a checkpoint the replay gate could never admit.
Make defaultCompactReserveTokens the single owner of the reserve default,
bounded by the KNOWN window: min(16384, floor(window / 4)), so both the
turn-boundary budget and the mid_turn high water derive from one value (8K
window -> 2048 reserve / 6144 budget; >= 64K windows byte-identical to before;
unknown window keeps the classic constant). An explicit
MAKA_CONTEXT_HISTORY_COMPACT_RESERVE_TOKENS is respected verbatim. Peers bound
the same way: opencode caps its compaction buffer by the model's output limit,
gemini-cli triggers at a window fraction; pi shares the flat-16384 blind spot.
Strengthen the shipped-defaults journey to fully-default derivation with
durable-persistence assertions (checkpoint recorded, compact block present,
replaced raw span gone), and add the unrescuable path: under the derived
default a turn with no safe span ends with the explicit
context_budget_exhausted outcome. A small-window guard locks the P2 behavior:
on an 8K model the default is completely inert when under budget — never a
pointless summarizer call.
Refs #882
* docs(runtime): align midTurn policy comments with the shipped runtime default
@likun666661
likun666661 merged commit 01fc111 into apache:mainJul 16, 2026
2 of 3 checks passed
Astro-Han added a commit that referenced this pull request Jul 16, 2026
…c CLI color tests) (#1077)
* fix(runtime): let active prune cover the newest step under mid-turn capacity
#986 unconditionally excluded the newest completed step from active
tool-result pruning, breaking three reviewed capacity-rescue invariants
(mid-turn finding C, capacity-replacement re-convergence, overflow
round-3 P1) and CI on main. Include the newest step exactly when
mid-turn capacity compaction is active, keeping the keep-newest-raw
default for every other configuration.
* test(cli): pin color level so ANSI assertions are hermetic
tui-ansi detects color capability from TERM/COLORTERM at module load, so
the pi-transcript and pi-tui-runner ANSI assertions inherited ambient
terminal capability: truecolor locally, level 0 on CI runners with no
TERM/COLORTERM, where every color function is a no-op and the #1064/#1066
color assertions fail deterministically. Pin level 3 via the existing
_setColorLevelForTesting seam, matching tui-ansi.test.ts.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@likun666661
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

feat(runtime): add attention-first semantic compaction - #986

Merged
likun666661 merged 6 commits into
apache:mainfrom
likun666661:feat/attention-first-semantic-compaction
Jul 16, 2026
Merged

feat(runtime): add attention-first semantic compaction#986
likun666661 merged 6 commits into
apache:mainfrom
likun666661:feat/attention-first-semantic-compaction

Conversation

@likun666661

@likun666661likun666661 commented Jul 14, 2026

Copy link
Copy Markdown
Member

Summary

  • preserve prior replay and the exact current-turn user message as an immutable head anchor
  • replace only the structurally complete middle span before the open provider-protocol tail
  • use the 8K total-context / 4K completed-span attention trigger and separate generation / accepted-projection budgets
  • add bounded structured and complete-sentence/line fallback validation, V2 rolling lineage, and shared attention/emergency scheduling
  • keep the provider-visible continuation strictly exact head anchor + LLM-authored projection + exact protocol tail
  • keep Runtime coverage, source refs, archive refs, preserved-tail indexes, and legacy state cards in the diagnostic side channel only

Safety properties

  • the head anchor is signature-checked on every prepareStep and is never semantic coverage
  • consecutive assistant reasoning/text/tool-call messages form one provider episode; incomplete multi-call episodes remain verbatim in the tail
  • empty, truncated, source-invalid, private-verifier-invalid, protocol-invalid, or over-budget projections do not replace source messages
  • token savings, compact-call cost, and cache loss remain diagnostics in attention mode
  • one prepareStep accepts at most one attention or emergency replacement
  • emergency compaction reports explicit head_anchor_exceeds_capacity / no_safe_completed_span outcomes instead of rewriting the user instruction
  • semanticCompact no longer generates, prompts with, records on new blocks, or renders deterministic state cards
  • rolling successors re-render the predecessor's LLM summary and ignore legacy state cards instead of replaying the old provider message verbatim

MIPS smoke finding

A GLM-5.2 make-mips-interpreter smoke on the earlier version was stopped after it exceeded 135 steps without creating /app/vm.js.

The retained trace showed:

  • 24 accepted semantic compactions and 228 tool calls
  • 19 later summaries whose next action was to write /app/vm.js
  • provider-visible state cards that instead repeated assistant planning such as “re-read” or “re-verify”
  • zero creation of /app/vm.js

The deterministic VM matcher treated any assistant text containing vm as operational state. Those inferred cards were rendered after every LLM summary and then copied into rolling successors, creating a positive feedback loop. The fix in 4eebe953 removes that semantic path rather than attempting to repair it with more regexes.

This smoke is mechanism evidence, not an A/B quality result. The old main-branch arm separately failed to activate semantic compaction because of tool_pair_split and eventually timed out near provider capacity, so it is not a usable success baseline.

Validation

After merging current upstream/main:

  • focused semantic + backend + mid-turn capacity integration: 187 passed, 0 failed
  • npm --workspace @maka/headless test: 964 passed, 0 failed, 1 skipped
  • npm run typecheck: passes across all workspaces after rebuilding changed workspace dependencies
  • Runtime full suite: 1943 passed, 7 skipped, with one unrelated local-only failure in main's macOS sandbox executable-root index assertion because this machine runs Node from /usr/local/bin/node
  • git diff --check: passes

No attention benefit is claimed yet, and this PR does not expand any surface default. A fresh same-commit A/B is still required after the mechanism correction.

Refs #981

Astro-Han added a commit that referenced this pull request Jul 15, 2026
* feat(runtime): default mid-turn capacity compaction on, runtime-owned
The mid_turn capacity invariant becomes a runtime-owned default: whenever
history compaction is enabled the runtime derives midTurn from the selected
model window (contextWindow - reserve), so Desktop, TUI, and headless inherit
it without copying config. MAKA_CONTEXT_HISTORY_COMPACT_MID_TURN=off stays as
the explicit escape hatch.
The backend still gates activation on the checkpoint seams, the persisted head
anchor, and a KNOWN context window, so a session without model metadata or a
child session with no anchor seam never misfires even though the default is on
(covered by new backend safety guards).
Refs #882
* feat: default semantic compaction off, runtime-owned across surfaces
Semantic compaction is the #981/#986 attention-first experiment, not the
standard historyCompact mechanism. Make the runtime the owner of its default:
it stays off unless a surface explicitly opts in (MAKA_CONTEXT_SEMANTIC_COMPACT
truthy, or a non-off MAKA_CONTEXT_SEMANTIC_COMPACT_MODE).
This lets the CLI drop its local semanticCompact strip and eat the runtime
default directly, so the surfaces stop duplicating the same off decision.
Desktop and headless inherit the off default without any per-surface wiring.
Refs #882
* test(runtime): lock the shipped default's proactive long-turn journey
Wire buildDefaultContextBudgetPolicy — the exact default every surface now
inherits — into the streaming backend and prove a long turn near the window
engages mid-turn compaction and continues to a clean end rather than being
truncated or surfacing a raw provider error.
The fake-backend Playwright suite cannot exercise this path: it is a separate
canned-event backend that never runs AiSdkBackend's prepareStep machinery, so a
desktop E2E could only fake the very compaction under test. This deterministic
runtime-layer test is the honest representative journey for the proactive path.
Refs #882
* fix(runtime): bound the derived compaction reserve by the model window
The flat 16384 reserve assumed large-window models: on gpt-4 (8192 window)
the default policy derived maxHistoryEstimatedTokens = max(1, 8192 - 16384) = 1
— a pre-existing turn-boundary pathology on main — and with mid_turn now on by
default the 1-token high water made every multi-step turn run the summarizer
for a checkpoint the replay gate could never admit.
Make defaultCompactReserveTokens the single owner of the reserve default,
bounded by the KNOWN window: min(16384, floor(window / 4)), so both the
turn-boundary budget and the mid_turn high water derive from one value (8K
window -> 2048 reserve / 6144 budget; >= 64K windows byte-identical to before;
unknown window keeps the classic constant). An explicit
MAKA_CONTEXT_HISTORY_COMPACT_RESERVE_TOKENS is respected verbatim. Peers bound
the same way: opencode caps its compaction buffer by the model's output limit,
gemini-cli triggers at a window fraction; pi shares the flat-16384 blind spot.
Strengthen the shipped-defaults journey to fully-default derivation with
durable-persistence assertions (checkpoint recorded, compact block present,
replaced raw span gone), and add the unrescuable path: under the derived
default a turn with no safe span ends with the explicit
context_budget_exhausted outcome. A small-window guard locks the P2 behavior:
on an 8K model the default is completely inert when under budget — never a
pointless summarizer call.
Refs #882
* docs(runtime): align midTurn policy comments with the shipped runtime default
@likun666661
likun666661 merged commit 01fc111 into apache:mainJul 16, 2026
2 of 3 checks passed
Astro-Han added a commit that referenced this pull request Jul 16, 2026
…c CLI color tests) (#1077)
* fix(runtime): let active prune cover the newest step under mid-turn capacity
#986 unconditionally excluded the newest completed step from active
tool-result pruning, breaking three reviewed capacity-rescue invariants
(mid-turn finding C, capacity-replacement re-convergence, overflow
round-3 P1) and CI on main. Include the newest step exactly when
mid-turn capacity compaction is active, keeping the keep-newest-raw
default for every other configuration.
* test(cli): pin color level so ANSI assertions are hermetic
tui-ansi detects color capability from TERM/COLORTERM at module load, so
the pi-transcript and pi-tui-runner ANSI assertions inherited ambient
terminal capability: truecolor locally, level 0 on CI runners with no
TERM/COLORTERM, where every color function is a no-op and the #1064/#1066
color assertions fail deterministically. Pin level 3 via the existing
_setColorLevelForTesting seam, matching tui-ansi.test.ts.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@likun666661