Skip to content

feat(headless): report active prune evidence - #323

Merged
Astro-Han merged 4 commits into
mainfrom
opencode/issue293-prune-ab-run
Jun 27, 2026
Merged

feat(headless): report active prune evidence#323
Astro-Han merged 4 commits into
mainfrom
opencode/issue293-prune-ab-run

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Jun 27, 2026

Copy link
Copy Markdown
Contributor

Summary

Add active tool-result prune diagnostics to Harbor cell output, A/B summaries, and markdown reports, including a B-active paired comparison slice.

Why

This covers part of the #293 phase-1 reporting gap for active tool-result pruning. Runtime already emits active prune diagnostics, but headless A/B reports previously only recognized stale prunedToolResults. That made active-prune runs hard to inspect and could mix untriggered or unpaired attempts into safety conclusions.

Scope

Changed:

  • Summarize and validate activePrunedToolResults, activeEstimatedTokensSaved, and activeArchiveFailures in Harbor cell output.
  • Include activeToolResultPrune in Harbor context budget policy snapshots.
  • Add active-prune reporting keyed by activePrunedToolResults > 0.
  • Report the active-prune subset as a B-active paired slice: the B attempts that activated active prune, plus the matching A (taskId, rep) attempts.
  • Report active subset task count, attempt count, observed/missing/coverage, pass rate, full token/cost, failure diagnostics, and context/archive diagnostics.
  • Keep activated attempt refs active-only so stale-only pruning does not appear as active-prune evidence.
  • Render existing archive diagnostics: archive placeholders, archive write failures, retrieved count/tokens, retrieval skipped, retrieval failures, and active archive failures.
  • Add/adjust headless tests for active diagnostics, active-only subset reporting, stale-only exclusion, paired active subset reporting, archive metric rendering, missing-pair coverage, and full token/cost rendering.

Not included:

  • No stale prior-context prune validation.
  • No multi-turn or autonomous continuation harness.
  • No archive reason-count schema expansion.
  • No field-list-driven context-budget metric refactor; that maintenance cleanup can be a follow-up PR.
  • No formal non-inferiority claim.

Verification

  • RED: npm --workspace @maka/headless test -- --test-name-pattern "renders active prune subset pair coverage" failed because the active subset line omitted attempts/observed/missing/coverage and full token/cost fields.
  • GREEN: npm --workspace @maka/headless test -- --test-name-pattern "active prune subset|context budget activation": 490 pass, 0 fail. The package test script still ran the full @maka/headless suite.
  • Re-rendered active-prune smoke report from existing results.jsonl: /Users/yuhan/.local/maka-eval/runs/issue293-prune-ab/issue293-active-smoke-001/runtime-policy-ab-report.md
  • Active-prune smoke run: /Users/yuhan/.local/maka-eval/runs/issue293-prune-ab/issue293-active-smoke-001
  • Tasks: count-dataset-tokens, extract-elf
  • A: context budget off
  • B: active prune + archive retrieval
  • B activation: 2/2 attempts, 2 tasks
  • activePrunedToolResults=1091
  • activeEstimatedTokensSaved=1400611
  • activeArchiveFailures=0
  • Infra/plumbing failures: 0
  • Clash/Mihomo observed traffic delta: 535.92 MiB

User-facing impact

None. This is headless benchmark reporting only.

Reviewer notes

The smoke run showed active pruning activates and reports correctly, but both B attempts ended with incomplete_tool_calls after hitting the single-turn 50-step cap. That is a harness limitation for longer tasks, not a non-inferiority result. Formal #293 evidence should wait for either shorter stable tasks or a benchmark-safe multi-turn continuation harness.

@Astro-HanAstro-Han changed the title feat(headless): report active prune diagnosticsfeat(headless): report active prune evidenceJun 27, 2026
@Astro-Han
Astro-Han merged commit 6ec15cf into mainJun 27, 2026
@Astro-Han
Astro-Han deleted the opencode/issue293-prune-ab-run branch June 27, 2026 06:45
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
feat(headless): report active prune evidence by Astro-Han · Pull Request #323 · apache/maka · GitHub
Skip to content

feat(headless): report active prune evidence - #323

Merged
Astro-Han merged 4 commits into
mainfrom
opencode/issue293-prune-ab-run
Jun 27, 2026
Merged

feat(headless): report active prune evidence#323
Astro-Han merged 4 commits into
mainfrom
opencode/issue293-prune-ab-run

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Jun 27, 2026

Copy link
Copy Markdown
Contributor

Summary

Add active tool-result prune diagnostics to Harbor cell output, A/B summaries, and markdown reports, including a B-active paired comparison slice.

Why

This covers part of the #293 phase-1 reporting gap for active tool-result pruning. Runtime already emits active prune diagnostics, but headless A/B reports previously only recognized stale prunedToolResults. That made active-prune runs hard to inspect and could mix untriggered or unpaired attempts into safety conclusions.

Scope

Changed:

  • Summarize and validate activePrunedToolResults, activeEstimatedTokensSaved, and activeArchiveFailures in Harbor cell output.
  • Include activeToolResultPrune in Harbor context budget policy snapshots.
  • Add active-prune reporting keyed by activePrunedToolResults > 0.
  • Report the active-prune subset as a B-active paired slice: the B attempts that activated active prune, plus the matching A (taskId, rep) attempts.
  • Report active subset task count, attempt count, observed/missing/coverage, pass rate, full token/cost, failure diagnostics, and context/archive diagnostics.
  • Keep activated attempt refs active-only so stale-only pruning does not appear as active-prune evidence.
  • Render existing archive diagnostics: archive placeholders, archive write failures, retrieved count/tokens, retrieval skipped, retrieval failures, and active archive failures.
  • Add/adjust headless tests for active diagnostics, active-only subset reporting, stale-only exclusion, paired active subset reporting, archive metric rendering, missing-pair coverage, and full token/cost rendering.

Not included:

  • No stale prior-context prune validation.
  • No multi-turn or autonomous continuation harness.
  • No archive reason-count schema expansion.
  • No field-list-driven context-budget metric refactor; that maintenance cleanup can be a follow-up PR.
  • No formal non-inferiority claim.

Verification

  • RED: npm --workspace @maka/headless test -- --test-name-pattern "renders active prune subset pair coverage" failed because the active subset line omitted attempts/observed/missing/coverage and full token/cost fields.
  • GREEN: npm --workspace @maka/headless test -- --test-name-pattern "active prune subset|context budget activation": 490 pass, 0 fail. The package test script still ran the full @maka/headless suite.
  • Re-rendered active-prune smoke report from existing results.jsonl: /Users/yuhan/.local/maka-eval/runs/issue293-prune-ab/issue293-active-smoke-001/runtime-policy-ab-report.md
  • Active-prune smoke run: /Users/yuhan/.local/maka-eval/runs/issue293-prune-ab/issue293-active-smoke-001
  • Tasks: count-dataset-tokens, extract-elf
  • A: context budget off
  • B: active prune + archive retrieval
  • B activation: 2/2 attempts, 2 tasks
  • activePrunedToolResults=1091
  • activeEstimatedTokensSaved=1400611
  • activeArchiveFailures=0
  • Infra/plumbing failures: 0
  • Clash/Mihomo observed traffic delta: 535.92 MiB

User-facing impact

None. This is headless benchmark reporting only.

Reviewer notes

The smoke run showed active pruning activates and reports correctly, but both B attempts ended with incomplete_tool_calls after hitting the single-turn 50-step cap. That is a harness limitation for longer tasks, not a non-inferiority result. Formal #293 evidence should wait for either shorter stable tasks or a benchmark-safe multi-turn continuation harness.

@Astro-HanAstro-Han changed the title feat(headless): report active prune diagnosticsfeat(headless): report active prune evidenceJun 27, 2026
@Astro-Han
Astro-Han merged commit 6ec15cf into mainJun 27, 2026
@Astro-Han
Astro-Han deleted the opencode/issue293-prune-ab-run branch June 27, 2026 06:45
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' feat(headless): report active prune evidence by Astro-Han · Pull Request #323 · apache/maka · GitHub
Skip to content

feat(headless): report active prune evidence - #323

Merged
Astro-Han merged 4 commits into
mainfrom
opencode/issue293-prune-ab-run
Jun 27, 2026
Merged

feat(headless): report active prune evidence#323
Astro-Han merged 4 commits into
mainfrom
opencode/issue293-prune-ab-run

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Jun 27, 2026

Copy link
Copy Markdown
Contributor

Summary

Add active tool-result prune diagnostics to Harbor cell output, A/B summaries, and markdown reports, including a B-active paired comparison slice.

Why

This covers part of the #293 phase-1 reporting gap for active tool-result pruning. Runtime already emits active prune diagnostics, but headless A/B reports previously only recognized stale prunedToolResults. That made active-prune runs hard to inspect and could mix untriggered or unpaired attempts into safety conclusions.

Scope

Changed:

  • Summarize and validate activePrunedToolResults, activeEstimatedTokensSaved, and activeArchiveFailures in Harbor cell output.
  • Include activeToolResultPrune in Harbor context budget policy snapshots.
  • Add active-prune reporting keyed by activePrunedToolResults > 0.
  • Report the active-prune subset as a B-active paired slice: the B attempts that activated active prune, plus the matching A (taskId, rep) attempts.
  • Report active subset task count, attempt count, observed/missing/coverage, pass rate, full token/cost, failure diagnostics, and context/archive diagnostics.
  • Keep activated attempt refs active-only so stale-only pruning does not appear as active-prune evidence.
  • Render existing archive diagnostics: archive placeholders, archive write failures, retrieved count/tokens, retrieval skipped, retrieval failures, and active archive failures.
  • Add/adjust headless tests for active diagnostics, active-only subset reporting, stale-only exclusion, paired active subset reporting, archive metric rendering, missing-pair coverage, and full token/cost rendering.

Not included:

  • No stale prior-context prune validation.
  • No multi-turn or autonomous continuation harness.
  • No archive reason-count schema expansion.
  • No field-list-driven context-budget metric refactor; that maintenance cleanup can be a follow-up PR.
  • No formal non-inferiority claim.

Verification

  • RED: npm --workspace @maka/headless test -- --test-name-pattern "renders active prune subset pair coverage" failed because the active subset line omitted attempts/observed/missing/coverage and full token/cost fields.
  • GREEN: npm --workspace @maka/headless test -- --test-name-pattern "active prune subset|context budget activation": 490 pass, 0 fail. The package test script still ran the full @maka/headless suite.
  • Re-rendered active-prune smoke report from existing results.jsonl: /Users/yuhan/.local/maka-eval/runs/issue293-prune-ab/issue293-active-smoke-001/runtime-policy-ab-report.md
  • Active-prune smoke run: /Users/yuhan/.local/maka-eval/runs/issue293-prune-ab/issue293-active-smoke-001
  • Tasks: count-dataset-tokens, extract-elf
  • A: context budget off
  • B: active prune + archive retrieval
  • B activation: 2/2 attempts, 2 tasks
  • activePrunedToolResults=1091
  • activeEstimatedTokensSaved=1400611
  • activeArchiveFailures=0
  • Infra/plumbing failures: 0
  • Clash/Mihomo observed traffic delta: 535.92 MiB

User-facing impact

None. This is headless benchmark reporting only.

Reviewer notes

The smoke run showed active pruning activates and reports correctly, but both B attempts ended with incomplete_tool_calls after hitting the single-turn 50-step cap. That is a harness limitation for longer tasks, not a non-inferiority result. Formal #293 evidence should wait for either shorter stable tasks or a benchmark-safe multi-turn continuation harness.

@Astro-HanAstro-Han changed the title feat(headless): report active prune diagnosticsfeat(headless): report active prune evidenceJun 27, 2026
@Astro-Han
Astro-Han merged commit 6ec15cf into mainJun 27, 2026
@Astro-Han
Astro-Han deleted the opencode/issue293-prune-ab-run branch June 27, 2026 06:45
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' feat(headless): report active prune evidence by Astro-Han · Pull Request #323 · apache/maka · GitHub
Skip to content

feat(headless): report active prune evidence - #323

Merged
Astro-Han merged 4 commits into
mainfrom
opencode/issue293-prune-ab-run
Jun 27, 2026
Merged

feat(headless): report active prune evidence#323
Astro-Han merged 4 commits into
mainfrom
opencode/issue293-prune-ab-run

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Jun 27, 2026

Copy link
Copy Markdown
Contributor

Summary

Add active tool-result prune diagnostics to Harbor cell output, A/B summaries, and markdown reports, including a B-active paired comparison slice.

Why

This covers part of the #293 phase-1 reporting gap for active tool-result pruning. Runtime already emits active prune diagnostics, but headless A/B reports previously only recognized stale prunedToolResults. That made active-prune runs hard to inspect and could mix untriggered or unpaired attempts into safety conclusions.

Scope

Changed:

  • Summarize and validate activePrunedToolResults, activeEstimatedTokensSaved, and activeArchiveFailures in Harbor cell output.
  • Include activeToolResultPrune in Harbor context budget policy snapshots.
  • Add active-prune reporting keyed by activePrunedToolResults > 0.
  • Report the active-prune subset as a B-active paired slice: the B attempts that activated active prune, plus the matching A (taskId, rep) attempts.
  • Report active subset task count, attempt count, observed/missing/coverage, pass rate, full token/cost, failure diagnostics, and context/archive diagnostics.
  • Keep activated attempt refs active-only so stale-only pruning does not appear as active-prune evidence.
  • Render existing archive diagnostics: archive placeholders, archive write failures, retrieved count/tokens, retrieval skipped, retrieval failures, and active archive failures.
  • Add/adjust headless tests for active diagnostics, active-only subset reporting, stale-only exclusion, paired active subset reporting, archive metric rendering, missing-pair coverage, and full token/cost rendering.

Not included:

  • No stale prior-context prune validation.
  • No multi-turn or autonomous continuation harness.
  • No archive reason-count schema expansion.
  • No field-list-driven context-budget metric refactor; that maintenance cleanup can be a follow-up PR.
  • No formal non-inferiority claim.

Verification

  • RED: npm --workspace @maka/headless test -- --test-name-pattern "renders active prune subset pair coverage" failed because the active subset line omitted attempts/observed/missing/coverage and full token/cost fields.
  • GREEN: npm --workspace @maka/headless test -- --test-name-pattern "active prune subset|context budget activation": 490 pass, 0 fail. The package test script still ran the full @maka/headless suite.
  • Re-rendered active-prune smoke report from existing results.jsonl: /Users/yuhan/.local/maka-eval/runs/issue293-prune-ab/issue293-active-smoke-001/runtime-policy-ab-report.md
  • Active-prune smoke run: /Users/yuhan/.local/maka-eval/runs/issue293-prune-ab/issue293-active-smoke-001
  • Tasks: count-dataset-tokens, extract-elf
  • A: context budget off
  • B: active prune + archive retrieval
  • B activation: 2/2 attempts, 2 tasks
  • activePrunedToolResults=1091
  • activeEstimatedTokensSaved=1400611
  • activeArchiveFailures=0
  • Infra/plumbing failures: 0
  • Clash/Mihomo observed traffic delta: 535.92 MiB

User-facing impact

None. This is headless benchmark reporting only.

Reviewer notes

The smoke run showed active pruning activates and reports correctly, but both B attempts ended with incomplete_tool_calls after hitting the single-turn 50-step cap. That is a harness limitation for longer tasks, not a non-inferiority result. Formal #293 evidence should wait for either shorter stable tasks or a benchmark-safe multi-turn continuation harness.

@Astro-HanAstro-Han changed the title feat(headless): report active prune diagnosticsfeat(headless): report active prune evidenceJun 27, 2026
@Astro-Han
Astro-Han merged commit 6ec15cf into mainJun 27, 2026
@Astro-Han
Astro-Han deleted the opencode/issue293-prune-ab-run branch June 27, 2026 06:45
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' feat(headless): report active prune evidence by Astro-Han · Pull Request #323 · apache/maka · GitHub
Skip to content

feat(headless): report active prune evidence - #323

Merged
Astro-Han merged 4 commits into
mainfrom
opencode/issue293-prune-ab-run
Jun 27, 2026
Merged

feat(headless): report active prune evidence#323
Astro-Han merged 4 commits into
mainfrom
opencode/issue293-prune-ab-run

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Jun 27, 2026

Copy link
Copy Markdown
Contributor

Summary

Add active tool-result prune diagnostics to Harbor cell output, A/B summaries, and markdown reports, including a B-active paired comparison slice.

Why

This covers part of the #293 phase-1 reporting gap for active tool-result pruning. Runtime already emits active prune diagnostics, but headless A/B reports previously only recognized stale prunedToolResults. That made active-prune runs hard to inspect and could mix untriggered or unpaired attempts into safety conclusions.

Scope

Changed:

  • Summarize and validate activePrunedToolResults, activeEstimatedTokensSaved, and activeArchiveFailures in Harbor cell output.
  • Include activeToolResultPrune in Harbor context budget policy snapshots.
  • Add active-prune reporting keyed by activePrunedToolResults > 0.
  • Report the active-prune subset as a B-active paired slice: the B attempts that activated active prune, plus the matching A (taskId, rep) attempts.
  • Report active subset task count, attempt count, observed/missing/coverage, pass rate, full token/cost, failure diagnostics, and context/archive diagnostics.
  • Keep activated attempt refs active-only so stale-only pruning does not appear as active-prune evidence.
  • Render existing archive diagnostics: archive placeholders, archive write failures, retrieved count/tokens, retrieval skipped, retrieval failures, and active archive failures.
  • Add/adjust headless tests for active diagnostics, active-only subset reporting, stale-only exclusion, paired active subset reporting, archive metric rendering, missing-pair coverage, and full token/cost rendering.

Not included:

  • No stale prior-context prune validation.
  • No multi-turn or autonomous continuation harness.
  • No archive reason-count schema expansion.
  • No field-list-driven context-budget metric refactor; that maintenance cleanup can be a follow-up PR.
  • No formal non-inferiority claim.

Verification

  • RED: npm --workspace @maka/headless test -- --test-name-pattern "renders active prune subset pair coverage" failed because the active subset line omitted attempts/observed/missing/coverage and full token/cost fields.
  • GREEN: npm --workspace @maka/headless test -- --test-name-pattern "active prune subset|context budget activation": 490 pass, 0 fail. The package test script still ran the full @maka/headless suite.
  • Re-rendered active-prune smoke report from existing results.jsonl: /Users/yuhan/.local/maka-eval/runs/issue293-prune-ab/issue293-active-smoke-001/runtime-policy-ab-report.md
  • Active-prune smoke run: /Users/yuhan/.local/maka-eval/runs/issue293-prune-ab/issue293-active-smoke-001
  • Tasks: count-dataset-tokens, extract-elf
  • A: context budget off
  • B: active prune + archive retrieval
  • B activation: 2/2 attempts, 2 tasks
  • activePrunedToolResults=1091
  • activeEstimatedTokensSaved=1400611
  • activeArchiveFailures=0
  • Infra/plumbing failures: 0
  • Clash/Mihomo observed traffic delta: 535.92 MiB

User-facing impact

None. This is headless benchmark reporting only.

Reviewer notes

The smoke run showed active pruning activates and reports correctly, but both B attempts ended with incomplete_tool_calls after hitting the single-turn 50-step cap. That is a harness limitation for longer tasks, not a non-inferiority result. Formal #293 evidence should wait for either shorter stable tasks or a benchmark-safe multi-turn continuation harness.

@Astro-HanAstro-Han changed the title feat(headless): report active prune diagnosticsfeat(headless): report active prune evidenceJun 27, 2026
@Astro-Han
Astro-Han merged commit 6ec15cf into mainJun 27, 2026
@Astro-Han
Astro-Han deleted the opencode/issue293-prune-ab-run branch June 27, 2026 06:45
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' feat(headless): report active prune evidence by Astro-Han · Pull Request #323 · apache/maka · GitHub
Skip to content

feat(headless): report active prune evidence - #323

Merged
Astro-Han merged 4 commits into
mainfrom
opencode/issue293-prune-ab-run
Jun 27, 2026
Merged

feat(headless): report active prune evidence#323
Astro-Han merged 4 commits into
mainfrom
opencode/issue293-prune-ab-run

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Jun 27, 2026

Copy link
Copy Markdown
Contributor

Summary

Add active tool-result prune diagnostics to Harbor cell output, A/B summaries, and markdown reports, including a B-active paired comparison slice.

Why

This covers part of the #293 phase-1 reporting gap for active tool-result pruning. Runtime already emits active prune diagnostics, but headless A/B reports previously only recognized stale prunedToolResults. That made active-prune runs hard to inspect and could mix untriggered or unpaired attempts into safety conclusions.

Scope

Changed:

  • Summarize and validate activePrunedToolResults, activeEstimatedTokensSaved, and activeArchiveFailures in Harbor cell output.
  • Include activeToolResultPrune in Harbor context budget policy snapshots.
  • Add active-prune reporting keyed by activePrunedToolResults > 0.
  • Report the active-prune subset as a B-active paired slice: the B attempts that activated active prune, plus the matching A (taskId, rep) attempts.
  • Report active subset task count, attempt count, observed/missing/coverage, pass rate, full token/cost, failure diagnostics, and context/archive diagnostics.
  • Keep activated attempt refs active-only so stale-only pruning does not appear as active-prune evidence.
  • Render existing archive diagnostics: archive placeholders, archive write failures, retrieved count/tokens, retrieval skipped, retrieval failures, and active archive failures.
  • Add/adjust headless tests for active diagnostics, active-only subset reporting, stale-only exclusion, paired active subset reporting, archive metric rendering, missing-pair coverage, and full token/cost rendering.

Not included:

  • No stale prior-context prune validation.
  • No multi-turn or autonomous continuation harness.
  • No archive reason-count schema expansion.
  • No field-list-driven context-budget metric refactor; that maintenance cleanup can be a follow-up PR.
  • No formal non-inferiority claim.

Verification

  • RED: npm --workspace @maka/headless test -- --test-name-pattern "renders active prune subset pair coverage" failed because the active subset line omitted attempts/observed/missing/coverage and full token/cost fields.
  • GREEN: npm --workspace @maka/headless test -- --test-name-pattern "active prune subset|context budget activation": 490 pass, 0 fail. The package test script still ran the full @maka/headless suite.
  • Re-rendered active-prune smoke report from existing results.jsonl: /Users/yuhan/.local/maka-eval/runs/issue293-prune-ab/issue293-active-smoke-001/runtime-policy-ab-report.md
  • Active-prune smoke run: /Users/yuhan/.local/maka-eval/runs/issue293-prune-ab/issue293-active-smoke-001
  • Tasks: count-dataset-tokens, extract-elf
  • A: context budget off
  • B: active prune + archive retrieval
  • B activation: 2/2 attempts, 2 tasks
  • activePrunedToolResults=1091
  • activeEstimatedTokensSaved=1400611
  • activeArchiveFailures=0
  • Infra/plumbing failures: 0
  • Clash/Mihomo observed traffic delta: 535.92 MiB

User-facing impact

None. This is headless benchmark reporting only.

Reviewer notes

The smoke run showed active pruning activates and reports correctly, but both B attempts ended with incomplete_tool_calls after hitting the single-turn 50-step cap. That is a harness limitation for longer tasks, not a non-inferiority result. Formal #293 evidence should wait for either shorter stable tasks or a benchmark-safe multi-turn continuation harness.

@Astro-HanAstro-Han changed the title feat(headless): report active prune diagnosticsfeat(headless): report active prune evidenceJun 27, 2026
@Astro-Han
Astro-Han merged commit 6ec15cf into mainJun 27, 2026
@Astro-Han
Astro-Han deleted the opencode/issue293-prune-ab-run branch June 27, 2026 06:45
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' feat(headless): report active prune evidence by Astro-Han · Pull Request #323 · apache/maka · GitHub
Skip to content

feat(headless): report active prune evidence - #323

Merged
Astro-Han merged 4 commits into
mainfrom
opencode/issue293-prune-ab-run
Jun 27, 2026
Merged

feat(headless): report active prune evidence#323
Astro-Han merged 4 commits into
mainfrom
opencode/issue293-prune-ab-run

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Jun 27, 2026

Copy link
Copy Markdown
Contributor

Summary

Add active tool-result prune diagnostics to Harbor cell output, A/B summaries, and markdown reports, including a B-active paired comparison slice.

Why

This covers part of the #293 phase-1 reporting gap for active tool-result pruning. Runtime already emits active prune diagnostics, but headless A/B reports previously only recognized stale prunedToolResults. That made active-prune runs hard to inspect and could mix untriggered or unpaired attempts into safety conclusions.

Scope

Changed:

  • Summarize and validate activePrunedToolResults, activeEstimatedTokensSaved, and activeArchiveFailures in Harbor cell output.
  • Include activeToolResultPrune in Harbor context budget policy snapshots.
  • Add active-prune reporting keyed by activePrunedToolResults > 0.
  • Report the active-prune subset as a B-active paired slice: the B attempts that activated active prune, plus the matching A (taskId, rep) attempts.
  • Report active subset task count, attempt count, observed/missing/coverage, pass rate, full token/cost, failure diagnostics, and context/archive diagnostics.
  • Keep activated attempt refs active-only so stale-only pruning does not appear as active-prune evidence.
  • Render existing archive diagnostics: archive placeholders, archive write failures, retrieved count/tokens, retrieval skipped, retrieval failures, and active archive failures.
  • Add/adjust headless tests for active diagnostics, active-only subset reporting, stale-only exclusion, paired active subset reporting, archive metric rendering, missing-pair coverage, and full token/cost rendering.

Not included:

  • No stale prior-context prune validation.
  • No multi-turn or autonomous continuation harness.
  • No archive reason-count schema expansion.
  • No field-list-driven context-budget metric refactor; that maintenance cleanup can be a follow-up PR.
  • No formal non-inferiority claim.

Verification

  • RED: npm --workspace @maka/headless test -- --test-name-pattern "renders active prune subset pair coverage" failed because the active subset line omitted attempts/observed/missing/coverage and full token/cost fields.
  • GREEN: npm --workspace @maka/headless test -- --test-name-pattern "active prune subset|context budget activation": 490 pass, 0 fail. The package test script still ran the full @maka/headless suite.
  • Re-rendered active-prune smoke report from existing results.jsonl: /Users/yuhan/.local/maka-eval/runs/issue293-prune-ab/issue293-active-smoke-001/runtime-policy-ab-report.md
  • Active-prune smoke run: /Users/yuhan/.local/maka-eval/runs/issue293-prune-ab/issue293-active-smoke-001
  • Tasks: count-dataset-tokens, extract-elf
  • A: context budget off
  • B: active prune + archive retrieval
  • B activation: 2/2 attempts, 2 tasks
  • activePrunedToolResults=1091
  • activeEstimatedTokensSaved=1400611
  • activeArchiveFailures=0
  • Infra/plumbing failures: 0
  • Clash/Mihomo observed traffic delta: 535.92 MiB

User-facing impact

None. This is headless benchmark reporting only.

Reviewer notes

The smoke run showed active pruning activates and reports correctly, but both B attempts ended with incomplete_tool_calls after hitting the single-turn 50-step cap. That is a harness limitation for longer tasks, not a non-inferiority result. Formal #293 evidence should wait for either shorter stable tasks or a benchmark-safe multi-turn continuation harness.

@Astro-HanAstro-Han changed the title feat(headless): report active prune diagnosticsfeat(headless): report active prune evidenceJun 27, 2026
@Astro-Han
Astro-Han merged commit 6ec15cf into mainJun 27, 2026
@Astro-Han
Astro-Han deleted the opencode/issue293-prune-ab-run branch June 27, 2026 06:45
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); feat(headless): report active prune evidence by Astro-Han · Pull Request #323 · apache/maka · GitHub
Skip to content

feat(headless): report active prune evidence - #323

Merged
Astro-Han merged 4 commits into
mainfrom
opencode/issue293-prune-ab-run
Jun 27, 2026
Merged

feat(headless): report active prune evidence#323
Astro-Han merged 4 commits into
mainfrom
opencode/issue293-prune-ab-run

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Jun 27, 2026

Copy link
Copy Markdown
Contributor

Summary

Add active tool-result prune diagnostics to Harbor cell output, A/B summaries, and markdown reports, including a B-active paired comparison slice.

Why

This covers part of the #293 phase-1 reporting gap for active tool-result pruning. Runtime already emits active prune diagnostics, but headless A/B reports previously only recognized stale prunedToolResults. That made active-prune runs hard to inspect and could mix untriggered or unpaired attempts into safety conclusions.

Scope

Changed:

  • Summarize and validate activePrunedToolResults, activeEstimatedTokensSaved, and activeArchiveFailures in Harbor cell output.
  • Include activeToolResultPrune in Harbor context budget policy snapshots.
  • Add active-prune reporting keyed by activePrunedToolResults > 0.
  • Report the active-prune subset as a B-active paired slice: the B attempts that activated active prune, plus the matching A (taskId, rep) attempts.
  • Report active subset task count, attempt count, observed/missing/coverage, pass rate, full token/cost, failure diagnostics, and context/archive diagnostics.
  • Keep activated attempt refs active-only so stale-only pruning does not appear as active-prune evidence.
  • Render existing archive diagnostics: archive placeholders, archive write failures, retrieved count/tokens, retrieval skipped, retrieval failures, and active archive failures.
  • Add/adjust headless tests for active diagnostics, active-only subset reporting, stale-only exclusion, paired active subset reporting, archive metric rendering, missing-pair coverage, and full token/cost rendering.

Not included:

  • No stale prior-context prune validation.
  • No multi-turn or autonomous continuation harness.
  • No archive reason-count schema expansion.
  • No field-list-driven context-budget metric refactor; that maintenance cleanup can be a follow-up PR.
  • No formal non-inferiority claim.

Verification

  • RED: npm --workspace @maka/headless test -- --test-name-pattern "renders active prune subset pair coverage" failed because the active subset line omitted attempts/observed/missing/coverage and full token/cost fields.
  • GREEN: npm --workspace @maka/headless test -- --test-name-pattern "active prune subset|context budget activation": 490 pass, 0 fail. The package test script still ran the full @maka/headless suite.
  • Re-rendered active-prune smoke report from existing results.jsonl: /Users/yuhan/.local/maka-eval/runs/issue293-prune-ab/issue293-active-smoke-001/runtime-policy-ab-report.md
  • Active-prune smoke run: /Users/yuhan/.local/maka-eval/runs/issue293-prune-ab/issue293-active-smoke-001
  • Tasks: count-dataset-tokens, extract-elf
  • A: context budget off
  • B: active prune + archive retrieval
  • B activation: 2/2 attempts, 2 tasks
  • activePrunedToolResults=1091
  • activeEstimatedTokensSaved=1400611
  • activeArchiveFailures=0
  • Infra/plumbing failures: 0
  • Clash/Mihomo observed traffic delta: 535.92 MiB

User-facing impact

None. This is headless benchmark reporting only.

Reviewer notes

The smoke run showed active pruning activates and reports correctly, but both B attempts ended with incomplete_tool_calls after hitting the single-turn 50-step cap. That is a harness limitation for longer tasks, not a non-inferiority result. Formal #293 evidence should wait for either shorter stable tasks or a benchmark-safe multi-turn continuation harness.

@Astro-HanAstro-Han changed the title feat(headless): report active prune diagnosticsfeat(headless): report active prune evidenceJun 27, 2026
@Astro-Han
Astro-Han merged commit 6ec15cf into mainJun 27, 2026
@Astro-Han
Astro-Han deleted the opencode/issue293-prune-ab-run branch June 27, 2026 06:45
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Astro-Han