Skip to content

fix(headless): lock valid structured verifier failures as scored results - #1591

Merged
Astro-Han merged 2 commits into
apache:mainfrom
me2seeks:fix/1585-structured-verifier-failure
Jul 30, 2026
Merged

fix(headless): lock valid structured verifier failures as scored results#1591
Astro-Han merged 2 commits into
apache:mainfrom
me2seeks:fix/1585-structured-verifier-failure

Conversation

@me2seeks

Copy link
Copy Markdown
Contributor

Summary

  • treat valid structured verifier pass and failure outcomes as authoritative even when the agent cell exits unsuccessfully
  • require the structured outcome, aggregate reward, and final verifier attempt to agree before overriding an infrastructure classification
  • project stored structured failures as scored results so resumed runs do not resample them
  • cover Harbor/Pier parity, infrastructure-only verifier outcomes, and WAL resume behavior

Closes#1585.

Verification

  • npm run build:test
  • npm --workspace @maka/headless test — 1,427 passed, 4 skipped because Harbor/Pier Python is unavailable, 0 failed
  • npm run typecheck
  • npm run lint
  • npm run format:check
  • git diff --check

Review focus

#1261 is a mechanical WAL-type extraction that also touches fixed-prompt-controller.ts. This PR keeps the scoring change and its regression coverage separate; I will rebase it onto the extracted type seam after #1261 lands.

@Astro-HanAstro-Han left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving current head 5765174 with a non-blocking P2 note.

Malformed or older verifier WAL data with a missing/non-array attempts field can throw a raw TypeError during projection and block resume. The shared verifier classification boundary should validate the runtime shape and map malformed data to the existing ungraded/infra path. The required test check is still pending.

Treat valid structured pass and failure outcomes as authoritative when the agent cell exits unsuccessfully. Validate reward and final-attempt agreement, and project stored failures as scored so they remain locked on resume.
Refs apache#1585
Treat verifier data from runners and persisted WAL as untrusted runtime input. Missing or non-array attempts now stay on the existing ungraded path instead of throwing during resume.
@me2seeks
me2seeksforce-pushed the fix/1585-structured-verifier-failure branch from 5765174 to 3a8026fCompareJuly 30, 2026 01:47
@me2seeks

Copy link
Copy Markdown
ContributorAuthor

Addressed the P2 in 3a8026f. The shared verifier classifier now treats runner and WAL payloads as untrusted runtime input, validates the harbor/verifier/attempt shapes before reading them, and returns the existing ungraded result for malformed data. The regression test covers both a missing attempts field and a non-array value; each resumes without resampling and remains scored=false / eligible=false.\n\nThe branch is rebased onto current main. Headless typecheck, lint, format, the focused 76-test controller suite, and the full Headless suite (1435 tests, 0 failures, 4 environment skips) pass locally.

@Astro-HanAstro-Han left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: taskCompletedEvent still treats any present harbor.verifier as scoring authority, even when structuredVerifierGrade() rejects it. On this head, { outcome: 'failed' } and non-array attempts both turn a max_tokens failure into task_completed with scored and eligible set to true. Sparse arrays also pass because Array#some skips holes. Please make the validated grade the only scoring input and cover malformed and sparse runner output.

@Astro-Han
Astro-Han merged commit 73b629f into apache:mainJul 30, 2026
3 checks passed
@me2seeks

Copy link
Copy Markdown
ContributorAuthor

Followed up in #1644. I removed the verifier-presence fallback, so runner data only affects scoring after structuredVerifierGrade() accepts the complete outcome/reward/attempt contract. The classifier now checks every array index explicitly, which rejects sparse attempts instead of letting Array#some skip holes.

The runner regressions cover missing and non-array attempts, a sparse array, and an internally inconsistent failed grade with a positive aggregate reward. Invalid verifier data remains ungraded, and a provider network failure still stays on the infrastructure path. The branch is rebased onto current main; build, lint, format, typecheck, and the full Headless suite pass locally.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(headless): lock valid structured verifier failures as scored results

2 participants

@me2seeks@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
fix(headless): lock valid structured verifier failures as scored results by me2seeks · Pull Request #1591 · apache/maka · GitHub
Skip to content

fix(headless): lock valid structured verifier failures as scored results - #1591

Merged
Astro-Han merged 2 commits into
apache:mainfrom
me2seeks:fix/1585-structured-verifier-failure
Jul 30, 2026
Merged

fix(headless): lock valid structured verifier failures as scored results#1591
Astro-Han merged 2 commits into
apache:mainfrom
me2seeks:fix/1585-structured-verifier-failure

Conversation

@me2seeks

Copy link
Copy Markdown
Contributor

Summary

  • treat valid structured verifier pass and failure outcomes as authoritative even when the agent cell exits unsuccessfully
  • require the structured outcome, aggregate reward, and final verifier attempt to agree before overriding an infrastructure classification
  • project stored structured failures as scored results so resumed runs do not resample them
  • cover Harbor/Pier parity, infrastructure-only verifier outcomes, and WAL resume behavior

Closes#1585.

Verification

  • npm run build:test
  • npm --workspace @maka/headless test — 1,427 passed, 4 skipped because Harbor/Pier Python is unavailable, 0 failed
  • npm run typecheck
  • npm run lint
  • npm run format:check
  • git diff --check

Review focus

#1261 is a mechanical WAL-type extraction that also touches fixed-prompt-controller.ts. This PR keeps the scoring change and its regression coverage separate; I will rebase it onto the extracted type seam after #1261 lands.

@Astro-HanAstro-Han left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving current head 5765174 with a non-blocking P2 note.

Malformed or older verifier WAL data with a missing/non-array attempts field can throw a raw TypeError during projection and block resume. The shared verifier classification boundary should validate the runtime shape and map malformed data to the existing ungraded/infra path. The required test check is still pending.

Treat valid structured pass and failure outcomes as authoritative when the agent cell exits unsuccessfully. Validate reward and final-attempt agreement, and project stored failures as scored so they remain locked on resume.
Refs apache#1585
Treat verifier data from runners and persisted WAL as untrusted runtime input. Missing or non-array attempts now stay on the existing ungraded path instead of throwing during resume.
@me2seeks
me2seeksforce-pushed the fix/1585-structured-verifier-failure branch from 5765174 to 3a8026fCompareJuly 30, 2026 01:47
@me2seeks

Copy link
Copy Markdown
ContributorAuthor

Addressed the P2 in 3a8026f. The shared verifier classifier now treats runner and WAL payloads as untrusted runtime input, validates the harbor/verifier/attempt shapes before reading them, and returns the existing ungraded result for malformed data. The regression test covers both a missing attempts field and a non-array value; each resumes without resampling and remains scored=false / eligible=false.\n\nThe branch is rebased onto current main. Headless typecheck, lint, format, the focused 76-test controller suite, and the full Headless suite (1435 tests, 0 failures, 4 environment skips) pass locally.

@Astro-HanAstro-Han left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: taskCompletedEvent still treats any present harbor.verifier as scoring authority, even when structuredVerifierGrade() rejects it. On this head, { outcome: 'failed' } and non-array attempts both turn a max_tokens failure into task_completed with scored and eligible set to true. Sparse arrays also pass because Array#some skips holes. Please make the validated grade the only scoring input and cover malformed and sparse runner output.

@Astro-Han
Astro-Han merged commit 73b629f into apache:mainJul 30, 2026
3 checks passed
@me2seeks

Copy link
Copy Markdown
ContributorAuthor

Followed up in #1644. I removed the verifier-presence fallback, so runner data only affects scoring after structuredVerifierGrade() accepts the complete outcome/reward/attempt contract. The classifier now checks every array index explicitly, which rejects sparse attempts instead of letting Array#some skip holes.

The runner regressions cover missing and non-array attempts, a sparse array, and an internally inconsistent failed grade with a positive aggregate reward. Invalid verifier data remains ungraded, and a provider network failure still stays on the infrastructure path. The branch is rebased onto current main; build, lint, format, typecheck, and the full Headless suite pass locally.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(headless): lock valid structured verifier failures as scored results

2 participants

@me2seeks@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix(headless): lock valid structured verifier failures as scored results by me2seeks · Pull Request #1591 · apache/maka · GitHub
Skip to content

fix(headless): lock valid structured verifier failures as scored results - #1591

Merged
Astro-Han merged 2 commits into
apache:mainfrom
me2seeks:fix/1585-structured-verifier-failure
Jul 30, 2026
Merged

fix(headless): lock valid structured verifier failures as scored results#1591
Astro-Han merged 2 commits into
apache:mainfrom
me2seeks:fix/1585-structured-verifier-failure

Conversation

@me2seeks

Copy link
Copy Markdown
Contributor

Summary

  • treat valid structured verifier pass and failure outcomes as authoritative even when the agent cell exits unsuccessfully
  • require the structured outcome, aggregate reward, and final verifier attempt to agree before overriding an infrastructure classification
  • project stored structured failures as scored results so resumed runs do not resample them
  • cover Harbor/Pier parity, infrastructure-only verifier outcomes, and WAL resume behavior

Closes#1585.

Verification

  • npm run build:test
  • npm --workspace @maka/headless test — 1,427 passed, 4 skipped because Harbor/Pier Python is unavailable, 0 failed
  • npm run typecheck
  • npm run lint
  • npm run format:check
  • git diff --check

Review focus

#1261 is a mechanical WAL-type extraction that also touches fixed-prompt-controller.ts. This PR keeps the scoring change and its regression coverage separate; I will rebase it onto the extracted type seam after #1261 lands.

@Astro-HanAstro-Han left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving current head 5765174 with a non-blocking P2 note.

Malformed or older verifier WAL data with a missing/non-array attempts field can throw a raw TypeError during projection and block resume. The shared verifier classification boundary should validate the runtime shape and map malformed data to the existing ungraded/infra path. The required test check is still pending.

Treat valid structured pass and failure outcomes as authoritative when the agent cell exits unsuccessfully. Validate reward and final-attempt agreement, and project stored failures as scored so they remain locked on resume.
Refs apache#1585
Treat verifier data from runners and persisted WAL as untrusted runtime input. Missing or non-array attempts now stay on the existing ungraded path instead of throwing during resume.
@me2seeks
me2seeksforce-pushed the fix/1585-structured-verifier-failure branch from 5765174 to 3a8026fCompareJuly 30, 2026 01:47
@me2seeks

Copy link
Copy Markdown
ContributorAuthor

Addressed the P2 in 3a8026f. The shared verifier classifier now treats runner and WAL payloads as untrusted runtime input, validates the harbor/verifier/attempt shapes before reading them, and returns the existing ungraded result for malformed data. The regression test covers both a missing attempts field and a non-array value; each resumes without resampling and remains scored=false / eligible=false.\n\nThe branch is rebased onto current main. Headless typecheck, lint, format, the focused 76-test controller suite, and the full Headless suite (1435 tests, 0 failures, 4 environment skips) pass locally.

@Astro-HanAstro-Han left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: taskCompletedEvent still treats any present harbor.verifier as scoring authority, even when structuredVerifierGrade() rejects it. On this head, { outcome: 'failed' } and non-array attempts both turn a max_tokens failure into task_completed with scored and eligible set to true. Sparse arrays also pass because Array#some skips holes. Please make the validated grade the only scoring input and cover malformed and sparse runner output.

@Astro-Han
Astro-Han merged commit 73b629f into apache:mainJul 30, 2026
3 checks passed
@me2seeks

Copy link
Copy Markdown
ContributorAuthor

Followed up in #1644. I removed the verifier-presence fallback, so runner data only affects scoring after structuredVerifierGrade() accepts the complete outcome/reward/attempt contract. The classifier now checks every array index explicitly, which rejects sparse attempts instead of letting Array#some skip holes.

The runner regressions cover missing and non-array attempts, a sparse array, and an internally inconsistent failed grade with a positive aggregate reward. Invalid verifier data remains ungraded, and a provider network failure still stays on the infrastructure path. The branch is rebased onto current main; build, lint, format, typecheck, and the full Headless suite pass locally.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(headless): lock valid structured verifier failures as scored results

2 participants

@me2seeks@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix(headless): lock valid structured verifier failures as scored results by me2seeks · Pull Request #1591 · apache/maka · GitHub
Skip to content

fix(headless): lock valid structured verifier failures as scored results - #1591

Merged
Astro-Han merged 2 commits into
apache:mainfrom
me2seeks:fix/1585-structured-verifier-failure
Jul 30, 2026
Merged

fix(headless): lock valid structured verifier failures as scored results#1591
Astro-Han merged 2 commits into
apache:mainfrom
me2seeks:fix/1585-structured-verifier-failure

Conversation

@me2seeks

Copy link
Copy Markdown
Contributor

Summary

  • treat valid structured verifier pass and failure outcomes as authoritative even when the agent cell exits unsuccessfully
  • require the structured outcome, aggregate reward, and final verifier attempt to agree before overriding an infrastructure classification
  • project stored structured failures as scored results so resumed runs do not resample them
  • cover Harbor/Pier parity, infrastructure-only verifier outcomes, and WAL resume behavior

Closes#1585.

Verification

  • npm run build:test
  • npm --workspace @maka/headless test — 1,427 passed, 4 skipped because Harbor/Pier Python is unavailable, 0 failed
  • npm run typecheck
  • npm run lint
  • npm run format:check
  • git diff --check

Review focus

#1261 is a mechanical WAL-type extraction that also touches fixed-prompt-controller.ts. This PR keeps the scoring change and its regression coverage separate; I will rebase it onto the extracted type seam after #1261 lands.

@Astro-HanAstro-Han left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving current head 5765174 with a non-blocking P2 note.

Malformed or older verifier WAL data with a missing/non-array attempts field can throw a raw TypeError during projection and block resume. The shared verifier classification boundary should validate the runtime shape and map malformed data to the existing ungraded/infra path. The required test check is still pending.

Treat valid structured pass and failure outcomes as authoritative when the agent cell exits unsuccessfully. Validate reward and final-attempt agreement, and project stored failures as scored so they remain locked on resume.
Refs apache#1585
Treat verifier data from runners and persisted WAL as untrusted runtime input. Missing or non-array attempts now stay on the existing ungraded path instead of throwing during resume.
@me2seeks
me2seeksforce-pushed the fix/1585-structured-verifier-failure branch from 5765174 to 3a8026fCompareJuly 30, 2026 01:47
@me2seeks

Copy link
Copy Markdown
ContributorAuthor

Addressed the P2 in 3a8026f. The shared verifier classifier now treats runner and WAL payloads as untrusted runtime input, validates the harbor/verifier/attempt shapes before reading them, and returns the existing ungraded result for malformed data. The regression test covers both a missing attempts field and a non-array value; each resumes without resampling and remains scored=false / eligible=false.\n\nThe branch is rebased onto current main. Headless typecheck, lint, format, the focused 76-test controller suite, and the full Headless suite (1435 tests, 0 failures, 4 environment skips) pass locally.

@Astro-HanAstro-Han left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: taskCompletedEvent still treats any present harbor.verifier as scoring authority, even when structuredVerifierGrade() rejects it. On this head, { outcome: 'failed' } and non-array attempts both turn a max_tokens failure into task_completed with scored and eligible set to true. Sparse arrays also pass because Array#some skips holes. Please make the validated grade the only scoring input and cover malformed and sparse runner output.

@Astro-Han
Astro-Han merged commit 73b629f into apache:mainJul 30, 2026
3 checks passed
@me2seeks

Copy link
Copy Markdown
ContributorAuthor

Followed up in #1644. I removed the verifier-presence fallback, so runner data only affects scoring after structuredVerifierGrade() accepts the complete outcome/reward/attempt contract. The classifier now checks every array index explicitly, which rejects sparse attempts instead of letting Array#some skip holes.

The runner regressions cover missing and non-array attempts, a sparse array, and an internally inconsistent failed grade with a positive aggregate reward. Invalid verifier data remains ungraded, and a provider network failure still stays on the infrastructure path. The branch is rebased onto current main; build, lint, format, typecheck, and the full Headless suite pass locally.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(headless): lock valid structured verifier failures as scored results

2 participants

@me2seeks@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' fix(headless): lock valid structured verifier failures as scored results by me2seeks · Pull Request #1591 · apache/maka · GitHub
Skip to content

fix(headless): lock valid structured verifier failures as scored results - #1591

Merged
Astro-Han merged 2 commits into
apache:mainfrom
me2seeks:fix/1585-structured-verifier-failure
Jul 30, 2026
Merged

fix(headless): lock valid structured verifier failures as scored results#1591
Astro-Han merged 2 commits into
apache:mainfrom
me2seeks:fix/1585-structured-verifier-failure

Conversation

@me2seeks

Copy link
Copy Markdown
Contributor

Summary

  • treat valid structured verifier pass and failure outcomes as authoritative even when the agent cell exits unsuccessfully
  • require the structured outcome, aggregate reward, and final verifier attempt to agree before overriding an infrastructure classification
  • project stored structured failures as scored results so resumed runs do not resample them
  • cover Harbor/Pier parity, infrastructure-only verifier outcomes, and WAL resume behavior

Closes#1585.

Verification

  • npm run build:test
  • npm --workspace @maka/headless test — 1,427 passed, 4 skipped because Harbor/Pier Python is unavailable, 0 failed
  • npm run typecheck
  • npm run lint
  • npm run format:check
  • git diff --check

Review focus

#1261 is a mechanical WAL-type extraction that also touches fixed-prompt-controller.ts. This PR keeps the scoring change and its regression coverage separate; I will rebase it onto the extracted type seam after #1261 lands.

@Astro-HanAstro-Han left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving current head 5765174 with a non-blocking P2 note.

Malformed or older verifier WAL data with a missing/non-array attempts field can throw a raw TypeError during projection and block resume. The shared verifier classification boundary should validate the runtime shape and map malformed data to the existing ungraded/infra path. The required test check is still pending.

Treat valid structured pass and failure outcomes as authoritative when the agent cell exits unsuccessfully. Validate reward and final-attempt agreement, and project stored failures as scored so they remain locked on resume.
Refs apache#1585
Treat verifier data from runners and persisted WAL as untrusted runtime input. Missing or non-array attempts now stay on the existing ungraded path instead of throwing during resume.
@me2seeks
me2seeksforce-pushed the fix/1585-structured-verifier-failure branch from 5765174 to 3a8026fCompareJuly 30, 2026 01:47
@me2seeks

Copy link
Copy Markdown
ContributorAuthor

Addressed the P2 in 3a8026f. The shared verifier classifier now treats runner and WAL payloads as untrusted runtime input, validates the harbor/verifier/attempt shapes before reading them, and returns the existing ungraded result for malformed data. The regression test covers both a missing attempts field and a non-array value; each resumes without resampling and remains scored=false / eligible=false.\n\nThe branch is rebased onto current main. Headless typecheck, lint, format, the focused 76-test controller suite, and the full Headless suite (1435 tests, 0 failures, 4 environment skips) pass locally.

@Astro-HanAstro-Han left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: taskCompletedEvent still treats any present harbor.verifier as scoring authority, even when structuredVerifierGrade() rejects it. On this head, { outcome: 'failed' } and non-array attempts both turn a max_tokens failure into task_completed with scored and eligible set to true. Sparse arrays also pass because Array#some skips holes. Please make the validated grade the only scoring input and cover malformed and sparse runner output.

@Astro-Han
Astro-Han merged commit 73b629f into apache:mainJul 30, 2026
3 checks passed
@me2seeks

Copy link
Copy Markdown
ContributorAuthor

Followed up in #1644. I removed the verifier-presence fallback, so runner data only affects scoring after structuredVerifierGrade() accepts the complete outcome/reward/attempt contract. The classifier now checks every array index explicitly, which rejects sparse attempts instead of letting Array#some skip holes.

The runner regressions cover missing and non-array attempts, a sparse array, and an internally inconsistent failed grade with a positive aggregate reward. Invalid verifier data remains ungraded, and a provider network failure still stays on the infrastructure path. The branch is rebased onto current main; build, lint, format, typecheck, and the full Headless suite pass locally.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(headless): lock valid structured verifier failures as scored results

2 participants

@me2seeks@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix(headless): lock valid structured verifier failures as scored results by me2seeks · Pull Request #1591 · apache/maka · GitHub
Skip to content

fix(headless): lock valid structured verifier failures as scored results - #1591

Merged
Astro-Han merged 2 commits into
apache:mainfrom
me2seeks:fix/1585-structured-verifier-failure
Jul 30, 2026
Merged

fix(headless): lock valid structured verifier failures as scored results#1591
Astro-Han merged 2 commits into
apache:mainfrom
me2seeks:fix/1585-structured-verifier-failure

Conversation

@me2seeks

Copy link
Copy Markdown
Contributor

Summary

  • treat valid structured verifier pass and failure outcomes as authoritative even when the agent cell exits unsuccessfully
  • require the structured outcome, aggregate reward, and final verifier attempt to agree before overriding an infrastructure classification
  • project stored structured failures as scored results so resumed runs do not resample them
  • cover Harbor/Pier parity, infrastructure-only verifier outcomes, and WAL resume behavior

Closes#1585.

Verification

  • npm run build:test
  • npm --workspace @maka/headless test — 1,427 passed, 4 skipped because Harbor/Pier Python is unavailable, 0 failed
  • npm run typecheck
  • npm run lint
  • npm run format:check
  • git diff --check

Review focus

#1261 is a mechanical WAL-type extraction that also touches fixed-prompt-controller.ts. This PR keeps the scoring change and its regression coverage separate; I will rebase it onto the extracted type seam after #1261 lands.

@Astro-HanAstro-Han left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving current head 5765174 with a non-blocking P2 note.

Malformed or older verifier WAL data with a missing/non-array attempts field can throw a raw TypeError during projection and block resume. The shared verifier classification boundary should validate the runtime shape and map malformed data to the existing ungraded/infra path. The required test check is still pending.

Treat valid structured pass and failure outcomes as authoritative when the agent cell exits unsuccessfully. Validate reward and final-attempt agreement, and project stored failures as scored so they remain locked on resume.
Refs apache#1585
Treat verifier data from runners and persisted WAL as untrusted runtime input. Missing or non-array attempts now stay on the existing ungraded path instead of throwing during resume.
@me2seeks
me2seeksforce-pushed the fix/1585-structured-verifier-failure branch from 5765174 to 3a8026fCompareJuly 30, 2026 01:47
@me2seeks

Copy link
Copy Markdown
ContributorAuthor

Addressed the P2 in 3a8026f. The shared verifier classifier now treats runner and WAL payloads as untrusted runtime input, validates the harbor/verifier/attempt shapes before reading them, and returns the existing ungraded result for malformed data. The regression test covers both a missing attempts field and a non-array value; each resumes without resampling and remains scored=false / eligible=false.\n\nThe branch is rebased onto current main. Headless typecheck, lint, format, the focused 76-test controller suite, and the full Headless suite (1435 tests, 0 failures, 4 environment skips) pass locally.

@Astro-HanAstro-Han left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: taskCompletedEvent still treats any present harbor.verifier as scoring authority, even when structuredVerifierGrade() rejects it. On this head, { outcome: 'failed' } and non-array attempts both turn a max_tokens failure into task_completed with scored and eligible set to true. Sparse arrays also pass because Array#some skips holes. Please make the validated grade the only scoring input and cover malformed and sparse runner output.

@Astro-Han
Astro-Han merged commit 73b629f into apache:mainJul 30, 2026
3 checks passed
@me2seeks

Copy link
Copy Markdown
ContributorAuthor

Followed up in #1644. I removed the verifier-presence fallback, so runner data only affects scoring after structuredVerifierGrade() accepts the complete outcome/reward/attempt contract. The classifier now checks every array index explicitly, which rejects sparse attempts instead of letting Array#some skip holes.

The runner regressions cover missing and non-array attempts, a sparse array, and an internally inconsistent failed grade with a positive aggregate reward. Invalid verifier data remains ungraded, and a provider network failure still stays on the infrastructure path. The branch is rebased onto current main; build, lint, format, typecheck, and the full Headless suite pass locally.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(headless): lock valid structured verifier failures as scored results

2 participants

@me2seeks@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix(headless): lock valid structured verifier failures as scored results by me2seeks · Pull Request #1591 · apache/maka · GitHub
Skip to content

fix(headless): lock valid structured verifier failures as scored results - #1591

Merged
Astro-Han merged 2 commits into
apache:mainfrom
me2seeks:fix/1585-structured-verifier-failure
Jul 30, 2026
Merged

fix(headless): lock valid structured verifier failures as scored results#1591
Astro-Han merged 2 commits into
apache:mainfrom
me2seeks:fix/1585-structured-verifier-failure

Conversation

@me2seeks

Copy link
Copy Markdown
Contributor

Summary

  • treat valid structured verifier pass and failure outcomes as authoritative even when the agent cell exits unsuccessfully
  • require the structured outcome, aggregate reward, and final verifier attempt to agree before overriding an infrastructure classification
  • project stored structured failures as scored results so resumed runs do not resample them
  • cover Harbor/Pier parity, infrastructure-only verifier outcomes, and WAL resume behavior

Closes#1585.

Verification

  • npm run build:test
  • npm --workspace @maka/headless test — 1,427 passed, 4 skipped because Harbor/Pier Python is unavailable, 0 failed
  • npm run typecheck
  • npm run lint
  • npm run format:check
  • git diff --check

Review focus

#1261 is a mechanical WAL-type extraction that also touches fixed-prompt-controller.ts. This PR keeps the scoring change and its regression coverage separate; I will rebase it onto the extracted type seam after #1261 lands.

@Astro-HanAstro-Han left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving current head 5765174 with a non-blocking P2 note.

Malformed or older verifier WAL data with a missing/non-array attempts field can throw a raw TypeError during projection and block resume. The shared verifier classification boundary should validate the runtime shape and map malformed data to the existing ungraded/infra path. The required test check is still pending.

Treat valid structured pass and failure outcomes as authoritative when the agent cell exits unsuccessfully. Validate reward and final-attempt agreement, and project stored failures as scored so they remain locked on resume.
Refs apache#1585
Treat verifier data from runners and persisted WAL as untrusted runtime input. Missing or non-array attempts now stay on the existing ungraded path instead of throwing during resume.
@me2seeks
me2seeksforce-pushed the fix/1585-structured-verifier-failure branch from 5765174 to 3a8026fCompareJuly 30, 2026 01:47
@me2seeks

Copy link
Copy Markdown
ContributorAuthor

Addressed the P2 in 3a8026f. The shared verifier classifier now treats runner and WAL payloads as untrusted runtime input, validates the harbor/verifier/attempt shapes before reading them, and returns the existing ungraded result for malformed data. The regression test covers both a missing attempts field and a non-array value; each resumes without resampling and remains scored=false / eligible=false.\n\nThe branch is rebased onto current main. Headless typecheck, lint, format, the focused 76-test controller suite, and the full Headless suite (1435 tests, 0 failures, 4 environment skips) pass locally.

@Astro-HanAstro-Han left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: taskCompletedEvent still treats any present harbor.verifier as scoring authority, even when structuredVerifierGrade() rejects it. On this head, { outcome: 'failed' } and non-array attempts both turn a max_tokens failure into task_completed with scored and eligible set to true. Sparse arrays also pass because Array#some skips holes. Please make the validated grade the only scoring input and cover malformed and sparse runner output.

@Astro-Han
Astro-Han merged commit 73b629f into apache:mainJul 30, 2026
3 checks passed
@me2seeks

Copy link
Copy Markdown
ContributorAuthor

Followed up in #1644. I removed the verifier-presence fallback, so runner data only affects scoring after structuredVerifierGrade() accepts the complete outcome/reward/attempt contract. The classifier now checks every array index explicitly, which rejects sparse attempts instead of letting Array#some skip holes.

The runner regressions cover missing and non-array attempts, a sparse array, and an internally inconsistent failed grade with a positive aggregate reward. Invalid verifier data remains ungraded, and a provider network failure still stays on the infrastructure path. The branch is rebased onto current main; build, lint, format, typecheck, and the full Headless suite pass locally.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(headless): lock valid structured verifier failures as scored results

2 participants

@me2seeks@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); fix(headless): lock valid structured verifier failures as scored results by me2seeks · Pull Request #1591 · apache/maka · GitHub
Skip to content

fix(headless): lock valid structured verifier failures as scored results - #1591

Merged
Astro-Han merged 2 commits into
apache:mainfrom
me2seeks:fix/1585-structured-verifier-failure
Jul 30, 2026
Merged

fix(headless): lock valid structured verifier failures as scored results#1591
Astro-Han merged 2 commits into
apache:mainfrom
me2seeks:fix/1585-structured-verifier-failure

Conversation

@me2seeks

Copy link
Copy Markdown
Contributor

Summary

  • treat valid structured verifier pass and failure outcomes as authoritative even when the agent cell exits unsuccessfully
  • require the structured outcome, aggregate reward, and final verifier attempt to agree before overriding an infrastructure classification
  • project stored structured failures as scored results so resumed runs do not resample them
  • cover Harbor/Pier parity, infrastructure-only verifier outcomes, and WAL resume behavior

Closes#1585.

Verification

  • npm run build:test
  • npm --workspace @maka/headless test — 1,427 passed, 4 skipped because Harbor/Pier Python is unavailable, 0 failed
  • npm run typecheck
  • npm run lint
  • npm run format:check
  • git diff --check

Review focus

#1261 is a mechanical WAL-type extraction that also touches fixed-prompt-controller.ts. This PR keeps the scoring change and its regression coverage separate; I will rebase it onto the extracted type seam after #1261 lands.

@Astro-HanAstro-Han left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving current head 5765174 with a non-blocking P2 note.

Malformed or older verifier WAL data with a missing/non-array attempts field can throw a raw TypeError during projection and block resume. The shared verifier classification boundary should validate the runtime shape and map malformed data to the existing ungraded/infra path. The required test check is still pending.

Treat valid structured pass and failure outcomes as authoritative when the agent cell exits unsuccessfully. Validate reward and final-attempt agreement, and project stored failures as scored so they remain locked on resume.
Refs apache#1585
Treat verifier data from runners and persisted WAL as untrusted runtime input. Missing or non-array attempts now stay on the existing ungraded path instead of throwing during resume.
@me2seeks
me2seeksforce-pushed the fix/1585-structured-verifier-failure branch from 5765174 to 3a8026fCompareJuly 30, 2026 01:47
@me2seeks

Copy link
Copy Markdown
ContributorAuthor

Addressed the P2 in 3a8026f. The shared verifier classifier now treats runner and WAL payloads as untrusted runtime input, validates the harbor/verifier/attempt shapes before reading them, and returns the existing ungraded result for malformed data. The regression test covers both a missing attempts field and a non-array value; each resumes without resampling and remains scored=false / eligible=false.\n\nThe branch is rebased onto current main. Headless typecheck, lint, format, the focused 76-test controller suite, and the full Headless suite (1435 tests, 0 failures, 4 environment skips) pass locally.

@Astro-HanAstro-Han left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: taskCompletedEvent still treats any present harbor.verifier as scoring authority, even when structuredVerifierGrade() rejects it. On this head, { outcome: 'failed' } and non-array attempts both turn a max_tokens failure into task_completed with scored and eligible set to true. Sparse arrays also pass because Array#some skips holes. Please make the validated grade the only scoring input and cover malformed and sparse runner output.

@Astro-Han
Astro-Han merged commit 73b629f into apache:mainJul 30, 2026
3 checks passed
@me2seeks

Copy link
Copy Markdown
ContributorAuthor

Followed up in #1644. I removed the verifier-presence fallback, so runner data only affects scoring after structuredVerifierGrade() accepts the complete outcome/reward/attempt contract. The classifier now checks every array index explicitly, which rejects sparse attempts instead of letting Array#some skip holes.

The runner regressions cover missing and non-array attempts, a sparse array, and an internally inconsistent failed grade with a positive aggregate reward. Invalid verifier data remains ungraded, and a provider network failure still stays on the infrastructure path. The branch is rebased onto current main; build, lint, format, typecheck, and the full Headless suite pass locally.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(headless): lock valid structured verifier failures as scored results

2 participants

@me2seeks@Astro-Han