Skip to content

fix(headless): lock valid structured verifier failures as scored results #1585

Description

@Astro-Han

What happened

The fixed-prompt headless controller treats a structured verifier pass as authoritative, but not a structured verifier failure.

When a Harbor cell has:

  • cell.status === "failed"
  • cell.errorClass === "runtime_error" or "infra_failed"
  • harbor.reward === 0
  • harbor.verifier.outcome === "failed"
  • a valid final verifier attempt with classification === "failed" and reward === 0

taskCompletedEvent() records the cell with scored: false.

The asymmetric control case is already accepted: when the same failed cell has a structured verifier pass with positive reward, structuredVerifierPassed makes it scored and passed.

This leaves a valid structured verifier failure replacement-eligible even though the verifier already produced a terminal pass/fail decision. It can cause unnecessary paid replacement attempts and can bias Pass@1 by retrying a result that should have been locked.

Two observed examples:

  • An OpenCode compile-compcert cell ended with errorClass: "infra_failed" and a structured verifier failure, but was recorded unscored.
  • A Maka qemu-startup cell ended with errorClass: "runtime_error" and a structured verifier failure, but was recorded unscored.

This is separate from the runtime trace-serialization defect addressed by #1558.

How to reproduce

  1. Exercise the fixed-prompt controller with a TaskRunOutput whose cell has:
    • status: "failed"
    • errorClass: "runtime_error" (repeat with "infra_failed")
  2. Give the output a valid Harbor result:
    • reward: 0
    • verifier.outcome: "failed"
    • final attempt { classification: "failed", reward: 0 }
  3. Observe that the emitted task_completed event has scored: false.
  4. Change the verifier outcome to a valid pass with positive reward.
  5. Observe that the failed cell is now accepted as scored: true.

The relevant logic is taskCompletedEvent() in packages/headless/src/fixed-prompt-controller.ts: structuredVerifierPassed participates in verifierGraded, but the corresponding structured verifier failure does not.

Expected behavior:

Once the structured verifier produces a valid passed or failed outcome, that result should be scored and locked regardless of the agent cell's terminal status. Only missing, malformed, or infrastructure-failed verifier outcomes should remain replacement-eligible.

Regression coverage should include:

  • structured pass on a failed agent cell → scored pass;
  • structured failure on a failed agent cell → scored fail;
  • verifier infrastructure failure → unscored;
  • a scored verifier failure is not eligible for replacement.

Environment

  • Maka observed subject commit: ee7e4ba5d
  • Confirmed affected logic remains at current commit: c561fed4e1
  • Surface: Headless / Harbor Terminal-Bench harness
  • macOS: 26.5.2
  • Node.js: v26.5.0

Logs, screenshots, or additional context

No provider HTTP 429 responses were involved in the observed examples. The issue concerns verifier authority and result classification, not provider availability.

Metadata

Metadata

Assignees

Labels

bugSomething isn't working

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions

    , 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
     blocks
    (function() {
    function addCopyButtons() {
    document.querySelectorAll('pre code').forEach(function(codeBlock) {
    if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
    codeBlock.parentElement.setAttribute('data-copy-added', 'true');
    var btn = document.createElement('button');
    btn.textContent = 'Copy';
    btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
    btn.onmouseover = function() { this.style.opacity = '1'; };
    btn.onmouseout = function() { this.style.opacity = '0.7'; };
    btn.onclick = function() {
    navigator.clipboard.writeText(codeBlock.textContent).then(function() {
    btn.textContent = 'Copied!';
    setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
    });
    };
    codeBlock.parentElement.style.position = 'relative';
    codeBlock.parentElement.appendChild(btn);
    });
    }
    addCopyButtons();
    // Re-run on dynamic content
    var observer = new MutationObserver(addCopyButtons);
    observer.observe(document.body, { childList: true, subtree: true });
    })();
    }
    } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
    })();
    (function(){
    try {
    var __m = "github.com";
    var __re = new RegExp('^' + "github\\.com" + '
    fix(headless): lock valid structured verifier failures as scored results · Issue #1585 · apache/maka · GitHub
    Skip to content

    fix(headless): lock valid structured verifier failures as scored results #1585

    Description

    @Astro-Han

    What happened

    The fixed-prompt headless controller treats a structured verifier pass as authoritative, but not a structured verifier failure.

    When a Harbor cell has:

    • cell.status === "failed"
    • cell.errorClass === "runtime_error" or "infra_failed"
    • harbor.reward === 0
    • harbor.verifier.outcome === "failed"
    • a valid final verifier attempt with classification === "failed" and reward === 0

    taskCompletedEvent() records the cell with scored: false.

    The asymmetric control case is already accepted: when the same failed cell has a structured verifier pass with positive reward, structuredVerifierPassed makes it scored and passed.

    This leaves a valid structured verifier failure replacement-eligible even though the verifier already produced a terminal pass/fail decision. It can cause unnecessary paid replacement attempts and can bias Pass@1 by retrying a result that should have been locked.

    Two observed examples:

    • An OpenCode compile-compcert cell ended with errorClass: "infra_failed" and a structured verifier failure, but was recorded unscored.
    • A Maka qemu-startup cell ended with errorClass: "runtime_error" and a structured verifier failure, but was recorded unscored.

    This is separate from the runtime trace-serialization defect addressed by #1558.

    How to reproduce

    1. Exercise the fixed-prompt controller with a TaskRunOutput whose cell has:
      • status: "failed"
      • errorClass: "runtime_error" (repeat with "infra_failed")
    2. Give the output a valid Harbor result:
      • reward: 0
      • verifier.outcome: "failed"
      • final attempt { classification: "failed", reward: 0 }
    3. Observe that the emitted task_completed event has scored: false.
    4. Change the verifier outcome to a valid pass with positive reward.
    5. Observe that the failed cell is now accepted as scored: true.

    The relevant logic is taskCompletedEvent() in packages/headless/src/fixed-prompt-controller.ts: structuredVerifierPassed participates in verifierGraded, but the corresponding structured verifier failure does not.

    Expected behavior:

    Once the structured verifier produces a valid passed or failed outcome, that result should be scored and locked regardless of the agent cell's terminal status. Only missing, malformed, or infrastructure-failed verifier outcomes should remain replacement-eligible.

    Regression coverage should include:

    • structured pass on a failed agent cell → scored pass;
    • structured failure on a failed agent cell → scored fail;
    • verifier infrastructure failure → unscored;
    • a scored verifier failure is not eligible for replacement.

    Environment

    • Maka observed subject commit: ee7e4ba5d
    • Confirmed affected logic remains at current commit: c561fed4e1
    • Surface: Headless / Harbor Terminal-Bench harness
    • macOS: 26.5.2
    • Node.js: v26.5.0

    Logs, screenshots, or additional context

    No provider HTTP 429 responses were involved in the observed examples. The issue concerns verifier authority and result classification, not provider availability.

    Metadata

    Metadata

    Assignees

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions

      , 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix(headless): lock valid structured verifier failures as scored results · Issue #1585 · apache/maka · GitHub
      Skip to content

      fix(headless): lock valid structured verifier failures as scored results #1585

      Description

      @Astro-Han

      What happened

      The fixed-prompt headless controller treats a structured verifier pass as authoritative, but not a structured verifier failure.

      When a Harbor cell has:

      • cell.status === "failed"
      • cell.errorClass === "runtime_error" or "infra_failed"
      • harbor.reward === 0
      • harbor.verifier.outcome === "failed"
      • a valid final verifier attempt with classification === "failed" and reward === 0

      taskCompletedEvent() records the cell with scored: false.

      The asymmetric control case is already accepted: when the same failed cell has a structured verifier pass with positive reward, structuredVerifierPassed makes it scored and passed.

      This leaves a valid structured verifier failure replacement-eligible even though the verifier already produced a terminal pass/fail decision. It can cause unnecessary paid replacement attempts and can bias Pass@1 by retrying a result that should have been locked.

      Two observed examples:

      • An OpenCode compile-compcert cell ended with errorClass: "infra_failed" and a structured verifier failure, but was recorded unscored.
      • A Maka qemu-startup cell ended with errorClass: "runtime_error" and a structured verifier failure, but was recorded unscored.

      This is separate from the runtime trace-serialization defect addressed by #1558.

      How to reproduce

      1. Exercise the fixed-prompt controller with a TaskRunOutput whose cell has:
        • status: "failed"
        • errorClass: "runtime_error" (repeat with "infra_failed")
      2. Give the output a valid Harbor result:
        • reward: 0
        • verifier.outcome: "failed"
        • final attempt { classification: "failed", reward: 0 }
      3. Observe that the emitted task_completed event has scored: false.
      4. Change the verifier outcome to a valid pass with positive reward.
      5. Observe that the failed cell is now accepted as scored: true.

      The relevant logic is taskCompletedEvent() in packages/headless/src/fixed-prompt-controller.ts: structuredVerifierPassed participates in verifierGraded, but the corresponding structured verifier failure does not.

      Expected behavior:

      Once the structured verifier produces a valid passed or failed outcome, that result should be scored and locked regardless of the agent cell's terminal status. Only missing, malformed, or infrastructure-failed verifier outcomes should remain replacement-eligible.

      Regression coverage should include:

      • structured pass on a failed agent cell → scored pass;
      • structured failure on a failed agent cell → scored fail;
      • verifier infrastructure failure → unscored;
      • a scored verifier failure is not eligible for replacement.

      Environment

      • Maka observed subject commit: ee7e4ba5d
      • Confirmed affected logic remains at current commit: c561fed4e1
      • Surface: Headless / Harbor Terminal-Bench harness
      • macOS: 26.5.2
      • Node.js: v26.5.0

      Logs, screenshots, or additional context

      No provider HTTP 429 responses were involved in the observed examples. The issue concerns verifier authority and result classification, not provider availability.

      Metadata

      Metadata

      Assignees

      Labels

      bugSomething isn't working

      Type

      No type

      Projects

      No projects

        Milestone

        No milestone

        Relationships

        None yet

        Development

        No branches or pull requests

        Issue actions

        , 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix(headless): lock valid structured verifier failures as scored results · Issue #1585 · apache/maka · GitHub
        Skip to content

        fix(headless): lock valid structured verifier failures as scored results #1585

        Description

        @Astro-Han

        What happened

        The fixed-prompt headless controller treats a structured verifier pass as authoritative, but not a structured verifier failure.

        When a Harbor cell has:

        • cell.status === "failed"
        • cell.errorClass === "runtime_error" or "infra_failed"
        • harbor.reward === 0
        • harbor.verifier.outcome === "failed"
        • a valid final verifier attempt with classification === "failed" and reward === 0

        taskCompletedEvent() records the cell with scored: false.

        The asymmetric control case is already accepted: when the same failed cell has a structured verifier pass with positive reward, structuredVerifierPassed makes it scored and passed.

        This leaves a valid structured verifier failure replacement-eligible even though the verifier already produced a terminal pass/fail decision. It can cause unnecessary paid replacement attempts and can bias Pass@1 by retrying a result that should have been locked.

        Two observed examples:

        • An OpenCode compile-compcert cell ended with errorClass: "infra_failed" and a structured verifier failure, but was recorded unscored.
        • A Maka qemu-startup cell ended with errorClass: "runtime_error" and a structured verifier failure, but was recorded unscored.

        This is separate from the runtime trace-serialization defect addressed by #1558.

        How to reproduce

        1. Exercise the fixed-prompt controller with a TaskRunOutput whose cell has:
          • status: "failed"
          • errorClass: "runtime_error" (repeat with "infra_failed")
        2. Give the output a valid Harbor result:
          • reward: 0
          • verifier.outcome: "failed"
          • final attempt { classification: "failed", reward: 0 }
        3. Observe that the emitted task_completed event has scored: false.
        4. Change the verifier outcome to a valid pass with positive reward.
        5. Observe that the failed cell is now accepted as scored: true.

        The relevant logic is taskCompletedEvent() in packages/headless/src/fixed-prompt-controller.ts: structuredVerifierPassed participates in verifierGraded, but the corresponding structured verifier failure does not.

        Expected behavior:

        Once the structured verifier produces a valid passed or failed outcome, that result should be scored and locked regardless of the agent cell's terminal status. Only missing, malformed, or infrastructure-failed verifier outcomes should remain replacement-eligible.

        Regression coverage should include:

        • structured pass on a failed agent cell → scored pass;
        • structured failure on a failed agent cell → scored fail;
        • verifier infrastructure failure → unscored;
        • a scored verifier failure is not eligible for replacement.

        Environment

        • Maka observed subject commit: ee7e4ba5d
        • Confirmed affected logic remains at current commit: c561fed4e1
        • Surface: Headless / Harbor Terminal-Bench harness
        • macOS: 26.5.2
        • Node.js: v26.5.0

        Logs, screenshots, or additional context

        No provider HTTP 429 responses were involved in the observed examples. The issue concerns verifier authority and result classification, not provider availability.

        Metadata

        Metadata

        Assignees

        Labels

        bugSomething isn't working

        Type

        No type

        Projects

        No projects

          Milestone

          No milestone

          Relationships

          None yet

          Development

          No branches or pull requests

          Issue actions

          , 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' fix(headless): lock valid structured verifier failures as scored results · Issue #1585 · apache/maka · GitHub
          Skip to content

          fix(headless): lock valid structured verifier failures as scored results #1585

          Description

          @Astro-Han

          What happened

          The fixed-prompt headless controller treats a structured verifier pass as authoritative, but not a structured verifier failure.

          When a Harbor cell has:

          • cell.status === "failed"
          • cell.errorClass === "runtime_error" or "infra_failed"
          • harbor.reward === 0
          • harbor.verifier.outcome === "failed"
          • a valid final verifier attempt with classification === "failed" and reward === 0

          taskCompletedEvent() records the cell with scored: false.

          The asymmetric control case is already accepted: when the same failed cell has a structured verifier pass with positive reward, structuredVerifierPassed makes it scored and passed.

          This leaves a valid structured verifier failure replacement-eligible even though the verifier already produced a terminal pass/fail decision. It can cause unnecessary paid replacement attempts and can bias Pass@1 by retrying a result that should have been locked.

          Two observed examples:

          • An OpenCode compile-compcert cell ended with errorClass: "infra_failed" and a structured verifier failure, but was recorded unscored.
          • A Maka qemu-startup cell ended with errorClass: "runtime_error" and a structured verifier failure, but was recorded unscored.

          This is separate from the runtime trace-serialization defect addressed by #1558.

          How to reproduce

          1. Exercise the fixed-prompt controller with a TaskRunOutput whose cell has:
            • status: "failed"
            • errorClass: "runtime_error" (repeat with "infra_failed")
          2. Give the output a valid Harbor result:
            • reward: 0
            • verifier.outcome: "failed"
            • final attempt { classification: "failed", reward: 0 }
          3. Observe that the emitted task_completed event has scored: false.
          4. Change the verifier outcome to a valid pass with positive reward.
          5. Observe that the failed cell is now accepted as scored: true.

          The relevant logic is taskCompletedEvent() in packages/headless/src/fixed-prompt-controller.ts: structuredVerifierPassed participates in verifierGraded, but the corresponding structured verifier failure does not.

          Expected behavior:

          Once the structured verifier produces a valid passed or failed outcome, that result should be scored and locked regardless of the agent cell's terminal status. Only missing, malformed, or infrastructure-failed verifier outcomes should remain replacement-eligible.

          Regression coverage should include:

          • structured pass on a failed agent cell → scored pass;
          • structured failure on a failed agent cell → scored fail;
          • verifier infrastructure failure → unscored;
          • a scored verifier failure is not eligible for replacement.

          Environment

          • Maka observed subject commit: ee7e4ba5d
          • Confirmed affected logic remains at current commit: c561fed4e1
          • Surface: Headless / Harbor Terminal-Bench harness
          • macOS: 26.5.2
          • Node.js: v26.5.0

          Logs, screenshots, or additional context

          No provider HTTP 429 responses were involved in the observed examples. The issue concerns verifier authority and result classification, not provider availability.

          Metadata

          Metadata

          Assignees

          Labels

          bugSomething isn't working

          Type

          No type

          Projects

          No projects

            Milestone

            No milestone

            Relationships

            None yet

            Development

            No branches or pull requests

            Issue actions

            , 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix(headless): lock valid structured verifier failures as scored results · Issue #1585 · apache/maka · GitHub
            Skip to content

            fix(headless): lock valid structured verifier failures as scored results #1585

            Description

            @Astro-Han

            What happened

            The fixed-prompt headless controller treats a structured verifier pass as authoritative, but not a structured verifier failure.

            When a Harbor cell has:

            • cell.status === "failed"
            • cell.errorClass === "runtime_error" or "infra_failed"
            • harbor.reward === 0
            • harbor.verifier.outcome === "failed"
            • a valid final verifier attempt with classification === "failed" and reward === 0

            taskCompletedEvent() records the cell with scored: false.

            The asymmetric control case is already accepted: when the same failed cell has a structured verifier pass with positive reward, structuredVerifierPassed makes it scored and passed.

            This leaves a valid structured verifier failure replacement-eligible even though the verifier already produced a terminal pass/fail decision. It can cause unnecessary paid replacement attempts and can bias Pass@1 by retrying a result that should have been locked.

            Two observed examples:

            • An OpenCode compile-compcert cell ended with errorClass: "infra_failed" and a structured verifier failure, but was recorded unscored.
            • A Maka qemu-startup cell ended with errorClass: "runtime_error" and a structured verifier failure, but was recorded unscored.

            This is separate from the runtime trace-serialization defect addressed by #1558.

            How to reproduce

            1. Exercise the fixed-prompt controller with a TaskRunOutput whose cell has:
              • status: "failed"
              • errorClass: "runtime_error" (repeat with "infra_failed")
            2. Give the output a valid Harbor result:
              • reward: 0
              • verifier.outcome: "failed"
              • final attempt { classification: "failed", reward: 0 }
            3. Observe that the emitted task_completed event has scored: false.
            4. Change the verifier outcome to a valid pass with positive reward.
            5. Observe that the failed cell is now accepted as scored: true.

            The relevant logic is taskCompletedEvent() in packages/headless/src/fixed-prompt-controller.ts: structuredVerifierPassed participates in verifierGraded, but the corresponding structured verifier failure does not.

            Expected behavior:

            Once the structured verifier produces a valid passed or failed outcome, that result should be scored and locked regardless of the agent cell's terminal status. Only missing, malformed, or infrastructure-failed verifier outcomes should remain replacement-eligible.

            Regression coverage should include:

            • structured pass on a failed agent cell → scored pass;
            • structured failure on a failed agent cell → scored fail;
            • verifier infrastructure failure → unscored;
            • a scored verifier failure is not eligible for replacement.

            Environment

            • Maka observed subject commit: ee7e4ba5d
            • Confirmed affected logic remains at current commit: c561fed4e1
            • Surface: Headless / Harbor Terminal-Bench harness
            • macOS: 26.5.2
            • Node.js: v26.5.0

            Logs, screenshots, or additional context

            No provider HTTP 429 responses were involved in the observed examples. The issue concerns verifier authority and result classification, not provider availability.

            Metadata

            Metadata

            Assignees

            Labels

            bugSomething isn't working

            Type

            No type

            Projects

            No projects

              Milestone

              No milestone

              Relationships

              None yet

              Development

              No branches or pull requests

              Issue actions

              , 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix(headless): lock valid structured verifier failures as scored results · Issue #1585 · apache/maka · GitHub
              Skip to content

              fix(headless): lock valid structured verifier failures as scored results #1585

              Description

              @Astro-Han

              What happened

              The fixed-prompt headless controller treats a structured verifier pass as authoritative, but not a structured verifier failure.

              When a Harbor cell has:

              • cell.status === "failed"
              • cell.errorClass === "runtime_error" or "infra_failed"
              • harbor.reward === 0
              • harbor.verifier.outcome === "failed"
              • a valid final verifier attempt with classification === "failed" and reward === 0

              taskCompletedEvent() records the cell with scored: false.

              The asymmetric control case is already accepted: when the same failed cell has a structured verifier pass with positive reward, structuredVerifierPassed makes it scored and passed.

              This leaves a valid structured verifier failure replacement-eligible even though the verifier already produced a terminal pass/fail decision. It can cause unnecessary paid replacement attempts and can bias Pass@1 by retrying a result that should have been locked.

              Two observed examples:

              • An OpenCode compile-compcert cell ended with errorClass: "infra_failed" and a structured verifier failure, but was recorded unscored.
              • A Maka qemu-startup cell ended with errorClass: "runtime_error" and a structured verifier failure, but was recorded unscored.

              This is separate from the runtime trace-serialization defect addressed by #1558.

              How to reproduce

              1. Exercise the fixed-prompt controller with a TaskRunOutput whose cell has:
                • status: "failed"
                • errorClass: "runtime_error" (repeat with "infra_failed")
              2. Give the output a valid Harbor result:
                • reward: 0
                • verifier.outcome: "failed"
                • final attempt { classification: "failed", reward: 0 }
              3. Observe that the emitted task_completed event has scored: false.
              4. Change the verifier outcome to a valid pass with positive reward.
              5. Observe that the failed cell is now accepted as scored: true.

              The relevant logic is taskCompletedEvent() in packages/headless/src/fixed-prompt-controller.ts: structuredVerifierPassed participates in verifierGraded, but the corresponding structured verifier failure does not.

              Expected behavior:

              Once the structured verifier produces a valid passed or failed outcome, that result should be scored and locked regardless of the agent cell's terminal status. Only missing, malformed, or infrastructure-failed verifier outcomes should remain replacement-eligible.

              Regression coverage should include:

              • structured pass on a failed agent cell → scored pass;
              • structured failure on a failed agent cell → scored fail;
              • verifier infrastructure failure → unscored;
              • a scored verifier failure is not eligible for replacement.

              Environment

              • Maka observed subject commit: ee7e4ba5d
              • Confirmed affected logic remains at current commit: c561fed4e1
              • Surface: Headless / Harbor Terminal-Bench harness
              • macOS: 26.5.2
              • Node.js: v26.5.0

              Logs, screenshots, or additional context

              No provider HTTP 429 responses were involved in the observed examples. The issue concerns verifier authority and result classification, not provider availability.

              Metadata

              Metadata

              Assignees

              Labels

              bugSomething isn't working

              Type

              No type

              Projects

              No projects

                Milestone

                No milestone

                Relationships

                None yet

                Development

                No branches or pull requests

                Issue actions

                , 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); fix(headless): lock valid structured verifier failures as scored results · Issue #1585 · apache/maka · GitHub
                Skip to content

                fix(headless): lock valid structured verifier failures as scored results #1585

                Description

                @Astro-Han

                What happened

                The fixed-prompt headless controller treats a structured verifier pass as authoritative, but not a structured verifier failure.

                When a Harbor cell has:

                • cell.status === "failed"
                • cell.errorClass === "runtime_error" or "infra_failed"
                • harbor.reward === 0
                • harbor.verifier.outcome === "failed"
                • a valid final verifier attempt with classification === "failed" and reward === 0

                taskCompletedEvent() records the cell with scored: false.

                The asymmetric control case is already accepted: when the same failed cell has a structured verifier pass with positive reward, structuredVerifierPassed makes it scored and passed.

                This leaves a valid structured verifier failure replacement-eligible even though the verifier already produced a terminal pass/fail decision. It can cause unnecessary paid replacement attempts and can bias Pass@1 by retrying a result that should have been locked.

                Two observed examples:

                • An OpenCode compile-compcert cell ended with errorClass: "infra_failed" and a structured verifier failure, but was recorded unscored.
                • A Maka qemu-startup cell ended with errorClass: "runtime_error" and a structured verifier failure, but was recorded unscored.

                This is separate from the runtime trace-serialization defect addressed by #1558.

                How to reproduce

                1. Exercise the fixed-prompt controller with a TaskRunOutput whose cell has:
                  • status: "failed"
                  • errorClass: "runtime_error" (repeat with "infra_failed")
                2. Give the output a valid Harbor result:
                  • reward: 0
                  • verifier.outcome: "failed"
                  • final attempt { classification: "failed", reward: 0 }
                3. Observe that the emitted task_completed event has scored: false.
                4. Change the verifier outcome to a valid pass with positive reward.
                5. Observe that the failed cell is now accepted as scored: true.

                The relevant logic is taskCompletedEvent() in packages/headless/src/fixed-prompt-controller.ts: structuredVerifierPassed participates in verifierGraded, but the corresponding structured verifier failure does not.

                Expected behavior:

                Once the structured verifier produces a valid passed or failed outcome, that result should be scored and locked regardless of the agent cell's terminal status. Only missing, malformed, or infrastructure-failed verifier outcomes should remain replacement-eligible.

                Regression coverage should include:

                • structured pass on a failed agent cell → scored pass;
                • structured failure on a failed agent cell → scored fail;
                • verifier infrastructure failure → unscored;
                • a scored verifier failure is not eligible for replacement.

                Environment

                • Maka observed subject commit: ee7e4ba5d
                • Confirmed affected logic remains at current commit: c561fed4e1
                • Surface: Headless / Harbor Terminal-Bench harness
                • macOS: 26.5.2
                • Node.js: v26.5.0

                Logs, screenshots, or additional context

                No provider HTTP 429 responses were involved in the observed examples. The issue concerns verifier authority and result classification, not provider availability.

                Metadata

                Metadata

                Assignees

                Labels

                bugSomething isn't working

                Type

                No type

                Projects

                No projects

                  Milestone

                  No milestone

                  Relationships

                  None yet

                  Development

                  No branches or pull requests

                  Issue actions