fix(test): metering test used the renamed tool (await_next → await_event) — main was red - #322

Merged
drewstone merged 1 commit into
mainfrom
fix/driver-inference-metering
Jun 17, 2026
Merged

fix(test): metering test used the renamed tool (await_next → await_event) — main was red#322
drewstone merged 1 commit into
mainfrom
fix/driver-inference-metering

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

main has 2 failing tests in driver-inference-metering.test.ts — this fixes them.

Root cause

#318 renamed the coordination wait verb await_nextawait_event. #319's metering test (merged around the same time) still scripted await_next. Since the tool no longer exists, the scripted driver's 'collect the worker' turn got {error:'unknown tool: await_next'}, the worker was never drained, and the two winner-expecting tests returned no-winner → failed.

The driver itself is fine — it fails loud (byName.get(name){error:'unknown tool: …'}, folded back to the LLM). A real LLM would recover by calling await_event; only the fixed test script couldn't. So this is purely a stale test literal.

Fix

All 6 await_next refs → await_event. Swept the whole repo — no stale await_next remains in src/tests/bench/docs/skills.

Verification

  • the file: 9/9 pass (was 7/9)
  • full suite 1030 pass, 0 failing (main was red at 2)

🤖 Generated with Claude Code

…t → await_event)
#318 renamed the coordination wait verb await_next → await_event; #319's metering test
(merged around the same time) still scripted await_next. Since the tool no longer exists,
the scripted driver's 'collect the worker' turn returned {error:'unknown tool: await_next'}
and the worker was never drained → result 'no-winner' → the 2 winner-expecting tests failed
on main. Updated all 6 refs to await_event (the driver already fails loud on the unknown
tool — only the fixed script couldn't recover, unlike a real LLM).
Full suite 1030 pass (was 2 failing on main); no stale await_next remains anywhere.

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved PR — 933d5f39

Blanket team auto-approval is enabled for this reviewer service.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: blanket_auto_approve · 2026-06-17T10:26:42Z

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Value Audit — sound

Verdictsound
Concerns0 (none)
Heuristic0.0s
Duplication0.0s
Interrogation154.5s (2 bridge agents)
Total154.5s

💰 Value — sound

Replaces six stale await_next literals in a single test file with the current await_event tool name so the scripted driver can drain workers again and the metering tests pass.

  • What it does: In tests/loops/driver-inference-metering.test.ts it changes every occurrence of the obsolete coordination tool name await_next to await_event in the scripted meteredChat turns and in the observability assertion (6 replacements total).
  • Goals it achieves: Restores the two failing winner-expecting tests by making the scripted driver call a tool that actually exists, so the worker settlement is drained and the tests return winner instead of no-winner; also brings the test file into line with the post-#318 coordination toolbox vocabulary.
  • Assessment: Good. The change is minimal, surgical, and exactly matches the codebase's current grain: await_event is the single canonical wait verb defined in src/mcp/tools/coordination.ts:402 and used throughout tests, docs, and skills. The PR author also swept the repo and confirmed no stale await_next remains.
  • Better / existing approach: none — this is the right approach. Adding a backwards-compatibility alias for await_next would fight the codebase's one-verb design and create unnecessary maintenance surface.
  • Model: kimi-code/kimi-for-coding
  • Bridge attempts: 3
  • Bridge warning: opencode/deepseek/deepseek-v4-pro: Bridge returned 503: {"error":{"message":"cli-bridge admission timed out after 30000ms","type":"admission_rejected","reason":"queue_timeout","admission":{"active":24,"queued":1,"maxActive":24,"maxQueue":32}}}; opencode/zai-coding-plan/glm-5.1: Bridge returned 503: {"error":{"message":"cli-bridge admission timed out after 30000ms","type":"admission_rejected","reas

🎯 Usefulness — sound

Restores a stale test script to the live coordination-verb surface; no new API, no dead code, no competing pattern.

  • Integration: The change replaces stale await_next literals with await_event in tests/loops/driver-inference-metering.test.ts so the scripted driver-LLM calls a tool that actually exists. The real tool is registered at src/mcp/tools/coordination.ts:402 as name: 'await_event'. A repo-wide grep for await_next returned no matches, so no stale references remain. await_event is already reachable across
  • Fit with existing patterns: It fits the established pattern exactly. The coordination driver exposes the verbs spawn_worker, await_event, observe_worker, steer_worker, stop over a live scope, and the production prompt in src/runtime/supervise/authoring.ts:56 already instructs the driver to use await_event (kinds:['settled']). The test was just out of sync after the rename in #318; the fix brings it back into th
  • Real-world viability: The runtime behavior is unchanged; the driver already handles unknown tools correctly by folding {error: 'unknown tool: ...'} back to the LLM (src/runtime/supervise/coordination-driver.ts:200-203), which a real LLM would recover from. The fixed test now correctly drains spawned workers via the valid tool and asserts on the real metering, nested re-homing, crash recovery, budget bounds, and obs
  • Model: kimi-code/kimi-for-coding
  • Bridge attempts: 3
  • Bridge warning: opencode/deepseek/deepseek-v4-pro: Bridge returned 503: {"error":{"message":"cli-bridge admission timed out after 30000ms","type":"admission_rejected","reason":"queue_timeout","admission":{"active":24,"queued":2,"maxActive":24,"maxQueue":32}}}; opencode/zai-coding-plan/glm-5.1: Bridge returned 503: {"error":{"message":"cli-bridge admission timed out after 30000ms","type":"admission_rejected","reas

No concerns — sound change, no better or existing approach found. ✅


What this audit checks

It judges the change on its merits — not whether it was tasked out in an issue. Unticketed, fast-moving work is fine; the question is whether the change is good and whether a better or existing approach should be used instead.

PassWhat it asks
HeuristicVague title? Whitespace-only or cruft-bearing diff? (content signals only)
DuplicationDo added function/class names already exist elsewhere in the repo?
Value AuditWhat does it do? What goal does it achieve? Is it good? Better architecture or already-exists?
Usefulness AuditDoes it integrate and fit? Will it hold up in real use and actually get used?

Findings are concerns, not blocks — the human reviewer decides what to do with them.

value-audit · 20260617T103131Z

@tangletools

Copy link
Copy Markdown
Contributor

✅ No Blockers — 933d5f39

Readiness 95/100 · Confidence 65/100 · 0 findings (none)

deepseekglmaggregate
Readiness959595
Confidence656565
Correctness959595
Security959595
Testing959595
Architecture959595

Full multi-shot audit completed 1/1 planned shots over 1 changed files. Global verifier still owns final merge decision. | Full multi-shot audit completed 1/1 planned shots over 1 changed files. Global verifier still owns final merge decision.

No findings.


tangletools · 2026-06-17T10:33:00Z · trace

@drewstone
drewstone merged commit d95652d into mainJun 17, 2026
1 check passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

fix(test): metering test used the renamed tool (await_next → await_event) — main was red - #322

Merged
drewstone merged 1 commit into
mainfrom
fix/driver-inference-metering
Jun 17, 2026
Merged

fix(test): metering test used the renamed tool (await_next → await_event) — main was red#322
drewstone merged 1 commit into
mainfrom
fix/driver-inference-metering

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

main has 2 failing tests in driver-inference-metering.test.ts — this fixes them.

Root cause

#318 renamed the coordination wait verb await_nextawait_event. #319's metering test (merged around the same time) still scripted await_next. Since the tool no longer exists, the scripted driver's 'collect the worker' turn got {error:'unknown tool: await_next'}, the worker was never drained, and the two winner-expecting tests returned no-winner → failed.

The driver itself is fine — it fails loud (byName.get(name){error:'unknown tool: …'}, folded back to the LLM). A real LLM would recover by calling await_event; only the fixed test script couldn't. So this is purely a stale test literal.

Fix

All 6 await_next refs → await_event. Swept the whole repo — no stale await_next remains in src/tests/bench/docs/skills.

Verification

  • the file: 9/9 pass (was 7/9)
  • full suite 1030 pass, 0 failing (main was red at 2)

🤖 Generated with Claude Code

…t → await_event)
#318 renamed the coordination wait verb await_next → await_event; #319's metering test
(merged around the same time) still scripted await_next. Since the tool no longer exists,
the scripted driver's 'collect the worker' turn returned {error:'unknown tool: await_next'}
and the worker was never drained → result 'no-winner' → the 2 winner-expecting tests failed
on main. Updated all 6 refs to await_event (the driver already fails loud on the unknown
tool — only the fixed script couldn't recover, unlike a real LLM).
Full suite 1030 pass (was 2 failing on main); no stale await_next remains anywhere.

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved PR — 933d5f39

Blanket team auto-approval is enabled for this reviewer service.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: blanket_auto_approve · 2026-06-17T10:26:42Z

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Value Audit — sound

Verdictsound
Concerns0 (none)
Heuristic0.0s
Duplication0.0s
Interrogation154.5s (2 bridge agents)
Total154.5s

💰 Value — sound

Replaces six stale await_next literals in a single test file with the current await_event tool name so the scripted driver can drain workers again and the metering tests pass.

  • What it does: In tests/loops/driver-inference-metering.test.ts it changes every occurrence of the obsolete coordination tool name await_next to await_event in the scripted meteredChat turns and in the observability assertion (6 replacements total).
  • Goals it achieves: Restores the two failing winner-expecting tests by making the scripted driver call a tool that actually exists, so the worker settlement is drained and the tests return winner instead of no-winner; also brings the test file into line with the post-#318 coordination toolbox vocabulary.
  • Assessment: Good. The change is minimal, surgical, and exactly matches the codebase's current grain: await_event is the single canonical wait verb defined in src/mcp/tools/coordination.ts:402 and used throughout tests, docs, and skills. The PR author also swept the repo and confirmed no stale await_next remains.
  • Better / existing approach: none — this is the right approach. Adding a backwards-compatibility alias for await_next would fight the codebase's one-verb design and create unnecessary maintenance surface.
  • Model: kimi-code/kimi-for-coding
  • Bridge attempts: 3
  • Bridge warning: opencode/deepseek/deepseek-v4-pro: Bridge returned 503: {"error":{"message":"cli-bridge admission timed out after 30000ms","type":"admission_rejected","reason":"queue_timeout","admission":{"active":24,"queued":1,"maxActive":24,"maxQueue":32}}}; opencode/zai-coding-plan/glm-5.1: Bridge returned 503: {"error":{"message":"cli-bridge admission timed out after 30000ms","type":"admission_rejected","reas

🎯 Usefulness — sound

Restores a stale test script to the live coordination-verb surface; no new API, no dead code, no competing pattern.

  • Integration: The change replaces stale await_next literals with await_event in tests/loops/driver-inference-metering.test.ts so the scripted driver-LLM calls a tool that actually exists. The real tool is registered at src/mcp/tools/coordination.ts:402 as name: 'await_event'. A repo-wide grep for await_next returned no matches, so no stale references remain. await_event is already reachable across
  • Fit with existing patterns: It fits the established pattern exactly. The coordination driver exposes the verbs spawn_worker, await_event, observe_worker, steer_worker, stop over a live scope, and the production prompt in src/runtime/supervise/authoring.ts:56 already instructs the driver to use await_event (kinds:['settled']). The test was just out of sync after the rename in #318; the fix brings it back into th
  • Real-world viability: The runtime behavior is unchanged; the driver already handles unknown tools correctly by folding {error: 'unknown tool: ...'} back to the LLM (src/runtime/supervise/coordination-driver.ts:200-203), which a real LLM would recover from. The fixed test now correctly drains spawned workers via the valid tool and asserts on the real metering, nested re-homing, crash recovery, budget bounds, and obs
  • Model: kimi-code/kimi-for-coding
  • Bridge attempts: 3
  • Bridge warning: opencode/deepseek/deepseek-v4-pro: Bridge returned 503: {"error":{"message":"cli-bridge admission timed out after 30000ms","type":"admission_rejected","reason":"queue_timeout","admission":{"active":24,"queued":2,"maxActive":24,"maxQueue":32}}}; opencode/zai-coding-plan/glm-5.1: Bridge returned 503: {"error":{"message":"cli-bridge admission timed out after 30000ms","type":"admission_rejected","reas

No concerns — sound change, no better or existing approach found. ✅


What this audit checks

It judges the change on its merits — not whether it was tasked out in an issue. Unticketed, fast-moving work is fine; the question is whether the change is good and whether a better or existing approach should be used instead.

PassWhat it asks
HeuristicVague title? Whitespace-only or cruft-bearing diff? (content signals only)
DuplicationDo added function/class names already exist elsewhere in the repo?
Value AuditWhat does it do? What goal does it achieve? Is it good? Better architecture or already-exists?
Usefulness AuditDoes it integrate and fit? Will it hold up in real use and actually get used?

Findings are concerns, not blocks — the human reviewer decides what to do with them.

value-audit · 20260617T103131Z

@tangletools

Copy link
Copy Markdown
Contributor

✅ No Blockers — 933d5f39

Readiness 95/100 · Confidence 65/100 · 0 findings (none)

deepseekglmaggregate
Readiness959595
Confidence656565
Correctness959595
Security959595
Testing959595
Architecture959595

Full multi-shot audit completed 1/1 planned shots over 1 changed files. Global verifier still owns final merge decision. | Full multi-shot audit completed 1/1 planned shots over 1 changed files. Global verifier still owns final merge decision.

No findings.


tangletools · 2026-06-17T10:33:00Z · trace

@drewstone
drewstone merged commit d95652d into mainJun 17, 2026
1 check passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(test): metering test used the renamed tool (await_next → await_event) — main was red - #322

Merged
drewstone merged 1 commit into
mainfrom
fix/driver-inference-metering
Jun 17, 2026
Merged

fix(test): metering test used the renamed tool (await_next → await_event) — main was red#322
drewstone merged 1 commit into
mainfrom
fix/driver-inference-metering

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

main has 2 failing tests in driver-inference-metering.test.ts — this fixes them.

Root cause

#318 renamed the coordination wait verb await_nextawait_event. #319's metering test (merged around the same time) still scripted await_next. Since the tool no longer exists, the scripted driver's 'collect the worker' turn got {error:'unknown tool: await_next'}, the worker was never drained, and the two winner-expecting tests returned no-winner → failed.

The driver itself is fine — it fails loud (byName.get(name){error:'unknown tool: …'}, folded back to the LLM). A real LLM would recover by calling await_event; only the fixed test script couldn't. So this is purely a stale test literal.

Fix

All 6 await_next refs → await_event. Swept the whole repo — no stale await_next remains in src/tests/bench/docs/skills.

Verification

  • the file: 9/9 pass (was 7/9)
  • full suite 1030 pass, 0 failing (main was red at 2)

🤖 Generated with Claude Code

…t → await_event)
#318 renamed the coordination wait verb await_next → await_event; #319's metering test
(merged around the same time) still scripted await_next. Since the tool no longer exists,
the scripted driver's 'collect the worker' turn returned {error:'unknown tool: await_next'}
and the worker was never drained → result 'no-winner' → the 2 winner-expecting tests failed
on main. Updated all 6 refs to await_event (the driver already fails loud on the unknown
tool — only the fixed script couldn't recover, unlike a real LLM).
Full suite 1030 pass (was 2 failing on main); no stale await_next remains anywhere.

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved PR — 933d5f39

Blanket team auto-approval is enabled for this reviewer service.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: blanket_auto_approve · 2026-06-17T10:26:42Z

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Value Audit — sound

Verdictsound
Concerns0 (none)
Heuristic0.0s
Duplication0.0s
Interrogation154.5s (2 bridge agents)
Total154.5s

💰 Value — sound

Replaces six stale await_next literals in a single test file with the current await_event tool name so the scripted driver can drain workers again and the metering tests pass.

  • What it does: In tests/loops/driver-inference-metering.test.ts it changes every occurrence of the obsolete coordination tool name await_next to await_event in the scripted meteredChat turns and in the observability assertion (6 replacements total).
  • Goals it achieves: Restores the two failing winner-expecting tests by making the scripted driver call a tool that actually exists, so the worker settlement is drained and the tests return winner instead of no-winner; also brings the test file into line with the post-#318 coordination toolbox vocabulary.
  • Assessment: Good. The change is minimal, surgical, and exactly matches the codebase's current grain: await_event is the single canonical wait verb defined in src/mcp/tools/coordination.ts:402 and used throughout tests, docs, and skills. The PR author also swept the repo and confirmed no stale await_next remains.
  • Better / existing approach: none — this is the right approach. Adding a backwards-compatibility alias for await_next would fight the codebase's one-verb design and create unnecessary maintenance surface.
  • Model: kimi-code/kimi-for-coding
  • Bridge attempts: 3
  • Bridge warning: opencode/deepseek/deepseek-v4-pro: Bridge returned 503: {"error":{"message":"cli-bridge admission timed out after 30000ms","type":"admission_rejected","reason":"queue_timeout","admission":{"active":24,"queued":1,"maxActive":24,"maxQueue":32}}}; opencode/zai-coding-plan/glm-5.1: Bridge returned 503: {"error":{"message":"cli-bridge admission timed out after 30000ms","type":"admission_rejected","reas

🎯 Usefulness — sound

Restores a stale test script to the live coordination-verb surface; no new API, no dead code, no competing pattern.

  • Integration: The change replaces stale await_next literals with await_event in tests/loops/driver-inference-metering.test.ts so the scripted driver-LLM calls a tool that actually exists. The real tool is registered at src/mcp/tools/coordination.ts:402 as name: 'await_event'. A repo-wide grep for await_next returned no matches, so no stale references remain. await_event is already reachable across
  • Fit with existing patterns: It fits the established pattern exactly. The coordination driver exposes the verbs spawn_worker, await_event, observe_worker, steer_worker, stop over a live scope, and the production prompt in src/runtime/supervise/authoring.ts:56 already instructs the driver to use await_event (kinds:['settled']). The test was just out of sync after the rename in #318; the fix brings it back into th
  • Real-world viability: The runtime behavior is unchanged; the driver already handles unknown tools correctly by folding {error: 'unknown tool: ...'} back to the LLM (src/runtime/supervise/coordination-driver.ts:200-203), which a real LLM would recover from. The fixed test now correctly drains spawned workers via the valid tool and asserts on the real metering, nested re-homing, crash recovery, budget bounds, and obs
  • Model: kimi-code/kimi-for-coding
  • Bridge attempts: 3
  • Bridge warning: opencode/deepseek/deepseek-v4-pro: Bridge returned 503: {"error":{"message":"cli-bridge admission timed out after 30000ms","type":"admission_rejected","reason":"queue_timeout","admission":{"active":24,"queued":2,"maxActive":24,"maxQueue":32}}}; opencode/zai-coding-plan/glm-5.1: Bridge returned 503: {"error":{"message":"cli-bridge admission timed out after 30000ms","type":"admission_rejected","reas

No concerns — sound change, no better or existing approach found. ✅


What this audit checks

It judges the change on its merits — not whether it was tasked out in an issue. Unticketed, fast-moving work is fine; the question is whether the change is good and whether a better or existing approach should be used instead.

PassWhat it asks
HeuristicVague title? Whitespace-only or cruft-bearing diff? (content signals only)
DuplicationDo added function/class names already exist elsewhere in the repo?
Value AuditWhat does it do? What goal does it achieve? Is it good? Better architecture or already-exists?
Usefulness AuditDoes it integrate and fit? Will it hold up in real use and actually get used?

Findings are concerns, not blocks — the human reviewer decides what to do with them.

value-audit · 20260617T103131Z

@tangletools

Copy link
Copy Markdown
Contributor

✅ No Blockers — 933d5f39

Readiness 95/100 · Confidence 65/100 · 0 findings (none)

deepseekglmaggregate
Readiness959595
Confidence656565
Correctness959595
Security959595
Testing959595
Architecture959595

Full multi-shot audit completed 1/1 planned shots over 1 changed files. Global verifier still owns final merge decision. | Full multi-shot audit completed 1/1 planned shots over 1 changed files. Global verifier still owns final merge decision.

No findings.


tangletools · 2026-06-17T10:33:00Z · trace

@drewstone
drewstone merged commit d95652d into mainJun 17, 2026
1 check passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(test): metering test used the renamed tool (await_next → await_event) — main was red - #322

Merged
drewstone merged 1 commit into
mainfrom
fix/driver-inference-metering
Jun 17, 2026
Merged

fix(test): metering test used the renamed tool (await_next → await_event) — main was red#322
drewstone merged 1 commit into
mainfrom
fix/driver-inference-metering

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

main has 2 failing tests in driver-inference-metering.test.ts — this fixes them.

Root cause

#318 renamed the coordination wait verb await_nextawait_event. #319's metering test (merged around the same time) still scripted await_next. Since the tool no longer exists, the scripted driver's 'collect the worker' turn got {error:'unknown tool: await_next'}, the worker was never drained, and the two winner-expecting tests returned no-winner → failed.

The driver itself is fine — it fails loud (byName.get(name){error:'unknown tool: …'}, folded back to the LLM). A real LLM would recover by calling await_event; only the fixed test script couldn't. So this is purely a stale test literal.

Fix

All 6 await_next refs → await_event. Swept the whole repo — no stale await_next remains in src/tests/bench/docs/skills.

Verification

  • the file: 9/9 pass (was 7/9)
  • full suite 1030 pass, 0 failing (main was red at 2)

🤖 Generated with Claude Code

…t → await_event)
#318 renamed the coordination wait verb await_next → await_event; #319's metering test
(merged around the same time) still scripted await_next. Since the tool no longer exists,
the scripted driver's 'collect the worker' turn returned {error:'unknown tool: await_next'}
and the worker was never drained → result 'no-winner' → the 2 winner-expecting tests failed
on main. Updated all 6 refs to await_event (the driver already fails loud on the unknown
tool — only the fixed script couldn't recover, unlike a real LLM).
Full suite 1030 pass (was 2 failing on main); no stale await_next remains anywhere.

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved PR — 933d5f39

Blanket team auto-approval is enabled for this reviewer service.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: blanket_auto_approve · 2026-06-17T10:26:42Z

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Value Audit — sound

Verdictsound
Concerns0 (none)
Heuristic0.0s
Duplication0.0s
Interrogation154.5s (2 bridge agents)
Total154.5s

💰 Value — sound

Replaces six stale await_next literals in a single test file with the current await_event tool name so the scripted driver can drain workers again and the metering tests pass.

  • What it does: In tests/loops/driver-inference-metering.test.ts it changes every occurrence of the obsolete coordination tool name await_next to await_event in the scripted meteredChat turns and in the observability assertion (6 replacements total).
  • Goals it achieves: Restores the two failing winner-expecting tests by making the scripted driver call a tool that actually exists, so the worker settlement is drained and the tests return winner instead of no-winner; also brings the test file into line with the post-#318 coordination toolbox vocabulary.
  • Assessment: Good. The change is minimal, surgical, and exactly matches the codebase's current grain: await_event is the single canonical wait verb defined in src/mcp/tools/coordination.ts:402 and used throughout tests, docs, and skills. The PR author also swept the repo and confirmed no stale await_next remains.
  • Better / existing approach: none — this is the right approach. Adding a backwards-compatibility alias for await_next would fight the codebase's one-verb design and create unnecessary maintenance surface.
  • Model: kimi-code/kimi-for-coding
  • Bridge attempts: 3
  • Bridge warning: opencode/deepseek/deepseek-v4-pro: Bridge returned 503: {"error":{"message":"cli-bridge admission timed out after 30000ms","type":"admission_rejected","reason":"queue_timeout","admission":{"active":24,"queued":1,"maxActive":24,"maxQueue":32}}}; opencode/zai-coding-plan/glm-5.1: Bridge returned 503: {"error":{"message":"cli-bridge admission timed out after 30000ms","type":"admission_rejected","reas

🎯 Usefulness — sound

Restores a stale test script to the live coordination-verb surface; no new API, no dead code, no competing pattern.

  • Integration: The change replaces stale await_next literals with await_event in tests/loops/driver-inference-metering.test.ts so the scripted driver-LLM calls a tool that actually exists. The real tool is registered at src/mcp/tools/coordination.ts:402 as name: 'await_event'. A repo-wide grep for await_next returned no matches, so no stale references remain. await_event is already reachable across
  • Fit with existing patterns: It fits the established pattern exactly. The coordination driver exposes the verbs spawn_worker, await_event, observe_worker, steer_worker, stop over a live scope, and the production prompt in src/runtime/supervise/authoring.ts:56 already instructs the driver to use await_event (kinds:['settled']). The test was just out of sync after the rename in #318; the fix brings it back into th
  • Real-world viability: The runtime behavior is unchanged; the driver already handles unknown tools correctly by folding {error: 'unknown tool: ...'} back to the LLM (src/runtime/supervise/coordination-driver.ts:200-203), which a real LLM would recover from. The fixed test now correctly drains spawned workers via the valid tool and asserts on the real metering, nested re-homing, crash recovery, budget bounds, and obs
  • Model: kimi-code/kimi-for-coding
  • Bridge attempts: 3
  • Bridge warning: opencode/deepseek/deepseek-v4-pro: Bridge returned 503: {"error":{"message":"cli-bridge admission timed out after 30000ms","type":"admission_rejected","reason":"queue_timeout","admission":{"active":24,"queued":2,"maxActive":24,"maxQueue":32}}}; opencode/zai-coding-plan/glm-5.1: Bridge returned 503: {"error":{"message":"cli-bridge admission timed out after 30000ms","type":"admission_rejected","reas

No concerns — sound change, no better or existing approach found. ✅


What this audit checks

It judges the change on its merits — not whether it was tasked out in an issue. Unticketed, fast-moving work is fine; the question is whether the change is good and whether a better or existing approach should be used instead.

PassWhat it asks
HeuristicVague title? Whitespace-only or cruft-bearing diff? (content signals only)
DuplicationDo added function/class names already exist elsewhere in the repo?
Value AuditWhat does it do? What goal does it achieve? Is it good? Better architecture or already-exists?
Usefulness AuditDoes it integrate and fit? Will it hold up in real use and actually get used?

Findings are concerns, not blocks — the human reviewer decides what to do with them.

value-audit · 20260617T103131Z

@tangletools

Copy link
Copy Markdown
Contributor

✅ No Blockers — 933d5f39

Readiness 95/100 · Confidence 65/100 · 0 findings (none)

deepseekglmaggregate
Readiness959595
Confidence656565
Correctness959595
Security959595
Testing959595
Architecture959595

Full multi-shot audit completed 1/1 planned shots over 1 changed files. Global verifier still owns final merge decision. | Full multi-shot audit completed 1/1 planned shots over 1 changed files. Global verifier still owns final merge decision.

No findings.


tangletools · 2026-06-17T10:33:00Z · trace

@drewstone
drewstone merged commit d95652d into mainJun 17, 2026
1 check passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

fix(test): metering test used the renamed tool (await_next → await_event) — main was red - #322

Merged
drewstone merged 1 commit into
mainfrom
fix/driver-inference-metering
Jun 17, 2026
Merged

fix(test): metering test used the renamed tool (await_next → await_event) — main was red#322
drewstone merged 1 commit into
mainfrom
fix/driver-inference-metering

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

main has 2 failing tests in driver-inference-metering.test.ts — this fixes them.

Root cause

#318 renamed the coordination wait verb await_nextawait_event. #319's metering test (merged around the same time) still scripted await_next. Since the tool no longer exists, the scripted driver's 'collect the worker' turn got {error:'unknown tool: await_next'}, the worker was never drained, and the two winner-expecting tests returned no-winner → failed.

The driver itself is fine — it fails loud (byName.get(name){error:'unknown tool: …'}, folded back to the LLM). A real LLM would recover by calling await_event; only the fixed test script couldn't. So this is purely a stale test literal.

Fix

All 6 await_next refs → await_event. Swept the whole repo — no stale await_next remains in src/tests/bench/docs/skills.

Verification

  • the file: 9/9 pass (was 7/9)
  • full suite 1030 pass, 0 failing (main was red at 2)

🤖 Generated with Claude Code

…t → await_event)
#318 renamed the coordination wait verb await_next → await_event; #319's metering test
(merged around the same time) still scripted await_next. Since the tool no longer exists,
the scripted driver's 'collect the worker' turn returned {error:'unknown tool: await_next'}
and the worker was never drained → result 'no-winner' → the 2 winner-expecting tests failed
on main. Updated all 6 refs to await_event (the driver already fails loud on the unknown
tool — only the fixed script couldn't recover, unlike a real LLM).
Full suite 1030 pass (was 2 failing on main); no stale await_next remains anywhere.

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved PR — 933d5f39

Blanket team auto-approval is enabled for this reviewer service.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: blanket_auto_approve · 2026-06-17T10:26:42Z

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Value Audit — sound

Verdictsound
Concerns0 (none)
Heuristic0.0s
Duplication0.0s
Interrogation154.5s (2 bridge agents)
Total154.5s

💰 Value — sound

Replaces six stale await_next literals in a single test file with the current await_event tool name so the scripted driver can drain workers again and the metering tests pass.

  • What it does: In tests/loops/driver-inference-metering.test.ts it changes every occurrence of the obsolete coordination tool name await_next to await_event in the scripted meteredChat turns and in the observability assertion (6 replacements total).
  • Goals it achieves: Restores the two failing winner-expecting tests by making the scripted driver call a tool that actually exists, so the worker settlement is drained and the tests return winner instead of no-winner; also brings the test file into line with the post-#318 coordination toolbox vocabulary.
  • Assessment: Good. The change is minimal, surgical, and exactly matches the codebase's current grain: await_event is the single canonical wait verb defined in src/mcp/tools/coordination.ts:402 and used throughout tests, docs, and skills. The PR author also swept the repo and confirmed no stale await_next remains.
  • Better / existing approach: none — this is the right approach. Adding a backwards-compatibility alias for await_next would fight the codebase's one-verb design and create unnecessary maintenance surface.
  • Model: kimi-code/kimi-for-coding
  • Bridge attempts: 3
  • Bridge warning: opencode/deepseek/deepseek-v4-pro: Bridge returned 503: {"error":{"message":"cli-bridge admission timed out after 30000ms","type":"admission_rejected","reason":"queue_timeout","admission":{"active":24,"queued":1,"maxActive":24,"maxQueue":32}}}; opencode/zai-coding-plan/glm-5.1: Bridge returned 503: {"error":{"message":"cli-bridge admission timed out after 30000ms","type":"admission_rejected","reas

🎯 Usefulness — sound

Restores a stale test script to the live coordination-verb surface; no new API, no dead code, no competing pattern.

  • Integration: The change replaces stale await_next literals with await_event in tests/loops/driver-inference-metering.test.ts so the scripted driver-LLM calls a tool that actually exists. The real tool is registered at src/mcp/tools/coordination.ts:402 as name: 'await_event'. A repo-wide grep for await_next returned no matches, so no stale references remain. await_event is already reachable across
  • Fit with existing patterns: It fits the established pattern exactly. The coordination driver exposes the verbs spawn_worker, await_event, observe_worker, steer_worker, stop over a live scope, and the production prompt in src/runtime/supervise/authoring.ts:56 already instructs the driver to use await_event (kinds:['settled']). The test was just out of sync after the rename in #318; the fix brings it back into th
  • Real-world viability: The runtime behavior is unchanged; the driver already handles unknown tools correctly by folding {error: 'unknown tool: ...'} back to the LLM (src/runtime/supervise/coordination-driver.ts:200-203), which a real LLM would recover from. The fixed test now correctly drains spawned workers via the valid tool and asserts on the real metering, nested re-homing, crash recovery, budget bounds, and obs
  • Model: kimi-code/kimi-for-coding
  • Bridge attempts: 3
  • Bridge warning: opencode/deepseek/deepseek-v4-pro: Bridge returned 503: {"error":{"message":"cli-bridge admission timed out after 30000ms","type":"admission_rejected","reason":"queue_timeout","admission":{"active":24,"queued":2,"maxActive":24,"maxQueue":32}}}; opencode/zai-coding-plan/glm-5.1: Bridge returned 503: {"error":{"message":"cli-bridge admission timed out after 30000ms","type":"admission_rejected","reas

No concerns — sound change, no better or existing approach found. ✅


What this audit checks

It judges the change on its merits — not whether it was tasked out in an issue. Unticketed, fast-moving work is fine; the question is whether the change is good and whether a better or existing approach should be used instead.

PassWhat it asks
HeuristicVague title? Whitespace-only or cruft-bearing diff? (content signals only)
DuplicationDo added function/class names already exist elsewhere in the repo?
Value AuditWhat does it do? What goal does it achieve? Is it good? Better architecture or already-exists?
Usefulness AuditDoes it integrate and fit? Will it hold up in real use and actually get used?

Findings are concerns, not blocks — the human reviewer decides what to do with them.

value-audit · 20260617T103131Z

@tangletools

Copy link
Copy Markdown
Contributor

✅ No Blockers — 933d5f39

Readiness 95/100 · Confidence 65/100 · 0 findings (none)

deepseekglmaggregate
Readiness959595
Confidence656565
Correctness959595
Security959595
Testing959595
Architecture959595

Full multi-shot audit completed 1/1 planned shots over 1 changed files. Global verifier still owns final merge decision. | Full multi-shot audit completed 1/1 planned shots over 1 changed files. Global verifier still owns final merge decision.

No findings.


tangletools · 2026-06-17T10:33:00Z · trace

@drewstone
drewstone merged commit d95652d into mainJun 17, 2026
1 check passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(test): metering test used the renamed tool (await_next → await_event) — main was red - #322

Merged
drewstone merged 1 commit into
mainfrom
fix/driver-inference-metering
Jun 17, 2026
Merged

fix(test): metering test used the renamed tool (await_next → await_event) — main was red#322
drewstone merged 1 commit into
mainfrom
fix/driver-inference-metering

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

main has 2 failing tests in driver-inference-metering.test.ts — this fixes them.

Root cause

#318 renamed the coordination wait verb await_nextawait_event. #319's metering test (merged around the same time) still scripted await_next. Since the tool no longer exists, the scripted driver's 'collect the worker' turn got {error:'unknown tool: await_next'}, the worker was never drained, and the two winner-expecting tests returned no-winner → failed.

The driver itself is fine — it fails loud (byName.get(name){error:'unknown tool: …'}, folded back to the LLM). A real LLM would recover by calling await_event; only the fixed test script couldn't. So this is purely a stale test literal.

Fix

All 6 await_next refs → await_event. Swept the whole repo — no stale await_next remains in src/tests/bench/docs/skills.

Verification

  • the file: 9/9 pass (was 7/9)
  • full suite 1030 pass, 0 failing (main was red at 2)

🤖 Generated with Claude Code

…t → await_event)
#318 renamed the coordination wait verb await_next → await_event; #319's metering test
(merged around the same time) still scripted await_next. Since the tool no longer exists,
the scripted driver's 'collect the worker' turn returned {error:'unknown tool: await_next'}
and the worker was never drained → result 'no-winner' → the 2 winner-expecting tests failed
on main. Updated all 6 refs to await_event (the driver already fails loud on the unknown
tool — only the fixed script couldn't recover, unlike a real LLM).
Full suite 1030 pass (was 2 failing on main); no stale await_next remains anywhere.

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved PR — 933d5f39

Blanket team auto-approval is enabled for this reviewer service.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: blanket_auto_approve · 2026-06-17T10:26:42Z

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Value Audit — sound

Verdictsound
Concerns0 (none)
Heuristic0.0s
Duplication0.0s
Interrogation154.5s (2 bridge agents)
Total154.5s

💰 Value — sound

Replaces six stale await_next literals in a single test file with the current await_event tool name so the scripted driver can drain workers again and the metering tests pass.

  • What it does: In tests/loops/driver-inference-metering.test.ts it changes every occurrence of the obsolete coordination tool name await_next to await_event in the scripted meteredChat turns and in the observability assertion (6 replacements total).
  • Goals it achieves: Restores the two failing winner-expecting tests by making the scripted driver call a tool that actually exists, so the worker settlement is drained and the tests return winner instead of no-winner; also brings the test file into line with the post-#318 coordination toolbox vocabulary.
  • Assessment: Good. The change is minimal, surgical, and exactly matches the codebase's current grain: await_event is the single canonical wait verb defined in src/mcp/tools/coordination.ts:402 and used throughout tests, docs, and skills. The PR author also swept the repo and confirmed no stale await_next remains.
  • Better / existing approach: none — this is the right approach. Adding a backwards-compatibility alias for await_next would fight the codebase's one-verb design and create unnecessary maintenance surface.
  • Model: kimi-code/kimi-for-coding
  • Bridge attempts: 3
  • Bridge warning: opencode/deepseek/deepseek-v4-pro: Bridge returned 503: {"error":{"message":"cli-bridge admission timed out after 30000ms","type":"admission_rejected","reason":"queue_timeout","admission":{"active":24,"queued":1,"maxActive":24,"maxQueue":32}}}; opencode/zai-coding-plan/glm-5.1: Bridge returned 503: {"error":{"message":"cli-bridge admission timed out after 30000ms","type":"admission_rejected","reas

🎯 Usefulness — sound

Restores a stale test script to the live coordination-verb surface; no new API, no dead code, no competing pattern.

  • Integration: The change replaces stale await_next literals with await_event in tests/loops/driver-inference-metering.test.ts so the scripted driver-LLM calls a tool that actually exists. The real tool is registered at src/mcp/tools/coordination.ts:402 as name: 'await_event'. A repo-wide grep for await_next returned no matches, so no stale references remain. await_event is already reachable across
  • Fit with existing patterns: It fits the established pattern exactly. The coordination driver exposes the verbs spawn_worker, await_event, observe_worker, steer_worker, stop over a live scope, and the production prompt in src/runtime/supervise/authoring.ts:56 already instructs the driver to use await_event (kinds:['settled']). The test was just out of sync after the rename in #318; the fix brings it back into th
  • Real-world viability: The runtime behavior is unchanged; the driver already handles unknown tools correctly by folding {error: 'unknown tool: ...'} back to the LLM (src/runtime/supervise/coordination-driver.ts:200-203), which a real LLM would recover from. The fixed test now correctly drains spawned workers via the valid tool and asserts on the real metering, nested re-homing, crash recovery, budget bounds, and obs
  • Model: kimi-code/kimi-for-coding
  • Bridge attempts: 3
  • Bridge warning: opencode/deepseek/deepseek-v4-pro: Bridge returned 503: {"error":{"message":"cli-bridge admission timed out after 30000ms","type":"admission_rejected","reason":"queue_timeout","admission":{"active":24,"queued":2,"maxActive":24,"maxQueue":32}}}; opencode/zai-coding-plan/glm-5.1: Bridge returned 503: {"error":{"message":"cli-bridge admission timed out after 30000ms","type":"admission_rejected","reas

No concerns — sound change, no better or existing approach found. ✅


What this audit checks

It judges the change on its merits — not whether it was tasked out in an issue. Unticketed, fast-moving work is fine; the question is whether the change is good and whether a better or existing approach should be used instead.

PassWhat it asks
HeuristicVague title? Whitespace-only or cruft-bearing diff? (content signals only)
DuplicationDo added function/class names already exist elsewhere in the repo?
Value AuditWhat does it do? What goal does it achieve? Is it good? Better architecture or already-exists?
Usefulness AuditDoes it integrate and fit? Will it hold up in real use and actually get used?

Findings are concerns, not blocks — the human reviewer decides what to do with them.

value-audit · 20260617T103131Z

@tangletools

Copy link
Copy Markdown
Contributor

✅ No Blockers — 933d5f39

Readiness 95/100 · Confidence 65/100 · 0 findings (none)

deepseekglmaggregate
Readiness959595
Confidence656565
Correctness959595
Security959595
Testing959595
Architecture959595

Full multi-shot audit completed 1/1 planned shots over 1 changed files. Global verifier still owns final merge decision. | Full multi-shot audit completed 1/1 planned shots over 1 changed files. Global verifier still owns final merge decision.

No findings.


tangletools · 2026-06-17T10:33:00Z · trace

@drewstone
drewstone merged commit d95652d into mainJun 17, 2026
1 check passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(test): metering test used the renamed tool (await_next → await_event) — main was red - #322

Merged
drewstone merged 1 commit into
mainfrom
fix/driver-inference-metering
Jun 17, 2026
Merged

fix(test): metering test used the renamed tool (await_next → await_event) — main was red#322
drewstone merged 1 commit into
mainfrom
fix/driver-inference-metering

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

main has 2 failing tests in driver-inference-metering.test.ts — this fixes them.

Root cause

#318 renamed the coordination wait verb await_nextawait_event. #319's metering test (merged around the same time) still scripted await_next. Since the tool no longer exists, the scripted driver's 'collect the worker' turn got {error:'unknown tool: await_next'}, the worker was never drained, and the two winner-expecting tests returned no-winner → failed.

The driver itself is fine — it fails loud (byName.get(name){error:'unknown tool: …'}, folded back to the LLM). A real LLM would recover by calling await_event; only the fixed test script couldn't. So this is purely a stale test literal.

Fix

All 6 await_next refs → await_event. Swept the whole repo — no stale await_next remains in src/tests/bench/docs/skills.

Verification

  • the file: 9/9 pass (was 7/9)
  • full suite 1030 pass, 0 failing (main was red at 2)

🤖 Generated with Claude Code

…t → await_event)
#318 renamed the coordination wait verb await_next → await_event; #319's metering test
(merged around the same time) still scripted await_next. Since the tool no longer exists,
the scripted driver's 'collect the worker' turn returned {error:'unknown tool: await_next'}
and the worker was never drained → result 'no-winner' → the 2 winner-expecting tests failed
on main. Updated all 6 refs to await_event (the driver already fails loud on the unknown
tool — only the fixed script couldn't recover, unlike a real LLM).
Full suite 1030 pass (was 2 failing on main); no stale await_next remains anywhere.

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved PR — 933d5f39

Blanket team auto-approval is enabled for this reviewer service.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: blanket_auto_approve · 2026-06-17T10:26:42Z

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Value Audit — sound

Verdictsound
Concerns0 (none)
Heuristic0.0s
Duplication0.0s
Interrogation154.5s (2 bridge agents)
Total154.5s

💰 Value — sound

Replaces six stale await_next literals in a single test file with the current await_event tool name so the scripted driver can drain workers again and the metering tests pass.

  • What it does: In tests/loops/driver-inference-metering.test.ts it changes every occurrence of the obsolete coordination tool name await_next to await_event in the scripted meteredChat turns and in the observability assertion (6 replacements total).
  • Goals it achieves: Restores the two failing winner-expecting tests by making the scripted driver call a tool that actually exists, so the worker settlement is drained and the tests return winner instead of no-winner; also brings the test file into line with the post-#318 coordination toolbox vocabulary.
  • Assessment: Good. The change is minimal, surgical, and exactly matches the codebase's current grain: await_event is the single canonical wait verb defined in src/mcp/tools/coordination.ts:402 and used throughout tests, docs, and skills. The PR author also swept the repo and confirmed no stale await_next remains.
  • Better / existing approach: none — this is the right approach. Adding a backwards-compatibility alias for await_next would fight the codebase's one-verb design and create unnecessary maintenance surface.
  • Model: kimi-code/kimi-for-coding
  • Bridge attempts: 3
  • Bridge warning: opencode/deepseek/deepseek-v4-pro: Bridge returned 503: {"error":{"message":"cli-bridge admission timed out after 30000ms","type":"admission_rejected","reason":"queue_timeout","admission":{"active":24,"queued":1,"maxActive":24,"maxQueue":32}}}; opencode/zai-coding-plan/glm-5.1: Bridge returned 503: {"error":{"message":"cli-bridge admission timed out after 30000ms","type":"admission_rejected","reas

🎯 Usefulness — sound

Restores a stale test script to the live coordination-verb surface; no new API, no dead code, no competing pattern.

  • Integration: The change replaces stale await_next literals with await_event in tests/loops/driver-inference-metering.test.ts so the scripted driver-LLM calls a tool that actually exists. The real tool is registered at src/mcp/tools/coordination.ts:402 as name: 'await_event'. A repo-wide grep for await_next returned no matches, so no stale references remain. await_event is already reachable across
  • Fit with existing patterns: It fits the established pattern exactly. The coordination driver exposes the verbs spawn_worker, await_event, observe_worker, steer_worker, stop over a live scope, and the production prompt in src/runtime/supervise/authoring.ts:56 already instructs the driver to use await_event (kinds:['settled']). The test was just out of sync after the rename in #318; the fix brings it back into th
  • Real-world viability: The runtime behavior is unchanged; the driver already handles unknown tools correctly by folding {error: 'unknown tool: ...'} back to the LLM (src/runtime/supervise/coordination-driver.ts:200-203), which a real LLM would recover from. The fixed test now correctly drains spawned workers via the valid tool and asserts on the real metering, nested re-homing, crash recovery, budget bounds, and obs
  • Model: kimi-code/kimi-for-coding
  • Bridge attempts: 3
  • Bridge warning: opencode/deepseek/deepseek-v4-pro: Bridge returned 503: {"error":{"message":"cli-bridge admission timed out after 30000ms","type":"admission_rejected","reason":"queue_timeout","admission":{"active":24,"queued":2,"maxActive":24,"maxQueue":32}}}; opencode/zai-coding-plan/glm-5.1: Bridge returned 503: {"error":{"message":"cli-bridge admission timed out after 30000ms","type":"admission_rejected","reas

No concerns — sound change, no better or existing approach found. ✅


What this audit checks

It judges the change on its merits — not whether it was tasked out in an issue. Unticketed, fast-moving work is fine; the question is whether the change is good and whether a better or existing approach should be used instead.

PassWhat it asks
HeuristicVague title? Whitespace-only or cruft-bearing diff? (content signals only)
DuplicationDo added function/class names already exist elsewhere in the repo?
Value AuditWhat does it do? What goal does it achieve? Is it good? Better architecture or already-exists?
Usefulness AuditDoes it integrate and fit? Will it hold up in real use and actually get used?

Findings are concerns, not blocks — the human reviewer decides what to do with them.

value-audit · 20260617T103131Z

@tangletools

Copy link
Copy Markdown
Contributor

✅ No Blockers — 933d5f39

Readiness 95/100 · Confidence 65/100 · 0 findings (none)

deepseekglmaggregate
Readiness959595
Confidence656565
Correctness959595
Security959595
Testing959595
Architecture959595

Full multi-shot audit completed 1/1 planned shots over 1 changed files. Global verifier still owns final merge decision. | Full multi-shot audit completed 1/1 planned shots over 1 changed files. Global verifier still owns final merge decision.

No findings.


tangletools · 2026-06-17T10:33:00Z · trace

@drewstone
drewstone merged commit d95652d into mainJun 17, 2026
1 check passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

fix(test): metering test used the renamed tool (await_next → await_event) — main was red - #322

Merged
drewstone merged 1 commit into
mainfrom
fix/driver-inference-metering
Jun 17, 2026
Merged

fix(test): metering test used the renamed tool (await_next → await_event) — main was red#322
drewstone merged 1 commit into
mainfrom
fix/driver-inference-metering

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

main has 2 failing tests in driver-inference-metering.test.ts — this fixes them.

Root cause

#318 renamed the coordination wait verb await_nextawait_event. #319's metering test (merged around the same time) still scripted await_next. Since the tool no longer exists, the scripted driver's 'collect the worker' turn got {error:'unknown tool: await_next'}, the worker was never drained, and the two winner-expecting tests returned no-winner → failed.

The driver itself is fine — it fails loud (byName.get(name){error:'unknown tool: …'}, folded back to the LLM). A real LLM would recover by calling await_event; only the fixed test script couldn't. So this is purely a stale test literal.

Fix

All 6 await_next refs → await_event. Swept the whole repo — no stale await_next remains in src/tests/bench/docs/skills.

Verification

  • the file: 9/9 pass (was 7/9)
  • full suite 1030 pass, 0 failing (main was red at 2)

🤖 Generated with Claude Code

…t → await_event)
#318 renamed the coordination wait verb await_next → await_event; #319's metering test
(merged around the same time) still scripted await_next. Since the tool no longer exists,
the scripted driver's 'collect the worker' turn returned {error:'unknown tool: await_next'}
and the worker was never drained → result 'no-winner' → the 2 winner-expecting tests failed
on main. Updated all 6 refs to await_event (the driver already fails loud on the unknown
tool — only the fixed script couldn't recover, unlike a real LLM).
Full suite 1030 pass (was 2 failing on main); no stale await_next remains anywhere.

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved PR — 933d5f39

Blanket team auto-approval is enabled for this reviewer service.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: blanket_auto_approve · 2026-06-17T10:26:42Z

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Value Audit — sound

Verdictsound
Concerns0 (none)
Heuristic0.0s
Duplication0.0s
Interrogation154.5s (2 bridge agents)
Total154.5s

💰 Value — sound

Replaces six stale await_next literals in a single test file with the current await_event tool name so the scripted driver can drain workers again and the metering tests pass.

  • What it does: In tests/loops/driver-inference-metering.test.ts it changes every occurrence of the obsolete coordination tool name await_next to await_event in the scripted meteredChat turns and in the observability assertion (6 replacements total).
  • Goals it achieves: Restores the two failing winner-expecting tests by making the scripted driver call a tool that actually exists, so the worker settlement is drained and the tests return winner instead of no-winner; also brings the test file into line with the post-#318 coordination toolbox vocabulary.
  • Assessment: Good. The change is minimal, surgical, and exactly matches the codebase's current grain: await_event is the single canonical wait verb defined in src/mcp/tools/coordination.ts:402 and used throughout tests, docs, and skills. The PR author also swept the repo and confirmed no stale await_next remains.
  • Better / existing approach: none — this is the right approach. Adding a backwards-compatibility alias for await_next would fight the codebase's one-verb design and create unnecessary maintenance surface.
  • Model: kimi-code/kimi-for-coding
  • Bridge attempts: 3
  • Bridge warning: opencode/deepseek/deepseek-v4-pro: Bridge returned 503: {"error":{"message":"cli-bridge admission timed out after 30000ms","type":"admission_rejected","reason":"queue_timeout","admission":{"active":24,"queued":1,"maxActive":24,"maxQueue":32}}}; opencode/zai-coding-plan/glm-5.1: Bridge returned 503: {"error":{"message":"cli-bridge admission timed out after 30000ms","type":"admission_rejected","reas

🎯 Usefulness — sound

Restores a stale test script to the live coordination-verb surface; no new API, no dead code, no competing pattern.

  • Integration: The change replaces stale await_next literals with await_event in tests/loops/driver-inference-metering.test.ts so the scripted driver-LLM calls a tool that actually exists. The real tool is registered at src/mcp/tools/coordination.ts:402 as name: 'await_event'. A repo-wide grep for await_next returned no matches, so no stale references remain. await_event is already reachable across
  • Fit with existing patterns: It fits the established pattern exactly. The coordination driver exposes the verbs spawn_worker, await_event, observe_worker, steer_worker, stop over a live scope, and the production prompt in src/runtime/supervise/authoring.ts:56 already instructs the driver to use await_event (kinds:['settled']). The test was just out of sync after the rename in #318; the fix brings it back into th
  • Real-world viability: The runtime behavior is unchanged; the driver already handles unknown tools correctly by folding {error: 'unknown tool: ...'} back to the LLM (src/runtime/supervise/coordination-driver.ts:200-203), which a real LLM would recover from. The fixed test now correctly drains spawned workers via the valid tool and asserts on the real metering, nested re-homing, crash recovery, budget bounds, and obs
  • Model: kimi-code/kimi-for-coding
  • Bridge attempts: 3
  • Bridge warning: opencode/deepseek/deepseek-v4-pro: Bridge returned 503: {"error":{"message":"cli-bridge admission timed out after 30000ms","type":"admission_rejected","reason":"queue_timeout","admission":{"active":24,"queued":2,"maxActive":24,"maxQueue":32}}}; opencode/zai-coding-plan/glm-5.1: Bridge returned 503: {"error":{"message":"cli-bridge admission timed out after 30000ms","type":"admission_rejected","reas

No concerns — sound change, no better or existing approach found. ✅


What this audit checks

It judges the change on its merits — not whether it was tasked out in an issue. Unticketed, fast-moving work is fine; the question is whether the change is good and whether a better or existing approach should be used instead.

PassWhat it asks
HeuristicVague title? Whitespace-only or cruft-bearing diff? (content signals only)
DuplicationDo added function/class names already exist elsewhere in the repo?
Value AuditWhat does it do? What goal does it achieve? Is it good? Better architecture or already-exists?
Usefulness AuditDoes it integrate and fit? Will it hold up in real use and actually get used?

Findings are concerns, not blocks — the human reviewer decides what to do with them.

value-audit · 20260617T103131Z

@tangletools

Copy link
Copy Markdown
Contributor

✅ No Blockers — 933d5f39

Readiness 95/100 · Confidence 65/100 · 0 findings (none)

deepseekglmaggregate
Readiness959595
Confidence656565
Correctness959595
Security959595
Testing959595
Architecture959595

Full multi-shot audit completed 1/1 planned shots over 1 changed files. Global verifier still owns final merge decision. | Full multi-shot audit completed 1/1 planned shots over 1 changed files. Global verifier still owns final merge decision.

No findings.


tangletools · 2026-06-17T10:33:00Z · trace

@drewstone
drewstone merged commit d95652d into mainJun 17, 2026
1 check passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools