feat(bench): router-backed loop executor — stateful research through the real kernel - #188

Merged
drewstone merged 3 commits into
mainfrom
feat/research-loop-executor
Jun 7, 2026
Merged

feat(bench): router-backed loop executor — stateful research through the real kernel#188
drewstone merged 3 commits into
mainfrom
feat/research-loop-executor

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

What

Make the research benches run through the real stateful kernel (runLoop + createDynamicDriver) — multi-round, analyst-steered, off-sandbox — instead of the one-shot RAG pool.

  • router-executor.ts — a router-backed LoopSandboxClient, the "router" cost-dial the one-flow header already names ("backend = the injected LoopSandboxClient (router / local-bridge / sandbox)"). Each streamPrompt = one research shot, off-sandbox. The kernel never branches on backend kind, so this drops in and the full loop (rounds + steering) runs with search working and no sandbox — the in-box egress allowlist (ops-board Suspend and resume: a node that parks the run, with the host owning persistence and wake #976) is irrelevant to research.
  • research-shot.ts — extracts the retrieve→answer body (runResearchShot) into a shared primitive, so the flat RAG worker (research-gate) and the kernel loop (research-loop) score the identical body.
  • research-loop.mts — the stateful runner: runExperiment with blind / analyst-steered / aggressive arms over ROUNDS.

Why

The provider leaderboard was one-shot RAG (K=1, no rounds, no resume) — it silently abandoned the stateful loop that's proven on EOPS/commit0. Research is retrieval, not in-box code execution, so it never needed a box; routing the executor through the router lets the same kernel drive it. This is the kernel's own interface (the anticipated LeafExecutor), not a shim.

Verification

  • All four files tsc-clean (full bench tsc: 0 errors).
  • SimpleQA smoke (real kernel, 3 arms × rounds): search-backed answers resolve; the loop adaptively stops at round 0 on success — correct depth behavior, and it keeps blind and the steered arms compute-matched.
  • finsearch smoke (decisive) — all records run 2 real rounds (round 0 fails → loop continues). The analyst arm reshapes round-1 prompts (312→533 chars; findings spliced in), aggressive-push reshapes (312→452, 641→781), and blind stays byte-identical across rounds — the clean unsteered equal-compute control. searches=1/round confirms router search firing off-sandbox.

The steering-vs-blind delta on research is now the experiment to run at real n; this PR delivers the mechanism.

Stacking

Stacked on #186 (base feat/research-leaderboard) — it refactors that PR's research-gate.mts onto the shared research-shot. Merge #186 first.

…un env passthrough
research-gate.mts: off-sandbox research-bench leaderboard (model x web-search-provider x multi-shot) over the router -- provider-pinned /v1/search + web_fetch, then answer. Deep-cleaned onto the kernel primitives (routerChatWithUsage, runPool, appendRunRecord, adapter.judge); deleted the reinvented pool/corpus/sandbox backends. 424 -> 259 lines.
experiment.ts: sandboxAgentRun gains an optional env passthrough (merged onto OPENAI_*), letting a caller pin the in-box agent search provider (TANGLE_SEARCH_DEFAULT_PROVIDER). rsi.ts forwards SEARCH / EXA_API_KEY to the box via it.
Verified: tsc clean; SimpleQA you-arm reproduces 2/2 through the cleaned worker.
…the real kernel
Research benches run through runLoop + createDynamicDriver (multi-round, analyst-steered) instead of the one-shot RAG pool. A router-backed LoopSandboxClient serves each research shot off-sandbox, so the kernel drives full rounds with search working and no sandbox dependency (the in-box egress allowlist, ops-board 976, is irrelevant to research). Extract runResearchShot into a shared research-shot module so the flat RAG worker (research-gate) and the kernel loop (research-loop) score the identical retrieve-answer body. Verified: finsearch smoke runs 2 real rounds with the analyst reshaping round-1 prompts and blind held unsteered (clean equal-compute control).
@tangletools

Copy link
Copy Markdown
Contributor

✅ No Blockers — b6cf2d98

Readiness 83/100 · Confidence 65/100 · 4 findings (4 low)

deepseek: Correctness 83 · Security 83 · Testing 83 · Architecture 83

Full multi-shot audit completed 1/1 planned shots over 4 changed files. Global verifier still owns final merge decision.

🟡 LOW No tests for the 4 changed bench files — bench/src/research-shot.ts

None of the 4 files in this shot (router-executor.ts, research-shot.ts, research-gate.mts, research-loop.mts) have dedicated tests. The closest coverage is experiment.test.mts which exercises runExperiment with a mock LoopSandboxClient — this indirectly covers the kernel path but doesn't test runResearchShot's search/answer logic, the gate's Phase 1/Phase 2 pipeline, or the router-executor's event shape. Risk: behavioral regressions in runResearchShot (the shared primitive) would only be caught by manual bench runs.

🟡 LOW Token usage from routerChatWithUsage dropped; kernel/lab sees zero usage — bench/src/research-shot.ts

runResearchShot (line 112) destructures only content from routerChatWithUsage(), dropping the usage and costUsd fields. The Shot type (line 28) has no usage fields. router-executor.tsline 32-34 emits only finalText, success, searches — no done event or data.usage. The kernel's extractLlmCallEvent (src/runtime/sandbox-events.ts:46) looks for type === 'result' events with `data.usa

🟡 LOW Shot.taskId set to box ID (router-research-N) rather than benchmark task ID in kernel path — bench/src/router-executor.ts

Line 31: runResearchShot(message, id, 0, cfg) uses the sandbox instance ID (router-research-0, etc.) as taskId. For the research-loop.mts kernel path this is harmless — the kernel tracks tasks by benchmark instanceId, not the Shot's internal taskId. But any debug logging or trace that reads Shot.taskId will see box IDs instead of benchmark task IDs, making it harder to correlate shots to tasks during troubleshooting. The research-gate.mts path correctly passes u.task.id (the benchmark task ID).

🟡 LOWas unknown as SandboxInstance cast skips type-checking on missing fields — bench/src/router-executor.ts

Line 38: as unknown as SandboxInstance suppresses TS checking. The returned object has only id, streamPrompt, delete but SandboxInstance from @tangle-network/sandbox likely requires status, events(), refresh(), sendCommand(), etc. The existing bench test mock (experiment.test.mtsline 17) uses Promise<any> for the same reason — an established pattern. The kernel only calls id, streamPrompt, delete on the box, so this is safe TODAY. Risk: if the kernel or runExperiment ever calls status or refresh()


tangletools · 2026-06-06T23:28:27Z · trace

tangletools
tangletools previously approved these changes Jun 6, 2026

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Approved — 4 non-blocking findings — b6cf2d98

Full multi-shot audit completed 1/1 planned shots over 4 changed files. Global verifier still owns final merge decision.

Full immutable report for this review: trace

Summary comment for this run: full summary


tangletools · 2026-06-06T23:28:27Z · immutable trace

@drewstone
drewstone changed the base branch from feat/research-leaderboard to mainJune 7, 2026 12:25
@drewstone
drewstone dismissed tangletools’s stale reviewJune 7, 2026 12:25

The base branch was changed.

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Refreshed approval after new commits — 8c2b7acc

A previous trusted approval on this PR was invalidated by new commits.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: stale_approval_refresh · 2026-06-07T12:29:11Z

@drewstone
drewstone merged commit d8e1032 into mainJun 7, 2026
1 check passed
@drewstone
drewstone deleted the feat/research-loop-executor branch June 7, 2026 12:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

feat(bench): router-backed loop executor — stateful research through the real kernel - #188

Merged
drewstone merged 3 commits into
mainfrom
feat/research-loop-executor
Jun 7, 2026
Merged

feat(bench): router-backed loop executor — stateful research through the real kernel#188
drewstone merged 3 commits into
mainfrom
feat/research-loop-executor

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

What

Make the research benches run through the real stateful kernel (runLoop + createDynamicDriver) — multi-round, analyst-steered, off-sandbox — instead of the one-shot RAG pool.

  • router-executor.ts — a router-backed LoopSandboxClient, the "router" cost-dial the one-flow header already names ("backend = the injected LoopSandboxClient (router / local-bridge / sandbox)"). Each streamPrompt = one research shot, off-sandbox. The kernel never branches on backend kind, so this drops in and the full loop (rounds + steering) runs with search working and no sandbox — the in-box egress allowlist (ops-board Suspend and resume: a node that parks the run, with the host owning persistence and wake #976) is irrelevant to research.
  • research-shot.ts — extracts the retrieve→answer body (runResearchShot) into a shared primitive, so the flat RAG worker (research-gate) and the kernel loop (research-loop) score the identical body.
  • research-loop.mts — the stateful runner: runExperiment with blind / analyst-steered / aggressive arms over ROUNDS.

Why

The provider leaderboard was one-shot RAG (K=1, no rounds, no resume) — it silently abandoned the stateful loop that's proven on EOPS/commit0. Research is retrieval, not in-box code execution, so it never needed a box; routing the executor through the router lets the same kernel drive it. This is the kernel's own interface (the anticipated LeafExecutor), not a shim.

Verification

  • All four files tsc-clean (full bench tsc: 0 errors).
  • SimpleQA smoke (real kernel, 3 arms × rounds): search-backed answers resolve; the loop adaptively stops at round 0 on success — correct depth behavior, and it keeps blind and the steered arms compute-matched.
  • finsearch smoke (decisive) — all records run 2 real rounds (round 0 fails → loop continues). The analyst arm reshapes round-1 prompts (312→533 chars; findings spliced in), aggressive-push reshapes (312→452, 641→781), and blind stays byte-identical across rounds — the clean unsteered equal-compute control. searches=1/round confirms router search firing off-sandbox.

The steering-vs-blind delta on research is now the experiment to run at real n; this PR delivers the mechanism.

Stacking

Stacked on #186 (base feat/research-leaderboard) — it refactors that PR's research-gate.mts onto the shared research-shot. Merge #186 first.

…un env passthrough
research-gate.mts: off-sandbox research-bench leaderboard (model x web-search-provider x multi-shot) over the router -- provider-pinned /v1/search + web_fetch, then answer. Deep-cleaned onto the kernel primitives (routerChatWithUsage, runPool, appendRunRecord, adapter.judge); deleted the reinvented pool/corpus/sandbox backends. 424 -> 259 lines.
experiment.ts: sandboxAgentRun gains an optional env passthrough (merged onto OPENAI_*), letting a caller pin the in-box agent search provider (TANGLE_SEARCH_DEFAULT_PROVIDER). rsi.ts forwards SEARCH / EXA_API_KEY to the box via it.
Verified: tsc clean; SimpleQA you-arm reproduces 2/2 through the cleaned worker.
…the real kernel
Research benches run through runLoop + createDynamicDriver (multi-round, analyst-steered) instead of the one-shot RAG pool. A router-backed LoopSandboxClient serves each research shot off-sandbox, so the kernel drives full rounds with search working and no sandbox dependency (the in-box egress allowlist, ops-board 976, is irrelevant to research). Extract runResearchShot into a shared research-shot module so the flat RAG worker (research-gate) and the kernel loop (research-loop) score the identical retrieve-answer body. Verified: finsearch smoke runs 2 real rounds with the analyst reshaping round-1 prompts and blind held unsteered (clean equal-compute control).
@tangletools

Copy link
Copy Markdown
Contributor

✅ No Blockers — b6cf2d98

Readiness 83/100 · Confidence 65/100 · 4 findings (4 low)

deepseek: Correctness 83 · Security 83 · Testing 83 · Architecture 83

Full multi-shot audit completed 1/1 planned shots over 4 changed files. Global verifier still owns final merge decision.

🟡 LOW No tests for the 4 changed bench files — bench/src/research-shot.ts

None of the 4 files in this shot (router-executor.ts, research-shot.ts, research-gate.mts, research-loop.mts) have dedicated tests. The closest coverage is experiment.test.mts which exercises runExperiment with a mock LoopSandboxClient — this indirectly covers the kernel path but doesn't test runResearchShot's search/answer logic, the gate's Phase 1/Phase 2 pipeline, or the router-executor's event shape. Risk: behavioral regressions in runResearchShot (the shared primitive) would only be caught by manual bench runs.

🟡 LOW Token usage from routerChatWithUsage dropped; kernel/lab sees zero usage — bench/src/research-shot.ts

runResearchShot (line 112) destructures only content from routerChatWithUsage(), dropping the usage and costUsd fields. The Shot type (line 28) has no usage fields. router-executor.tsline 32-34 emits only finalText, success, searches — no done event or data.usage. The kernel's extractLlmCallEvent (src/runtime/sandbox-events.ts:46) looks for type === 'result' events with `data.usa

🟡 LOW Shot.taskId set to box ID (router-research-N) rather than benchmark task ID in kernel path — bench/src/router-executor.ts

Line 31: runResearchShot(message, id, 0, cfg) uses the sandbox instance ID (router-research-0, etc.) as taskId. For the research-loop.mts kernel path this is harmless — the kernel tracks tasks by benchmark instanceId, not the Shot's internal taskId. But any debug logging or trace that reads Shot.taskId will see box IDs instead of benchmark task IDs, making it harder to correlate shots to tasks during troubleshooting. The research-gate.mts path correctly passes u.task.id (the benchmark task ID).

🟡 LOWas unknown as SandboxInstance cast skips type-checking on missing fields — bench/src/router-executor.ts

Line 38: as unknown as SandboxInstance suppresses TS checking. The returned object has only id, streamPrompt, delete but SandboxInstance from @tangle-network/sandbox likely requires status, events(), refresh(), sendCommand(), etc. The existing bench test mock (experiment.test.mtsline 17) uses Promise<any> for the same reason — an established pattern. The kernel only calls id, streamPrompt, delete on the box, so this is safe TODAY. Risk: if the kernel or runExperiment ever calls status or refresh()


tangletools · 2026-06-06T23:28:27Z · trace

tangletools
tangletools previously approved these changes Jun 6, 2026

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Approved — 4 non-blocking findings — b6cf2d98

Full multi-shot audit completed 1/1 planned shots over 4 changed files. Global verifier still owns final merge decision.

Full immutable report for this review: trace

Summary comment for this run: full summary


tangletools · 2026-06-06T23:28:27Z · immutable trace

@drewstone
drewstone changed the base branch from feat/research-leaderboard to mainJune 7, 2026 12:25
@drewstone
drewstone dismissed tangletools’s stale reviewJune 7, 2026 12:25

The base branch was changed.

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Refreshed approval after new commits — 8c2b7acc

A previous trusted approval on this PR was invalidated by new commits.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: stale_approval_refresh · 2026-06-07T12:29:11Z

@drewstone
drewstone merged commit d8e1032 into mainJun 7, 2026
1 check passed
@drewstone
drewstone deleted the feat/research-loop-executor branch June 7, 2026 12:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(bench): router-backed loop executor — stateful research through the real kernel - #188

Merged
drewstone merged 3 commits into
mainfrom
feat/research-loop-executor
Jun 7, 2026
Merged

feat(bench): router-backed loop executor — stateful research through the real kernel#188
drewstone merged 3 commits into
mainfrom
feat/research-loop-executor

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

What

Make the research benches run through the real stateful kernel (runLoop + createDynamicDriver) — multi-round, analyst-steered, off-sandbox — instead of the one-shot RAG pool.

  • router-executor.ts — a router-backed LoopSandboxClient, the "router" cost-dial the one-flow header already names ("backend = the injected LoopSandboxClient (router / local-bridge / sandbox)"). Each streamPrompt = one research shot, off-sandbox. The kernel never branches on backend kind, so this drops in and the full loop (rounds + steering) runs with search working and no sandbox — the in-box egress allowlist (ops-board Suspend and resume: a node that parks the run, with the host owning persistence and wake #976) is irrelevant to research.
  • research-shot.ts — extracts the retrieve→answer body (runResearchShot) into a shared primitive, so the flat RAG worker (research-gate) and the kernel loop (research-loop) score the identical body.
  • research-loop.mts — the stateful runner: runExperiment with blind / analyst-steered / aggressive arms over ROUNDS.

Why

The provider leaderboard was one-shot RAG (K=1, no rounds, no resume) — it silently abandoned the stateful loop that's proven on EOPS/commit0. Research is retrieval, not in-box code execution, so it never needed a box; routing the executor through the router lets the same kernel drive it. This is the kernel's own interface (the anticipated LeafExecutor), not a shim.

Verification

  • All four files tsc-clean (full bench tsc: 0 errors).
  • SimpleQA smoke (real kernel, 3 arms × rounds): search-backed answers resolve; the loop adaptively stops at round 0 on success — correct depth behavior, and it keeps blind and the steered arms compute-matched.
  • finsearch smoke (decisive) — all records run 2 real rounds (round 0 fails → loop continues). The analyst arm reshapes round-1 prompts (312→533 chars; findings spliced in), aggressive-push reshapes (312→452, 641→781), and blind stays byte-identical across rounds — the clean unsteered equal-compute control. searches=1/round confirms router search firing off-sandbox.

The steering-vs-blind delta on research is now the experiment to run at real n; this PR delivers the mechanism.

Stacking

Stacked on #186 (base feat/research-leaderboard) — it refactors that PR's research-gate.mts onto the shared research-shot. Merge #186 first.

…un env passthrough
research-gate.mts: off-sandbox research-bench leaderboard (model x web-search-provider x multi-shot) over the router -- provider-pinned /v1/search + web_fetch, then answer. Deep-cleaned onto the kernel primitives (routerChatWithUsage, runPool, appendRunRecord, adapter.judge); deleted the reinvented pool/corpus/sandbox backends. 424 -> 259 lines.
experiment.ts: sandboxAgentRun gains an optional env passthrough (merged onto OPENAI_*), letting a caller pin the in-box agent search provider (TANGLE_SEARCH_DEFAULT_PROVIDER). rsi.ts forwards SEARCH / EXA_API_KEY to the box via it.
Verified: tsc clean; SimpleQA you-arm reproduces 2/2 through the cleaned worker.
…the real kernel
Research benches run through runLoop + createDynamicDriver (multi-round, analyst-steered) instead of the one-shot RAG pool. A router-backed LoopSandboxClient serves each research shot off-sandbox, so the kernel drives full rounds with search working and no sandbox dependency (the in-box egress allowlist, ops-board 976, is irrelevant to research). Extract runResearchShot into a shared research-shot module so the flat RAG worker (research-gate) and the kernel loop (research-loop) score the identical retrieve-answer body. Verified: finsearch smoke runs 2 real rounds with the analyst reshaping round-1 prompts and blind held unsteered (clean equal-compute control).
@tangletools

Copy link
Copy Markdown
Contributor

✅ No Blockers — b6cf2d98

Readiness 83/100 · Confidence 65/100 · 4 findings (4 low)

deepseek: Correctness 83 · Security 83 · Testing 83 · Architecture 83

Full multi-shot audit completed 1/1 planned shots over 4 changed files. Global verifier still owns final merge decision.

🟡 LOW No tests for the 4 changed bench files — bench/src/research-shot.ts

None of the 4 files in this shot (router-executor.ts, research-shot.ts, research-gate.mts, research-loop.mts) have dedicated tests. The closest coverage is experiment.test.mts which exercises runExperiment with a mock LoopSandboxClient — this indirectly covers the kernel path but doesn't test runResearchShot's search/answer logic, the gate's Phase 1/Phase 2 pipeline, or the router-executor's event shape. Risk: behavioral regressions in runResearchShot (the shared primitive) would only be caught by manual bench runs.

🟡 LOW Token usage from routerChatWithUsage dropped; kernel/lab sees zero usage — bench/src/research-shot.ts

runResearchShot (line 112) destructures only content from routerChatWithUsage(), dropping the usage and costUsd fields. The Shot type (line 28) has no usage fields. router-executor.tsline 32-34 emits only finalText, success, searches — no done event or data.usage. The kernel's extractLlmCallEvent (src/runtime/sandbox-events.ts:46) looks for type === 'result' events with `data.usa

🟡 LOW Shot.taskId set to box ID (router-research-N) rather than benchmark task ID in kernel path — bench/src/router-executor.ts

Line 31: runResearchShot(message, id, 0, cfg) uses the sandbox instance ID (router-research-0, etc.) as taskId. For the research-loop.mts kernel path this is harmless — the kernel tracks tasks by benchmark instanceId, not the Shot's internal taskId. But any debug logging or trace that reads Shot.taskId will see box IDs instead of benchmark task IDs, making it harder to correlate shots to tasks during troubleshooting. The research-gate.mts path correctly passes u.task.id (the benchmark task ID).

🟡 LOWas unknown as SandboxInstance cast skips type-checking on missing fields — bench/src/router-executor.ts

Line 38: as unknown as SandboxInstance suppresses TS checking. The returned object has only id, streamPrompt, delete but SandboxInstance from @tangle-network/sandbox likely requires status, events(), refresh(), sendCommand(), etc. The existing bench test mock (experiment.test.mtsline 17) uses Promise<any> for the same reason — an established pattern. The kernel only calls id, streamPrompt, delete on the box, so this is safe TODAY. Risk: if the kernel or runExperiment ever calls status or refresh()


tangletools · 2026-06-06T23:28:27Z · trace

tangletools
tangletools previously approved these changes Jun 6, 2026

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Approved — 4 non-blocking findings — b6cf2d98

Full multi-shot audit completed 1/1 planned shots over 4 changed files. Global verifier still owns final merge decision.

Full immutable report for this review: trace

Summary comment for this run: full summary


tangletools · 2026-06-06T23:28:27Z · immutable trace

@drewstone
drewstone changed the base branch from feat/research-leaderboard to mainJune 7, 2026 12:25
@drewstone
drewstone dismissed tangletools’s stale reviewJune 7, 2026 12:25

The base branch was changed.

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Refreshed approval after new commits — 8c2b7acc

A previous trusted approval on this PR was invalidated by new commits.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: stale_approval_refresh · 2026-06-07T12:29:11Z

@drewstone
drewstone merged commit d8e1032 into mainJun 7, 2026
1 check passed
@drewstone
drewstone deleted the feat/research-loop-executor branch June 7, 2026 12:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(bench): router-backed loop executor — stateful research through the real kernel - #188

Merged
drewstone merged 3 commits into
mainfrom
feat/research-loop-executor
Jun 7, 2026
Merged

feat(bench): router-backed loop executor — stateful research through the real kernel#188
drewstone merged 3 commits into
mainfrom
feat/research-loop-executor

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

What

Make the research benches run through the real stateful kernel (runLoop + createDynamicDriver) — multi-round, analyst-steered, off-sandbox — instead of the one-shot RAG pool.

  • router-executor.ts — a router-backed LoopSandboxClient, the "router" cost-dial the one-flow header already names ("backend = the injected LoopSandboxClient (router / local-bridge / sandbox)"). Each streamPrompt = one research shot, off-sandbox. The kernel never branches on backend kind, so this drops in and the full loop (rounds + steering) runs with search working and no sandbox — the in-box egress allowlist (ops-board Suspend and resume: a node that parks the run, with the host owning persistence and wake #976) is irrelevant to research.
  • research-shot.ts — extracts the retrieve→answer body (runResearchShot) into a shared primitive, so the flat RAG worker (research-gate) and the kernel loop (research-loop) score the identical body.
  • research-loop.mts — the stateful runner: runExperiment with blind / analyst-steered / aggressive arms over ROUNDS.

Why

The provider leaderboard was one-shot RAG (K=1, no rounds, no resume) — it silently abandoned the stateful loop that's proven on EOPS/commit0. Research is retrieval, not in-box code execution, so it never needed a box; routing the executor through the router lets the same kernel drive it. This is the kernel's own interface (the anticipated LeafExecutor), not a shim.

Verification

  • All four files tsc-clean (full bench tsc: 0 errors).
  • SimpleQA smoke (real kernel, 3 arms × rounds): search-backed answers resolve; the loop adaptively stops at round 0 on success — correct depth behavior, and it keeps blind and the steered arms compute-matched.
  • finsearch smoke (decisive) — all records run 2 real rounds (round 0 fails → loop continues). The analyst arm reshapes round-1 prompts (312→533 chars; findings spliced in), aggressive-push reshapes (312→452, 641→781), and blind stays byte-identical across rounds — the clean unsteered equal-compute control. searches=1/round confirms router search firing off-sandbox.

The steering-vs-blind delta on research is now the experiment to run at real n; this PR delivers the mechanism.

Stacking

Stacked on #186 (base feat/research-leaderboard) — it refactors that PR's research-gate.mts onto the shared research-shot. Merge #186 first.

…un env passthrough
research-gate.mts: off-sandbox research-bench leaderboard (model x web-search-provider x multi-shot) over the router -- provider-pinned /v1/search + web_fetch, then answer. Deep-cleaned onto the kernel primitives (routerChatWithUsage, runPool, appendRunRecord, adapter.judge); deleted the reinvented pool/corpus/sandbox backends. 424 -> 259 lines.
experiment.ts: sandboxAgentRun gains an optional env passthrough (merged onto OPENAI_*), letting a caller pin the in-box agent search provider (TANGLE_SEARCH_DEFAULT_PROVIDER). rsi.ts forwards SEARCH / EXA_API_KEY to the box via it.
Verified: tsc clean; SimpleQA you-arm reproduces 2/2 through the cleaned worker.
…the real kernel
Research benches run through runLoop + createDynamicDriver (multi-round, analyst-steered) instead of the one-shot RAG pool. A router-backed LoopSandboxClient serves each research shot off-sandbox, so the kernel drives full rounds with search working and no sandbox dependency (the in-box egress allowlist, ops-board 976, is irrelevant to research). Extract runResearchShot into a shared research-shot module so the flat RAG worker (research-gate) and the kernel loop (research-loop) score the identical retrieve-answer body. Verified: finsearch smoke runs 2 real rounds with the analyst reshaping round-1 prompts and blind held unsteered (clean equal-compute control).
@tangletools

Copy link
Copy Markdown
Contributor

✅ No Blockers — b6cf2d98

Readiness 83/100 · Confidence 65/100 · 4 findings (4 low)

deepseek: Correctness 83 · Security 83 · Testing 83 · Architecture 83

Full multi-shot audit completed 1/1 planned shots over 4 changed files. Global verifier still owns final merge decision.

🟡 LOW No tests for the 4 changed bench files — bench/src/research-shot.ts

None of the 4 files in this shot (router-executor.ts, research-shot.ts, research-gate.mts, research-loop.mts) have dedicated tests. The closest coverage is experiment.test.mts which exercises runExperiment with a mock LoopSandboxClient — this indirectly covers the kernel path but doesn't test runResearchShot's search/answer logic, the gate's Phase 1/Phase 2 pipeline, or the router-executor's event shape. Risk: behavioral regressions in runResearchShot (the shared primitive) would only be caught by manual bench runs.

🟡 LOW Token usage from routerChatWithUsage dropped; kernel/lab sees zero usage — bench/src/research-shot.ts

runResearchShot (line 112) destructures only content from routerChatWithUsage(), dropping the usage and costUsd fields. The Shot type (line 28) has no usage fields. router-executor.tsline 32-34 emits only finalText, success, searches — no done event or data.usage. The kernel's extractLlmCallEvent (src/runtime/sandbox-events.ts:46) looks for type === 'result' events with `data.usa

🟡 LOW Shot.taskId set to box ID (router-research-N) rather than benchmark task ID in kernel path — bench/src/router-executor.ts

Line 31: runResearchShot(message, id, 0, cfg) uses the sandbox instance ID (router-research-0, etc.) as taskId. For the research-loop.mts kernel path this is harmless — the kernel tracks tasks by benchmark instanceId, not the Shot's internal taskId. But any debug logging or trace that reads Shot.taskId will see box IDs instead of benchmark task IDs, making it harder to correlate shots to tasks during troubleshooting. The research-gate.mts path correctly passes u.task.id (the benchmark task ID).

🟡 LOWas unknown as SandboxInstance cast skips type-checking on missing fields — bench/src/router-executor.ts

Line 38: as unknown as SandboxInstance suppresses TS checking. The returned object has only id, streamPrompt, delete but SandboxInstance from @tangle-network/sandbox likely requires status, events(), refresh(), sendCommand(), etc. The existing bench test mock (experiment.test.mtsline 17) uses Promise<any> for the same reason — an established pattern. The kernel only calls id, streamPrompt, delete on the box, so this is safe TODAY. Risk: if the kernel or runExperiment ever calls status or refresh()


tangletools · 2026-06-06T23:28:27Z · trace

tangletools
tangletools previously approved these changes Jun 6, 2026

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Approved — 4 non-blocking findings — b6cf2d98

Full multi-shot audit completed 1/1 planned shots over 4 changed files. Global verifier still owns final merge decision.

Full immutable report for this review: trace

Summary comment for this run: full summary


tangletools · 2026-06-06T23:28:27Z · immutable trace

@drewstone
drewstone changed the base branch from feat/research-leaderboard to mainJune 7, 2026 12:25
@drewstone
drewstone dismissed tangletools’s stale reviewJune 7, 2026 12:25

The base branch was changed.

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Refreshed approval after new commits — 8c2b7acc

A previous trusted approval on this PR was invalidated by new commits.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: stale_approval_refresh · 2026-06-07T12:29:11Z

@drewstone
drewstone merged commit d8e1032 into mainJun 7, 2026
1 check passed
@drewstone
drewstone deleted the feat/research-loop-executor branch June 7, 2026 12:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

feat(bench): router-backed loop executor — stateful research through the real kernel - #188

Merged
drewstone merged 3 commits into
mainfrom
feat/research-loop-executor
Jun 7, 2026
Merged

feat(bench): router-backed loop executor — stateful research through the real kernel#188
drewstone merged 3 commits into
mainfrom
feat/research-loop-executor

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

What

Make the research benches run through the real stateful kernel (runLoop + createDynamicDriver) — multi-round, analyst-steered, off-sandbox — instead of the one-shot RAG pool.

  • router-executor.ts — a router-backed LoopSandboxClient, the "router" cost-dial the one-flow header already names ("backend = the injected LoopSandboxClient (router / local-bridge / sandbox)"). Each streamPrompt = one research shot, off-sandbox. The kernel never branches on backend kind, so this drops in and the full loop (rounds + steering) runs with search working and no sandbox — the in-box egress allowlist (ops-board Suspend and resume: a node that parks the run, with the host owning persistence and wake #976) is irrelevant to research.
  • research-shot.ts — extracts the retrieve→answer body (runResearchShot) into a shared primitive, so the flat RAG worker (research-gate) and the kernel loop (research-loop) score the identical body.
  • research-loop.mts — the stateful runner: runExperiment with blind / analyst-steered / aggressive arms over ROUNDS.

Why

The provider leaderboard was one-shot RAG (K=1, no rounds, no resume) — it silently abandoned the stateful loop that's proven on EOPS/commit0. Research is retrieval, not in-box code execution, so it never needed a box; routing the executor through the router lets the same kernel drive it. This is the kernel's own interface (the anticipated LeafExecutor), not a shim.

Verification

  • All four files tsc-clean (full bench tsc: 0 errors).
  • SimpleQA smoke (real kernel, 3 arms × rounds): search-backed answers resolve; the loop adaptively stops at round 0 on success — correct depth behavior, and it keeps blind and the steered arms compute-matched.
  • finsearch smoke (decisive) — all records run 2 real rounds (round 0 fails → loop continues). The analyst arm reshapes round-1 prompts (312→533 chars; findings spliced in), aggressive-push reshapes (312→452, 641→781), and blind stays byte-identical across rounds — the clean unsteered equal-compute control. searches=1/round confirms router search firing off-sandbox.

The steering-vs-blind delta on research is now the experiment to run at real n; this PR delivers the mechanism.

Stacking

Stacked on #186 (base feat/research-leaderboard) — it refactors that PR's research-gate.mts onto the shared research-shot. Merge #186 first.

…un env passthrough
research-gate.mts: off-sandbox research-bench leaderboard (model x web-search-provider x multi-shot) over the router -- provider-pinned /v1/search + web_fetch, then answer. Deep-cleaned onto the kernel primitives (routerChatWithUsage, runPool, appendRunRecord, adapter.judge); deleted the reinvented pool/corpus/sandbox backends. 424 -> 259 lines.
experiment.ts: sandboxAgentRun gains an optional env passthrough (merged onto OPENAI_*), letting a caller pin the in-box agent search provider (TANGLE_SEARCH_DEFAULT_PROVIDER). rsi.ts forwards SEARCH / EXA_API_KEY to the box via it.
Verified: tsc clean; SimpleQA you-arm reproduces 2/2 through the cleaned worker.
…the real kernel
Research benches run through runLoop + createDynamicDriver (multi-round, analyst-steered) instead of the one-shot RAG pool. A router-backed LoopSandboxClient serves each research shot off-sandbox, so the kernel drives full rounds with search working and no sandbox dependency (the in-box egress allowlist, ops-board 976, is irrelevant to research). Extract runResearchShot into a shared research-shot module so the flat RAG worker (research-gate) and the kernel loop (research-loop) score the identical retrieve-answer body. Verified: finsearch smoke runs 2 real rounds with the analyst reshaping round-1 prompts and blind held unsteered (clean equal-compute control).
@tangletools

Copy link
Copy Markdown
Contributor

✅ No Blockers — b6cf2d98

Readiness 83/100 · Confidence 65/100 · 4 findings (4 low)

deepseek: Correctness 83 · Security 83 · Testing 83 · Architecture 83

Full multi-shot audit completed 1/1 planned shots over 4 changed files. Global verifier still owns final merge decision.

🟡 LOW No tests for the 4 changed bench files — bench/src/research-shot.ts

None of the 4 files in this shot (router-executor.ts, research-shot.ts, research-gate.mts, research-loop.mts) have dedicated tests. The closest coverage is experiment.test.mts which exercises runExperiment with a mock LoopSandboxClient — this indirectly covers the kernel path but doesn't test runResearchShot's search/answer logic, the gate's Phase 1/Phase 2 pipeline, or the router-executor's event shape. Risk: behavioral regressions in runResearchShot (the shared primitive) would only be caught by manual bench runs.

🟡 LOW Token usage from routerChatWithUsage dropped; kernel/lab sees zero usage — bench/src/research-shot.ts

runResearchShot (line 112) destructures only content from routerChatWithUsage(), dropping the usage and costUsd fields. The Shot type (line 28) has no usage fields. router-executor.tsline 32-34 emits only finalText, success, searches — no done event or data.usage. The kernel's extractLlmCallEvent (src/runtime/sandbox-events.ts:46) looks for type === 'result' events with `data.usa

🟡 LOW Shot.taskId set to box ID (router-research-N) rather than benchmark task ID in kernel path — bench/src/router-executor.ts

Line 31: runResearchShot(message, id, 0, cfg) uses the sandbox instance ID (router-research-0, etc.) as taskId. For the research-loop.mts kernel path this is harmless — the kernel tracks tasks by benchmark instanceId, not the Shot's internal taskId. But any debug logging or trace that reads Shot.taskId will see box IDs instead of benchmark task IDs, making it harder to correlate shots to tasks during troubleshooting. The research-gate.mts path correctly passes u.task.id (the benchmark task ID).

🟡 LOWas unknown as SandboxInstance cast skips type-checking on missing fields — bench/src/router-executor.ts

Line 38: as unknown as SandboxInstance suppresses TS checking. The returned object has only id, streamPrompt, delete but SandboxInstance from @tangle-network/sandbox likely requires status, events(), refresh(), sendCommand(), etc. The existing bench test mock (experiment.test.mtsline 17) uses Promise<any> for the same reason — an established pattern. The kernel only calls id, streamPrompt, delete on the box, so this is safe TODAY. Risk: if the kernel or runExperiment ever calls status or refresh()


tangletools · 2026-06-06T23:28:27Z · trace

tangletools
tangletools previously approved these changes Jun 6, 2026

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Approved — 4 non-blocking findings — b6cf2d98

Full multi-shot audit completed 1/1 planned shots over 4 changed files. Global verifier still owns final merge decision.

Full immutable report for this review: trace

Summary comment for this run: full summary


tangletools · 2026-06-06T23:28:27Z · immutable trace

@drewstone
drewstone changed the base branch from feat/research-leaderboard to mainJune 7, 2026 12:25
@drewstone
drewstone dismissed tangletools’s stale reviewJune 7, 2026 12:25

The base branch was changed.

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Refreshed approval after new commits — 8c2b7acc

A previous trusted approval on this PR was invalidated by new commits.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: stale_approval_refresh · 2026-06-07T12:29:11Z

@drewstone
drewstone merged commit d8e1032 into mainJun 7, 2026
1 check passed
@drewstone
drewstone deleted the feat/research-loop-executor branch June 7, 2026 12:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(bench): router-backed loop executor — stateful research through the real kernel - #188

Merged
drewstone merged 3 commits into
mainfrom
feat/research-loop-executor
Jun 7, 2026
Merged

feat(bench): router-backed loop executor — stateful research through the real kernel#188
drewstone merged 3 commits into
mainfrom
feat/research-loop-executor

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

What

Make the research benches run through the real stateful kernel (runLoop + createDynamicDriver) — multi-round, analyst-steered, off-sandbox — instead of the one-shot RAG pool.

  • router-executor.ts — a router-backed LoopSandboxClient, the "router" cost-dial the one-flow header already names ("backend = the injected LoopSandboxClient (router / local-bridge / sandbox)"). Each streamPrompt = one research shot, off-sandbox. The kernel never branches on backend kind, so this drops in and the full loop (rounds + steering) runs with search working and no sandbox — the in-box egress allowlist (ops-board Suspend and resume: a node that parks the run, with the host owning persistence and wake #976) is irrelevant to research.
  • research-shot.ts — extracts the retrieve→answer body (runResearchShot) into a shared primitive, so the flat RAG worker (research-gate) and the kernel loop (research-loop) score the identical body.
  • research-loop.mts — the stateful runner: runExperiment with blind / analyst-steered / aggressive arms over ROUNDS.

Why

The provider leaderboard was one-shot RAG (K=1, no rounds, no resume) — it silently abandoned the stateful loop that's proven on EOPS/commit0. Research is retrieval, not in-box code execution, so it never needed a box; routing the executor through the router lets the same kernel drive it. This is the kernel's own interface (the anticipated LeafExecutor), not a shim.

Verification

  • All four files tsc-clean (full bench tsc: 0 errors).
  • SimpleQA smoke (real kernel, 3 arms × rounds): search-backed answers resolve; the loop adaptively stops at round 0 on success — correct depth behavior, and it keeps blind and the steered arms compute-matched.
  • finsearch smoke (decisive) — all records run 2 real rounds (round 0 fails → loop continues). The analyst arm reshapes round-1 prompts (312→533 chars; findings spliced in), aggressive-push reshapes (312→452, 641→781), and blind stays byte-identical across rounds — the clean unsteered equal-compute control. searches=1/round confirms router search firing off-sandbox.

The steering-vs-blind delta on research is now the experiment to run at real n; this PR delivers the mechanism.

Stacking

Stacked on #186 (base feat/research-leaderboard) — it refactors that PR's research-gate.mts onto the shared research-shot. Merge #186 first.

…un env passthrough
research-gate.mts: off-sandbox research-bench leaderboard (model x web-search-provider x multi-shot) over the router -- provider-pinned /v1/search + web_fetch, then answer. Deep-cleaned onto the kernel primitives (routerChatWithUsage, runPool, appendRunRecord, adapter.judge); deleted the reinvented pool/corpus/sandbox backends. 424 -> 259 lines.
experiment.ts: sandboxAgentRun gains an optional env passthrough (merged onto OPENAI_*), letting a caller pin the in-box agent search provider (TANGLE_SEARCH_DEFAULT_PROVIDER). rsi.ts forwards SEARCH / EXA_API_KEY to the box via it.
Verified: tsc clean; SimpleQA you-arm reproduces 2/2 through the cleaned worker.
…the real kernel
Research benches run through runLoop + createDynamicDriver (multi-round, analyst-steered) instead of the one-shot RAG pool. A router-backed LoopSandboxClient serves each research shot off-sandbox, so the kernel drives full rounds with search working and no sandbox dependency (the in-box egress allowlist, ops-board 976, is irrelevant to research). Extract runResearchShot into a shared research-shot module so the flat RAG worker (research-gate) and the kernel loop (research-loop) score the identical retrieve-answer body. Verified: finsearch smoke runs 2 real rounds with the analyst reshaping round-1 prompts and blind held unsteered (clean equal-compute control).
@tangletools

Copy link
Copy Markdown
Contributor

✅ No Blockers — b6cf2d98

Readiness 83/100 · Confidence 65/100 · 4 findings (4 low)

deepseek: Correctness 83 · Security 83 · Testing 83 · Architecture 83

Full multi-shot audit completed 1/1 planned shots over 4 changed files. Global verifier still owns final merge decision.

🟡 LOW No tests for the 4 changed bench files — bench/src/research-shot.ts

None of the 4 files in this shot (router-executor.ts, research-shot.ts, research-gate.mts, research-loop.mts) have dedicated tests. The closest coverage is experiment.test.mts which exercises runExperiment with a mock LoopSandboxClient — this indirectly covers the kernel path but doesn't test runResearchShot's search/answer logic, the gate's Phase 1/Phase 2 pipeline, or the router-executor's event shape. Risk: behavioral regressions in runResearchShot (the shared primitive) would only be caught by manual bench runs.

🟡 LOW Token usage from routerChatWithUsage dropped; kernel/lab sees zero usage — bench/src/research-shot.ts

runResearchShot (line 112) destructures only content from routerChatWithUsage(), dropping the usage and costUsd fields. The Shot type (line 28) has no usage fields. router-executor.tsline 32-34 emits only finalText, success, searches — no done event or data.usage. The kernel's extractLlmCallEvent (src/runtime/sandbox-events.ts:46) looks for type === 'result' events with `data.usa

🟡 LOW Shot.taskId set to box ID (router-research-N) rather than benchmark task ID in kernel path — bench/src/router-executor.ts

Line 31: runResearchShot(message, id, 0, cfg) uses the sandbox instance ID (router-research-0, etc.) as taskId. For the research-loop.mts kernel path this is harmless — the kernel tracks tasks by benchmark instanceId, not the Shot's internal taskId. But any debug logging or trace that reads Shot.taskId will see box IDs instead of benchmark task IDs, making it harder to correlate shots to tasks during troubleshooting. The research-gate.mts path correctly passes u.task.id (the benchmark task ID).

🟡 LOWas unknown as SandboxInstance cast skips type-checking on missing fields — bench/src/router-executor.ts

Line 38: as unknown as SandboxInstance suppresses TS checking. The returned object has only id, streamPrompt, delete but SandboxInstance from @tangle-network/sandbox likely requires status, events(), refresh(), sendCommand(), etc. The existing bench test mock (experiment.test.mtsline 17) uses Promise<any> for the same reason — an established pattern. The kernel only calls id, streamPrompt, delete on the box, so this is safe TODAY. Risk: if the kernel or runExperiment ever calls status or refresh()


tangletools · 2026-06-06T23:28:27Z · trace

tangletools
tangletools previously approved these changes Jun 6, 2026

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Approved — 4 non-blocking findings — b6cf2d98

Full multi-shot audit completed 1/1 planned shots over 4 changed files. Global verifier still owns final merge decision.

Full immutable report for this review: trace

Summary comment for this run: full summary


tangletools · 2026-06-06T23:28:27Z · immutable trace

@drewstone
drewstone changed the base branch from feat/research-leaderboard to mainJune 7, 2026 12:25
@drewstone
drewstone dismissed tangletools’s stale reviewJune 7, 2026 12:25

The base branch was changed.

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Refreshed approval after new commits — 8c2b7acc

A previous trusted approval on this PR was invalidated by new commits.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: stale_approval_refresh · 2026-06-07T12:29:11Z

@drewstone
drewstone merged commit d8e1032 into mainJun 7, 2026
1 check passed
@drewstone
drewstone deleted the feat/research-loop-executor branch June 7, 2026 12:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(bench): router-backed loop executor — stateful research through the real kernel - #188

Merged
drewstone merged 3 commits into
mainfrom
feat/research-loop-executor
Jun 7, 2026
Merged

feat(bench): router-backed loop executor — stateful research through the real kernel#188
drewstone merged 3 commits into
mainfrom
feat/research-loop-executor

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

What

Make the research benches run through the real stateful kernel (runLoop + createDynamicDriver) — multi-round, analyst-steered, off-sandbox — instead of the one-shot RAG pool.

  • router-executor.ts — a router-backed LoopSandboxClient, the "router" cost-dial the one-flow header already names ("backend = the injected LoopSandboxClient (router / local-bridge / sandbox)"). Each streamPrompt = one research shot, off-sandbox. The kernel never branches on backend kind, so this drops in and the full loop (rounds + steering) runs with search working and no sandbox — the in-box egress allowlist (ops-board Suspend and resume: a node that parks the run, with the host owning persistence and wake #976) is irrelevant to research.
  • research-shot.ts — extracts the retrieve→answer body (runResearchShot) into a shared primitive, so the flat RAG worker (research-gate) and the kernel loop (research-loop) score the identical body.
  • research-loop.mts — the stateful runner: runExperiment with blind / analyst-steered / aggressive arms over ROUNDS.

Why

The provider leaderboard was one-shot RAG (K=1, no rounds, no resume) — it silently abandoned the stateful loop that's proven on EOPS/commit0. Research is retrieval, not in-box code execution, so it never needed a box; routing the executor through the router lets the same kernel drive it. This is the kernel's own interface (the anticipated LeafExecutor), not a shim.

Verification

  • All four files tsc-clean (full bench tsc: 0 errors).
  • SimpleQA smoke (real kernel, 3 arms × rounds): search-backed answers resolve; the loop adaptively stops at round 0 on success — correct depth behavior, and it keeps blind and the steered arms compute-matched.
  • finsearch smoke (decisive) — all records run 2 real rounds (round 0 fails → loop continues). The analyst arm reshapes round-1 prompts (312→533 chars; findings spliced in), aggressive-push reshapes (312→452, 641→781), and blind stays byte-identical across rounds — the clean unsteered equal-compute control. searches=1/round confirms router search firing off-sandbox.

The steering-vs-blind delta on research is now the experiment to run at real n; this PR delivers the mechanism.

Stacking

Stacked on #186 (base feat/research-leaderboard) — it refactors that PR's research-gate.mts onto the shared research-shot. Merge #186 first.

…un env passthrough
research-gate.mts: off-sandbox research-bench leaderboard (model x web-search-provider x multi-shot) over the router -- provider-pinned /v1/search + web_fetch, then answer. Deep-cleaned onto the kernel primitives (routerChatWithUsage, runPool, appendRunRecord, adapter.judge); deleted the reinvented pool/corpus/sandbox backends. 424 -> 259 lines.
experiment.ts: sandboxAgentRun gains an optional env passthrough (merged onto OPENAI_*), letting a caller pin the in-box agent search provider (TANGLE_SEARCH_DEFAULT_PROVIDER). rsi.ts forwards SEARCH / EXA_API_KEY to the box via it.
Verified: tsc clean; SimpleQA you-arm reproduces 2/2 through the cleaned worker.
…the real kernel
Research benches run through runLoop + createDynamicDriver (multi-round, analyst-steered) instead of the one-shot RAG pool. A router-backed LoopSandboxClient serves each research shot off-sandbox, so the kernel drives full rounds with search working and no sandbox dependency (the in-box egress allowlist, ops-board 976, is irrelevant to research). Extract runResearchShot into a shared research-shot module so the flat RAG worker (research-gate) and the kernel loop (research-loop) score the identical retrieve-answer body. Verified: finsearch smoke runs 2 real rounds with the analyst reshaping round-1 prompts and blind held unsteered (clean equal-compute control).
@tangletools

Copy link
Copy Markdown
Contributor

✅ No Blockers — b6cf2d98

Readiness 83/100 · Confidence 65/100 · 4 findings (4 low)

deepseek: Correctness 83 · Security 83 · Testing 83 · Architecture 83

Full multi-shot audit completed 1/1 planned shots over 4 changed files. Global verifier still owns final merge decision.

🟡 LOW No tests for the 4 changed bench files — bench/src/research-shot.ts

None of the 4 files in this shot (router-executor.ts, research-shot.ts, research-gate.mts, research-loop.mts) have dedicated tests. The closest coverage is experiment.test.mts which exercises runExperiment with a mock LoopSandboxClient — this indirectly covers the kernel path but doesn't test runResearchShot's search/answer logic, the gate's Phase 1/Phase 2 pipeline, or the router-executor's event shape. Risk: behavioral regressions in runResearchShot (the shared primitive) would only be caught by manual bench runs.

🟡 LOW Token usage from routerChatWithUsage dropped; kernel/lab sees zero usage — bench/src/research-shot.ts

runResearchShot (line 112) destructures only content from routerChatWithUsage(), dropping the usage and costUsd fields. The Shot type (line 28) has no usage fields. router-executor.tsline 32-34 emits only finalText, success, searches — no done event or data.usage. The kernel's extractLlmCallEvent (src/runtime/sandbox-events.ts:46) looks for type === 'result' events with `data.usa

🟡 LOW Shot.taskId set to box ID (router-research-N) rather than benchmark task ID in kernel path — bench/src/router-executor.ts

Line 31: runResearchShot(message, id, 0, cfg) uses the sandbox instance ID (router-research-0, etc.) as taskId. For the research-loop.mts kernel path this is harmless — the kernel tracks tasks by benchmark instanceId, not the Shot's internal taskId. But any debug logging or trace that reads Shot.taskId will see box IDs instead of benchmark task IDs, making it harder to correlate shots to tasks during troubleshooting. The research-gate.mts path correctly passes u.task.id (the benchmark task ID).

🟡 LOWas unknown as SandboxInstance cast skips type-checking on missing fields — bench/src/router-executor.ts

Line 38: as unknown as SandboxInstance suppresses TS checking. The returned object has only id, streamPrompt, delete but SandboxInstance from @tangle-network/sandbox likely requires status, events(), refresh(), sendCommand(), etc. The existing bench test mock (experiment.test.mtsline 17) uses Promise<any> for the same reason — an established pattern. The kernel only calls id, streamPrompt, delete on the box, so this is safe TODAY. Risk: if the kernel or runExperiment ever calls status or refresh()


tangletools · 2026-06-06T23:28:27Z · trace

tangletools
tangletools previously approved these changes Jun 6, 2026

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Approved — 4 non-blocking findings — b6cf2d98

Full multi-shot audit completed 1/1 planned shots over 4 changed files. Global verifier still owns final merge decision.

Full immutable report for this review: trace

Summary comment for this run: full summary


tangletools · 2026-06-06T23:28:27Z · immutable trace

@drewstone
drewstone changed the base branch from feat/research-leaderboard to mainJune 7, 2026 12:25
@drewstone
drewstone dismissed tangletools’s stale reviewJune 7, 2026 12:25

The base branch was changed.

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Refreshed approval after new commits — 8c2b7acc

A previous trusted approval on this PR was invalidated by new commits.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: stale_approval_refresh · 2026-06-07T12:29:11Z

@drewstone
drewstone merged commit d8e1032 into mainJun 7, 2026
1 check passed
@drewstone
drewstone deleted the feat/research-loop-executor branch June 7, 2026 12:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

feat(bench): router-backed loop executor — stateful research through the real kernel - #188

Merged
drewstone merged 3 commits into
mainfrom
feat/research-loop-executor
Jun 7, 2026
Merged

feat(bench): router-backed loop executor — stateful research through the real kernel#188
drewstone merged 3 commits into
mainfrom
feat/research-loop-executor

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

What

Make the research benches run through the real stateful kernel (runLoop + createDynamicDriver) — multi-round, analyst-steered, off-sandbox — instead of the one-shot RAG pool.

  • router-executor.ts — a router-backed LoopSandboxClient, the "router" cost-dial the one-flow header already names ("backend = the injected LoopSandboxClient (router / local-bridge / sandbox)"). Each streamPrompt = one research shot, off-sandbox. The kernel never branches on backend kind, so this drops in and the full loop (rounds + steering) runs with search working and no sandbox — the in-box egress allowlist (ops-board Suspend and resume: a node that parks the run, with the host owning persistence and wake #976) is irrelevant to research.
  • research-shot.ts — extracts the retrieve→answer body (runResearchShot) into a shared primitive, so the flat RAG worker (research-gate) and the kernel loop (research-loop) score the identical body.
  • research-loop.mts — the stateful runner: runExperiment with blind / analyst-steered / aggressive arms over ROUNDS.

Why

The provider leaderboard was one-shot RAG (K=1, no rounds, no resume) — it silently abandoned the stateful loop that's proven on EOPS/commit0. Research is retrieval, not in-box code execution, so it never needed a box; routing the executor through the router lets the same kernel drive it. This is the kernel's own interface (the anticipated LeafExecutor), not a shim.

Verification

  • All four files tsc-clean (full bench tsc: 0 errors).
  • SimpleQA smoke (real kernel, 3 arms × rounds): search-backed answers resolve; the loop adaptively stops at round 0 on success — correct depth behavior, and it keeps blind and the steered arms compute-matched.
  • finsearch smoke (decisive) — all records run 2 real rounds (round 0 fails → loop continues). The analyst arm reshapes round-1 prompts (312→533 chars; findings spliced in), aggressive-push reshapes (312→452, 641→781), and blind stays byte-identical across rounds — the clean unsteered equal-compute control. searches=1/round confirms router search firing off-sandbox.

The steering-vs-blind delta on research is now the experiment to run at real n; this PR delivers the mechanism.

Stacking

Stacked on #186 (base feat/research-leaderboard) — it refactors that PR's research-gate.mts onto the shared research-shot. Merge #186 first.

…un env passthrough
research-gate.mts: off-sandbox research-bench leaderboard (model x web-search-provider x multi-shot) over the router -- provider-pinned /v1/search + web_fetch, then answer. Deep-cleaned onto the kernel primitives (routerChatWithUsage, runPool, appendRunRecord, adapter.judge); deleted the reinvented pool/corpus/sandbox backends. 424 -> 259 lines.
experiment.ts: sandboxAgentRun gains an optional env passthrough (merged onto OPENAI_*), letting a caller pin the in-box agent search provider (TANGLE_SEARCH_DEFAULT_PROVIDER). rsi.ts forwards SEARCH / EXA_API_KEY to the box via it.
Verified: tsc clean; SimpleQA you-arm reproduces 2/2 through the cleaned worker.
…the real kernel
Research benches run through runLoop + createDynamicDriver (multi-round, analyst-steered) instead of the one-shot RAG pool. A router-backed LoopSandboxClient serves each research shot off-sandbox, so the kernel drives full rounds with search working and no sandbox dependency (the in-box egress allowlist, ops-board 976, is irrelevant to research). Extract runResearchShot into a shared research-shot module so the flat RAG worker (research-gate) and the kernel loop (research-loop) score the identical retrieve-answer body. Verified: finsearch smoke runs 2 real rounds with the analyst reshaping round-1 prompts and blind held unsteered (clean equal-compute control).
@tangletools

Copy link
Copy Markdown
Contributor

✅ No Blockers — b6cf2d98

Readiness 83/100 · Confidence 65/100 · 4 findings (4 low)

deepseek: Correctness 83 · Security 83 · Testing 83 · Architecture 83

Full multi-shot audit completed 1/1 planned shots over 4 changed files. Global verifier still owns final merge decision.

🟡 LOW No tests for the 4 changed bench files — bench/src/research-shot.ts

None of the 4 files in this shot (router-executor.ts, research-shot.ts, research-gate.mts, research-loop.mts) have dedicated tests. The closest coverage is experiment.test.mts which exercises runExperiment with a mock LoopSandboxClient — this indirectly covers the kernel path but doesn't test runResearchShot's search/answer logic, the gate's Phase 1/Phase 2 pipeline, or the router-executor's event shape. Risk: behavioral regressions in runResearchShot (the shared primitive) would only be caught by manual bench runs.

🟡 LOW Token usage from routerChatWithUsage dropped; kernel/lab sees zero usage — bench/src/research-shot.ts

runResearchShot (line 112) destructures only content from routerChatWithUsage(), dropping the usage and costUsd fields. The Shot type (line 28) has no usage fields. router-executor.tsline 32-34 emits only finalText, success, searches — no done event or data.usage. The kernel's extractLlmCallEvent (src/runtime/sandbox-events.ts:46) looks for type === 'result' events with `data.usa

🟡 LOW Shot.taskId set to box ID (router-research-N) rather than benchmark task ID in kernel path — bench/src/router-executor.ts

Line 31: runResearchShot(message, id, 0, cfg) uses the sandbox instance ID (router-research-0, etc.) as taskId. For the research-loop.mts kernel path this is harmless — the kernel tracks tasks by benchmark instanceId, not the Shot's internal taskId. But any debug logging or trace that reads Shot.taskId will see box IDs instead of benchmark task IDs, making it harder to correlate shots to tasks during troubleshooting. The research-gate.mts path correctly passes u.task.id (the benchmark task ID).

🟡 LOWas unknown as SandboxInstance cast skips type-checking on missing fields — bench/src/router-executor.ts

Line 38: as unknown as SandboxInstance suppresses TS checking. The returned object has only id, streamPrompt, delete but SandboxInstance from @tangle-network/sandbox likely requires status, events(), refresh(), sendCommand(), etc. The existing bench test mock (experiment.test.mtsline 17) uses Promise<any> for the same reason — an established pattern. The kernel only calls id, streamPrompt, delete on the box, so this is safe TODAY. Risk: if the kernel or runExperiment ever calls status or refresh()


tangletools · 2026-06-06T23:28:27Z · trace

tangletools
tangletools previously approved these changes Jun 6, 2026

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Approved — 4 non-blocking findings — b6cf2d98

Full multi-shot audit completed 1/1 planned shots over 4 changed files. Global verifier still owns final merge decision.

Full immutable report for this review: trace

Summary comment for this run: full summary


tangletools · 2026-06-06T23:28:27Z · immutable trace

@drewstone
drewstone changed the base branch from feat/research-leaderboard to mainJune 7, 2026 12:25
@drewstone
drewstone dismissed tangletools’s stale reviewJune 7, 2026 12:25

The base branch was changed.

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Refreshed approval after new commits — 8c2b7acc

A previous trusted approval on this PR was invalidated by new commits.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: stale_approval_refresh · 2026-06-07T12:29:11Z

@drewstone
drewstone merged commit d8e1032 into mainJun 7, 2026
1 check passed
@drewstone
drewstone deleted the feat/research-loop-executor branch June 7, 2026 12:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools