Uh oh!
There was an error while loading. Please reload this page.
feat(runtime): execute and grade immutable agent candidates - #507
Conversation
Extend the existing TraceStore, durable content addressing, candidate interface, and shared profile materializer into one verify-prepare-finalize execution path.
drewstone
commented
Jul 11, 2026
@tangletools review the current head for merge. |
tangletools
left a comment
There was a problem hiding this comment.
✅ Auto-approved drewstone PR — eedc44c8
This PR was opened by the trusted drewstone account.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.
tangletools · auto-approval · reason: drewstone_author · 2026-07-11T06:00:22Z
tangletools
left a comment
There was a problem hiding this comment.
🟢 Value Audit — sound
| Verdict | sound |
| Concerns | 0 (none) |
| Heuristic | 0.0s |
| Duplication | 0.0s |
| Interrogation | 291.5s (2 bridge agents) |
| Total | 291.5s |
💰 Value — sound
Adds a durable, cross-process, exactly-once pipeline that turns a verified immutable candidate bundle into a tamper-evident V2 run receipt — a genuinely new layer that consumes the published substrate contracts rather than reinventing the in-process kernels, and it is in the repo's grain.
- What it does: A new
src/candidate-execution/subsystem (~27 files) with one atomic path:verifyAgentCandidateBundle(verify.ts:37, consumesagentCandidateBundleSchemafrom@tangle-network/agent-interface) →prepareAgentCandidateExecution(prepare.ts:95; materializes task/candidate/profile workspaces, delegates profile→harness to@tangle-network/agent-profile-materializeat prepare.ts:167-174, reserv - Goals it achieves: Exactly-once, crash-recoverable execution of sealed benchmark candidates across processes/machines without double-grading, secret leakage, or unproven/late result writes (separate frozen execution/cleanup/result budgets, lease fencing, two-phase terminal commit, zero-model-call retry eligibility — claim.ts:76-91,215-222). It binds the full evidence chain (patch, task tree, executable grader bytes,
- Assessment: Good on its merits and in-grain. It is a real capability gap the existing kernels do not fill, and it follows the repo's hard rules: domain-clean via injected adapters (CLAUDE.md layering), profile realization delegated to substrate (
materializeCandidateProfile/applyAgentCandidateWorkspacePlan, prepare.ts:25-28,167-174 — honoring the §1.5 'never write a profile→harness realizer' law), neutral - Better / existing approach: None — this is the right approach. I checked the obvious existing equivalents and none cover durable cross-process candidate-bundle execution + receipts:
src/runtime/superviseSpawnJournal/budget.ts:80anddriver-executor.tsare an in-process, single-tree content-addressed event log for equal-k replay/resume of the recursive atom (no cross-process claim/lease/terminal-CAS, no protected gra - Model: opencode/kimi-for-coding/k2p7
- Bridge attempts: 1
🎯 Usefulness — sound
Adds a coherent verify→prepare→claim→execute→finalize spine that turns a sealed candidate bundle into a tamper-evident V2 run receipt, consuming real published contracts with no existing equivalent in the codebase.
- Integration: Fully reachable: exported via
export * from './candidate-execution'at src/index.ts:35, and its symbols are already in the generated API catalog (docs/api/primitive-catalog.md:110-111,142) with the docs-freshness gate green. It consumes the published @tangle-network/agent-interface 0.22.0 contract — I confirmed `AgentCandidateRunReceiptV2 extends Omit<AgentCandidateRunReceiptV1, 'schemaVersion'> - Fit with existing patterns: Fits the codebase grain precisely. Respects the layering law (agent-runtime → agent-eval, never reverse): imports TraceStore/BenchmarkEvaluation/isLlmSpan/REDACTION_VERSION from agent-eval and contract types from agent-interface, never imports agent-knowledge. Does NOT compete with the improvement/ module's
CandidateGenerator/improvementDriver(src/improvement/improvement-driver.ts:36) — that - Real-world viability: Rigorously handles the non-happy paths. Tests cover exactly-one-concurrent-invocation (tests/candidate-execution-execute.test.ts:710), hanging process stop leaving the claim for recovery (:993), unknown settlement blocking zero-spend guessing (:1061), zero-call retry only for pre-model failures (:1126), lease expiry, cross-process claim races, and replay terminal writes — 1,420 tests pass. The exe
- Model: opencode/zai-coding-plan/glm-5.2
- Bridge attempts: 1
No concerns — sound change, no better or existing approach found. ✅
What this audit checks
It judges the change on its merits — not whether it was tasked out in an issue. Unticketed, fast-moving work is fine; the question is whether the change is good and whether a better or existing approach should be used instead.
| Pass | What it asks |
|---|---|
| Heuristic | Vague title? Whitespace-only or cruft-bearing diff? (content signals only) |
| Duplication | Do added function/class names already exist elsewhere in the repo? |
| Value Audit | What does it do? What goal does it achieve? Is it good? Better architecture or already-exists? |
| Usefulness Audit | Does it integrate and fit? Will it hold up in real use and actually get used? |
Findings are concerns, not blocks — the human reviewer decides what to do with them.
tangletools
left a comment
There was a problem hiding this comment.
✅ Auto-approved drewstone PR — a4fec22e
This PR was opened by the trusted drewstone account.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.
tangletools · auto-approval · reason: drewstone_author · 2026-07-11T06:15:17Z
tangletools
left a comment
There was a problem hiding this comment.
✅ Auto-approved drewstone PR — fa6db037
This PR was opened by the trusted drewstone account.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.
tangletools · auto-approval · reason: drewstone_author · 2026-07-11T06:19:52Z
Uh oh!
There was an error while loading. Please reload this page.
Summary
Verification
pnpm lint— 426 files cleanpnpm run typecheck— runtime and examples cleanpnpm test— 146 files, 1,420 passed, 2 skippedpnpm build— JavaScript and declarations builtpnpm verify:package— public exports loadpnpm docs:freshness— version, peer, symbol, and generated-catalog checks cleanNo benchmark model calls were made in this change.