Skip to content

feat(runtime): execute and grade immutable agent candidates - #507

Merged
drewstone merged 6 commits into
mainfrom
feat/candidate-execution-plan
Jul 11, 2026
Merged

feat(runtime): execute and grade immutable agent candidates#507
drewstone merged 6 commits into
mainfrom
feat/candidate-execution-plan

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

Summary

  • add one atomic path from a verified candidate bundle to a durable V2 run receipt
  • execute detached task/profile/code bytes behind a one-shot cross-process claim with crash recovery
  • bind the exact verified patch, task tree, executable grader bytes, raw grader output, protected trace, model ledger, and isolated memory evidence
  • enforce separate frozen execution, cleanup, and result budgets with cancellation and lease fencing
  • reject secret-bearing archives/manifests before persistence and keep recovery trace access read-only
  • consume the published @tangle-network/agent-interface 0.22.0 receipt contract and profile materializer 0.3.0

Verification

  • pnpm lint — 426 files clean
  • pnpm run typecheck — runtime and examples clean
  • pnpm test — 146 files, 1,420 passed, 2 skipped
  • pnpm build — JavaScript and declarations built
  • pnpm verify:package — public exports load
  • pnpm docs:freshness — version, peer, symbol, and generated-catalog checks clean
  • hostile re-audit — 4/4 reproduced timing/persistence failures fixed; 0 late result writes

No benchmark model calls were made in this change.

Extend the existing TraceStore, durable content addressing, candidate interface, and shared profile materializer into one verify-prepare-finalize execution path.
@drewstone

Copy link
Copy Markdown
ContributorAuthor

@tangletools review the current head for merge.

tangletools
tangletools previously approved these changes Jul 11, 2026

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved drewstone PR — eedc44c8

This PR was opened by the trusted drewstone account.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: drewstone_author · 2026-07-11T06:00:22Z

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Value Audit — sound

Verdictsound
Concerns0 (none)
Heuristic0.0s
Duplication0.0s
Interrogation291.5s (2 bridge agents)
Total291.5s

💰 Value — sound

Adds a durable, cross-process, exactly-once pipeline that turns a verified immutable candidate bundle into a tamper-evident V2 run receipt — a genuinely new layer that consumes the published substrate contracts rather than reinventing the in-process kernels, and it is in the repo's grain.

  • What it does: A new src/candidate-execution/ subsystem (~27 files) with one atomic path: verifyAgentCandidateBundle (verify.ts:37, consumes agentCandidateBundleSchema from @tangle-network/agent-interface) → prepareAgentCandidateExecution (prepare.ts:95; materializes task/candidate/profile workspaces, delegates profile→harness to @tangle-network/agent-profile-materialize at prepare.ts:167-174, reserv
  • Goals it achieves: Exactly-once, crash-recoverable execution of sealed benchmark candidates across processes/machines without double-grading, secret leakage, or unproven/late result writes (separate frozen execution/cleanup/result budgets, lease fencing, two-phase terminal commit, zero-model-call retry eligibility — claim.ts:76-91,215-222). It binds the full evidence chain (patch, task tree, executable grader bytes,
  • Assessment: Good on its merits and in-grain. It is a real capability gap the existing kernels do not fill, and it follows the repo's hard rules: domain-clean via injected adapters (CLAUDE.md layering), profile realization delegated to substrate (materializeCandidateProfile/applyAgentCandidateWorkspacePlan, prepare.ts:25-28,167-174 — honoring the §1.5 'never write a profile→harness realizer' law), neutral
  • Better / existing approach: None — this is the right approach. I checked the obvious existing equivalents and none cover durable cross-process candidate-bundle execution + receipts: src/runtime/superviseSpawnJournal/budget.ts:80 and driver-executor.ts are an in-process, single-tree content-addressed event log for equal-k replay/resume of the recursive atom (no cross-process claim/lease/terminal-CAS, no protected gra
  • Model: opencode/kimi-for-coding/k2p7
  • Bridge attempts: 1

🎯 Usefulness — sound

Adds a coherent verify→prepare→claim→execute→finalize spine that turns a sealed candidate bundle into a tamper-evident V2 run receipt, consuming real published contracts with no existing equivalent in the codebase.

  • Integration: Fully reachable: exported via export * from './candidate-execution' at src/index.ts:35, and its symbols are already in the generated API catalog (docs/api/primitive-catalog.md:110-111,142) with the docs-freshness gate green. It consumes the published @tangle-network/agent-interface 0.22.0 contract — I confirmed `AgentCandidateRunReceiptV2 extends Omit<AgentCandidateRunReceiptV1, 'schemaVersion'>
  • Fit with existing patterns: Fits the codebase grain precisely. Respects the layering law (agent-runtime → agent-eval, never reverse): imports TraceStore/BenchmarkEvaluation/isLlmSpan/REDACTION_VERSION from agent-eval and contract types from agent-interface, never imports agent-knowledge. Does NOT compete with the improvement/ module's CandidateGenerator/improvementDriver (src/improvement/improvement-driver.ts:36) — that
  • Real-world viability: Rigorously handles the non-happy paths. Tests cover exactly-one-concurrent-invocation (tests/candidate-execution-execute.test.ts:710), hanging process stop leaving the claim for recovery (:993), unknown settlement blocking zero-spend guessing (:1061), zero-call retry only for pre-model failures (:1126), lease expiry, cross-process claim races, and replay terminal writes — 1,420 tests pass. The exe
  • Model: opencode/zai-coding-plan/glm-5.2
  • Bridge attempts: 1

No concerns — sound change, no better or existing approach found. ✅


What this audit checks

It judges the change on its merits — not whether it was tasked out in an issue. Unticketed, fast-moving work is fine; the question is whether the change is good and whether a better or existing approach should be used instead.

PassWhat it asks
HeuristicVague title? Whitespace-only or cruft-bearing diff? (content signals only)
DuplicationDo added function/class names already exist elsewhere in the repo?
Value AuditWhat does it do? What goal does it achieve? Is it good? Better architecture or already-exists?
Usefulness AuditDoes it integrate and fit? Will it hold up in real use and actually get used?

Findings are concerns, not blocks — the human reviewer decides what to do with them.

value-audit · 20260711T060700Z

tangletools
tangletools previously approved these changes Jul 11, 2026

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved drewstone PR — a4fec22e

This PR was opened by the trusted drewstone account.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: drewstone_author · 2026-07-11T06:15:17Z

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved drewstone PR — fa6db037

This PR was opened by the trusted drewstone account.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: drewstone_author · 2026-07-11T06:19:52Z

@drewstone
drewstone merged commit 164da02 into mainJul 11, 2026
2 checks passed
@drewstonedrewstone mentioned this pull request Jul 11, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
feat(runtime): execute and grade immutable agent candidates by drewstone · Pull Request #507 · tangle-network/agent-runtime · GitHub
Skip to content

feat(runtime): execute and grade immutable agent candidates - #507

Merged
drewstone merged 6 commits into
mainfrom
feat/candidate-execution-plan
Jul 11, 2026
Merged

feat(runtime): execute and grade immutable agent candidates#507
drewstone merged 6 commits into
mainfrom
feat/candidate-execution-plan

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

Summary

  • add one atomic path from a verified candidate bundle to a durable V2 run receipt
  • execute detached task/profile/code bytes behind a one-shot cross-process claim with crash recovery
  • bind the exact verified patch, task tree, executable grader bytes, raw grader output, protected trace, model ledger, and isolated memory evidence
  • enforce separate frozen execution, cleanup, and result budgets with cancellation and lease fencing
  • reject secret-bearing archives/manifests before persistence and keep recovery trace access read-only
  • consume the published @tangle-network/agent-interface 0.22.0 receipt contract and profile materializer 0.3.0

Verification

  • pnpm lint — 426 files clean
  • pnpm run typecheck — runtime and examples clean
  • pnpm test — 146 files, 1,420 passed, 2 skipped
  • pnpm build — JavaScript and declarations built
  • pnpm verify:package — public exports load
  • pnpm docs:freshness — version, peer, symbol, and generated-catalog checks clean
  • hostile re-audit — 4/4 reproduced timing/persistence failures fixed; 0 late result writes

No benchmark model calls were made in this change.

Extend the existing TraceStore, durable content addressing, candidate interface, and shared profile materializer into one verify-prepare-finalize execution path.
@drewstone

Copy link
Copy Markdown
ContributorAuthor

@tangletools review the current head for merge.

tangletools
tangletools previously approved these changes Jul 11, 2026

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved drewstone PR — eedc44c8

This PR was opened by the trusted drewstone account.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: drewstone_author · 2026-07-11T06:00:22Z

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Value Audit — sound

Verdictsound
Concerns0 (none)
Heuristic0.0s
Duplication0.0s
Interrogation291.5s (2 bridge agents)
Total291.5s

💰 Value — sound

Adds a durable, cross-process, exactly-once pipeline that turns a verified immutable candidate bundle into a tamper-evident V2 run receipt — a genuinely new layer that consumes the published substrate contracts rather than reinventing the in-process kernels, and it is in the repo's grain.

  • What it does: A new src/candidate-execution/ subsystem (~27 files) with one atomic path: verifyAgentCandidateBundle (verify.ts:37, consumes agentCandidateBundleSchema from @tangle-network/agent-interface) → prepareAgentCandidateExecution (prepare.ts:95; materializes task/candidate/profile workspaces, delegates profile→harness to @tangle-network/agent-profile-materialize at prepare.ts:167-174, reserv
  • Goals it achieves: Exactly-once, crash-recoverable execution of sealed benchmark candidates across processes/machines without double-grading, secret leakage, or unproven/late result writes (separate frozen execution/cleanup/result budgets, lease fencing, two-phase terminal commit, zero-model-call retry eligibility — claim.ts:76-91,215-222). It binds the full evidence chain (patch, task tree, executable grader bytes,
  • Assessment: Good on its merits and in-grain. It is a real capability gap the existing kernels do not fill, and it follows the repo's hard rules: domain-clean via injected adapters (CLAUDE.md layering), profile realization delegated to substrate (materializeCandidateProfile/applyAgentCandidateWorkspacePlan, prepare.ts:25-28,167-174 — honoring the §1.5 'never write a profile→harness realizer' law), neutral
  • Better / existing approach: None — this is the right approach. I checked the obvious existing equivalents and none cover durable cross-process candidate-bundle execution + receipts: src/runtime/superviseSpawnJournal/budget.ts:80 and driver-executor.ts are an in-process, single-tree content-addressed event log for equal-k replay/resume of the recursive atom (no cross-process claim/lease/terminal-CAS, no protected gra
  • Model: opencode/kimi-for-coding/k2p7
  • Bridge attempts: 1

🎯 Usefulness — sound

Adds a coherent verify→prepare→claim→execute→finalize spine that turns a sealed candidate bundle into a tamper-evident V2 run receipt, consuming real published contracts with no existing equivalent in the codebase.

  • Integration: Fully reachable: exported via export * from './candidate-execution' at src/index.ts:35, and its symbols are already in the generated API catalog (docs/api/primitive-catalog.md:110-111,142) with the docs-freshness gate green. It consumes the published @tangle-network/agent-interface 0.22.0 contract — I confirmed `AgentCandidateRunReceiptV2 extends Omit<AgentCandidateRunReceiptV1, 'schemaVersion'>
  • Fit with existing patterns: Fits the codebase grain precisely. Respects the layering law (agent-runtime → agent-eval, never reverse): imports TraceStore/BenchmarkEvaluation/isLlmSpan/REDACTION_VERSION from agent-eval and contract types from agent-interface, never imports agent-knowledge. Does NOT compete with the improvement/ module's CandidateGenerator/improvementDriver (src/improvement/improvement-driver.ts:36) — that
  • Real-world viability: Rigorously handles the non-happy paths. Tests cover exactly-one-concurrent-invocation (tests/candidate-execution-execute.test.ts:710), hanging process stop leaving the claim for recovery (:993), unknown settlement blocking zero-spend guessing (:1061), zero-call retry only for pre-model failures (:1126), lease expiry, cross-process claim races, and replay terminal writes — 1,420 tests pass. The exe
  • Model: opencode/zai-coding-plan/glm-5.2
  • Bridge attempts: 1

No concerns — sound change, no better or existing approach found. ✅


What this audit checks

It judges the change on its merits — not whether it was tasked out in an issue. Unticketed, fast-moving work is fine; the question is whether the change is good and whether a better or existing approach should be used instead.

PassWhat it asks
HeuristicVague title? Whitespace-only or cruft-bearing diff? (content signals only)
DuplicationDo added function/class names already exist elsewhere in the repo?
Value AuditWhat does it do? What goal does it achieve? Is it good? Better architecture or already-exists?
Usefulness AuditDoes it integrate and fit? Will it hold up in real use and actually get used?

Findings are concerns, not blocks — the human reviewer decides what to do with them.

value-audit · 20260711T060700Z

tangletools
tangletools previously approved these changes Jul 11, 2026

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved drewstone PR — a4fec22e

This PR was opened by the trusted drewstone account.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: drewstone_author · 2026-07-11T06:15:17Z

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved drewstone PR — fa6db037

This PR was opened by the trusted drewstone account.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: drewstone_author · 2026-07-11T06:19:52Z

@drewstone
drewstone merged commit 164da02 into mainJul 11, 2026
2 checks passed
@drewstonedrewstone mentioned this pull request Jul 11, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' feat(runtime): execute and grade immutable agent candidates by drewstone · Pull Request #507 · tangle-network/agent-runtime · GitHub
Skip to content

feat(runtime): execute and grade immutable agent candidates - #507

Merged
drewstone merged 6 commits into
mainfrom
feat/candidate-execution-plan
Jul 11, 2026
Merged

feat(runtime): execute and grade immutable agent candidates#507
drewstone merged 6 commits into
mainfrom
feat/candidate-execution-plan

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

Summary

  • add one atomic path from a verified candidate bundle to a durable V2 run receipt
  • execute detached task/profile/code bytes behind a one-shot cross-process claim with crash recovery
  • bind the exact verified patch, task tree, executable grader bytes, raw grader output, protected trace, model ledger, and isolated memory evidence
  • enforce separate frozen execution, cleanup, and result budgets with cancellation and lease fencing
  • reject secret-bearing archives/manifests before persistence and keep recovery trace access read-only
  • consume the published @tangle-network/agent-interface 0.22.0 receipt contract and profile materializer 0.3.0

Verification

  • pnpm lint — 426 files clean
  • pnpm run typecheck — runtime and examples clean
  • pnpm test — 146 files, 1,420 passed, 2 skipped
  • pnpm build — JavaScript and declarations built
  • pnpm verify:package — public exports load
  • pnpm docs:freshness — version, peer, symbol, and generated-catalog checks clean
  • hostile re-audit — 4/4 reproduced timing/persistence failures fixed; 0 late result writes

No benchmark model calls were made in this change.

Extend the existing TraceStore, durable content addressing, candidate interface, and shared profile materializer into one verify-prepare-finalize execution path.
@drewstone

Copy link
Copy Markdown
ContributorAuthor

@tangletools review the current head for merge.

tangletools
tangletools previously approved these changes Jul 11, 2026

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved drewstone PR — eedc44c8

This PR was opened by the trusted drewstone account.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: drewstone_author · 2026-07-11T06:00:22Z

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Value Audit — sound

Verdictsound
Concerns0 (none)
Heuristic0.0s
Duplication0.0s
Interrogation291.5s (2 bridge agents)
Total291.5s

💰 Value — sound

Adds a durable, cross-process, exactly-once pipeline that turns a verified immutable candidate bundle into a tamper-evident V2 run receipt — a genuinely new layer that consumes the published substrate contracts rather than reinventing the in-process kernels, and it is in the repo's grain.

  • What it does: A new src/candidate-execution/ subsystem (~27 files) with one atomic path: verifyAgentCandidateBundle (verify.ts:37, consumes agentCandidateBundleSchema from @tangle-network/agent-interface) → prepareAgentCandidateExecution (prepare.ts:95; materializes task/candidate/profile workspaces, delegates profile→harness to @tangle-network/agent-profile-materialize at prepare.ts:167-174, reserv
  • Goals it achieves: Exactly-once, crash-recoverable execution of sealed benchmark candidates across processes/machines without double-grading, secret leakage, or unproven/late result writes (separate frozen execution/cleanup/result budgets, lease fencing, two-phase terminal commit, zero-model-call retry eligibility — claim.ts:76-91,215-222). It binds the full evidence chain (patch, task tree, executable grader bytes,
  • Assessment: Good on its merits and in-grain. It is a real capability gap the existing kernels do not fill, and it follows the repo's hard rules: domain-clean via injected adapters (CLAUDE.md layering), profile realization delegated to substrate (materializeCandidateProfile/applyAgentCandidateWorkspacePlan, prepare.ts:25-28,167-174 — honoring the §1.5 'never write a profile→harness realizer' law), neutral
  • Better / existing approach: None — this is the right approach. I checked the obvious existing equivalents and none cover durable cross-process candidate-bundle execution + receipts: src/runtime/superviseSpawnJournal/budget.ts:80 and driver-executor.ts are an in-process, single-tree content-addressed event log for equal-k replay/resume of the recursive atom (no cross-process claim/lease/terminal-CAS, no protected gra
  • Model: opencode/kimi-for-coding/k2p7
  • Bridge attempts: 1

🎯 Usefulness — sound

Adds a coherent verify→prepare→claim→execute→finalize spine that turns a sealed candidate bundle into a tamper-evident V2 run receipt, consuming real published contracts with no existing equivalent in the codebase.

  • Integration: Fully reachable: exported via export * from './candidate-execution' at src/index.ts:35, and its symbols are already in the generated API catalog (docs/api/primitive-catalog.md:110-111,142) with the docs-freshness gate green. It consumes the published @tangle-network/agent-interface 0.22.0 contract — I confirmed `AgentCandidateRunReceiptV2 extends Omit<AgentCandidateRunReceiptV1, 'schemaVersion'>
  • Fit with existing patterns: Fits the codebase grain precisely. Respects the layering law (agent-runtime → agent-eval, never reverse): imports TraceStore/BenchmarkEvaluation/isLlmSpan/REDACTION_VERSION from agent-eval and contract types from agent-interface, never imports agent-knowledge. Does NOT compete with the improvement/ module's CandidateGenerator/improvementDriver (src/improvement/improvement-driver.ts:36) — that
  • Real-world viability: Rigorously handles the non-happy paths. Tests cover exactly-one-concurrent-invocation (tests/candidate-execution-execute.test.ts:710), hanging process stop leaving the claim for recovery (:993), unknown settlement blocking zero-spend guessing (:1061), zero-call retry only for pre-model failures (:1126), lease expiry, cross-process claim races, and replay terminal writes — 1,420 tests pass. The exe
  • Model: opencode/zai-coding-plan/glm-5.2
  • Bridge attempts: 1

No concerns — sound change, no better or existing approach found. ✅


What this audit checks

It judges the change on its merits — not whether it was tasked out in an issue. Unticketed, fast-moving work is fine; the question is whether the change is good and whether a better or existing approach should be used instead.

PassWhat it asks
HeuristicVague title? Whitespace-only or cruft-bearing diff? (content signals only)
DuplicationDo added function/class names already exist elsewhere in the repo?
Value AuditWhat does it do? What goal does it achieve? Is it good? Better architecture or already-exists?
Usefulness AuditDoes it integrate and fit? Will it hold up in real use and actually get used?

Findings are concerns, not blocks — the human reviewer decides what to do with them.

value-audit · 20260711T060700Z

tangletools
tangletools previously approved these changes Jul 11, 2026

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved drewstone PR — a4fec22e

This PR was opened by the trusted drewstone account.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: drewstone_author · 2026-07-11T06:15:17Z

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved drewstone PR — fa6db037

This PR was opened by the trusted drewstone account.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: drewstone_author · 2026-07-11T06:19:52Z

@drewstone
drewstone merged commit 164da02 into mainJul 11, 2026
2 checks passed
@drewstonedrewstone mentioned this pull request Jul 11, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' feat(runtime): execute and grade immutable agent candidates by drewstone · Pull Request #507 · tangle-network/agent-runtime · GitHub
Skip to content

feat(runtime): execute and grade immutable agent candidates - #507

Merged
drewstone merged 6 commits into
mainfrom
feat/candidate-execution-plan
Jul 11, 2026
Merged

feat(runtime): execute and grade immutable agent candidates#507
drewstone merged 6 commits into
mainfrom
feat/candidate-execution-plan

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

Summary

  • add one atomic path from a verified candidate bundle to a durable V2 run receipt
  • execute detached task/profile/code bytes behind a one-shot cross-process claim with crash recovery
  • bind the exact verified patch, task tree, executable grader bytes, raw grader output, protected trace, model ledger, and isolated memory evidence
  • enforce separate frozen execution, cleanup, and result budgets with cancellation and lease fencing
  • reject secret-bearing archives/manifests before persistence and keep recovery trace access read-only
  • consume the published @tangle-network/agent-interface 0.22.0 receipt contract and profile materializer 0.3.0

Verification

  • pnpm lint — 426 files clean
  • pnpm run typecheck — runtime and examples clean
  • pnpm test — 146 files, 1,420 passed, 2 skipped
  • pnpm build — JavaScript and declarations built
  • pnpm verify:package — public exports load
  • pnpm docs:freshness — version, peer, symbol, and generated-catalog checks clean
  • hostile re-audit — 4/4 reproduced timing/persistence failures fixed; 0 late result writes

No benchmark model calls were made in this change.

Extend the existing TraceStore, durable content addressing, candidate interface, and shared profile materializer into one verify-prepare-finalize execution path.
@drewstone

Copy link
Copy Markdown
ContributorAuthor

@tangletools review the current head for merge.

tangletools
tangletools previously approved these changes Jul 11, 2026

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved drewstone PR — eedc44c8

This PR was opened by the trusted drewstone account.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: drewstone_author · 2026-07-11T06:00:22Z

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Value Audit — sound

Verdictsound
Concerns0 (none)
Heuristic0.0s
Duplication0.0s
Interrogation291.5s (2 bridge agents)
Total291.5s

💰 Value — sound

Adds a durable, cross-process, exactly-once pipeline that turns a verified immutable candidate bundle into a tamper-evident V2 run receipt — a genuinely new layer that consumes the published substrate contracts rather than reinventing the in-process kernels, and it is in the repo's grain.

  • What it does: A new src/candidate-execution/ subsystem (~27 files) with one atomic path: verifyAgentCandidateBundle (verify.ts:37, consumes agentCandidateBundleSchema from @tangle-network/agent-interface) → prepareAgentCandidateExecution (prepare.ts:95; materializes task/candidate/profile workspaces, delegates profile→harness to @tangle-network/agent-profile-materialize at prepare.ts:167-174, reserv
  • Goals it achieves: Exactly-once, crash-recoverable execution of sealed benchmark candidates across processes/machines without double-grading, secret leakage, or unproven/late result writes (separate frozen execution/cleanup/result budgets, lease fencing, two-phase terminal commit, zero-model-call retry eligibility — claim.ts:76-91,215-222). It binds the full evidence chain (patch, task tree, executable grader bytes,
  • Assessment: Good on its merits and in-grain. It is a real capability gap the existing kernels do not fill, and it follows the repo's hard rules: domain-clean via injected adapters (CLAUDE.md layering), profile realization delegated to substrate (materializeCandidateProfile/applyAgentCandidateWorkspacePlan, prepare.ts:25-28,167-174 — honoring the §1.5 'never write a profile→harness realizer' law), neutral
  • Better / existing approach: None — this is the right approach. I checked the obvious existing equivalents and none cover durable cross-process candidate-bundle execution + receipts: src/runtime/superviseSpawnJournal/budget.ts:80 and driver-executor.ts are an in-process, single-tree content-addressed event log for equal-k replay/resume of the recursive atom (no cross-process claim/lease/terminal-CAS, no protected gra
  • Model: opencode/kimi-for-coding/k2p7
  • Bridge attempts: 1

🎯 Usefulness — sound

Adds a coherent verify→prepare→claim→execute→finalize spine that turns a sealed candidate bundle into a tamper-evident V2 run receipt, consuming real published contracts with no existing equivalent in the codebase.

  • Integration: Fully reachable: exported via export * from './candidate-execution' at src/index.ts:35, and its symbols are already in the generated API catalog (docs/api/primitive-catalog.md:110-111,142) with the docs-freshness gate green. It consumes the published @tangle-network/agent-interface 0.22.0 contract — I confirmed `AgentCandidateRunReceiptV2 extends Omit<AgentCandidateRunReceiptV1, 'schemaVersion'>
  • Fit with existing patterns: Fits the codebase grain precisely. Respects the layering law (agent-runtime → agent-eval, never reverse): imports TraceStore/BenchmarkEvaluation/isLlmSpan/REDACTION_VERSION from agent-eval and contract types from agent-interface, never imports agent-knowledge. Does NOT compete with the improvement/ module's CandidateGenerator/improvementDriver (src/improvement/improvement-driver.ts:36) — that
  • Real-world viability: Rigorously handles the non-happy paths. Tests cover exactly-one-concurrent-invocation (tests/candidate-execution-execute.test.ts:710), hanging process stop leaving the claim for recovery (:993), unknown settlement blocking zero-spend guessing (:1061), zero-call retry only for pre-model failures (:1126), lease expiry, cross-process claim races, and replay terminal writes — 1,420 tests pass. The exe
  • Model: opencode/zai-coding-plan/glm-5.2
  • Bridge attempts: 1

No concerns — sound change, no better or existing approach found. ✅


What this audit checks

It judges the change on its merits — not whether it was tasked out in an issue. Unticketed, fast-moving work is fine; the question is whether the change is good and whether a better or existing approach should be used instead.

PassWhat it asks
HeuristicVague title? Whitespace-only or cruft-bearing diff? (content signals only)
DuplicationDo added function/class names already exist elsewhere in the repo?
Value AuditWhat does it do? What goal does it achieve? Is it good? Better architecture or already-exists?
Usefulness AuditDoes it integrate and fit? Will it hold up in real use and actually get used?

Findings are concerns, not blocks — the human reviewer decides what to do with them.

value-audit · 20260711T060700Z

tangletools
tangletools previously approved these changes Jul 11, 2026

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved drewstone PR — a4fec22e

This PR was opened by the trusted drewstone account.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: drewstone_author · 2026-07-11T06:15:17Z

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved drewstone PR — fa6db037

This PR was opened by the trusted drewstone account.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: drewstone_author · 2026-07-11T06:19:52Z

@drewstone
drewstone merged commit 164da02 into mainJul 11, 2026
2 checks passed
@drewstonedrewstone mentioned this pull request Jul 11, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' feat(runtime): execute and grade immutable agent candidates by drewstone · Pull Request #507 · tangle-network/agent-runtime · GitHub
Skip to content

feat(runtime): execute and grade immutable agent candidates - #507

Merged
drewstone merged 6 commits into
mainfrom
feat/candidate-execution-plan
Jul 11, 2026
Merged

feat(runtime): execute and grade immutable agent candidates#507
drewstone merged 6 commits into
mainfrom
feat/candidate-execution-plan

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

Summary

  • add one atomic path from a verified candidate bundle to a durable V2 run receipt
  • execute detached task/profile/code bytes behind a one-shot cross-process claim with crash recovery
  • bind the exact verified patch, task tree, executable grader bytes, raw grader output, protected trace, model ledger, and isolated memory evidence
  • enforce separate frozen execution, cleanup, and result budgets with cancellation and lease fencing
  • reject secret-bearing archives/manifests before persistence and keep recovery trace access read-only
  • consume the published @tangle-network/agent-interface 0.22.0 receipt contract and profile materializer 0.3.0

Verification

  • pnpm lint — 426 files clean
  • pnpm run typecheck — runtime and examples clean
  • pnpm test — 146 files, 1,420 passed, 2 skipped
  • pnpm build — JavaScript and declarations built
  • pnpm verify:package — public exports load
  • pnpm docs:freshness — version, peer, symbol, and generated-catalog checks clean
  • hostile re-audit — 4/4 reproduced timing/persistence failures fixed; 0 late result writes

No benchmark model calls were made in this change.

Extend the existing TraceStore, durable content addressing, candidate interface, and shared profile materializer into one verify-prepare-finalize execution path.
@drewstone

Copy link
Copy Markdown
ContributorAuthor

@tangletools review the current head for merge.

tangletools
tangletools previously approved these changes Jul 11, 2026

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved drewstone PR — eedc44c8

This PR was opened by the trusted drewstone account.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: drewstone_author · 2026-07-11T06:00:22Z

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Value Audit — sound

Verdictsound
Concerns0 (none)
Heuristic0.0s
Duplication0.0s
Interrogation291.5s (2 bridge agents)
Total291.5s

💰 Value — sound

Adds a durable, cross-process, exactly-once pipeline that turns a verified immutable candidate bundle into a tamper-evident V2 run receipt — a genuinely new layer that consumes the published substrate contracts rather than reinventing the in-process kernels, and it is in the repo's grain.

  • What it does: A new src/candidate-execution/ subsystem (~27 files) with one atomic path: verifyAgentCandidateBundle (verify.ts:37, consumes agentCandidateBundleSchema from @tangle-network/agent-interface) → prepareAgentCandidateExecution (prepare.ts:95; materializes task/candidate/profile workspaces, delegates profile→harness to @tangle-network/agent-profile-materialize at prepare.ts:167-174, reserv
  • Goals it achieves: Exactly-once, crash-recoverable execution of sealed benchmark candidates across processes/machines without double-grading, secret leakage, or unproven/late result writes (separate frozen execution/cleanup/result budgets, lease fencing, two-phase terminal commit, zero-model-call retry eligibility — claim.ts:76-91,215-222). It binds the full evidence chain (patch, task tree, executable grader bytes,
  • Assessment: Good on its merits and in-grain. It is a real capability gap the existing kernels do not fill, and it follows the repo's hard rules: domain-clean via injected adapters (CLAUDE.md layering), profile realization delegated to substrate (materializeCandidateProfile/applyAgentCandidateWorkspacePlan, prepare.ts:25-28,167-174 — honoring the §1.5 'never write a profile→harness realizer' law), neutral
  • Better / existing approach: None — this is the right approach. I checked the obvious existing equivalents and none cover durable cross-process candidate-bundle execution + receipts: src/runtime/superviseSpawnJournal/budget.ts:80 and driver-executor.ts are an in-process, single-tree content-addressed event log for equal-k replay/resume of the recursive atom (no cross-process claim/lease/terminal-CAS, no protected gra
  • Model: opencode/kimi-for-coding/k2p7
  • Bridge attempts: 1

🎯 Usefulness — sound

Adds a coherent verify→prepare→claim→execute→finalize spine that turns a sealed candidate bundle into a tamper-evident V2 run receipt, consuming real published contracts with no existing equivalent in the codebase.

  • Integration: Fully reachable: exported via export * from './candidate-execution' at src/index.ts:35, and its symbols are already in the generated API catalog (docs/api/primitive-catalog.md:110-111,142) with the docs-freshness gate green. It consumes the published @tangle-network/agent-interface 0.22.0 contract — I confirmed `AgentCandidateRunReceiptV2 extends Omit<AgentCandidateRunReceiptV1, 'schemaVersion'>
  • Fit with existing patterns: Fits the codebase grain precisely. Respects the layering law (agent-runtime → agent-eval, never reverse): imports TraceStore/BenchmarkEvaluation/isLlmSpan/REDACTION_VERSION from agent-eval and contract types from agent-interface, never imports agent-knowledge. Does NOT compete with the improvement/ module's CandidateGenerator/improvementDriver (src/improvement/improvement-driver.ts:36) — that
  • Real-world viability: Rigorously handles the non-happy paths. Tests cover exactly-one-concurrent-invocation (tests/candidate-execution-execute.test.ts:710), hanging process stop leaving the claim for recovery (:993), unknown settlement blocking zero-spend guessing (:1061), zero-call retry only for pre-model failures (:1126), lease expiry, cross-process claim races, and replay terminal writes — 1,420 tests pass. The exe
  • Model: opencode/zai-coding-plan/glm-5.2
  • Bridge attempts: 1

No concerns — sound change, no better or existing approach found. ✅


What this audit checks

It judges the change on its merits — not whether it was tasked out in an issue. Unticketed, fast-moving work is fine; the question is whether the change is good and whether a better or existing approach should be used instead.

PassWhat it asks
HeuristicVague title? Whitespace-only or cruft-bearing diff? (content signals only)
DuplicationDo added function/class names already exist elsewhere in the repo?
Value AuditWhat does it do? What goal does it achieve? Is it good? Better architecture or already-exists?
Usefulness AuditDoes it integrate and fit? Will it hold up in real use and actually get used?

Findings are concerns, not blocks — the human reviewer decides what to do with them.

value-audit · 20260711T060700Z

tangletools
tangletools previously approved these changes Jul 11, 2026

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved drewstone PR — a4fec22e

This PR was opened by the trusted drewstone account.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: drewstone_author · 2026-07-11T06:15:17Z

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved drewstone PR — fa6db037

This PR was opened by the trusted drewstone account.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: drewstone_author · 2026-07-11T06:19:52Z

@drewstone
drewstone merged commit 164da02 into mainJul 11, 2026
2 checks passed
@drewstonedrewstone mentioned this pull request Jul 11, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' feat(runtime): execute and grade immutable agent candidates by drewstone · Pull Request #507 · tangle-network/agent-runtime · GitHub
Skip to content

feat(runtime): execute and grade immutable agent candidates - #507

Merged
drewstone merged 6 commits into
mainfrom
feat/candidate-execution-plan
Jul 11, 2026
Merged

feat(runtime): execute and grade immutable agent candidates#507
drewstone merged 6 commits into
mainfrom
feat/candidate-execution-plan

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

Summary

  • add one atomic path from a verified candidate bundle to a durable V2 run receipt
  • execute detached task/profile/code bytes behind a one-shot cross-process claim with crash recovery
  • bind the exact verified patch, task tree, executable grader bytes, raw grader output, protected trace, model ledger, and isolated memory evidence
  • enforce separate frozen execution, cleanup, and result budgets with cancellation and lease fencing
  • reject secret-bearing archives/manifests before persistence and keep recovery trace access read-only
  • consume the published @tangle-network/agent-interface 0.22.0 receipt contract and profile materializer 0.3.0

Verification

  • pnpm lint — 426 files clean
  • pnpm run typecheck — runtime and examples clean
  • pnpm test — 146 files, 1,420 passed, 2 skipped
  • pnpm build — JavaScript and declarations built
  • pnpm verify:package — public exports load
  • pnpm docs:freshness — version, peer, symbol, and generated-catalog checks clean
  • hostile re-audit — 4/4 reproduced timing/persistence failures fixed; 0 late result writes

No benchmark model calls were made in this change.

Extend the existing TraceStore, durable content addressing, candidate interface, and shared profile materializer into one verify-prepare-finalize execution path.
@drewstone

Copy link
Copy Markdown
ContributorAuthor

@tangletools review the current head for merge.

tangletools
tangletools previously approved these changes Jul 11, 2026

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved drewstone PR — eedc44c8

This PR was opened by the trusted drewstone account.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: drewstone_author · 2026-07-11T06:00:22Z

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Value Audit — sound

Verdictsound
Concerns0 (none)
Heuristic0.0s
Duplication0.0s
Interrogation291.5s (2 bridge agents)
Total291.5s

💰 Value — sound

Adds a durable, cross-process, exactly-once pipeline that turns a verified immutable candidate bundle into a tamper-evident V2 run receipt — a genuinely new layer that consumes the published substrate contracts rather than reinventing the in-process kernels, and it is in the repo's grain.

  • What it does: A new src/candidate-execution/ subsystem (~27 files) with one atomic path: verifyAgentCandidateBundle (verify.ts:37, consumes agentCandidateBundleSchema from @tangle-network/agent-interface) → prepareAgentCandidateExecution (prepare.ts:95; materializes task/candidate/profile workspaces, delegates profile→harness to @tangle-network/agent-profile-materialize at prepare.ts:167-174, reserv
  • Goals it achieves: Exactly-once, crash-recoverable execution of sealed benchmark candidates across processes/machines without double-grading, secret leakage, or unproven/late result writes (separate frozen execution/cleanup/result budgets, lease fencing, two-phase terminal commit, zero-model-call retry eligibility — claim.ts:76-91,215-222). It binds the full evidence chain (patch, task tree, executable grader bytes,
  • Assessment: Good on its merits and in-grain. It is a real capability gap the existing kernels do not fill, and it follows the repo's hard rules: domain-clean via injected adapters (CLAUDE.md layering), profile realization delegated to substrate (materializeCandidateProfile/applyAgentCandidateWorkspacePlan, prepare.ts:25-28,167-174 — honoring the §1.5 'never write a profile→harness realizer' law), neutral
  • Better / existing approach: None — this is the right approach. I checked the obvious existing equivalents and none cover durable cross-process candidate-bundle execution + receipts: src/runtime/superviseSpawnJournal/budget.ts:80 and driver-executor.ts are an in-process, single-tree content-addressed event log for equal-k replay/resume of the recursive atom (no cross-process claim/lease/terminal-CAS, no protected gra
  • Model: opencode/kimi-for-coding/k2p7
  • Bridge attempts: 1

🎯 Usefulness — sound

Adds a coherent verify→prepare→claim→execute→finalize spine that turns a sealed candidate bundle into a tamper-evident V2 run receipt, consuming real published contracts with no existing equivalent in the codebase.

  • Integration: Fully reachable: exported via export * from './candidate-execution' at src/index.ts:35, and its symbols are already in the generated API catalog (docs/api/primitive-catalog.md:110-111,142) with the docs-freshness gate green. It consumes the published @tangle-network/agent-interface 0.22.0 contract — I confirmed `AgentCandidateRunReceiptV2 extends Omit<AgentCandidateRunReceiptV1, 'schemaVersion'>
  • Fit with existing patterns: Fits the codebase grain precisely. Respects the layering law (agent-runtime → agent-eval, never reverse): imports TraceStore/BenchmarkEvaluation/isLlmSpan/REDACTION_VERSION from agent-eval and contract types from agent-interface, never imports agent-knowledge. Does NOT compete with the improvement/ module's CandidateGenerator/improvementDriver (src/improvement/improvement-driver.ts:36) — that
  • Real-world viability: Rigorously handles the non-happy paths. Tests cover exactly-one-concurrent-invocation (tests/candidate-execution-execute.test.ts:710), hanging process stop leaving the claim for recovery (:993), unknown settlement blocking zero-spend guessing (:1061), zero-call retry only for pre-model failures (:1126), lease expiry, cross-process claim races, and replay terminal writes — 1,420 tests pass. The exe
  • Model: opencode/zai-coding-plan/glm-5.2
  • Bridge attempts: 1

No concerns — sound change, no better or existing approach found. ✅


What this audit checks

It judges the change on its merits — not whether it was tasked out in an issue. Unticketed, fast-moving work is fine; the question is whether the change is good and whether a better or existing approach should be used instead.

PassWhat it asks
HeuristicVague title? Whitespace-only or cruft-bearing diff? (content signals only)
DuplicationDo added function/class names already exist elsewhere in the repo?
Value AuditWhat does it do? What goal does it achieve? Is it good? Better architecture or already-exists?
Usefulness AuditDoes it integrate and fit? Will it hold up in real use and actually get used?

Findings are concerns, not blocks — the human reviewer decides what to do with them.

value-audit · 20260711T060700Z

tangletools
tangletools previously approved these changes Jul 11, 2026

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved drewstone PR — a4fec22e

This PR was opened by the trusted drewstone account.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: drewstone_author · 2026-07-11T06:15:17Z

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved drewstone PR — fa6db037

This PR was opened by the trusted drewstone account.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: drewstone_author · 2026-07-11T06:19:52Z

@drewstone
drewstone merged commit 164da02 into mainJul 11, 2026
2 checks passed
@drewstonedrewstone mentioned this pull request Jul 11, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' feat(runtime): execute and grade immutable agent candidates by drewstone · Pull Request #507 · tangle-network/agent-runtime · GitHub
Skip to content

feat(runtime): execute and grade immutable agent candidates - #507

Merged
drewstone merged 6 commits into
mainfrom
feat/candidate-execution-plan
Jul 11, 2026
Merged

feat(runtime): execute and grade immutable agent candidates#507
drewstone merged 6 commits into
mainfrom
feat/candidate-execution-plan

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

Summary

  • add one atomic path from a verified candidate bundle to a durable V2 run receipt
  • execute detached task/profile/code bytes behind a one-shot cross-process claim with crash recovery
  • bind the exact verified patch, task tree, executable grader bytes, raw grader output, protected trace, model ledger, and isolated memory evidence
  • enforce separate frozen execution, cleanup, and result budgets with cancellation and lease fencing
  • reject secret-bearing archives/manifests before persistence and keep recovery trace access read-only
  • consume the published @tangle-network/agent-interface 0.22.0 receipt contract and profile materializer 0.3.0

Verification

  • pnpm lint — 426 files clean
  • pnpm run typecheck — runtime and examples clean
  • pnpm test — 146 files, 1,420 passed, 2 skipped
  • pnpm build — JavaScript and declarations built
  • pnpm verify:package — public exports load
  • pnpm docs:freshness — version, peer, symbol, and generated-catalog checks clean
  • hostile re-audit — 4/4 reproduced timing/persistence failures fixed; 0 late result writes

No benchmark model calls were made in this change.

Extend the existing TraceStore, durable content addressing, candidate interface, and shared profile materializer into one verify-prepare-finalize execution path.
@drewstone

Copy link
Copy Markdown
ContributorAuthor

@tangletools review the current head for merge.

tangletools
tangletools previously approved these changes Jul 11, 2026

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved drewstone PR — eedc44c8

This PR was opened by the trusted drewstone account.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: drewstone_author · 2026-07-11T06:00:22Z

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Value Audit — sound

Verdictsound
Concerns0 (none)
Heuristic0.0s
Duplication0.0s
Interrogation291.5s (2 bridge agents)
Total291.5s

💰 Value — sound

Adds a durable, cross-process, exactly-once pipeline that turns a verified immutable candidate bundle into a tamper-evident V2 run receipt — a genuinely new layer that consumes the published substrate contracts rather than reinventing the in-process kernels, and it is in the repo's grain.

  • What it does: A new src/candidate-execution/ subsystem (~27 files) with one atomic path: verifyAgentCandidateBundle (verify.ts:37, consumes agentCandidateBundleSchema from @tangle-network/agent-interface) → prepareAgentCandidateExecution (prepare.ts:95; materializes task/candidate/profile workspaces, delegates profile→harness to @tangle-network/agent-profile-materialize at prepare.ts:167-174, reserv
  • Goals it achieves: Exactly-once, crash-recoverable execution of sealed benchmark candidates across processes/machines without double-grading, secret leakage, or unproven/late result writes (separate frozen execution/cleanup/result budgets, lease fencing, two-phase terminal commit, zero-model-call retry eligibility — claim.ts:76-91,215-222). It binds the full evidence chain (patch, task tree, executable grader bytes,
  • Assessment: Good on its merits and in-grain. It is a real capability gap the existing kernels do not fill, and it follows the repo's hard rules: domain-clean via injected adapters (CLAUDE.md layering), profile realization delegated to substrate (materializeCandidateProfile/applyAgentCandidateWorkspacePlan, prepare.ts:25-28,167-174 — honoring the §1.5 'never write a profile→harness realizer' law), neutral
  • Better / existing approach: None — this is the right approach. I checked the obvious existing equivalents and none cover durable cross-process candidate-bundle execution + receipts: src/runtime/superviseSpawnJournal/budget.ts:80 and driver-executor.ts are an in-process, single-tree content-addressed event log for equal-k replay/resume of the recursive atom (no cross-process claim/lease/terminal-CAS, no protected gra
  • Model: opencode/kimi-for-coding/k2p7
  • Bridge attempts: 1

🎯 Usefulness — sound

Adds a coherent verify→prepare→claim→execute→finalize spine that turns a sealed candidate bundle into a tamper-evident V2 run receipt, consuming real published contracts with no existing equivalent in the codebase.

  • Integration: Fully reachable: exported via export * from './candidate-execution' at src/index.ts:35, and its symbols are already in the generated API catalog (docs/api/primitive-catalog.md:110-111,142) with the docs-freshness gate green. It consumes the published @tangle-network/agent-interface 0.22.0 contract — I confirmed `AgentCandidateRunReceiptV2 extends Omit<AgentCandidateRunReceiptV1, 'schemaVersion'>
  • Fit with existing patterns: Fits the codebase grain precisely. Respects the layering law (agent-runtime → agent-eval, never reverse): imports TraceStore/BenchmarkEvaluation/isLlmSpan/REDACTION_VERSION from agent-eval and contract types from agent-interface, never imports agent-knowledge. Does NOT compete with the improvement/ module's CandidateGenerator/improvementDriver (src/improvement/improvement-driver.ts:36) — that
  • Real-world viability: Rigorously handles the non-happy paths. Tests cover exactly-one-concurrent-invocation (tests/candidate-execution-execute.test.ts:710), hanging process stop leaving the claim for recovery (:993), unknown settlement blocking zero-spend guessing (:1061), zero-call retry only for pre-model failures (:1126), lease expiry, cross-process claim races, and replay terminal writes — 1,420 tests pass. The exe
  • Model: opencode/zai-coding-plan/glm-5.2
  • Bridge attempts: 1

No concerns — sound change, no better or existing approach found. ✅


What this audit checks

It judges the change on its merits — not whether it was tasked out in an issue. Unticketed, fast-moving work is fine; the question is whether the change is good and whether a better or existing approach should be used instead.

PassWhat it asks
HeuristicVague title? Whitespace-only or cruft-bearing diff? (content signals only)
DuplicationDo added function/class names already exist elsewhere in the repo?
Value AuditWhat does it do? What goal does it achieve? Is it good? Better architecture or already-exists?
Usefulness AuditDoes it integrate and fit? Will it hold up in real use and actually get used?

Findings are concerns, not blocks — the human reviewer decides what to do with them.

value-audit · 20260711T060700Z

tangletools
tangletools previously approved these changes Jul 11, 2026

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved drewstone PR — a4fec22e

This PR was opened by the trusted drewstone account.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: drewstone_author · 2026-07-11T06:15:17Z

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved drewstone PR — fa6db037

This PR was opened by the trusted drewstone account.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: drewstone_author · 2026-07-11T06:19:52Z

@drewstone
drewstone merged commit 164da02 into mainJul 11, 2026
2 checks passed
@drewstonedrewstone mentioned this pull request Jul 11, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); feat(runtime): execute and grade immutable agent candidates by drewstone · Pull Request #507 · tangle-network/agent-runtime · GitHub
Skip to content

feat(runtime): execute and grade immutable agent candidates - #507

Merged
drewstone merged 6 commits into
mainfrom
feat/candidate-execution-plan
Jul 11, 2026
Merged

feat(runtime): execute and grade immutable agent candidates#507
drewstone merged 6 commits into
mainfrom
feat/candidate-execution-plan

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

Summary

  • add one atomic path from a verified candidate bundle to a durable V2 run receipt
  • execute detached task/profile/code bytes behind a one-shot cross-process claim with crash recovery
  • bind the exact verified patch, task tree, executable grader bytes, raw grader output, protected trace, model ledger, and isolated memory evidence
  • enforce separate frozen execution, cleanup, and result budgets with cancellation and lease fencing
  • reject secret-bearing archives/manifests before persistence and keep recovery trace access read-only
  • consume the published @tangle-network/agent-interface 0.22.0 receipt contract and profile materializer 0.3.0

Verification

  • pnpm lint — 426 files clean
  • pnpm run typecheck — runtime and examples clean
  • pnpm test — 146 files, 1,420 passed, 2 skipped
  • pnpm build — JavaScript and declarations built
  • pnpm verify:package — public exports load
  • pnpm docs:freshness — version, peer, symbol, and generated-catalog checks clean
  • hostile re-audit — 4/4 reproduced timing/persistence failures fixed; 0 late result writes

No benchmark model calls were made in this change.

Extend the existing TraceStore, durable content addressing, candidate interface, and shared profile materializer into one verify-prepare-finalize execution path.
@drewstone

Copy link
Copy Markdown
ContributorAuthor

@tangletools review the current head for merge.

tangletools
tangletools previously approved these changes Jul 11, 2026

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved drewstone PR — eedc44c8

This PR was opened by the trusted drewstone account.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: drewstone_author · 2026-07-11T06:00:22Z

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Value Audit — sound

Verdictsound
Concerns0 (none)
Heuristic0.0s
Duplication0.0s
Interrogation291.5s (2 bridge agents)
Total291.5s

💰 Value — sound

Adds a durable, cross-process, exactly-once pipeline that turns a verified immutable candidate bundle into a tamper-evident V2 run receipt — a genuinely new layer that consumes the published substrate contracts rather than reinventing the in-process kernels, and it is in the repo's grain.

  • What it does: A new src/candidate-execution/ subsystem (~27 files) with one atomic path: verifyAgentCandidateBundle (verify.ts:37, consumes agentCandidateBundleSchema from @tangle-network/agent-interface) → prepareAgentCandidateExecution (prepare.ts:95; materializes task/candidate/profile workspaces, delegates profile→harness to @tangle-network/agent-profile-materialize at prepare.ts:167-174, reserv
  • Goals it achieves: Exactly-once, crash-recoverable execution of sealed benchmark candidates across processes/machines without double-grading, secret leakage, or unproven/late result writes (separate frozen execution/cleanup/result budgets, lease fencing, two-phase terminal commit, zero-model-call retry eligibility — claim.ts:76-91,215-222). It binds the full evidence chain (patch, task tree, executable grader bytes,
  • Assessment: Good on its merits and in-grain. It is a real capability gap the existing kernels do not fill, and it follows the repo's hard rules: domain-clean via injected adapters (CLAUDE.md layering), profile realization delegated to substrate (materializeCandidateProfile/applyAgentCandidateWorkspacePlan, prepare.ts:25-28,167-174 — honoring the §1.5 'never write a profile→harness realizer' law), neutral
  • Better / existing approach: None — this is the right approach. I checked the obvious existing equivalents and none cover durable cross-process candidate-bundle execution + receipts: src/runtime/superviseSpawnJournal/budget.ts:80 and driver-executor.ts are an in-process, single-tree content-addressed event log for equal-k replay/resume of the recursive atom (no cross-process claim/lease/terminal-CAS, no protected gra
  • Model: opencode/kimi-for-coding/k2p7
  • Bridge attempts: 1

🎯 Usefulness — sound

Adds a coherent verify→prepare→claim→execute→finalize spine that turns a sealed candidate bundle into a tamper-evident V2 run receipt, consuming real published contracts with no existing equivalent in the codebase.

  • Integration: Fully reachable: exported via export * from './candidate-execution' at src/index.ts:35, and its symbols are already in the generated API catalog (docs/api/primitive-catalog.md:110-111,142) with the docs-freshness gate green. It consumes the published @tangle-network/agent-interface 0.22.0 contract — I confirmed `AgentCandidateRunReceiptV2 extends Omit<AgentCandidateRunReceiptV1, 'schemaVersion'>
  • Fit with existing patterns: Fits the codebase grain precisely. Respects the layering law (agent-runtime → agent-eval, never reverse): imports TraceStore/BenchmarkEvaluation/isLlmSpan/REDACTION_VERSION from agent-eval and contract types from agent-interface, never imports agent-knowledge. Does NOT compete with the improvement/ module's CandidateGenerator/improvementDriver (src/improvement/improvement-driver.ts:36) — that
  • Real-world viability: Rigorously handles the non-happy paths. Tests cover exactly-one-concurrent-invocation (tests/candidate-execution-execute.test.ts:710), hanging process stop leaving the claim for recovery (:993), unknown settlement blocking zero-spend guessing (:1061), zero-call retry only for pre-model failures (:1126), lease expiry, cross-process claim races, and replay terminal writes — 1,420 tests pass. The exe
  • Model: opencode/zai-coding-plan/glm-5.2
  • Bridge attempts: 1

No concerns — sound change, no better or existing approach found. ✅


What this audit checks

It judges the change on its merits — not whether it was tasked out in an issue. Unticketed, fast-moving work is fine; the question is whether the change is good and whether a better or existing approach should be used instead.

PassWhat it asks
HeuristicVague title? Whitespace-only or cruft-bearing diff? (content signals only)
DuplicationDo added function/class names already exist elsewhere in the repo?
Value AuditWhat does it do? What goal does it achieve? Is it good? Better architecture or already-exists?
Usefulness AuditDoes it integrate and fit? Will it hold up in real use and actually get used?

Findings are concerns, not blocks — the human reviewer decides what to do with them.

value-audit · 20260711T060700Z

tangletools
tangletools previously approved these changes Jul 11, 2026

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved drewstone PR — a4fec22e

This PR was opened by the trusted drewstone account.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: drewstone_author · 2026-07-11T06:15:17Z

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved drewstone PR — fa6db037

This PR was opened by the trusted drewstone account.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: drewstone_author · 2026-07-11T06:19:52Z

@drewstone
drewstone merged commit 164da02 into mainJul 11, 2026
2 checks passed
@drewstonedrewstone mentioned this pull request Jul 11, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools