docs(examples): three-persona cleanup — newcomer/senior/junior - #376

Merged
drewstone merged 1 commit into
mainfrom
examples/three-persona-cleanup
Jun 24, 2026
Merged

docs(examples): three-persona cleanup — newcomer/senior/junior#376
drewstone merged 1 commit into
mainfrom
examples/three-persona-cleanup

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

Applies the genuine-defect fixes from a three-persona (newcomer / senior / junior) review of the 22 examples. Scope: fix the things that actively mislead a reader, and clean the rough edges that teach the wrong habit — without churning the examples judged good/great. Every example still typechecks (typecheck:examples), the held-out coding-benchmark anti-cheat + its smoke test stay intact, and docs:check stays green.

Genuine defects fixed (these mislead a reader)

  • self-improving-loop — the demo gates at n=3 yet called it "the production held-out gate's statistical core" flatly, re-teaching this repo's documented fix: persist final runtime stream failures #1 failure mode (the small-n mirage). Added a loud minimum-evidence-floor caveat at the gate() definition AND in the README: the production gate floors the evidence (heldoutSignificance won't report a pair under minSamples, default 8; HeldOutGate rejects below minProductiveRuns with few_runs) — never ship a real change on n=3. (helps the junior most — the persona most likely to lift gate() verbatim).
  • delegate — was a Tier-1 example with no README, and the file was a CI test (E2E PASSED, process.exit) wearing an example's clothes, with internal history in the header. Added README.md; split a lean teaching delegate.ts + reusable shared.ts from the regression proof, which moved to tests/delegate-example.test.ts (env-gated: a paid live e2e when TANGLE_API_KEY is set, an always-on offline fail-loud assertion otherwise); stripped the history per the repo's no-history-in-source rule. (newcomer + senior).
  • improve — README named ImprovementDriver / gepaDriver; improve() actually builds SurfaceProposer / gepaProposer. Corrected both, plus the stale test path (src/improvement/improve.test.ts). (a README grep now resolves).
  • knowledge-gating — README featured adapter.onKnowledgeBlocked as the headline hook but the adapter never defined it (doc/code drift). Wired the hook into the adapter so the blocked run demonstrates it (converts the gap into a "would ask the user" decision that flows through as the stop reason).

Quality wins (remove ceremony / amateur tells, no behavior change)

  • coding-benchmark — renamed offlineSolutionsofflineAgentScripts with a clear 2-line header; kept the rate-limiter cheat/real pair inline (the one anti-cheat teaching moment), moved the round-invariant csv-parser / lru-cache real impls to fixtures.ts as readable template literals (no + '\n' + escaped-string ceremony). The held-out anti-cheat, the firewall test, the reps-don't-fake-n regression, and the BH-corrected stats are all unchanged and still pass.
  • supervise — wrapped the flagship in main().catch (matching the sibling examples) and uncommented the completion-oracle deliverable so the headline models the safe path.
  • ui-auditLENSES_TO_RUNlensesToRun (the publish-safe module-global convention; an UPPERCASE module-global trips the Tangle obfuscator).
  • driver-loop / researcher-loop — one-line justification at the offline-box as unknown as SandboxInstance casts (the other casts already carried one).
  • strategy-evolution README — one line noting promoted: false at toy scale is the gate working, not a break.

Left alone (already good/great — no churn)

driver-loop, strategy-suite, supervisor-loop, chat-handler, recursive-supervisor, runtime-run, stream-backends, sanitized-telemetry-streaming, mcp-delegation, fleet-delegation, intelligence-recommend, intelligence-drop-in, agents-of-all-shapes, product-eval. The verdict's older snapshot flagged a few of these (mcp-delegation's delegate_ui_audit, the fleet-delegation casts) but they already carry the right framing/justification on current main; forcing fake-complete SandboxInstance helpers would add ceremony against the cleanup goal.

Verification

  • pnpm run build — clean
  • pnpm run typecheck (src + examples) — clean
  • pnpm run lint — clean (333 files)
  • pnpm run docs:check — green (0 errors; freshness OK)
  • pnpm test — 115 files / 1120 passed, 2 skipped (the env-gated live e2es)
  • ran each materially-changed example offline: knowledge-gating (hook fires), self-improving-loop, ui-audit, coding-benchmark (identical 0.944 leaderboard)

Apply the genuine-defect fixes from the three-persona example verdict, leaving
good/great examples untouched.
- self-improving-loop: add a loud minimum-evidence-floor caveat at the gate and
in the README — the demo gates at n=3 for runnability, but the production gate
floors at minSamples (8 in heldoutSignificance) / minProductiveRuns; never
ship a real change on n=3 (the small-n mirage).
- delegate: add a README, split a lean teaching delegate.ts + shared.ts from the
regression proof (moved to tests/delegate-example.test.ts — env-gated live e2e
+ an always-on offline fail-loud assertion); drop the test-in-example clothing
(E2E PASSED / process.exit) and the internal history from the header.
- improve: README symbol drift — ImprovementDriver → SurfaceProposer,
gepaDriver → gepaProposer (match what improve() actually builds); fix the test
path to src/improvement/improve.test.ts.
- knowledge-gating: wire the headline onKnowledgeBlocked hook into the adapter so
the README's documented hook is demonstrated (the blocked run now converts the
gap into a "would ask the user" decision).
- coding-benchmark: simplify offlineSolutions → offlineAgentScripts; keep the
rate-limiter cheat/real pair inline (the one anti-cheat teaching moment), move
csv/lru real impls to fixtures.ts as readable template literals (no escaped
strings). The held-out anti-cheat + smoke + firewall tests stay intact.
- supervise: wrap the flagship in main().catch (match the siblings) and uncomment
the completion-oracle deliverable so the headline models the safe path.
- ui-audit: LENSES_TO_RUN → lensesToRun (the publish-safe module-global convention).
- driver-loop / researcher-loop: one-line justification at the offline-box casts.
- strategy-evolution README: note promoted:false at toy scale is the gate working.

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved PR — efb6b428

Blanket team auto-approval is enabled for this reviewer service.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: blanket_auto_approve · 2026-06-24T15:09:11Z

@drewstone
drewstone merged commit 322211d into mainJun 24, 2026
1 check passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

docs(examples): three-persona cleanup — newcomer/senior/junior - #376

Merged
drewstone merged 1 commit into
mainfrom
examples/three-persona-cleanup
Jun 24, 2026
Merged

docs(examples): three-persona cleanup — newcomer/senior/junior#376
drewstone merged 1 commit into
mainfrom
examples/three-persona-cleanup

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

Applies the genuine-defect fixes from a three-persona (newcomer / senior / junior) review of the 22 examples. Scope: fix the things that actively mislead a reader, and clean the rough edges that teach the wrong habit — without churning the examples judged good/great. Every example still typechecks (typecheck:examples), the held-out coding-benchmark anti-cheat + its smoke test stay intact, and docs:check stays green.

Genuine defects fixed (these mislead a reader)

  • self-improving-loop — the demo gates at n=3 yet called it "the production held-out gate's statistical core" flatly, re-teaching this repo's documented fix: persist final runtime stream failures #1 failure mode (the small-n mirage). Added a loud minimum-evidence-floor caveat at the gate() definition AND in the README: the production gate floors the evidence (heldoutSignificance won't report a pair under minSamples, default 8; HeldOutGate rejects below minProductiveRuns with few_runs) — never ship a real change on n=3. (helps the junior most — the persona most likely to lift gate() verbatim).
  • delegate — was a Tier-1 example with no README, and the file was a CI test (E2E PASSED, process.exit) wearing an example's clothes, with internal history in the header. Added README.md; split a lean teaching delegate.ts + reusable shared.ts from the regression proof, which moved to tests/delegate-example.test.ts (env-gated: a paid live e2e when TANGLE_API_KEY is set, an always-on offline fail-loud assertion otherwise); stripped the history per the repo's no-history-in-source rule. (newcomer + senior).
  • improve — README named ImprovementDriver / gepaDriver; improve() actually builds SurfaceProposer / gepaProposer. Corrected both, plus the stale test path (src/improvement/improve.test.ts). (a README grep now resolves).
  • knowledge-gating — README featured adapter.onKnowledgeBlocked as the headline hook but the adapter never defined it (doc/code drift). Wired the hook into the adapter so the blocked run demonstrates it (converts the gap into a "would ask the user" decision that flows through as the stop reason).

Quality wins (remove ceremony / amateur tells, no behavior change)

  • coding-benchmark — renamed offlineSolutionsofflineAgentScripts with a clear 2-line header; kept the rate-limiter cheat/real pair inline (the one anti-cheat teaching moment), moved the round-invariant csv-parser / lru-cache real impls to fixtures.ts as readable template literals (no + '\n' + escaped-string ceremony). The held-out anti-cheat, the firewall test, the reps-don't-fake-n regression, and the BH-corrected stats are all unchanged and still pass.
  • supervise — wrapped the flagship in main().catch (matching the sibling examples) and uncommented the completion-oracle deliverable so the headline models the safe path.
  • ui-auditLENSES_TO_RUNlensesToRun (the publish-safe module-global convention; an UPPERCASE module-global trips the Tangle obfuscator).
  • driver-loop / researcher-loop — one-line justification at the offline-box as unknown as SandboxInstance casts (the other casts already carried one).
  • strategy-evolution README — one line noting promoted: false at toy scale is the gate working, not a break.

Left alone (already good/great — no churn)

driver-loop, strategy-suite, supervisor-loop, chat-handler, recursive-supervisor, runtime-run, stream-backends, sanitized-telemetry-streaming, mcp-delegation, fleet-delegation, intelligence-recommend, intelligence-drop-in, agents-of-all-shapes, product-eval. The verdict's older snapshot flagged a few of these (mcp-delegation's delegate_ui_audit, the fleet-delegation casts) but they already carry the right framing/justification on current main; forcing fake-complete SandboxInstance helpers would add ceremony against the cleanup goal.

Verification

  • pnpm run build — clean
  • pnpm run typecheck (src + examples) — clean
  • pnpm run lint — clean (333 files)
  • pnpm run docs:check — green (0 errors; freshness OK)
  • pnpm test — 115 files / 1120 passed, 2 skipped (the env-gated live e2es)
  • ran each materially-changed example offline: knowledge-gating (hook fires), self-improving-loop, ui-audit, coding-benchmark (identical 0.944 leaderboard)

Apply the genuine-defect fixes from the three-persona example verdict, leaving
good/great examples untouched.
- self-improving-loop: add a loud minimum-evidence-floor caveat at the gate and
in the README — the demo gates at n=3 for runnability, but the production gate
floors at minSamples (8 in heldoutSignificance) / minProductiveRuns; never
ship a real change on n=3 (the small-n mirage).
- delegate: add a README, split a lean teaching delegate.ts + shared.ts from the
regression proof (moved to tests/delegate-example.test.ts — env-gated live e2e
+ an always-on offline fail-loud assertion); drop the test-in-example clothing
(E2E PASSED / process.exit) and the internal history from the header.
- improve: README symbol drift — ImprovementDriver → SurfaceProposer,
gepaDriver → gepaProposer (match what improve() actually builds); fix the test
path to src/improvement/improve.test.ts.
- knowledge-gating: wire the headline onKnowledgeBlocked hook into the adapter so
the README's documented hook is demonstrated (the blocked run now converts the
gap into a "would ask the user" decision).
- coding-benchmark: simplify offlineSolutions → offlineAgentScripts; keep the
rate-limiter cheat/real pair inline (the one anti-cheat teaching moment), move
csv/lru real impls to fixtures.ts as readable template literals (no escaped
strings). The held-out anti-cheat + smoke + firewall tests stay intact.
- supervise: wrap the flagship in main().catch (match the siblings) and uncomment
the completion-oracle deliverable so the headline models the safe path.
- ui-audit: LENSES_TO_RUN → lensesToRun (the publish-safe module-global convention).
- driver-loop / researcher-loop: one-line justification at the offline-box casts.
- strategy-evolution README: note promoted:false at toy scale is the gate working.

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved PR — efb6b428

Blanket team auto-approval is enabled for this reviewer service.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: blanket_auto_approve · 2026-06-24T15:09:11Z

@drewstone
drewstone merged commit 322211d into mainJun 24, 2026
1 check passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

docs(examples): three-persona cleanup — newcomer/senior/junior - #376

Merged
drewstone merged 1 commit into
mainfrom
examples/three-persona-cleanup
Jun 24, 2026
Merged

docs(examples): three-persona cleanup — newcomer/senior/junior#376
drewstone merged 1 commit into
mainfrom
examples/three-persona-cleanup

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

Applies the genuine-defect fixes from a three-persona (newcomer / senior / junior) review of the 22 examples. Scope: fix the things that actively mislead a reader, and clean the rough edges that teach the wrong habit — without churning the examples judged good/great. Every example still typechecks (typecheck:examples), the held-out coding-benchmark anti-cheat + its smoke test stay intact, and docs:check stays green.

Genuine defects fixed (these mislead a reader)

  • self-improving-loop — the demo gates at n=3 yet called it "the production held-out gate's statistical core" flatly, re-teaching this repo's documented fix: persist final runtime stream failures #1 failure mode (the small-n mirage). Added a loud minimum-evidence-floor caveat at the gate() definition AND in the README: the production gate floors the evidence (heldoutSignificance won't report a pair under minSamples, default 8; HeldOutGate rejects below minProductiveRuns with few_runs) — never ship a real change on n=3. (helps the junior most — the persona most likely to lift gate() verbatim).
  • delegate — was a Tier-1 example with no README, and the file was a CI test (E2E PASSED, process.exit) wearing an example's clothes, with internal history in the header. Added README.md; split a lean teaching delegate.ts + reusable shared.ts from the regression proof, which moved to tests/delegate-example.test.ts (env-gated: a paid live e2e when TANGLE_API_KEY is set, an always-on offline fail-loud assertion otherwise); stripped the history per the repo's no-history-in-source rule. (newcomer + senior).
  • improve — README named ImprovementDriver / gepaDriver; improve() actually builds SurfaceProposer / gepaProposer. Corrected both, plus the stale test path (src/improvement/improve.test.ts). (a README grep now resolves).
  • knowledge-gating — README featured adapter.onKnowledgeBlocked as the headline hook but the adapter never defined it (doc/code drift). Wired the hook into the adapter so the blocked run demonstrates it (converts the gap into a "would ask the user" decision that flows through as the stop reason).

Quality wins (remove ceremony / amateur tells, no behavior change)

  • coding-benchmark — renamed offlineSolutionsofflineAgentScripts with a clear 2-line header; kept the rate-limiter cheat/real pair inline (the one anti-cheat teaching moment), moved the round-invariant csv-parser / lru-cache real impls to fixtures.ts as readable template literals (no + '\n' + escaped-string ceremony). The held-out anti-cheat, the firewall test, the reps-don't-fake-n regression, and the BH-corrected stats are all unchanged and still pass.
  • supervise — wrapped the flagship in main().catch (matching the sibling examples) and uncommented the completion-oracle deliverable so the headline models the safe path.
  • ui-auditLENSES_TO_RUNlensesToRun (the publish-safe module-global convention; an UPPERCASE module-global trips the Tangle obfuscator).
  • driver-loop / researcher-loop — one-line justification at the offline-box as unknown as SandboxInstance casts (the other casts already carried one).
  • strategy-evolution README — one line noting promoted: false at toy scale is the gate working, not a break.

Left alone (already good/great — no churn)

driver-loop, strategy-suite, supervisor-loop, chat-handler, recursive-supervisor, runtime-run, stream-backends, sanitized-telemetry-streaming, mcp-delegation, fleet-delegation, intelligence-recommend, intelligence-drop-in, agents-of-all-shapes, product-eval. The verdict's older snapshot flagged a few of these (mcp-delegation's delegate_ui_audit, the fleet-delegation casts) but they already carry the right framing/justification on current main; forcing fake-complete SandboxInstance helpers would add ceremony against the cleanup goal.

Verification

  • pnpm run build — clean
  • pnpm run typecheck (src + examples) — clean
  • pnpm run lint — clean (333 files)
  • pnpm run docs:check — green (0 errors; freshness OK)
  • pnpm test — 115 files / 1120 passed, 2 skipped (the env-gated live e2es)
  • ran each materially-changed example offline: knowledge-gating (hook fires), self-improving-loop, ui-audit, coding-benchmark (identical 0.944 leaderboard)

Apply the genuine-defect fixes from the three-persona example verdict, leaving
good/great examples untouched.
- self-improving-loop: add a loud minimum-evidence-floor caveat at the gate and
in the README — the demo gates at n=3 for runnability, but the production gate
floors at minSamples (8 in heldoutSignificance) / minProductiveRuns; never
ship a real change on n=3 (the small-n mirage).
- delegate: add a README, split a lean teaching delegate.ts + shared.ts from the
regression proof (moved to tests/delegate-example.test.ts — env-gated live e2e
+ an always-on offline fail-loud assertion); drop the test-in-example clothing
(E2E PASSED / process.exit) and the internal history from the header.
- improve: README symbol drift — ImprovementDriver → SurfaceProposer,
gepaDriver → gepaProposer (match what improve() actually builds); fix the test
path to src/improvement/improve.test.ts.
- knowledge-gating: wire the headline onKnowledgeBlocked hook into the adapter so
the README's documented hook is demonstrated (the blocked run now converts the
gap into a "would ask the user" decision).
- coding-benchmark: simplify offlineSolutions → offlineAgentScripts; keep the
rate-limiter cheat/real pair inline (the one anti-cheat teaching moment), move
csv/lru real impls to fixtures.ts as readable template literals (no escaped
strings). The held-out anti-cheat + smoke + firewall tests stay intact.
- supervise: wrap the flagship in main().catch (match the siblings) and uncomment
the completion-oracle deliverable so the headline models the safe path.
- ui-audit: LENSES_TO_RUN → lensesToRun (the publish-safe module-global convention).
- driver-loop / researcher-loop: one-line justification at the offline-box casts.
- strategy-evolution README: note promoted:false at toy scale is the gate working.

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved PR — efb6b428

Blanket team auto-approval is enabled for this reviewer service.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: blanket_auto_approve · 2026-06-24T15:09:11Z

@drewstone
drewstone merged commit 322211d into mainJun 24, 2026
1 check passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

docs(examples): three-persona cleanup — newcomer/senior/junior - #376

Merged
drewstone merged 1 commit into
mainfrom
examples/three-persona-cleanup
Jun 24, 2026
Merged

docs(examples): three-persona cleanup — newcomer/senior/junior#376
drewstone merged 1 commit into
mainfrom
examples/three-persona-cleanup

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

Applies the genuine-defect fixes from a three-persona (newcomer / senior / junior) review of the 22 examples. Scope: fix the things that actively mislead a reader, and clean the rough edges that teach the wrong habit — without churning the examples judged good/great. Every example still typechecks (typecheck:examples), the held-out coding-benchmark anti-cheat + its smoke test stay intact, and docs:check stays green.

Genuine defects fixed (these mislead a reader)

  • self-improving-loop — the demo gates at n=3 yet called it "the production held-out gate's statistical core" flatly, re-teaching this repo's documented fix: persist final runtime stream failures #1 failure mode (the small-n mirage). Added a loud minimum-evidence-floor caveat at the gate() definition AND in the README: the production gate floors the evidence (heldoutSignificance won't report a pair under minSamples, default 8; HeldOutGate rejects below minProductiveRuns with few_runs) — never ship a real change on n=3. (helps the junior most — the persona most likely to lift gate() verbatim).
  • delegate — was a Tier-1 example with no README, and the file was a CI test (E2E PASSED, process.exit) wearing an example's clothes, with internal history in the header. Added README.md; split a lean teaching delegate.ts + reusable shared.ts from the regression proof, which moved to tests/delegate-example.test.ts (env-gated: a paid live e2e when TANGLE_API_KEY is set, an always-on offline fail-loud assertion otherwise); stripped the history per the repo's no-history-in-source rule. (newcomer + senior).
  • improve — README named ImprovementDriver / gepaDriver; improve() actually builds SurfaceProposer / gepaProposer. Corrected both, plus the stale test path (src/improvement/improve.test.ts). (a README grep now resolves).
  • knowledge-gating — README featured adapter.onKnowledgeBlocked as the headline hook but the adapter never defined it (doc/code drift). Wired the hook into the adapter so the blocked run demonstrates it (converts the gap into a "would ask the user" decision that flows through as the stop reason).

Quality wins (remove ceremony / amateur tells, no behavior change)

  • coding-benchmark — renamed offlineSolutionsofflineAgentScripts with a clear 2-line header; kept the rate-limiter cheat/real pair inline (the one anti-cheat teaching moment), moved the round-invariant csv-parser / lru-cache real impls to fixtures.ts as readable template literals (no + '\n' + escaped-string ceremony). The held-out anti-cheat, the firewall test, the reps-don't-fake-n regression, and the BH-corrected stats are all unchanged and still pass.
  • supervise — wrapped the flagship in main().catch (matching the sibling examples) and uncommented the completion-oracle deliverable so the headline models the safe path.
  • ui-auditLENSES_TO_RUNlensesToRun (the publish-safe module-global convention; an UPPERCASE module-global trips the Tangle obfuscator).
  • driver-loop / researcher-loop — one-line justification at the offline-box as unknown as SandboxInstance casts (the other casts already carried one).
  • strategy-evolution README — one line noting promoted: false at toy scale is the gate working, not a break.

Left alone (already good/great — no churn)

driver-loop, strategy-suite, supervisor-loop, chat-handler, recursive-supervisor, runtime-run, stream-backends, sanitized-telemetry-streaming, mcp-delegation, fleet-delegation, intelligence-recommend, intelligence-drop-in, agents-of-all-shapes, product-eval. The verdict's older snapshot flagged a few of these (mcp-delegation's delegate_ui_audit, the fleet-delegation casts) but they already carry the right framing/justification on current main; forcing fake-complete SandboxInstance helpers would add ceremony against the cleanup goal.

Verification

  • pnpm run build — clean
  • pnpm run typecheck (src + examples) — clean
  • pnpm run lint — clean (333 files)
  • pnpm run docs:check — green (0 errors; freshness OK)
  • pnpm test — 115 files / 1120 passed, 2 skipped (the env-gated live e2es)
  • ran each materially-changed example offline: knowledge-gating (hook fires), self-improving-loop, ui-audit, coding-benchmark (identical 0.944 leaderboard)

Apply the genuine-defect fixes from the three-persona example verdict, leaving
good/great examples untouched.
- self-improving-loop: add a loud minimum-evidence-floor caveat at the gate and
in the README — the demo gates at n=3 for runnability, but the production gate
floors at minSamples (8 in heldoutSignificance) / minProductiveRuns; never
ship a real change on n=3 (the small-n mirage).
- delegate: add a README, split a lean teaching delegate.ts + shared.ts from the
regression proof (moved to tests/delegate-example.test.ts — env-gated live e2e
+ an always-on offline fail-loud assertion); drop the test-in-example clothing
(E2E PASSED / process.exit) and the internal history from the header.
- improve: README symbol drift — ImprovementDriver → SurfaceProposer,
gepaDriver → gepaProposer (match what improve() actually builds); fix the test
path to src/improvement/improve.test.ts.
- knowledge-gating: wire the headline onKnowledgeBlocked hook into the adapter so
the README's documented hook is demonstrated (the blocked run now converts the
gap into a "would ask the user" decision).
- coding-benchmark: simplify offlineSolutions → offlineAgentScripts; keep the
rate-limiter cheat/real pair inline (the one anti-cheat teaching moment), move
csv/lru real impls to fixtures.ts as readable template literals (no escaped
strings). The held-out anti-cheat + smoke + firewall tests stay intact.
- supervise: wrap the flagship in main().catch (match the siblings) and uncomment
the completion-oracle deliverable so the headline models the safe path.
- ui-audit: LENSES_TO_RUN → lensesToRun (the publish-safe module-global convention).
- driver-loop / researcher-loop: one-line justification at the offline-box casts.
- strategy-evolution README: note promoted:false at toy scale is the gate working.

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved PR — efb6b428

Blanket team auto-approval is enabled for this reviewer service.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: blanket_auto_approve · 2026-06-24T15:09:11Z

@drewstone
drewstone merged commit 322211d into mainJun 24, 2026
1 check passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

docs(examples): three-persona cleanup — newcomer/senior/junior - #376

Merged
drewstone merged 1 commit into
mainfrom
examples/three-persona-cleanup
Jun 24, 2026
Merged

docs(examples): three-persona cleanup — newcomer/senior/junior#376
drewstone merged 1 commit into
mainfrom
examples/three-persona-cleanup

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

Applies the genuine-defect fixes from a three-persona (newcomer / senior / junior) review of the 22 examples. Scope: fix the things that actively mislead a reader, and clean the rough edges that teach the wrong habit — without churning the examples judged good/great. Every example still typechecks (typecheck:examples), the held-out coding-benchmark anti-cheat + its smoke test stay intact, and docs:check stays green.

Genuine defects fixed (these mislead a reader)

  • self-improving-loop — the demo gates at n=3 yet called it "the production held-out gate's statistical core" flatly, re-teaching this repo's documented fix: persist final runtime stream failures #1 failure mode (the small-n mirage). Added a loud minimum-evidence-floor caveat at the gate() definition AND in the README: the production gate floors the evidence (heldoutSignificance won't report a pair under minSamples, default 8; HeldOutGate rejects below minProductiveRuns with few_runs) — never ship a real change on n=3. (helps the junior most — the persona most likely to lift gate() verbatim).
  • delegate — was a Tier-1 example with no README, and the file was a CI test (E2E PASSED, process.exit) wearing an example's clothes, with internal history in the header. Added README.md; split a lean teaching delegate.ts + reusable shared.ts from the regression proof, which moved to tests/delegate-example.test.ts (env-gated: a paid live e2e when TANGLE_API_KEY is set, an always-on offline fail-loud assertion otherwise); stripped the history per the repo's no-history-in-source rule. (newcomer + senior).
  • improve — README named ImprovementDriver / gepaDriver; improve() actually builds SurfaceProposer / gepaProposer. Corrected both, plus the stale test path (src/improvement/improve.test.ts). (a README grep now resolves).
  • knowledge-gating — README featured adapter.onKnowledgeBlocked as the headline hook but the adapter never defined it (doc/code drift). Wired the hook into the adapter so the blocked run demonstrates it (converts the gap into a "would ask the user" decision that flows through as the stop reason).

Quality wins (remove ceremony / amateur tells, no behavior change)

  • coding-benchmark — renamed offlineSolutionsofflineAgentScripts with a clear 2-line header; kept the rate-limiter cheat/real pair inline (the one anti-cheat teaching moment), moved the round-invariant csv-parser / lru-cache real impls to fixtures.ts as readable template literals (no + '\n' + escaped-string ceremony). The held-out anti-cheat, the firewall test, the reps-don't-fake-n regression, and the BH-corrected stats are all unchanged and still pass.
  • supervise — wrapped the flagship in main().catch (matching the sibling examples) and uncommented the completion-oracle deliverable so the headline models the safe path.
  • ui-auditLENSES_TO_RUNlensesToRun (the publish-safe module-global convention; an UPPERCASE module-global trips the Tangle obfuscator).
  • driver-loop / researcher-loop — one-line justification at the offline-box as unknown as SandboxInstance casts (the other casts already carried one).
  • strategy-evolution README — one line noting promoted: false at toy scale is the gate working, not a break.

Left alone (already good/great — no churn)

driver-loop, strategy-suite, supervisor-loop, chat-handler, recursive-supervisor, runtime-run, stream-backends, sanitized-telemetry-streaming, mcp-delegation, fleet-delegation, intelligence-recommend, intelligence-drop-in, agents-of-all-shapes, product-eval. The verdict's older snapshot flagged a few of these (mcp-delegation's delegate_ui_audit, the fleet-delegation casts) but they already carry the right framing/justification on current main; forcing fake-complete SandboxInstance helpers would add ceremony against the cleanup goal.

Verification

  • pnpm run build — clean
  • pnpm run typecheck (src + examples) — clean
  • pnpm run lint — clean (333 files)
  • pnpm run docs:check — green (0 errors; freshness OK)
  • pnpm test — 115 files / 1120 passed, 2 skipped (the env-gated live e2es)
  • ran each materially-changed example offline: knowledge-gating (hook fires), self-improving-loop, ui-audit, coding-benchmark (identical 0.944 leaderboard)

Apply the genuine-defect fixes from the three-persona example verdict, leaving
good/great examples untouched.
- self-improving-loop: add a loud minimum-evidence-floor caveat at the gate and
in the README — the demo gates at n=3 for runnability, but the production gate
floors at minSamples (8 in heldoutSignificance) / minProductiveRuns; never
ship a real change on n=3 (the small-n mirage).
- delegate: add a README, split a lean teaching delegate.ts + shared.ts from the
regression proof (moved to tests/delegate-example.test.ts — env-gated live e2e
+ an always-on offline fail-loud assertion); drop the test-in-example clothing
(E2E PASSED / process.exit) and the internal history from the header.
- improve: README symbol drift — ImprovementDriver → SurfaceProposer,
gepaDriver → gepaProposer (match what improve() actually builds); fix the test
path to src/improvement/improve.test.ts.
- knowledge-gating: wire the headline onKnowledgeBlocked hook into the adapter so
the README's documented hook is demonstrated (the blocked run now converts the
gap into a "would ask the user" decision).
- coding-benchmark: simplify offlineSolutions → offlineAgentScripts; keep the
rate-limiter cheat/real pair inline (the one anti-cheat teaching moment), move
csv/lru real impls to fixtures.ts as readable template literals (no escaped
strings). The held-out anti-cheat + smoke + firewall tests stay intact.
- supervise: wrap the flagship in main().catch (match the siblings) and uncomment
the completion-oracle deliverable so the headline models the safe path.
- ui-audit: LENSES_TO_RUN → lensesToRun (the publish-safe module-global convention).
- driver-loop / researcher-loop: one-line justification at the offline-box casts.
- strategy-evolution README: note promoted:false at toy scale is the gate working.

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved PR — efb6b428

Blanket team auto-approval is enabled for this reviewer service.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: blanket_auto_approve · 2026-06-24T15:09:11Z

@drewstone
drewstone merged commit 322211d into mainJun 24, 2026
1 check passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

docs(examples): three-persona cleanup — newcomer/senior/junior - #376

Merged
drewstone merged 1 commit into
mainfrom
examples/three-persona-cleanup
Jun 24, 2026
Merged

docs(examples): three-persona cleanup — newcomer/senior/junior#376
drewstone merged 1 commit into
mainfrom
examples/three-persona-cleanup

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

Applies the genuine-defect fixes from a three-persona (newcomer / senior / junior) review of the 22 examples. Scope: fix the things that actively mislead a reader, and clean the rough edges that teach the wrong habit — without churning the examples judged good/great. Every example still typechecks (typecheck:examples), the held-out coding-benchmark anti-cheat + its smoke test stay intact, and docs:check stays green.

Genuine defects fixed (these mislead a reader)

  • self-improving-loop — the demo gates at n=3 yet called it "the production held-out gate's statistical core" flatly, re-teaching this repo's documented fix: persist final runtime stream failures #1 failure mode (the small-n mirage). Added a loud minimum-evidence-floor caveat at the gate() definition AND in the README: the production gate floors the evidence (heldoutSignificance won't report a pair under minSamples, default 8; HeldOutGate rejects below minProductiveRuns with few_runs) — never ship a real change on n=3. (helps the junior most — the persona most likely to lift gate() verbatim).
  • delegate — was a Tier-1 example with no README, and the file was a CI test (E2E PASSED, process.exit) wearing an example's clothes, with internal history in the header. Added README.md; split a lean teaching delegate.ts + reusable shared.ts from the regression proof, which moved to tests/delegate-example.test.ts (env-gated: a paid live e2e when TANGLE_API_KEY is set, an always-on offline fail-loud assertion otherwise); stripped the history per the repo's no-history-in-source rule. (newcomer + senior).
  • improve — README named ImprovementDriver / gepaDriver; improve() actually builds SurfaceProposer / gepaProposer. Corrected both, plus the stale test path (src/improvement/improve.test.ts). (a README grep now resolves).
  • knowledge-gating — README featured adapter.onKnowledgeBlocked as the headline hook but the adapter never defined it (doc/code drift). Wired the hook into the adapter so the blocked run demonstrates it (converts the gap into a "would ask the user" decision that flows through as the stop reason).

Quality wins (remove ceremony / amateur tells, no behavior change)

  • coding-benchmark — renamed offlineSolutionsofflineAgentScripts with a clear 2-line header; kept the rate-limiter cheat/real pair inline (the one anti-cheat teaching moment), moved the round-invariant csv-parser / lru-cache real impls to fixtures.ts as readable template literals (no + '\n' + escaped-string ceremony). The held-out anti-cheat, the firewall test, the reps-don't-fake-n regression, and the BH-corrected stats are all unchanged and still pass.
  • supervise — wrapped the flagship in main().catch (matching the sibling examples) and uncommented the completion-oracle deliverable so the headline models the safe path.
  • ui-auditLENSES_TO_RUNlensesToRun (the publish-safe module-global convention; an UPPERCASE module-global trips the Tangle obfuscator).
  • driver-loop / researcher-loop — one-line justification at the offline-box as unknown as SandboxInstance casts (the other casts already carried one).
  • strategy-evolution README — one line noting promoted: false at toy scale is the gate working, not a break.

Left alone (already good/great — no churn)

driver-loop, strategy-suite, supervisor-loop, chat-handler, recursive-supervisor, runtime-run, stream-backends, sanitized-telemetry-streaming, mcp-delegation, fleet-delegation, intelligence-recommend, intelligence-drop-in, agents-of-all-shapes, product-eval. The verdict's older snapshot flagged a few of these (mcp-delegation's delegate_ui_audit, the fleet-delegation casts) but they already carry the right framing/justification on current main; forcing fake-complete SandboxInstance helpers would add ceremony against the cleanup goal.

Verification

  • pnpm run build — clean
  • pnpm run typecheck (src + examples) — clean
  • pnpm run lint — clean (333 files)
  • pnpm run docs:check — green (0 errors; freshness OK)
  • pnpm test — 115 files / 1120 passed, 2 skipped (the env-gated live e2es)
  • ran each materially-changed example offline: knowledge-gating (hook fires), self-improving-loop, ui-audit, coding-benchmark (identical 0.944 leaderboard)

Apply the genuine-defect fixes from the three-persona example verdict, leaving
good/great examples untouched.
- self-improving-loop: add a loud minimum-evidence-floor caveat at the gate and
in the README — the demo gates at n=3 for runnability, but the production gate
floors at minSamples (8 in heldoutSignificance) / minProductiveRuns; never
ship a real change on n=3 (the small-n mirage).
- delegate: add a README, split a lean teaching delegate.ts + shared.ts from the
regression proof (moved to tests/delegate-example.test.ts — env-gated live e2e
+ an always-on offline fail-loud assertion); drop the test-in-example clothing
(E2E PASSED / process.exit) and the internal history from the header.
- improve: README symbol drift — ImprovementDriver → SurfaceProposer,
gepaDriver → gepaProposer (match what improve() actually builds); fix the test
path to src/improvement/improve.test.ts.
- knowledge-gating: wire the headline onKnowledgeBlocked hook into the adapter so
the README's documented hook is demonstrated (the blocked run now converts the
gap into a "would ask the user" decision).
- coding-benchmark: simplify offlineSolutions → offlineAgentScripts; keep the
rate-limiter cheat/real pair inline (the one anti-cheat teaching moment), move
csv/lru real impls to fixtures.ts as readable template literals (no escaped
strings). The held-out anti-cheat + smoke + firewall tests stay intact.
- supervise: wrap the flagship in main().catch (match the siblings) and uncomment
the completion-oracle deliverable so the headline models the safe path.
- ui-audit: LENSES_TO_RUN → lensesToRun (the publish-safe module-global convention).
- driver-loop / researcher-loop: one-line justification at the offline-box casts.
- strategy-evolution README: note promoted:false at toy scale is the gate working.

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved PR — efb6b428

Blanket team auto-approval is enabled for this reviewer service.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: blanket_auto_approve · 2026-06-24T15:09:11Z

@drewstone
drewstone merged commit 322211d into mainJun 24, 2026
1 check passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

docs(examples): three-persona cleanup — newcomer/senior/junior - #376

Merged
drewstone merged 1 commit into
mainfrom
examples/three-persona-cleanup
Jun 24, 2026
Merged

docs(examples): three-persona cleanup — newcomer/senior/junior#376
drewstone merged 1 commit into
mainfrom
examples/three-persona-cleanup

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

Applies the genuine-defect fixes from a three-persona (newcomer / senior / junior) review of the 22 examples. Scope: fix the things that actively mislead a reader, and clean the rough edges that teach the wrong habit — without churning the examples judged good/great. Every example still typechecks (typecheck:examples), the held-out coding-benchmark anti-cheat + its smoke test stay intact, and docs:check stays green.

Genuine defects fixed (these mislead a reader)

  • self-improving-loop — the demo gates at n=3 yet called it "the production held-out gate's statistical core" flatly, re-teaching this repo's documented fix: persist final runtime stream failures #1 failure mode (the small-n mirage). Added a loud minimum-evidence-floor caveat at the gate() definition AND in the README: the production gate floors the evidence (heldoutSignificance won't report a pair under minSamples, default 8; HeldOutGate rejects below minProductiveRuns with few_runs) — never ship a real change on n=3. (helps the junior most — the persona most likely to lift gate() verbatim).
  • delegate — was a Tier-1 example with no README, and the file was a CI test (E2E PASSED, process.exit) wearing an example's clothes, with internal history in the header. Added README.md; split a lean teaching delegate.ts + reusable shared.ts from the regression proof, which moved to tests/delegate-example.test.ts (env-gated: a paid live e2e when TANGLE_API_KEY is set, an always-on offline fail-loud assertion otherwise); stripped the history per the repo's no-history-in-source rule. (newcomer + senior).
  • improve — README named ImprovementDriver / gepaDriver; improve() actually builds SurfaceProposer / gepaProposer. Corrected both, plus the stale test path (src/improvement/improve.test.ts). (a README grep now resolves).
  • knowledge-gating — README featured adapter.onKnowledgeBlocked as the headline hook but the adapter never defined it (doc/code drift). Wired the hook into the adapter so the blocked run demonstrates it (converts the gap into a "would ask the user" decision that flows through as the stop reason).

Quality wins (remove ceremony / amateur tells, no behavior change)

  • coding-benchmark — renamed offlineSolutionsofflineAgentScripts with a clear 2-line header; kept the rate-limiter cheat/real pair inline (the one anti-cheat teaching moment), moved the round-invariant csv-parser / lru-cache real impls to fixtures.ts as readable template literals (no + '\n' + escaped-string ceremony). The held-out anti-cheat, the firewall test, the reps-don't-fake-n regression, and the BH-corrected stats are all unchanged and still pass.
  • supervise — wrapped the flagship in main().catch (matching the sibling examples) and uncommented the completion-oracle deliverable so the headline models the safe path.
  • ui-auditLENSES_TO_RUNlensesToRun (the publish-safe module-global convention; an UPPERCASE module-global trips the Tangle obfuscator).
  • driver-loop / researcher-loop — one-line justification at the offline-box as unknown as SandboxInstance casts (the other casts already carried one).
  • strategy-evolution README — one line noting promoted: false at toy scale is the gate working, not a break.

Left alone (already good/great — no churn)

driver-loop, strategy-suite, supervisor-loop, chat-handler, recursive-supervisor, runtime-run, stream-backends, sanitized-telemetry-streaming, mcp-delegation, fleet-delegation, intelligence-recommend, intelligence-drop-in, agents-of-all-shapes, product-eval. The verdict's older snapshot flagged a few of these (mcp-delegation's delegate_ui_audit, the fleet-delegation casts) but they already carry the right framing/justification on current main; forcing fake-complete SandboxInstance helpers would add ceremony against the cleanup goal.

Verification

  • pnpm run build — clean
  • pnpm run typecheck (src + examples) — clean
  • pnpm run lint — clean (333 files)
  • pnpm run docs:check — green (0 errors; freshness OK)
  • pnpm test — 115 files / 1120 passed, 2 skipped (the env-gated live e2es)
  • ran each materially-changed example offline: knowledge-gating (hook fires), self-improving-loop, ui-audit, coding-benchmark (identical 0.944 leaderboard)

Apply the genuine-defect fixes from the three-persona example verdict, leaving
good/great examples untouched.
- self-improving-loop: add a loud minimum-evidence-floor caveat at the gate and
in the README — the demo gates at n=3 for runnability, but the production gate
floors at minSamples (8 in heldoutSignificance) / minProductiveRuns; never
ship a real change on n=3 (the small-n mirage).
- delegate: add a README, split a lean teaching delegate.ts + shared.ts from the
regression proof (moved to tests/delegate-example.test.ts — env-gated live e2e
+ an always-on offline fail-loud assertion); drop the test-in-example clothing
(E2E PASSED / process.exit) and the internal history from the header.
- improve: README symbol drift — ImprovementDriver → SurfaceProposer,
gepaDriver → gepaProposer (match what improve() actually builds); fix the test
path to src/improvement/improve.test.ts.
- knowledge-gating: wire the headline onKnowledgeBlocked hook into the adapter so
the README's documented hook is demonstrated (the blocked run now converts the
gap into a "would ask the user" decision).
- coding-benchmark: simplify offlineSolutions → offlineAgentScripts; keep the
rate-limiter cheat/real pair inline (the one anti-cheat teaching moment), move
csv/lru real impls to fixtures.ts as readable template literals (no escaped
strings). The held-out anti-cheat + smoke + firewall tests stay intact.
- supervise: wrap the flagship in main().catch (match the siblings) and uncomment
the completion-oracle deliverable so the headline models the safe path.
- ui-audit: LENSES_TO_RUN → lensesToRun (the publish-safe module-global convention).
- driver-loop / researcher-loop: one-line justification at the offline-box casts.
- strategy-evolution README: note promoted:false at toy scale is the gate working.

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved PR — efb6b428

Blanket team auto-approval is enabled for this reviewer service.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: blanket_auto_approve · 2026-06-24T15:09:11Z

@drewstone
drewstone merged commit 322211d into mainJun 24, 2026
1 check passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

docs(examples): three-persona cleanup — newcomer/senior/junior - #376

Merged
drewstone merged 1 commit into
mainfrom
examples/three-persona-cleanup
Jun 24, 2026
Merged

docs(examples): three-persona cleanup — newcomer/senior/junior#376
drewstone merged 1 commit into
mainfrom
examples/three-persona-cleanup

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

Applies the genuine-defect fixes from a three-persona (newcomer / senior / junior) review of the 22 examples. Scope: fix the things that actively mislead a reader, and clean the rough edges that teach the wrong habit — without churning the examples judged good/great. Every example still typechecks (typecheck:examples), the held-out coding-benchmark anti-cheat + its smoke test stay intact, and docs:check stays green.

Genuine defects fixed (these mislead a reader)

  • self-improving-loop — the demo gates at n=3 yet called it "the production held-out gate's statistical core" flatly, re-teaching this repo's documented fix: persist final runtime stream failures #1 failure mode (the small-n mirage). Added a loud minimum-evidence-floor caveat at the gate() definition AND in the README: the production gate floors the evidence (heldoutSignificance won't report a pair under minSamples, default 8; HeldOutGate rejects below minProductiveRuns with few_runs) — never ship a real change on n=3. (helps the junior most — the persona most likely to lift gate() verbatim).
  • delegate — was a Tier-1 example with no README, and the file was a CI test (E2E PASSED, process.exit) wearing an example's clothes, with internal history in the header. Added README.md; split a lean teaching delegate.ts + reusable shared.ts from the regression proof, which moved to tests/delegate-example.test.ts (env-gated: a paid live e2e when TANGLE_API_KEY is set, an always-on offline fail-loud assertion otherwise); stripped the history per the repo's no-history-in-source rule. (newcomer + senior).
  • improve — README named ImprovementDriver / gepaDriver; improve() actually builds SurfaceProposer / gepaProposer. Corrected both, plus the stale test path (src/improvement/improve.test.ts). (a README grep now resolves).
  • knowledge-gating — README featured adapter.onKnowledgeBlocked as the headline hook but the adapter never defined it (doc/code drift). Wired the hook into the adapter so the blocked run demonstrates it (converts the gap into a "would ask the user" decision that flows through as the stop reason).

Quality wins (remove ceremony / amateur tells, no behavior change)

  • coding-benchmark — renamed offlineSolutionsofflineAgentScripts with a clear 2-line header; kept the rate-limiter cheat/real pair inline (the one anti-cheat teaching moment), moved the round-invariant csv-parser / lru-cache real impls to fixtures.ts as readable template literals (no + '\n' + escaped-string ceremony). The held-out anti-cheat, the firewall test, the reps-don't-fake-n regression, and the BH-corrected stats are all unchanged and still pass.
  • supervise — wrapped the flagship in main().catch (matching the sibling examples) and uncommented the completion-oracle deliverable so the headline models the safe path.
  • ui-auditLENSES_TO_RUNlensesToRun (the publish-safe module-global convention; an UPPERCASE module-global trips the Tangle obfuscator).
  • driver-loop / researcher-loop — one-line justification at the offline-box as unknown as SandboxInstance casts (the other casts already carried one).
  • strategy-evolution README — one line noting promoted: false at toy scale is the gate working, not a break.

Left alone (already good/great — no churn)

driver-loop, strategy-suite, supervisor-loop, chat-handler, recursive-supervisor, runtime-run, stream-backends, sanitized-telemetry-streaming, mcp-delegation, fleet-delegation, intelligence-recommend, intelligence-drop-in, agents-of-all-shapes, product-eval. The verdict's older snapshot flagged a few of these (mcp-delegation's delegate_ui_audit, the fleet-delegation casts) but they already carry the right framing/justification on current main; forcing fake-complete SandboxInstance helpers would add ceremony against the cleanup goal.

Verification

  • pnpm run build — clean
  • pnpm run typecheck (src + examples) — clean
  • pnpm run lint — clean (333 files)
  • pnpm run docs:check — green (0 errors; freshness OK)
  • pnpm test — 115 files / 1120 passed, 2 skipped (the env-gated live e2es)
  • ran each materially-changed example offline: knowledge-gating (hook fires), self-improving-loop, ui-audit, coding-benchmark (identical 0.944 leaderboard)

Apply the genuine-defect fixes from the three-persona example verdict, leaving
good/great examples untouched.
- self-improving-loop: add a loud minimum-evidence-floor caveat at the gate and
in the README — the demo gates at n=3 for runnability, but the production gate
floors at minSamples (8 in heldoutSignificance) / minProductiveRuns; never
ship a real change on n=3 (the small-n mirage).
- delegate: add a README, split a lean teaching delegate.ts + shared.ts from the
regression proof (moved to tests/delegate-example.test.ts — env-gated live e2e
+ an always-on offline fail-loud assertion); drop the test-in-example clothing
(E2E PASSED / process.exit) and the internal history from the header.
- improve: README symbol drift — ImprovementDriver → SurfaceProposer,
gepaDriver → gepaProposer (match what improve() actually builds); fix the test
path to src/improvement/improve.test.ts.
- knowledge-gating: wire the headline onKnowledgeBlocked hook into the adapter so
the README's documented hook is demonstrated (the blocked run now converts the
gap into a "would ask the user" decision).
- coding-benchmark: simplify offlineSolutions → offlineAgentScripts; keep the
rate-limiter cheat/real pair inline (the one anti-cheat teaching moment), move
csv/lru real impls to fixtures.ts as readable template literals (no escaped
strings). The held-out anti-cheat + smoke + firewall tests stay intact.
- supervise: wrap the flagship in main().catch (match the siblings) and uncomment
the completion-oracle deliverable so the headline models the safe path.
- ui-audit: LENSES_TO_RUN → lensesToRun (the publish-safe module-global convention).
- driver-loop / researcher-loop: one-line justification at the offline-box casts.
- strategy-evolution README: note promoted:false at toy scale is the gate working.

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved PR — efb6b428

Blanket team auto-approval is enabled for this reviewer service.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: blanket_auto_approve · 2026-06-24T15:09:11Z

@drewstone
drewstone merged commit 322211d into mainJun 24, 2026
1 check passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools