docs(canonical-api): close the anti-reinvention gaps + de-reinvent the examples - #370

Closed
drewstone wants to merge 1 commit into
mainfrom
docs/kill-example-reinvention
Closed

docs(canonical-api): close the anti-reinvention gaps + de-reinvent the examples#370
drewstone wants to merge 1 commit into
mainfrom
docs/kill-example-reinvention

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

What

The anti-reinvention decision table in docs/canonical-api.md (§2 — the doc agents are told to read before writing orchestration/measurement code) had zero rows for whole families of live exports, so agents kept hand-rolling them. This adds them, and de-reinvents the one example that hand-rolled primitives.

1. Reference gaps closed (docs/canonical-api.md §2)

Six grouped rows added, in the existing "I want to ___ → use ___ → NOT ___" voice:

FamilyUse (the real primitive)Was being hand-rolled as
statisticsevery test on the agent-eval main barrelpairedBootstrap, wilson, pairedTTest, wilcoxonSignedRank, mcnemar/mcnemarPower/mcnemarRequiredN, mannWhitneyU, passAtK, confidenceInterval, benjaminiHochberg/bonferroni, cohensD/cliffsDelta, requiredSampleSize/pairedMde, eProcess, weightedComposite, corpusInterRaterAgreement, pairedRiskDifference, pearsonR/spearmanR, paretoFrontier/dominatesa bootstrap-with-PRNG, a Wilson/McNemar/power calc, a Pareto sweep, a Cohen's-d per gate (the #1 measurement reinvention)
judgea campaign JudgeConfig, or ensembleJudge for a multi-model panela hand-built judge prompt loop, or the @deprecatedJudgeFn factories
authenticityscoreAuthenticitygateRealness over an AuthenticitySignals (agent-eval/authenticity subpath)a regex realness scorer, or trusting a buildability score that rewards a polished fake
verificationnew MultiLayerVerifier(layers).run(...) + gradeSemanticStatus(...)a hand-chained compile→lint→test→semantic pipeline
reference-replayrunReferenceReplay/scoreReferenceReplay + decideReferenceReplayPromotion over jsonlReferenceReplayStorea hand-rolled "does the candidate reproduce the reference" matcher
token/usage seamextractLlmCallEvent / reportLoopUsage (UsageSink) — /loopsre-walking sandbox events to tally tokens yourself

The judge row also kills a recurring confusion: there is no llmJudge export in this version — that name appears only in a campaign-types docstring; the live judge surface is JudgeConfig + ensembleJudge. A future agent reading this doc can no longer conclude "no judge primitive exists."

The freshness gate (scripts/check-docs-freshness.mjs) now loads the agent-eval/authenticity subpath barrel into the §2 export universe, so a row that recommends importing from that subpath resolves — exactly parallel to how contract/campaign/index are already loaded. This strengthens coverage (a real published export surface the gate previously couldn't see), it does not weaken the gate.

2. Example de-reinvented

examples/self-improving-loop was the only example (of 21 audited) that hand-rolled substrate primitives:

  • analyst — local interface AnalystFinding { rootCause; proposedMutation } + runAnalyst() → now the canonical AnalystFinding stamped by makeFinding (the exact shape improve(profile, findings, opts) reflects on); mutation rides recommended_action.
  • gate — bare v1Mean - v0Mean >= 0.5 point comparison → now pairedBootstrap(v0Scores, v1Scores, { seed }), shipping only when the paired-bootstrap CI lower bound clears 0 — the statistical core the production held-out gate (HeldOutGate / improve() over selfImprove) is built on.

Stays offline + deterministic; the demo still ships v1, now on a +5.00 paired median, 95% CI [5.00, 6.00], n=3. README + comments mirror the already-clean companion examples (improve/, intelligence-recommend/). Only the analyst body, proposer, and LLM remain scripted (for reproducibility) — the finding type and gate statistic are now the real substrate.

Verify (all green)

  • pnpm docs:checkdocs freshness: OK (§2 table symbols against 4134 public exports; prose symbols against 5795 resolvable; docs/api regen diff clean)
  • pnpm run build — clean
  • pnpm run typecheck + typecheck:examples — clean (the fixed example compiles)
  • pnpm run lint — Checked 319 files, no fixes applied
  • Example runs offline e2e: ships v1 on the paired CI [5.00, 6.00]

DO NOT MERGE — operator review.

…e examples
The anti-reinvention decision table (§2) had zero rows for whole families of
live exports, so agents kept hand-rolling them. Add grouped rows for:
- statistics — every test on the agent-eval main barrel (pairedBootstrap,
wilson, pairedTTest, wilcoxonSignedRank, mcnemar/mcnemarPower, mannWhitneyU,
passAtK, confidenceInterval, benjaminiHochberg/bonferroni, cohensD/cliffsDelta,
requiredSampleSize/pairedMde, eProcess, weightedComposite,
corpusInterRaterAgreement, pairedRiskDifference, pearsonR/spearmanR,
paretoFrontier/dominates)
- judge — JudgeConfig + ensembleJudge, with a warn-off for the @deprecated
JudgeFn factories and an explicit note that there is no llmJudge export
- authenticity — scoreAuthenticity / gateRealness / AuthenticitySignals
(agent-eval/authenticity subpath)
- verification — MultiLayerVerifier + gradeSemanticStatus
- reference-replay — runReferenceReplay / scoreReferenceReplay /
decideReferenceReplayPromotion over jsonlReferenceReplayStore
- token/usage seam — extractLlmCallEvent + reportLoopUsage / UsageSink (/loops)
Teach the freshness gate the agent-eval/authenticity subpath barrel so a §2 row
that recommends importing from it resolves (parallel to contract/campaign/index).
De-reinvent the one example that hand-rolled primitives:
examples/self-improving-loop replaces its local AnalystFinding interface +
runAnalyst with the canonical AnalystFinding stamped by makeFinding, and its
bare v1Mean-v0Mean>=0.5 point-comparison gate with pairedBootstrap (the
statistical core the production held-out gate is built on). Stays offline and
deterministic; ships v1 on a +5.00 [5.00, 6.00] paired CI. README + comments
mirror the clean companion examples (improve/, intelligence-recommend/).

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved PR — 70e098d8

Blanket team auto-approval is enabled for this reviewer service.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: blanket_auto_approve · 2026-06-24T08:54:31Z

drewstone added a commit that referenced this pull request Jun 24, 2026
…fImprove not a bare mean gate
Folds the one real example fix from #370 (otherwise superseded by the generated
catalog) into this PR: self-improving-loop hand-rolled the ship gate as a bare
mean comparison instead of the real HeldOutGate/selfImprove primitives.
@drewstone

Copy link
Copy Markdown
ContributorAuthor

Superseded by #371. The primitive inventory is now GENERATED (docs/api/primitive-catalog.md, freshness-gate-enforced) so it can't go stale by hand — the root cause this PR hand-patched. The ~30 primitives catalogued here are now auto-generated; the one real example fix (self-improving-loop → HeldOutGate/selfImprove) was folded into #371.

drewstone added a commit that referenced this pull request Jun 24, 2026
…erence cannot go stale (#371)
* docs(api): generate the primitive catalog so the anti-reinvention reference cannot go stale
The hand-listed primitive inventory in docs/canonical-api.md drifted from source:
it had zero mentions of live exports (scoreAuthenticity, gateRealness,
MultiLayerVerifier, wilson, pairedTTest, runProfileMatrix, extractUsage, …). Anything
derivable from source must be generated, not hand-written — only judgment stays curated.
- scripts/gen-primitive-catalog.mjs reads the LIVE exports of (a) this package's own
public subpaths (from package.json `exports`) and (b) a curated category->subpath map
of the @tangle-network/agent-eval substrate surfaces agents should reuse (judge,
authenticity, verification, statistics, campaign, token/usage). Extraction is via the
TypeScript compiler API over a virtual re-export entry, so it follows aliased
re-exports and content-hashed bundle files — the exact things that rot a hand list.
Emits docs/api/primitive-catalog.md with a GENERATED header (name, import path,
one-line summary per export, grouped by surface).
- Wired into `docs:api` (runs after TypeDoc). The freshness gate gains a seventh class
(CATALOG): it regenerates the catalog to a temp file and byte-compares to the committed
copy, so a new/removed/renamed live export absent from the catalog is a RED BUILD.
- Shrank canonical-api.md: removed the export-inventory enumeration from the banner and
the §2 preamble, replaced with pointers to docs/api/primitive-catalog.md. Kept all the
judgment — the decision gate, §1.5 AgentProfile law, the §2 "I want to -> use -> NOT"
table and every "Do NOT". The version + substrate-peer pins stay (gate-enforced).
- MAINTAINING.md documents the generated-inventory layer, CLASS 7, and its fix path.
* chore(deps): bump agent-eval to 0.99.0 + regenerate primitive catalog
agent-eval 0.99.0 adds llmJudge (+ the full current judge/auth/verify/stats
surface); regenerating the generated catalog picks it up with zero hand-work,
which is the point of the generator. Lockfile was pinned at 0.97.0 (pre-llmJudge)
despite agent-eval already being in minimumReleaseAgeExclude.
* fix(deps): keep agent-eval peer floor at >=0.97.0
Only the examples (devDependency) need 0.99.0 for llmJudge; agent-runtime's src
does not, so the peer floor must not force consumers onto 0.99.0. Catalog +
lockfile stay on the resolved 0.99.0 so the examples get llmJudge.
* docs(examples): de-reinvent self-improving-loop — use HeldOutGate/selfImprove not a bare mean gate
Folds the one real example fix from #370 (otherwise superseded by the generated
catalog) into this PR: self-improving-loop hand-rolled the ship gate as a bare
mean comparison instead of the real HeldOutGate/selfImprove primitives.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

docs(canonical-api): close the anti-reinvention gaps + de-reinvent the examples - #370

Closed
drewstone wants to merge 1 commit into
mainfrom
docs/kill-example-reinvention
Closed

docs(canonical-api): close the anti-reinvention gaps + de-reinvent the examples#370
drewstone wants to merge 1 commit into
mainfrom
docs/kill-example-reinvention

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

What

The anti-reinvention decision table in docs/canonical-api.md (§2 — the doc agents are told to read before writing orchestration/measurement code) had zero rows for whole families of live exports, so agents kept hand-rolling them. This adds them, and de-reinvents the one example that hand-rolled primitives.

1. Reference gaps closed (docs/canonical-api.md §2)

Six grouped rows added, in the existing "I want to ___ → use ___ → NOT ___" voice:

FamilyUse (the real primitive)Was being hand-rolled as
statisticsevery test on the agent-eval main barrelpairedBootstrap, wilson, pairedTTest, wilcoxonSignedRank, mcnemar/mcnemarPower/mcnemarRequiredN, mannWhitneyU, passAtK, confidenceInterval, benjaminiHochberg/bonferroni, cohensD/cliffsDelta, requiredSampleSize/pairedMde, eProcess, weightedComposite, corpusInterRaterAgreement, pairedRiskDifference, pearsonR/spearmanR, paretoFrontier/dominatesa bootstrap-with-PRNG, a Wilson/McNemar/power calc, a Pareto sweep, a Cohen's-d per gate (the #1 measurement reinvention)
judgea campaign JudgeConfig, or ensembleJudge for a multi-model panela hand-built judge prompt loop, or the @deprecatedJudgeFn factories
authenticityscoreAuthenticitygateRealness over an AuthenticitySignals (agent-eval/authenticity subpath)a regex realness scorer, or trusting a buildability score that rewards a polished fake
verificationnew MultiLayerVerifier(layers).run(...) + gradeSemanticStatus(...)a hand-chained compile→lint→test→semantic pipeline
reference-replayrunReferenceReplay/scoreReferenceReplay + decideReferenceReplayPromotion over jsonlReferenceReplayStorea hand-rolled "does the candidate reproduce the reference" matcher
token/usage seamextractLlmCallEvent / reportLoopUsage (UsageSink) — /loopsre-walking sandbox events to tally tokens yourself

The judge row also kills a recurring confusion: there is no llmJudge export in this version — that name appears only in a campaign-types docstring; the live judge surface is JudgeConfig + ensembleJudge. A future agent reading this doc can no longer conclude "no judge primitive exists."

The freshness gate (scripts/check-docs-freshness.mjs) now loads the agent-eval/authenticity subpath barrel into the §2 export universe, so a row that recommends importing from that subpath resolves — exactly parallel to how contract/campaign/index are already loaded. This strengthens coverage (a real published export surface the gate previously couldn't see), it does not weaken the gate.

2. Example de-reinvented

examples/self-improving-loop was the only example (of 21 audited) that hand-rolled substrate primitives:

  • analyst — local interface AnalystFinding { rootCause; proposedMutation } + runAnalyst() → now the canonical AnalystFinding stamped by makeFinding (the exact shape improve(profile, findings, opts) reflects on); mutation rides recommended_action.
  • gate — bare v1Mean - v0Mean >= 0.5 point comparison → now pairedBootstrap(v0Scores, v1Scores, { seed }), shipping only when the paired-bootstrap CI lower bound clears 0 — the statistical core the production held-out gate (HeldOutGate / improve() over selfImprove) is built on.

Stays offline + deterministic; the demo still ships v1, now on a +5.00 paired median, 95% CI [5.00, 6.00], n=3. README + comments mirror the already-clean companion examples (improve/, intelligence-recommend/). Only the analyst body, proposer, and LLM remain scripted (for reproducibility) — the finding type and gate statistic are now the real substrate.

Verify (all green)

  • pnpm docs:checkdocs freshness: OK (§2 table symbols against 4134 public exports; prose symbols against 5795 resolvable; docs/api regen diff clean)
  • pnpm run build — clean
  • pnpm run typecheck + typecheck:examples — clean (the fixed example compiles)
  • pnpm run lint — Checked 319 files, no fixes applied
  • Example runs offline e2e: ships v1 on the paired CI [5.00, 6.00]

DO NOT MERGE — operator review.

…e examples
The anti-reinvention decision table (§2) had zero rows for whole families of
live exports, so agents kept hand-rolling them. Add grouped rows for:
- statistics — every test on the agent-eval main barrel (pairedBootstrap,
wilson, pairedTTest, wilcoxonSignedRank, mcnemar/mcnemarPower, mannWhitneyU,
passAtK, confidenceInterval, benjaminiHochberg/bonferroni, cohensD/cliffsDelta,
requiredSampleSize/pairedMde, eProcess, weightedComposite,
corpusInterRaterAgreement, pairedRiskDifference, pearsonR/spearmanR,
paretoFrontier/dominates)
- judge — JudgeConfig + ensembleJudge, with a warn-off for the @deprecated
JudgeFn factories and an explicit note that there is no llmJudge export
- authenticity — scoreAuthenticity / gateRealness / AuthenticitySignals
(agent-eval/authenticity subpath)
- verification — MultiLayerVerifier + gradeSemanticStatus
- reference-replay — runReferenceReplay / scoreReferenceReplay /
decideReferenceReplayPromotion over jsonlReferenceReplayStore
- token/usage seam — extractLlmCallEvent + reportLoopUsage / UsageSink (/loops)
Teach the freshness gate the agent-eval/authenticity subpath barrel so a §2 row
that recommends importing from it resolves (parallel to contract/campaign/index).
De-reinvent the one example that hand-rolled primitives:
examples/self-improving-loop replaces its local AnalystFinding interface +
runAnalyst with the canonical AnalystFinding stamped by makeFinding, and its
bare v1Mean-v0Mean>=0.5 point-comparison gate with pairedBootstrap (the
statistical core the production held-out gate is built on). Stays offline and
deterministic; ships v1 on a +5.00 [5.00, 6.00] paired CI. README + comments
mirror the clean companion examples (improve/, intelligence-recommend/).

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved PR — 70e098d8

Blanket team auto-approval is enabled for this reviewer service.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: blanket_auto_approve · 2026-06-24T08:54:31Z

drewstone added a commit that referenced this pull request Jun 24, 2026
…fImprove not a bare mean gate
Folds the one real example fix from #370 (otherwise superseded by the generated
catalog) into this PR: self-improving-loop hand-rolled the ship gate as a bare
mean comparison instead of the real HeldOutGate/selfImprove primitives.
@drewstone

Copy link
Copy Markdown
ContributorAuthor

Superseded by #371. The primitive inventory is now GENERATED (docs/api/primitive-catalog.md, freshness-gate-enforced) so it can't go stale by hand — the root cause this PR hand-patched. The ~30 primitives catalogued here are now auto-generated; the one real example fix (self-improving-loop → HeldOutGate/selfImprove) was folded into #371.

drewstone added a commit that referenced this pull request Jun 24, 2026
…erence cannot go stale (#371)
* docs(api): generate the primitive catalog so the anti-reinvention reference cannot go stale
The hand-listed primitive inventory in docs/canonical-api.md drifted from source:
it had zero mentions of live exports (scoreAuthenticity, gateRealness,
MultiLayerVerifier, wilson, pairedTTest, runProfileMatrix, extractUsage, …). Anything
derivable from source must be generated, not hand-written — only judgment stays curated.
- scripts/gen-primitive-catalog.mjs reads the LIVE exports of (a) this package's own
public subpaths (from package.json `exports`) and (b) a curated category->subpath map
of the @tangle-network/agent-eval substrate surfaces agents should reuse (judge,
authenticity, verification, statistics, campaign, token/usage). Extraction is via the
TypeScript compiler API over a virtual re-export entry, so it follows aliased
re-exports and content-hashed bundle files — the exact things that rot a hand list.
Emits docs/api/primitive-catalog.md with a GENERATED header (name, import path,
one-line summary per export, grouped by surface).
- Wired into `docs:api` (runs after TypeDoc). The freshness gate gains a seventh class
(CATALOG): it regenerates the catalog to a temp file and byte-compares to the committed
copy, so a new/removed/renamed live export absent from the catalog is a RED BUILD.
- Shrank canonical-api.md: removed the export-inventory enumeration from the banner and
the §2 preamble, replaced with pointers to docs/api/primitive-catalog.md. Kept all the
judgment — the decision gate, §1.5 AgentProfile law, the §2 "I want to -> use -> NOT"
table and every "Do NOT". The version + substrate-peer pins stay (gate-enforced).
- MAINTAINING.md documents the generated-inventory layer, CLASS 7, and its fix path.
* chore(deps): bump agent-eval to 0.99.0 + regenerate primitive catalog
agent-eval 0.99.0 adds llmJudge (+ the full current judge/auth/verify/stats
surface); regenerating the generated catalog picks it up with zero hand-work,
which is the point of the generator. Lockfile was pinned at 0.97.0 (pre-llmJudge)
despite agent-eval already being in minimumReleaseAgeExclude.
* fix(deps): keep agent-eval peer floor at >=0.97.0
Only the examples (devDependency) need 0.99.0 for llmJudge; agent-runtime's src
does not, so the peer floor must not force consumers onto 0.99.0. Catalog +
lockfile stay on the resolved 0.99.0 so the examples get llmJudge.
* docs(examples): de-reinvent self-improving-loop — use HeldOutGate/selfImprove not a bare mean gate
Folds the one real example fix from #370 (otherwise superseded by the generated
catalog) into this PR: self-improving-loop hand-rolled the ship gate as a bare
mean comparison instead of the real HeldOutGate/selfImprove primitives.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

docs(canonical-api): close the anti-reinvention gaps + de-reinvent the examples - #370

Closed
drewstone wants to merge 1 commit into
mainfrom
docs/kill-example-reinvention
Closed

docs(canonical-api): close the anti-reinvention gaps + de-reinvent the examples#370
drewstone wants to merge 1 commit into
mainfrom
docs/kill-example-reinvention

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

What

The anti-reinvention decision table in docs/canonical-api.md (§2 — the doc agents are told to read before writing orchestration/measurement code) had zero rows for whole families of live exports, so agents kept hand-rolling them. This adds them, and de-reinvents the one example that hand-rolled primitives.

1. Reference gaps closed (docs/canonical-api.md §2)

Six grouped rows added, in the existing "I want to ___ → use ___ → NOT ___" voice:

FamilyUse (the real primitive)Was being hand-rolled as
statisticsevery test on the agent-eval main barrelpairedBootstrap, wilson, pairedTTest, wilcoxonSignedRank, mcnemar/mcnemarPower/mcnemarRequiredN, mannWhitneyU, passAtK, confidenceInterval, benjaminiHochberg/bonferroni, cohensD/cliffsDelta, requiredSampleSize/pairedMde, eProcess, weightedComposite, corpusInterRaterAgreement, pairedRiskDifference, pearsonR/spearmanR, paretoFrontier/dominatesa bootstrap-with-PRNG, a Wilson/McNemar/power calc, a Pareto sweep, a Cohen's-d per gate (the #1 measurement reinvention)
judgea campaign JudgeConfig, or ensembleJudge for a multi-model panela hand-built judge prompt loop, or the @deprecatedJudgeFn factories
authenticityscoreAuthenticitygateRealness over an AuthenticitySignals (agent-eval/authenticity subpath)a regex realness scorer, or trusting a buildability score that rewards a polished fake
verificationnew MultiLayerVerifier(layers).run(...) + gradeSemanticStatus(...)a hand-chained compile→lint→test→semantic pipeline
reference-replayrunReferenceReplay/scoreReferenceReplay + decideReferenceReplayPromotion over jsonlReferenceReplayStorea hand-rolled "does the candidate reproduce the reference" matcher
token/usage seamextractLlmCallEvent / reportLoopUsage (UsageSink) — /loopsre-walking sandbox events to tally tokens yourself

The judge row also kills a recurring confusion: there is no llmJudge export in this version — that name appears only in a campaign-types docstring; the live judge surface is JudgeConfig + ensembleJudge. A future agent reading this doc can no longer conclude "no judge primitive exists."

The freshness gate (scripts/check-docs-freshness.mjs) now loads the agent-eval/authenticity subpath barrel into the §2 export universe, so a row that recommends importing from that subpath resolves — exactly parallel to how contract/campaign/index are already loaded. This strengthens coverage (a real published export surface the gate previously couldn't see), it does not weaken the gate.

2. Example de-reinvented

examples/self-improving-loop was the only example (of 21 audited) that hand-rolled substrate primitives:

  • analyst — local interface AnalystFinding { rootCause; proposedMutation } + runAnalyst() → now the canonical AnalystFinding stamped by makeFinding (the exact shape improve(profile, findings, opts) reflects on); mutation rides recommended_action.
  • gate — bare v1Mean - v0Mean >= 0.5 point comparison → now pairedBootstrap(v0Scores, v1Scores, { seed }), shipping only when the paired-bootstrap CI lower bound clears 0 — the statistical core the production held-out gate (HeldOutGate / improve() over selfImprove) is built on.

Stays offline + deterministic; the demo still ships v1, now on a +5.00 paired median, 95% CI [5.00, 6.00], n=3. README + comments mirror the already-clean companion examples (improve/, intelligence-recommend/). Only the analyst body, proposer, and LLM remain scripted (for reproducibility) — the finding type and gate statistic are now the real substrate.

Verify (all green)

  • pnpm docs:checkdocs freshness: OK (§2 table symbols against 4134 public exports; prose symbols against 5795 resolvable; docs/api regen diff clean)
  • pnpm run build — clean
  • pnpm run typecheck + typecheck:examples — clean (the fixed example compiles)
  • pnpm run lint — Checked 319 files, no fixes applied
  • Example runs offline e2e: ships v1 on the paired CI [5.00, 6.00]

DO NOT MERGE — operator review.

…e examples
The anti-reinvention decision table (§2) had zero rows for whole families of
live exports, so agents kept hand-rolling them. Add grouped rows for:
- statistics — every test on the agent-eval main barrel (pairedBootstrap,
wilson, pairedTTest, wilcoxonSignedRank, mcnemar/mcnemarPower, mannWhitneyU,
passAtK, confidenceInterval, benjaminiHochberg/bonferroni, cohensD/cliffsDelta,
requiredSampleSize/pairedMde, eProcess, weightedComposite,
corpusInterRaterAgreement, pairedRiskDifference, pearsonR/spearmanR,
paretoFrontier/dominates)
- judge — JudgeConfig + ensembleJudge, with a warn-off for the @deprecated
JudgeFn factories and an explicit note that there is no llmJudge export
- authenticity — scoreAuthenticity / gateRealness / AuthenticitySignals
(agent-eval/authenticity subpath)
- verification — MultiLayerVerifier + gradeSemanticStatus
- reference-replay — runReferenceReplay / scoreReferenceReplay /
decideReferenceReplayPromotion over jsonlReferenceReplayStore
- token/usage seam — extractLlmCallEvent + reportLoopUsage / UsageSink (/loops)
Teach the freshness gate the agent-eval/authenticity subpath barrel so a §2 row
that recommends importing from it resolves (parallel to contract/campaign/index).
De-reinvent the one example that hand-rolled primitives:
examples/self-improving-loop replaces its local AnalystFinding interface +
runAnalyst with the canonical AnalystFinding stamped by makeFinding, and its
bare v1Mean-v0Mean>=0.5 point-comparison gate with pairedBootstrap (the
statistical core the production held-out gate is built on). Stays offline and
deterministic; ships v1 on a +5.00 [5.00, 6.00] paired CI. README + comments
mirror the clean companion examples (improve/, intelligence-recommend/).

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved PR — 70e098d8

Blanket team auto-approval is enabled for this reviewer service.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: blanket_auto_approve · 2026-06-24T08:54:31Z

drewstone added a commit that referenced this pull request Jun 24, 2026
…fImprove not a bare mean gate
Folds the one real example fix from #370 (otherwise superseded by the generated
catalog) into this PR: self-improving-loop hand-rolled the ship gate as a bare
mean comparison instead of the real HeldOutGate/selfImprove primitives.
@drewstone

Copy link
Copy Markdown
ContributorAuthor

Superseded by #371. The primitive inventory is now GENERATED (docs/api/primitive-catalog.md, freshness-gate-enforced) so it can't go stale by hand — the root cause this PR hand-patched. The ~30 primitives catalogued here are now auto-generated; the one real example fix (self-improving-loop → HeldOutGate/selfImprove) was folded into #371.

drewstone added a commit that referenced this pull request Jun 24, 2026
…erence cannot go stale (#371)
* docs(api): generate the primitive catalog so the anti-reinvention reference cannot go stale
The hand-listed primitive inventory in docs/canonical-api.md drifted from source:
it had zero mentions of live exports (scoreAuthenticity, gateRealness,
MultiLayerVerifier, wilson, pairedTTest, runProfileMatrix, extractUsage, …). Anything
derivable from source must be generated, not hand-written — only judgment stays curated.
- scripts/gen-primitive-catalog.mjs reads the LIVE exports of (a) this package's own
public subpaths (from package.json `exports`) and (b) a curated category->subpath map
of the @tangle-network/agent-eval substrate surfaces agents should reuse (judge,
authenticity, verification, statistics, campaign, token/usage). Extraction is via the
TypeScript compiler API over a virtual re-export entry, so it follows aliased
re-exports and content-hashed bundle files — the exact things that rot a hand list.
Emits docs/api/primitive-catalog.md with a GENERATED header (name, import path,
one-line summary per export, grouped by surface).
- Wired into `docs:api` (runs after TypeDoc). The freshness gate gains a seventh class
(CATALOG): it regenerates the catalog to a temp file and byte-compares to the committed
copy, so a new/removed/renamed live export absent from the catalog is a RED BUILD.
- Shrank canonical-api.md: removed the export-inventory enumeration from the banner and
the §2 preamble, replaced with pointers to docs/api/primitive-catalog.md. Kept all the
judgment — the decision gate, §1.5 AgentProfile law, the §2 "I want to -> use -> NOT"
table and every "Do NOT". The version + substrate-peer pins stay (gate-enforced).
- MAINTAINING.md documents the generated-inventory layer, CLASS 7, and its fix path.
* chore(deps): bump agent-eval to 0.99.0 + regenerate primitive catalog
agent-eval 0.99.0 adds llmJudge (+ the full current judge/auth/verify/stats
surface); regenerating the generated catalog picks it up with zero hand-work,
which is the point of the generator. Lockfile was pinned at 0.97.0 (pre-llmJudge)
despite agent-eval already being in minimumReleaseAgeExclude.
* fix(deps): keep agent-eval peer floor at >=0.97.0
Only the examples (devDependency) need 0.99.0 for llmJudge; agent-runtime's src
does not, so the peer floor must not force consumers onto 0.99.0. Catalog +
lockfile stay on the resolved 0.99.0 so the examples get llmJudge.
* docs(examples): de-reinvent self-improving-loop — use HeldOutGate/selfImprove not a bare mean gate
Folds the one real example fix from #370 (otherwise superseded by the generated
catalog) into this PR: self-improving-loop hand-rolled the ship gate as a bare
mean comparison instead of the real HeldOutGate/selfImprove primitives.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

docs(canonical-api): close the anti-reinvention gaps + de-reinvent the examples - #370

Closed
drewstone wants to merge 1 commit into
mainfrom
docs/kill-example-reinvention
Closed

docs(canonical-api): close the anti-reinvention gaps + de-reinvent the examples#370
drewstone wants to merge 1 commit into
mainfrom
docs/kill-example-reinvention

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

What

The anti-reinvention decision table in docs/canonical-api.md (§2 — the doc agents are told to read before writing orchestration/measurement code) had zero rows for whole families of live exports, so agents kept hand-rolling them. This adds them, and de-reinvents the one example that hand-rolled primitives.

1. Reference gaps closed (docs/canonical-api.md §2)

Six grouped rows added, in the existing "I want to ___ → use ___ → NOT ___" voice:

FamilyUse (the real primitive)Was being hand-rolled as
statisticsevery test on the agent-eval main barrelpairedBootstrap, wilson, pairedTTest, wilcoxonSignedRank, mcnemar/mcnemarPower/mcnemarRequiredN, mannWhitneyU, passAtK, confidenceInterval, benjaminiHochberg/bonferroni, cohensD/cliffsDelta, requiredSampleSize/pairedMde, eProcess, weightedComposite, corpusInterRaterAgreement, pairedRiskDifference, pearsonR/spearmanR, paretoFrontier/dominatesa bootstrap-with-PRNG, a Wilson/McNemar/power calc, a Pareto sweep, a Cohen's-d per gate (the #1 measurement reinvention)
judgea campaign JudgeConfig, or ensembleJudge for a multi-model panela hand-built judge prompt loop, or the @deprecatedJudgeFn factories
authenticityscoreAuthenticitygateRealness over an AuthenticitySignals (agent-eval/authenticity subpath)a regex realness scorer, or trusting a buildability score that rewards a polished fake
verificationnew MultiLayerVerifier(layers).run(...) + gradeSemanticStatus(...)a hand-chained compile→lint→test→semantic pipeline
reference-replayrunReferenceReplay/scoreReferenceReplay + decideReferenceReplayPromotion over jsonlReferenceReplayStorea hand-rolled "does the candidate reproduce the reference" matcher
token/usage seamextractLlmCallEvent / reportLoopUsage (UsageSink) — /loopsre-walking sandbox events to tally tokens yourself

The judge row also kills a recurring confusion: there is no llmJudge export in this version — that name appears only in a campaign-types docstring; the live judge surface is JudgeConfig + ensembleJudge. A future agent reading this doc can no longer conclude "no judge primitive exists."

The freshness gate (scripts/check-docs-freshness.mjs) now loads the agent-eval/authenticity subpath barrel into the §2 export universe, so a row that recommends importing from that subpath resolves — exactly parallel to how contract/campaign/index are already loaded. This strengthens coverage (a real published export surface the gate previously couldn't see), it does not weaken the gate.

2. Example de-reinvented

examples/self-improving-loop was the only example (of 21 audited) that hand-rolled substrate primitives:

  • analyst — local interface AnalystFinding { rootCause; proposedMutation } + runAnalyst() → now the canonical AnalystFinding stamped by makeFinding (the exact shape improve(profile, findings, opts) reflects on); mutation rides recommended_action.
  • gate — bare v1Mean - v0Mean >= 0.5 point comparison → now pairedBootstrap(v0Scores, v1Scores, { seed }), shipping only when the paired-bootstrap CI lower bound clears 0 — the statistical core the production held-out gate (HeldOutGate / improve() over selfImprove) is built on.

Stays offline + deterministic; the demo still ships v1, now on a +5.00 paired median, 95% CI [5.00, 6.00], n=3. README + comments mirror the already-clean companion examples (improve/, intelligence-recommend/). Only the analyst body, proposer, and LLM remain scripted (for reproducibility) — the finding type and gate statistic are now the real substrate.

Verify (all green)

  • pnpm docs:checkdocs freshness: OK (§2 table symbols against 4134 public exports; prose symbols against 5795 resolvable; docs/api regen diff clean)
  • pnpm run build — clean
  • pnpm run typecheck + typecheck:examples — clean (the fixed example compiles)
  • pnpm run lint — Checked 319 files, no fixes applied
  • Example runs offline e2e: ships v1 on the paired CI [5.00, 6.00]

DO NOT MERGE — operator review.

…e examples
The anti-reinvention decision table (§2) had zero rows for whole families of
live exports, so agents kept hand-rolling them. Add grouped rows for:
- statistics — every test on the agent-eval main barrel (pairedBootstrap,
wilson, pairedTTest, wilcoxonSignedRank, mcnemar/mcnemarPower, mannWhitneyU,
passAtK, confidenceInterval, benjaminiHochberg/bonferroni, cohensD/cliffsDelta,
requiredSampleSize/pairedMde, eProcess, weightedComposite,
corpusInterRaterAgreement, pairedRiskDifference, pearsonR/spearmanR,
paretoFrontier/dominates)
- judge — JudgeConfig + ensembleJudge, with a warn-off for the @deprecated
JudgeFn factories and an explicit note that there is no llmJudge export
- authenticity — scoreAuthenticity / gateRealness / AuthenticitySignals
(agent-eval/authenticity subpath)
- verification — MultiLayerVerifier + gradeSemanticStatus
- reference-replay — runReferenceReplay / scoreReferenceReplay /
decideReferenceReplayPromotion over jsonlReferenceReplayStore
- token/usage seam — extractLlmCallEvent + reportLoopUsage / UsageSink (/loops)
Teach the freshness gate the agent-eval/authenticity subpath barrel so a §2 row
that recommends importing from it resolves (parallel to contract/campaign/index).
De-reinvent the one example that hand-rolled primitives:
examples/self-improving-loop replaces its local AnalystFinding interface +
runAnalyst with the canonical AnalystFinding stamped by makeFinding, and its
bare v1Mean-v0Mean>=0.5 point-comparison gate with pairedBootstrap (the
statistical core the production held-out gate is built on). Stays offline and
deterministic; ships v1 on a +5.00 [5.00, 6.00] paired CI. README + comments
mirror the clean companion examples (improve/, intelligence-recommend/).

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved PR — 70e098d8

Blanket team auto-approval is enabled for this reviewer service.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: blanket_auto_approve · 2026-06-24T08:54:31Z

drewstone added a commit that referenced this pull request Jun 24, 2026
…fImprove not a bare mean gate
Folds the one real example fix from #370 (otherwise superseded by the generated
catalog) into this PR: self-improving-loop hand-rolled the ship gate as a bare
mean comparison instead of the real HeldOutGate/selfImprove primitives.
@drewstone

Copy link
Copy Markdown
ContributorAuthor

Superseded by #371. The primitive inventory is now GENERATED (docs/api/primitive-catalog.md, freshness-gate-enforced) so it can't go stale by hand — the root cause this PR hand-patched. The ~30 primitives catalogued here are now auto-generated; the one real example fix (self-improving-loop → HeldOutGate/selfImprove) was folded into #371.

drewstone added a commit that referenced this pull request Jun 24, 2026
…erence cannot go stale (#371)
* docs(api): generate the primitive catalog so the anti-reinvention reference cannot go stale
The hand-listed primitive inventory in docs/canonical-api.md drifted from source:
it had zero mentions of live exports (scoreAuthenticity, gateRealness,
MultiLayerVerifier, wilson, pairedTTest, runProfileMatrix, extractUsage, …). Anything
derivable from source must be generated, not hand-written — only judgment stays curated.
- scripts/gen-primitive-catalog.mjs reads the LIVE exports of (a) this package's own
public subpaths (from package.json `exports`) and (b) a curated category->subpath map
of the @tangle-network/agent-eval substrate surfaces agents should reuse (judge,
authenticity, verification, statistics, campaign, token/usage). Extraction is via the
TypeScript compiler API over a virtual re-export entry, so it follows aliased
re-exports and content-hashed bundle files — the exact things that rot a hand list.
Emits docs/api/primitive-catalog.md with a GENERATED header (name, import path,
one-line summary per export, grouped by surface).
- Wired into `docs:api` (runs after TypeDoc). The freshness gate gains a seventh class
(CATALOG): it regenerates the catalog to a temp file and byte-compares to the committed
copy, so a new/removed/renamed live export absent from the catalog is a RED BUILD.
- Shrank canonical-api.md: removed the export-inventory enumeration from the banner and
the §2 preamble, replaced with pointers to docs/api/primitive-catalog.md. Kept all the
judgment — the decision gate, §1.5 AgentProfile law, the §2 "I want to -> use -> NOT"
table and every "Do NOT". The version + substrate-peer pins stay (gate-enforced).
- MAINTAINING.md documents the generated-inventory layer, CLASS 7, and its fix path.
* chore(deps): bump agent-eval to 0.99.0 + regenerate primitive catalog
agent-eval 0.99.0 adds llmJudge (+ the full current judge/auth/verify/stats
surface); regenerating the generated catalog picks it up with zero hand-work,
which is the point of the generator. Lockfile was pinned at 0.97.0 (pre-llmJudge)
despite agent-eval already being in minimumReleaseAgeExclude.
* fix(deps): keep agent-eval peer floor at >=0.97.0
Only the examples (devDependency) need 0.99.0 for llmJudge; agent-runtime's src
does not, so the peer floor must not force consumers onto 0.99.0. Catalog +
lockfile stay on the resolved 0.99.0 so the examples get llmJudge.
* docs(examples): de-reinvent self-improving-loop — use HeldOutGate/selfImprove not a bare mean gate
Folds the one real example fix from #370 (otherwise superseded by the generated
catalog) into this PR: self-improving-loop hand-rolled the ship gate as a bare
mean comparison instead of the real HeldOutGate/selfImprove primitives.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

docs(canonical-api): close the anti-reinvention gaps + de-reinvent the examples - #370

Closed
drewstone wants to merge 1 commit into
mainfrom
docs/kill-example-reinvention
Closed

docs(canonical-api): close the anti-reinvention gaps + de-reinvent the examples#370
drewstone wants to merge 1 commit into
mainfrom
docs/kill-example-reinvention

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

What

The anti-reinvention decision table in docs/canonical-api.md (§2 — the doc agents are told to read before writing orchestration/measurement code) had zero rows for whole families of live exports, so agents kept hand-rolling them. This adds them, and de-reinvents the one example that hand-rolled primitives.

1. Reference gaps closed (docs/canonical-api.md §2)

Six grouped rows added, in the existing "I want to ___ → use ___ → NOT ___" voice:

FamilyUse (the real primitive)Was being hand-rolled as
statisticsevery test on the agent-eval main barrelpairedBootstrap, wilson, pairedTTest, wilcoxonSignedRank, mcnemar/mcnemarPower/mcnemarRequiredN, mannWhitneyU, passAtK, confidenceInterval, benjaminiHochberg/bonferroni, cohensD/cliffsDelta, requiredSampleSize/pairedMde, eProcess, weightedComposite, corpusInterRaterAgreement, pairedRiskDifference, pearsonR/spearmanR, paretoFrontier/dominatesa bootstrap-with-PRNG, a Wilson/McNemar/power calc, a Pareto sweep, a Cohen's-d per gate (the #1 measurement reinvention)
judgea campaign JudgeConfig, or ensembleJudge for a multi-model panela hand-built judge prompt loop, or the @deprecatedJudgeFn factories
authenticityscoreAuthenticitygateRealness over an AuthenticitySignals (agent-eval/authenticity subpath)a regex realness scorer, or trusting a buildability score that rewards a polished fake
verificationnew MultiLayerVerifier(layers).run(...) + gradeSemanticStatus(...)a hand-chained compile→lint→test→semantic pipeline
reference-replayrunReferenceReplay/scoreReferenceReplay + decideReferenceReplayPromotion over jsonlReferenceReplayStorea hand-rolled "does the candidate reproduce the reference" matcher
token/usage seamextractLlmCallEvent / reportLoopUsage (UsageSink) — /loopsre-walking sandbox events to tally tokens yourself

The judge row also kills a recurring confusion: there is no llmJudge export in this version — that name appears only in a campaign-types docstring; the live judge surface is JudgeConfig + ensembleJudge. A future agent reading this doc can no longer conclude "no judge primitive exists."

The freshness gate (scripts/check-docs-freshness.mjs) now loads the agent-eval/authenticity subpath barrel into the §2 export universe, so a row that recommends importing from that subpath resolves — exactly parallel to how contract/campaign/index are already loaded. This strengthens coverage (a real published export surface the gate previously couldn't see), it does not weaken the gate.

2. Example de-reinvented

examples/self-improving-loop was the only example (of 21 audited) that hand-rolled substrate primitives:

  • analyst — local interface AnalystFinding { rootCause; proposedMutation } + runAnalyst() → now the canonical AnalystFinding stamped by makeFinding (the exact shape improve(profile, findings, opts) reflects on); mutation rides recommended_action.
  • gate — bare v1Mean - v0Mean >= 0.5 point comparison → now pairedBootstrap(v0Scores, v1Scores, { seed }), shipping only when the paired-bootstrap CI lower bound clears 0 — the statistical core the production held-out gate (HeldOutGate / improve() over selfImprove) is built on.

Stays offline + deterministic; the demo still ships v1, now on a +5.00 paired median, 95% CI [5.00, 6.00], n=3. README + comments mirror the already-clean companion examples (improve/, intelligence-recommend/). Only the analyst body, proposer, and LLM remain scripted (for reproducibility) — the finding type and gate statistic are now the real substrate.

Verify (all green)

  • pnpm docs:checkdocs freshness: OK (§2 table symbols against 4134 public exports; prose symbols against 5795 resolvable; docs/api regen diff clean)
  • pnpm run build — clean
  • pnpm run typecheck + typecheck:examples — clean (the fixed example compiles)
  • pnpm run lint — Checked 319 files, no fixes applied
  • Example runs offline e2e: ships v1 on the paired CI [5.00, 6.00]

DO NOT MERGE — operator review.

…e examples
The anti-reinvention decision table (§2) had zero rows for whole families of
live exports, so agents kept hand-rolling them. Add grouped rows for:
- statistics — every test on the agent-eval main barrel (pairedBootstrap,
wilson, pairedTTest, wilcoxonSignedRank, mcnemar/mcnemarPower, mannWhitneyU,
passAtK, confidenceInterval, benjaminiHochberg/bonferroni, cohensD/cliffsDelta,
requiredSampleSize/pairedMde, eProcess, weightedComposite,
corpusInterRaterAgreement, pairedRiskDifference, pearsonR/spearmanR,
paretoFrontier/dominates)
- judge — JudgeConfig + ensembleJudge, with a warn-off for the @deprecated
JudgeFn factories and an explicit note that there is no llmJudge export
- authenticity — scoreAuthenticity / gateRealness / AuthenticitySignals
(agent-eval/authenticity subpath)
- verification — MultiLayerVerifier + gradeSemanticStatus
- reference-replay — runReferenceReplay / scoreReferenceReplay /
decideReferenceReplayPromotion over jsonlReferenceReplayStore
- token/usage seam — extractLlmCallEvent + reportLoopUsage / UsageSink (/loops)
Teach the freshness gate the agent-eval/authenticity subpath barrel so a §2 row
that recommends importing from it resolves (parallel to contract/campaign/index).
De-reinvent the one example that hand-rolled primitives:
examples/self-improving-loop replaces its local AnalystFinding interface +
runAnalyst with the canonical AnalystFinding stamped by makeFinding, and its
bare v1Mean-v0Mean>=0.5 point-comparison gate with pairedBootstrap (the
statistical core the production held-out gate is built on). Stays offline and
deterministic; ships v1 on a +5.00 [5.00, 6.00] paired CI. README + comments
mirror the clean companion examples (improve/, intelligence-recommend/).

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved PR — 70e098d8

Blanket team auto-approval is enabled for this reviewer service.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: blanket_auto_approve · 2026-06-24T08:54:31Z

drewstone added a commit that referenced this pull request Jun 24, 2026
…fImprove not a bare mean gate
Folds the one real example fix from #370 (otherwise superseded by the generated
catalog) into this PR: self-improving-loop hand-rolled the ship gate as a bare
mean comparison instead of the real HeldOutGate/selfImprove primitives.
@drewstone

Copy link
Copy Markdown
ContributorAuthor

Superseded by #371. The primitive inventory is now GENERATED (docs/api/primitive-catalog.md, freshness-gate-enforced) so it can't go stale by hand — the root cause this PR hand-patched. The ~30 primitives catalogued here are now auto-generated; the one real example fix (self-improving-loop → HeldOutGate/selfImprove) was folded into #371.

drewstone added a commit that referenced this pull request Jun 24, 2026
…erence cannot go stale (#371)
* docs(api): generate the primitive catalog so the anti-reinvention reference cannot go stale
The hand-listed primitive inventory in docs/canonical-api.md drifted from source:
it had zero mentions of live exports (scoreAuthenticity, gateRealness,
MultiLayerVerifier, wilson, pairedTTest, runProfileMatrix, extractUsage, …). Anything
derivable from source must be generated, not hand-written — only judgment stays curated.
- scripts/gen-primitive-catalog.mjs reads the LIVE exports of (a) this package's own
public subpaths (from package.json `exports`) and (b) a curated category->subpath map
of the @tangle-network/agent-eval substrate surfaces agents should reuse (judge,
authenticity, verification, statistics, campaign, token/usage). Extraction is via the
TypeScript compiler API over a virtual re-export entry, so it follows aliased
re-exports and content-hashed bundle files — the exact things that rot a hand list.
Emits docs/api/primitive-catalog.md with a GENERATED header (name, import path,
one-line summary per export, grouped by surface).
- Wired into `docs:api` (runs after TypeDoc). The freshness gate gains a seventh class
(CATALOG): it regenerates the catalog to a temp file and byte-compares to the committed
copy, so a new/removed/renamed live export absent from the catalog is a RED BUILD.
- Shrank canonical-api.md: removed the export-inventory enumeration from the banner and
the §2 preamble, replaced with pointers to docs/api/primitive-catalog.md. Kept all the
judgment — the decision gate, §1.5 AgentProfile law, the §2 "I want to -> use -> NOT"
table and every "Do NOT". The version + substrate-peer pins stay (gate-enforced).
- MAINTAINING.md documents the generated-inventory layer, CLASS 7, and its fix path.
* chore(deps): bump agent-eval to 0.99.0 + regenerate primitive catalog
agent-eval 0.99.0 adds llmJudge (+ the full current judge/auth/verify/stats
surface); regenerating the generated catalog picks it up with zero hand-work,
which is the point of the generator. Lockfile was pinned at 0.97.0 (pre-llmJudge)
despite agent-eval already being in minimumReleaseAgeExclude.
* fix(deps): keep agent-eval peer floor at >=0.97.0
Only the examples (devDependency) need 0.99.0 for llmJudge; agent-runtime's src
does not, so the peer floor must not force consumers onto 0.99.0. Catalog +
lockfile stay on the resolved 0.99.0 so the examples get llmJudge.
* docs(examples): de-reinvent self-improving-loop — use HeldOutGate/selfImprove not a bare mean gate
Folds the one real example fix from #370 (otherwise superseded by the generated
catalog) into this PR: self-improving-loop hand-rolled the ship gate as a bare
mean comparison instead of the real HeldOutGate/selfImprove primitives.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

docs(canonical-api): close the anti-reinvention gaps + de-reinvent the examples - #370

Closed
drewstone wants to merge 1 commit into
mainfrom
docs/kill-example-reinvention
Closed

docs(canonical-api): close the anti-reinvention gaps + de-reinvent the examples#370
drewstone wants to merge 1 commit into
mainfrom
docs/kill-example-reinvention

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

What

The anti-reinvention decision table in docs/canonical-api.md (§2 — the doc agents are told to read before writing orchestration/measurement code) had zero rows for whole families of live exports, so agents kept hand-rolling them. This adds them, and de-reinvents the one example that hand-rolled primitives.

1. Reference gaps closed (docs/canonical-api.md §2)

Six grouped rows added, in the existing "I want to ___ → use ___ → NOT ___" voice:

FamilyUse (the real primitive)Was being hand-rolled as
statisticsevery test on the agent-eval main barrelpairedBootstrap, wilson, pairedTTest, wilcoxonSignedRank, mcnemar/mcnemarPower/mcnemarRequiredN, mannWhitneyU, passAtK, confidenceInterval, benjaminiHochberg/bonferroni, cohensD/cliffsDelta, requiredSampleSize/pairedMde, eProcess, weightedComposite, corpusInterRaterAgreement, pairedRiskDifference, pearsonR/spearmanR, paretoFrontier/dominatesa bootstrap-with-PRNG, a Wilson/McNemar/power calc, a Pareto sweep, a Cohen's-d per gate (the #1 measurement reinvention)
judgea campaign JudgeConfig, or ensembleJudge for a multi-model panela hand-built judge prompt loop, or the @deprecatedJudgeFn factories
authenticityscoreAuthenticitygateRealness over an AuthenticitySignals (agent-eval/authenticity subpath)a regex realness scorer, or trusting a buildability score that rewards a polished fake
verificationnew MultiLayerVerifier(layers).run(...) + gradeSemanticStatus(...)a hand-chained compile→lint→test→semantic pipeline
reference-replayrunReferenceReplay/scoreReferenceReplay + decideReferenceReplayPromotion over jsonlReferenceReplayStorea hand-rolled "does the candidate reproduce the reference" matcher
token/usage seamextractLlmCallEvent / reportLoopUsage (UsageSink) — /loopsre-walking sandbox events to tally tokens yourself

The judge row also kills a recurring confusion: there is no llmJudge export in this version — that name appears only in a campaign-types docstring; the live judge surface is JudgeConfig + ensembleJudge. A future agent reading this doc can no longer conclude "no judge primitive exists."

The freshness gate (scripts/check-docs-freshness.mjs) now loads the agent-eval/authenticity subpath barrel into the §2 export universe, so a row that recommends importing from that subpath resolves — exactly parallel to how contract/campaign/index are already loaded. This strengthens coverage (a real published export surface the gate previously couldn't see), it does not weaken the gate.

2. Example de-reinvented

examples/self-improving-loop was the only example (of 21 audited) that hand-rolled substrate primitives:

  • analyst — local interface AnalystFinding { rootCause; proposedMutation } + runAnalyst() → now the canonical AnalystFinding stamped by makeFinding (the exact shape improve(profile, findings, opts) reflects on); mutation rides recommended_action.
  • gate — bare v1Mean - v0Mean >= 0.5 point comparison → now pairedBootstrap(v0Scores, v1Scores, { seed }), shipping only when the paired-bootstrap CI lower bound clears 0 — the statistical core the production held-out gate (HeldOutGate / improve() over selfImprove) is built on.

Stays offline + deterministic; the demo still ships v1, now on a +5.00 paired median, 95% CI [5.00, 6.00], n=3. README + comments mirror the already-clean companion examples (improve/, intelligence-recommend/). Only the analyst body, proposer, and LLM remain scripted (for reproducibility) — the finding type and gate statistic are now the real substrate.

Verify (all green)

  • pnpm docs:checkdocs freshness: OK (§2 table symbols against 4134 public exports; prose symbols against 5795 resolvable; docs/api regen diff clean)
  • pnpm run build — clean
  • pnpm run typecheck + typecheck:examples — clean (the fixed example compiles)
  • pnpm run lint — Checked 319 files, no fixes applied
  • Example runs offline e2e: ships v1 on the paired CI [5.00, 6.00]

DO NOT MERGE — operator review.

…e examples
The anti-reinvention decision table (§2) had zero rows for whole families of
live exports, so agents kept hand-rolling them. Add grouped rows for:
- statistics — every test on the agent-eval main barrel (pairedBootstrap,
wilson, pairedTTest, wilcoxonSignedRank, mcnemar/mcnemarPower, mannWhitneyU,
passAtK, confidenceInterval, benjaminiHochberg/bonferroni, cohensD/cliffsDelta,
requiredSampleSize/pairedMde, eProcess, weightedComposite,
corpusInterRaterAgreement, pairedRiskDifference, pearsonR/spearmanR,
paretoFrontier/dominates)
- judge — JudgeConfig + ensembleJudge, with a warn-off for the @deprecated
JudgeFn factories and an explicit note that there is no llmJudge export
- authenticity — scoreAuthenticity / gateRealness / AuthenticitySignals
(agent-eval/authenticity subpath)
- verification — MultiLayerVerifier + gradeSemanticStatus
- reference-replay — runReferenceReplay / scoreReferenceReplay /
decideReferenceReplayPromotion over jsonlReferenceReplayStore
- token/usage seam — extractLlmCallEvent + reportLoopUsage / UsageSink (/loops)
Teach the freshness gate the agent-eval/authenticity subpath barrel so a §2 row
that recommends importing from it resolves (parallel to contract/campaign/index).
De-reinvent the one example that hand-rolled primitives:
examples/self-improving-loop replaces its local AnalystFinding interface +
runAnalyst with the canonical AnalystFinding stamped by makeFinding, and its
bare v1Mean-v0Mean>=0.5 point-comparison gate with pairedBootstrap (the
statistical core the production held-out gate is built on). Stays offline and
deterministic; ships v1 on a +5.00 [5.00, 6.00] paired CI. README + comments
mirror the clean companion examples (improve/, intelligence-recommend/).

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved PR — 70e098d8

Blanket team auto-approval is enabled for this reviewer service.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: blanket_auto_approve · 2026-06-24T08:54:31Z

drewstone added a commit that referenced this pull request Jun 24, 2026
…fImprove not a bare mean gate
Folds the one real example fix from #370 (otherwise superseded by the generated
catalog) into this PR: self-improving-loop hand-rolled the ship gate as a bare
mean comparison instead of the real HeldOutGate/selfImprove primitives.
@drewstone

Copy link
Copy Markdown
ContributorAuthor

Superseded by #371. The primitive inventory is now GENERATED (docs/api/primitive-catalog.md, freshness-gate-enforced) so it can't go stale by hand — the root cause this PR hand-patched. The ~30 primitives catalogued here are now auto-generated; the one real example fix (self-improving-loop → HeldOutGate/selfImprove) was folded into #371.

drewstone added a commit that referenced this pull request Jun 24, 2026
…erence cannot go stale (#371)
* docs(api): generate the primitive catalog so the anti-reinvention reference cannot go stale
The hand-listed primitive inventory in docs/canonical-api.md drifted from source:
it had zero mentions of live exports (scoreAuthenticity, gateRealness,
MultiLayerVerifier, wilson, pairedTTest, runProfileMatrix, extractUsage, …). Anything
derivable from source must be generated, not hand-written — only judgment stays curated.
- scripts/gen-primitive-catalog.mjs reads the LIVE exports of (a) this package's own
public subpaths (from package.json `exports`) and (b) a curated category->subpath map
of the @tangle-network/agent-eval substrate surfaces agents should reuse (judge,
authenticity, verification, statistics, campaign, token/usage). Extraction is via the
TypeScript compiler API over a virtual re-export entry, so it follows aliased
re-exports and content-hashed bundle files — the exact things that rot a hand list.
Emits docs/api/primitive-catalog.md with a GENERATED header (name, import path,
one-line summary per export, grouped by surface).
- Wired into `docs:api` (runs after TypeDoc). The freshness gate gains a seventh class
(CATALOG): it regenerates the catalog to a temp file and byte-compares to the committed
copy, so a new/removed/renamed live export absent from the catalog is a RED BUILD.
- Shrank canonical-api.md: removed the export-inventory enumeration from the banner and
the §2 preamble, replaced with pointers to docs/api/primitive-catalog.md. Kept all the
judgment — the decision gate, §1.5 AgentProfile law, the §2 "I want to -> use -> NOT"
table and every "Do NOT". The version + substrate-peer pins stay (gate-enforced).
- MAINTAINING.md documents the generated-inventory layer, CLASS 7, and its fix path.
* chore(deps): bump agent-eval to 0.99.0 + regenerate primitive catalog
agent-eval 0.99.0 adds llmJudge (+ the full current judge/auth/verify/stats
surface); regenerating the generated catalog picks it up with zero hand-work,
which is the point of the generator. Lockfile was pinned at 0.97.0 (pre-llmJudge)
despite agent-eval already being in minimumReleaseAgeExclude.
* fix(deps): keep agent-eval peer floor at >=0.97.0
Only the examples (devDependency) need 0.99.0 for llmJudge; agent-runtime's src
does not, so the peer floor must not force consumers onto 0.99.0. Catalog +
lockfile stay on the resolved 0.99.0 so the examples get llmJudge.
* docs(examples): de-reinvent self-improving-loop — use HeldOutGate/selfImprove not a bare mean gate
Folds the one real example fix from #370 (otherwise superseded by the generated
catalog) into this PR: self-improving-loop hand-rolled the ship gate as a bare
mean comparison instead of the real HeldOutGate/selfImprove primitives.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

docs(canonical-api): close the anti-reinvention gaps + de-reinvent the examples - #370

Closed
drewstone wants to merge 1 commit into
mainfrom
docs/kill-example-reinvention
Closed

docs(canonical-api): close the anti-reinvention gaps + de-reinvent the examples#370
drewstone wants to merge 1 commit into
mainfrom
docs/kill-example-reinvention

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

What

The anti-reinvention decision table in docs/canonical-api.md (§2 — the doc agents are told to read before writing orchestration/measurement code) had zero rows for whole families of live exports, so agents kept hand-rolling them. This adds them, and de-reinvents the one example that hand-rolled primitives.

1. Reference gaps closed (docs/canonical-api.md §2)

Six grouped rows added, in the existing "I want to ___ → use ___ → NOT ___" voice:

FamilyUse (the real primitive)Was being hand-rolled as
statisticsevery test on the agent-eval main barrelpairedBootstrap, wilson, pairedTTest, wilcoxonSignedRank, mcnemar/mcnemarPower/mcnemarRequiredN, mannWhitneyU, passAtK, confidenceInterval, benjaminiHochberg/bonferroni, cohensD/cliffsDelta, requiredSampleSize/pairedMde, eProcess, weightedComposite, corpusInterRaterAgreement, pairedRiskDifference, pearsonR/spearmanR, paretoFrontier/dominatesa bootstrap-with-PRNG, a Wilson/McNemar/power calc, a Pareto sweep, a Cohen's-d per gate (the #1 measurement reinvention)
judgea campaign JudgeConfig, or ensembleJudge for a multi-model panela hand-built judge prompt loop, or the @deprecatedJudgeFn factories
authenticityscoreAuthenticitygateRealness over an AuthenticitySignals (agent-eval/authenticity subpath)a regex realness scorer, or trusting a buildability score that rewards a polished fake
verificationnew MultiLayerVerifier(layers).run(...) + gradeSemanticStatus(...)a hand-chained compile→lint→test→semantic pipeline
reference-replayrunReferenceReplay/scoreReferenceReplay + decideReferenceReplayPromotion over jsonlReferenceReplayStorea hand-rolled "does the candidate reproduce the reference" matcher
token/usage seamextractLlmCallEvent / reportLoopUsage (UsageSink) — /loopsre-walking sandbox events to tally tokens yourself

The judge row also kills a recurring confusion: there is no llmJudge export in this version — that name appears only in a campaign-types docstring; the live judge surface is JudgeConfig + ensembleJudge. A future agent reading this doc can no longer conclude "no judge primitive exists."

The freshness gate (scripts/check-docs-freshness.mjs) now loads the agent-eval/authenticity subpath barrel into the §2 export universe, so a row that recommends importing from that subpath resolves — exactly parallel to how contract/campaign/index are already loaded. This strengthens coverage (a real published export surface the gate previously couldn't see), it does not weaken the gate.

2. Example de-reinvented

examples/self-improving-loop was the only example (of 21 audited) that hand-rolled substrate primitives:

  • analyst — local interface AnalystFinding { rootCause; proposedMutation } + runAnalyst() → now the canonical AnalystFinding stamped by makeFinding (the exact shape improve(profile, findings, opts) reflects on); mutation rides recommended_action.
  • gate — bare v1Mean - v0Mean >= 0.5 point comparison → now pairedBootstrap(v0Scores, v1Scores, { seed }), shipping only when the paired-bootstrap CI lower bound clears 0 — the statistical core the production held-out gate (HeldOutGate / improve() over selfImprove) is built on.

Stays offline + deterministic; the demo still ships v1, now on a +5.00 paired median, 95% CI [5.00, 6.00], n=3. README + comments mirror the already-clean companion examples (improve/, intelligence-recommend/). Only the analyst body, proposer, and LLM remain scripted (for reproducibility) — the finding type and gate statistic are now the real substrate.

Verify (all green)

  • pnpm docs:checkdocs freshness: OK (§2 table symbols against 4134 public exports; prose symbols against 5795 resolvable; docs/api regen diff clean)
  • pnpm run build — clean
  • pnpm run typecheck + typecheck:examples — clean (the fixed example compiles)
  • pnpm run lint — Checked 319 files, no fixes applied
  • Example runs offline e2e: ships v1 on the paired CI [5.00, 6.00]

DO NOT MERGE — operator review.

…e examples
The anti-reinvention decision table (§2) had zero rows for whole families of
live exports, so agents kept hand-rolling them. Add grouped rows for:
- statistics — every test on the agent-eval main barrel (pairedBootstrap,
wilson, pairedTTest, wilcoxonSignedRank, mcnemar/mcnemarPower, mannWhitneyU,
passAtK, confidenceInterval, benjaminiHochberg/bonferroni, cohensD/cliffsDelta,
requiredSampleSize/pairedMde, eProcess, weightedComposite,
corpusInterRaterAgreement, pairedRiskDifference, pearsonR/spearmanR,
paretoFrontier/dominates)
- judge — JudgeConfig + ensembleJudge, with a warn-off for the @deprecated
JudgeFn factories and an explicit note that there is no llmJudge export
- authenticity — scoreAuthenticity / gateRealness / AuthenticitySignals
(agent-eval/authenticity subpath)
- verification — MultiLayerVerifier + gradeSemanticStatus
- reference-replay — runReferenceReplay / scoreReferenceReplay /
decideReferenceReplayPromotion over jsonlReferenceReplayStore
- token/usage seam — extractLlmCallEvent + reportLoopUsage / UsageSink (/loops)
Teach the freshness gate the agent-eval/authenticity subpath barrel so a §2 row
that recommends importing from it resolves (parallel to contract/campaign/index).
De-reinvent the one example that hand-rolled primitives:
examples/self-improving-loop replaces its local AnalystFinding interface +
runAnalyst with the canonical AnalystFinding stamped by makeFinding, and its
bare v1Mean-v0Mean>=0.5 point-comparison gate with pairedBootstrap (the
statistical core the production held-out gate is built on). Stays offline and
deterministic; ships v1 on a +5.00 [5.00, 6.00] paired CI. README + comments
mirror the clean companion examples (improve/, intelligence-recommend/).

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved PR — 70e098d8

Blanket team auto-approval is enabled for this reviewer service.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: blanket_auto_approve · 2026-06-24T08:54:31Z

drewstone added a commit that referenced this pull request Jun 24, 2026
…fImprove not a bare mean gate
Folds the one real example fix from #370 (otherwise superseded by the generated
catalog) into this PR: self-improving-loop hand-rolled the ship gate as a bare
mean comparison instead of the real HeldOutGate/selfImprove primitives.
@drewstone

Copy link
Copy Markdown
ContributorAuthor

Superseded by #371. The primitive inventory is now GENERATED (docs/api/primitive-catalog.md, freshness-gate-enforced) so it can't go stale by hand — the root cause this PR hand-patched. The ~30 primitives catalogued here are now auto-generated; the one real example fix (self-improving-loop → HeldOutGate/selfImprove) was folded into #371.

drewstone added a commit that referenced this pull request Jun 24, 2026
…erence cannot go stale (#371)
* docs(api): generate the primitive catalog so the anti-reinvention reference cannot go stale
The hand-listed primitive inventory in docs/canonical-api.md drifted from source:
it had zero mentions of live exports (scoreAuthenticity, gateRealness,
MultiLayerVerifier, wilson, pairedTTest, runProfileMatrix, extractUsage, …). Anything
derivable from source must be generated, not hand-written — only judgment stays curated.
- scripts/gen-primitive-catalog.mjs reads the LIVE exports of (a) this package's own
public subpaths (from package.json `exports`) and (b) a curated category->subpath map
of the @tangle-network/agent-eval substrate surfaces agents should reuse (judge,
authenticity, verification, statistics, campaign, token/usage). Extraction is via the
TypeScript compiler API over a virtual re-export entry, so it follows aliased
re-exports and content-hashed bundle files — the exact things that rot a hand list.
Emits docs/api/primitive-catalog.md with a GENERATED header (name, import path,
one-line summary per export, grouped by surface).
- Wired into `docs:api` (runs after TypeDoc). The freshness gate gains a seventh class
(CATALOG): it regenerates the catalog to a temp file and byte-compares to the committed
copy, so a new/removed/renamed live export absent from the catalog is a RED BUILD.
- Shrank canonical-api.md: removed the export-inventory enumeration from the banner and
the §2 preamble, replaced with pointers to docs/api/primitive-catalog.md. Kept all the
judgment — the decision gate, §1.5 AgentProfile law, the §2 "I want to -> use -> NOT"
table and every "Do NOT". The version + substrate-peer pins stay (gate-enforced).
- MAINTAINING.md documents the generated-inventory layer, CLASS 7, and its fix path.
* chore(deps): bump agent-eval to 0.99.0 + regenerate primitive catalog
agent-eval 0.99.0 adds llmJudge (+ the full current judge/auth/verify/stats
surface); regenerating the generated catalog picks it up with zero hand-work,
which is the point of the generator. Lockfile was pinned at 0.97.0 (pre-llmJudge)
despite agent-eval already being in minimumReleaseAgeExclude.
* fix(deps): keep agent-eval peer floor at >=0.97.0
Only the examples (devDependency) need 0.99.0 for llmJudge; agent-runtime's src
does not, so the peer floor must not force consumers onto 0.99.0. Catalog +
lockfile stay on the resolved 0.99.0 so the examples get llmJudge.
* docs(examples): de-reinvent self-improving-loop — use HeldOutGate/selfImprove not a bare mean gate
Folds the one real example fix from #370 (otherwise superseded by the generated
catalog) into this PR: self-improving-loop hand-rolled the ship gate as a bare
mean comparison instead of the real HeldOutGate/selfImprove primitives.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

docs(canonical-api): close the anti-reinvention gaps + de-reinvent the examples - #370

Closed
drewstone wants to merge 1 commit into
mainfrom
docs/kill-example-reinvention
Closed

docs(canonical-api): close the anti-reinvention gaps + de-reinvent the examples#370
drewstone wants to merge 1 commit into
mainfrom
docs/kill-example-reinvention

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

What

The anti-reinvention decision table in docs/canonical-api.md (§2 — the doc agents are told to read before writing orchestration/measurement code) had zero rows for whole families of live exports, so agents kept hand-rolling them. This adds them, and de-reinvents the one example that hand-rolled primitives.

1. Reference gaps closed (docs/canonical-api.md §2)

Six grouped rows added, in the existing "I want to ___ → use ___ → NOT ___" voice:

FamilyUse (the real primitive)Was being hand-rolled as
statisticsevery test on the agent-eval main barrelpairedBootstrap, wilson, pairedTTest, wilcoxonSignedRank, mcnemar/mcnemarPower/mcnemarRequiredN, mannWhitneyU, passAtK, confidenceInterval, benjaminiHochberg/bonferroni, cohensD/cliffsDelta, requiredSampleSize/pairedMde, eProcess, weightedComposite, corpusInterRaterAgreement, pairedRiskDifference, pearsonR/spearmanR, paretoFrontier/dominatesa bootstrap-with-PRNG, a Wilson/McNemar/power calc, a Pareto sweep, a Cohen's-d per gate (the #1 measurement reinvention)
judgea campaign JudgeConfig, or ensembleJudge for a multi-model panela hand-built judge prompt loop, or the @deprecatedJudgeFn factories
authenticityscoreAuthenticitygateRealness over an AuthenticitySignals (agent-eval/authenticity subpath)a regex realness scorer, or trusting a buildability score that rewards a polished fake
verificationnew MultiLayerVerifier(layers).run(...) + gradeSemanticStatus(...)a hand-chained compile→lint→test→semantic pipeline
reference-replayrunReferenceReplay/scoreReferenceReplay + decideReferenceReplayPromotion over jsonlReferenceReplayStorea hand-rolled "does the candidate reproduce the reference" matcher
token/usage seamextractLlmCallEvent / reportLoopUsage (UsageSink) — /loopsre-walking sandbox events to tally tokens yourself

The judge row also kills a recurring confusion: there is no llmJudge export in this version — that name appears only in a campaign-types docstring; the live judge surface is JudgeConfig + ensembleJudge. A future agent reading this doc can no longer conclude "no judge primitive exists."

The freshness gate (scripts/check-docs-freshness.mjs) now loads the agent-eval/authenticity subpath barrel into the §2 export universe, so a row that recommends importing from that subpath resolves — exactly parallel to how contract/campaign/index are already loaded. This strengthens coverage (a real published export surface the gate previously couldn't see), it does not weaken the gate.

2. Example de-reinvented

examples/self-improving-loop was the only example (of 21 audited) that hand-rolled substrate primitives:

  • analyst — local interface AnalystFinding { rootCause; proposedMutation } + runAnalyst() → now the canonical AnalystFinding stamped by makeFinding (the exact shape improve(profile, findings, opts) reflects on); mutation rides recommended_action.
  • gate — bare v1Mean - v0Mean >= 0.5 point comparison → now pairedBootstrap(v0Scores, v1Scores, { seed }), shipping only when the paired-bootstrap CI lower bound clears 0 — the statistical core the production held-out gate (HeldOutGate / improve() over selfImprove) is built on.

Stays offline + deterministic; the demo still ships v1, now on a +5.00 paired median, 95% CI [5.00, 6.00], n=3. README + comments mirror the already-clean companion examples (improve/, intelligence-recommend/). Only the analyst body, proposer, and LLM remain scripted (for reproducibility) — the finding type and gate statistic are now the real substrate.

Verify (all green)

  • pnpm docs:checkdocs freshness: OK (§2 table symbols against 4134 public exports; prose symbols against 5795 resolvable; docs/api regen diff clean)
  • pnpm run build — clean
  • pnpm run typecheck + typecheck:examples — clean (the fixed example compiles)
  • pnpm run lint — Checked 319 files, no fixes applied
  • Example runs offline e2e: ships v1 on the paired CI [5.00, 6.00]

DO NOT MERGE — operator review.

…e examples
The anti-reinvention decision table (§2) had zero rows for whole families of
live exports, so agents kept hand-rolling them. Add grouped rows for:
- statistics — every test on the agent-eval main barrel (pairedBootstrap,
wilson, pairedTTest, wilcoxonSignedRank, mcnemar/mcnemarPower, mannWhitneyU,
passAtK, confidenceInterval, benjaminiHochberg/bonferroni, cohensD/cliffsDelta,
requiredSampleSize/pairedMde, eProcess, weightedComposite,
corpusInterRaterAgreement, pairedRiskDifference, pearsonR/spearmanR,
paretoFrontier/dominates)
- judge — JudgeConfig + ensembleJudge, with a warn-off for the @deprecated
JudgeFn factories and an explicit note that there is no llmJudge export
- authenticity — scoreAuthenticity / gateRealness / AuthenticitySignals
(agent-eval/authenticity subpath)
- verification — MultiLayerVerifier + gradeSemanticStatus
- reference-replay — runReferenceReplay / scoreReferenceReplay /
decideReferenceReplayPromotion over jsonlReferenceReplayStore
- token/usage seam — extractLlmCallEvent + reportLoopUsage / UsageSink (/loops)
Teach the freshness gate the agent-eval/authenticity subpath barrel so a §2 row
that recommends importing from it resolves (parallel to contract/campaign/index).
De-reinvent the one example that hand-rolled primitives:
examples/self-improving-loop replaces its local AnalystFinding interface +
runAnalyst with the canonical AnalystFinding stamped by makeFinding, and its
bare v1Mean-v0Mean>=0.5 point-comparison gate with pairedBootstrap (the
statistical core the production held-out gate is built on). Stays offline and
deterministic; ships v1 on a +5.00 [5.00, 6.00] paired CI. README + comments
mirror the clean companion examples (improve/, intelligence-recommend/).

@tangletoolstangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved PR — 70e098d8

Blanket team auto-approval is enabled for this reviewer service.
The full PR reviewer audit still runs separately and will publish findings if it detects issues.

tangletools · auto-approval · reason: blanket_auto_approve · 2026-06-24T08:54:31Z

drewstone added a commit that referenced this pull request Jun 24, 2026
…fImprove not a bare mean gate
Folds the one real example fix from #370 (otherwise superseded by the generated
catalog) into this PR: self-improving-loop hand-rolled the ship gate as a bare
mean comparison instead of the real HeldOutGate/selfImprove primitives.
@drewstone

Copy link
Copy Markdown
ContributorAuthor

Superseded by #371. The primitive inventory is now GENERATED (docs/api/primitive-catalog.md, freshness-gate-enforced) so it can't go stale by hand — the root cause this PR hand-patched. The ~30 primitives catalogued here are now auto-generated; the one real example fix (self-improving-loop → HeldOutGate/selfImprove) was folded into #371.

drewstone added a commit that referenced this pull request Jun 24, 2026
…erence cannot go stale (#371)
* docs(api): generate the primitive catalog so the anti-reinvention reference cannot go stale
The hand-listed primitive inventory in docs/canonical-api.md drifted from source:
it had zero mentions of live exports (scoreAuthenticity, gateRealness,
MultiLayerVerifier, wilson, pairedTTest, runProfileMatrix, extractUsage, …). Anything
derivable from source must be generated, not hand-written — only judgment stays curated.
- scripts/gen-primitive-catalog.mjs reads the LIVE exports of (a) this package's own
public subpaths (from package.json `exports`) and (b) a curated category->subpath map
of the @tangle-network/agent-eval substrate surfaces agents should reuse (judge,
authenticity, verification, statistics, campaign, token/usage). Extraction is via the
TypeScript compiler API over a virtual re-export entry, so it follows aliased
re-exports and content-hashed bundle files — the exact things that rot a hand list.
Emits docs/api/primitive-catalog.md with a GENERATED header (name, import path,
one-line summary per export, grouped by surface).
- Wired into `docs:api` (runs after TypeDoc). The freshness gate gains a seventh class
(CATALOG): it regenerates the catalog to a temp file and byte-compares to the committed
copy, so a new/removed/renamed live export absent from the catalog is a RED BUILD.
- Shrank canonical-api.md: removed the export-inventory enumeration from the banner and
the §2 preamble, replaced with pointers to docs/api/primitive-catalog.md. Kept all the
judgment — the decision gate, §1.5 AgentProfile law, the §2 "I want to -> use -> NOT"
table and every "Do NOT". The version + substrate-peer pins stay (gate-enforced).
- MAINTAINING.md documents the generated-inventory layer, CLASS 7, and its fix path.
* chore(deps): bump agent-eval to 0.99.0 + regenerate primitive catalog
agent-eval 0.99.0 adds llmJudge (+ the full current judge/auth/verify/stats
surface); regenerating the generated catalog picks it up with zero hand-work,
which is the point of the generator. Lockfile was pinned at 0.97.0 (pre-llmJudge)
despite agent-eval already being in minimumReleaseAgeExclude.
* fix(deps): keep agent-eval peer floor at >=0.97.0
Only the examples (devDependency) need 0.99.0 for llmJudge; agent-runtime's src
does not, so the peer floor must not force consumers onto 0.99.0. Catalog +
lockfile stay on the resolved 0.99.0 so the examples get llmJudge.
* docs(examples): de-reinvent self-improving-loop — use HeldOutGate/selfImprove not a bare mean gate
Folds the one real example fix from #370 (otherwise superseded by the generated
catalog) into this PR: self-improving-loop hand-rolled the ship gate as a bare
mean comparison instead of the real HeldOutGate/selfImprove primitives.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@drewstone@tangletools