Skip to content

M19.3/M19.5/M19.7: the three training gates become functions - #40

Merged
codeitlikemiley merged 1 commit into
mainfrom
m19-executable-gates
Aug 24, 2026
Merged

M19.3/M19.5/M19.7: the three training gates become functions#40
codeitlikemiley merged 1 commit into
mainfrom
m19-executable-gates

Conversation

@codeitlikemiley

Copy link
Copy Markdown
Owner

docs/19 states three acceptance gates in prose. None existed as code, so none could be failed — which is exactly what M19.1 was written against. The models are still absent and that does not change; what changes is that when one arrives, there is something to judge it with.

No fabricated scores. The gates ship, not results.

M19.3 — Score::beats(&incumbent, 10.0)

The milestone's bar is relative ("beats heuristic by ≥10pt"); the existing meets_the_gate is an absolute floor answering a different question.

  • The margin is percentage points, not a fraction. Passing 0.10 and meaning ten points is the obvious way to get this wrong, both readings compile and both look right, so a test pins the unit.
  • A confidently-wrong challenger is refused regardless of margin. The dangerous quadrant is what route-bench exists to measure. A caller checking only the margin would ship a model that is better on average and catastrophic on the cases where the router acts on the answer.

M19.5 — Report::cost_saving / cheaper_than(0.25, …)

reduce-bench had no cost dimension at all. It scored retention and token ratio, which are not money — so "≥25% cheaper" was unmeasurable even with a model in hand.

Cost rather than tokens, because the two move independently: a reducer that keeps 40% of the tokens and hands them to a model at 3× the price is more expensive, and the old ratio-only scorecard would have called that a 60% win.

Prices are per token and supplied by the caller — panday-harness has no price table and must not grow one, or the gate starts drifting with a catalog. A saving against an incumbent that spends nothing returns None rather than a number. Both halves required: cheap-with-regressions is a reducer that saves money by losing the answer.

M19.7 — Scorecard::beats(&base)

On the artifact type every suite already emits, so it isn't agent-bench's alone.

It returns Option<bool>, and the None is the substance. Comparing a tuned run over 10 tasks against a base run over 41 is how this measurement gets faked without anyone lying: both numbers are real, the ratio is meaningless, and a plain bool would have hidden it. Refused on suite mismatch, on an empty run, and on differing case counts. .unwrap_or(false) — what a caller will actually write — fails closed.

Falsified

Injected defectResult
M19.3 margin treated as a fractionred — unit test
M19.5 prices both sides identicallyred — cost test
M19.7 case-count guard removedred — two tests, incl. fail-closed

Green after restoring: 34/34 router, 42/42 harness, 13/13 types.

Numbering

No new milestone numbers — these are the gates of M19.3/5/7. Each bullet in docs/19 now records that its gate is executable while its model is not.

Verification

fmt · clippy --workspace --all-targets -D warnings · 1114 workspace tests, 0 failed · schemas · ts-sdk · sbom (609 components) · deny.

https://claude.ai/code/session_017kFpYDqvz6sKGSkM4YKaRf

docs/19 states three acceptance gates in prose. None existed as code, so none
could be failed — which is precisely what M19.1 was written against. The models
are still absent and that is unchanged; what changes is that when one arrives
there is something to judge it with.
**M19.3 — `Score::beats(&incumbent, 10.0)`.** The milestone's bar is relative
("beats heuristic by ≥10pt"); `meets_the_gate` is an absolute floor and answers
a different question. The margin is percentage points, not a fraction: passing
`0.10` and meaning ten points is the obvious way to get this wrong, both
readings compile, so a test pins the unit. A confidently-wrong challenger is
refused regardless of margin — the dangerous quadrant is what route-bench exists
to measure, and a caller checking only the margin would ship a model that is
better on average and catastrophic where the router acts on it.
**M19.5 — `Report::cost_saving` and `cheaper_than(0.25, ..)`.** reduce-bench had
**no cost dimension at all**: it scored retention and token ratio, which are not
money, so "≥25% cheaper" was unmeasurable with a model in hand. Cost rather than
tokens because the two move independently — a reducer keeping 40% of the tokens
and handing them to a model at 3x the price is *more* expensive, and the old
scorecard would have called that a 60% win. Prices are per token and supplied by
the caller; this crate has no price table and must not grow one, or the gate
drifts with a catalog. A saving against an incumbent that spends nothing is
`None`, not a number. Both halves are required: cheap-with-regressions is a
reducer that saves money by losing the answer.
**M19.7 — `Scorecard::beats(&base)`**, on the artifact type every suite emits.
It returns `Option<bool>` and the `None` is the substance: comparing a tuned run
over ten tasks against a base over forty-one is how this gets faked without
anyone lying — both numbers real, ratio meaningless — and a plain `bool` would
have hidden it. Refused on suite mismatch, on an empty run, and on differing
case counts. `.unwrap_or(false)` fails closed.
No new milestone numbers: these are the gates *of* M19.3/5/7, so each bullet
records that its gate is now executable while its model is not.
All three falsified. Making the margin a fraction reddens the unit test; pricing
both sides of reduce-bench identically reddens the cost test; removing the
case-count guard reddens two, including the fail-closed one. Green after
restoring.
Verified: fmt, clippy -D warnings, 1114 workspace tests, schemas, ts-sdk, sbom,
deny.
Claude-Session: https://claude.ai/code/session_017kFpYDqvz6sKGSkM4YKaRf
@codeitlikemiley
codeitlikemiley marked this pull request as ready for review August 24, 2026 00:15
@codeitlikemiley
codeitlikemiley merged commit 124029f into mainAug 24, 2026
6 checks passed
@codeitlikemiley
codeitlikemiley deleted the m19-executable-gates branch August 24, 2026 00:15
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@codeitlikemiley
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
M19.3/M19.5/M19.7: the three training gates become functions by codeitlikemiley · Pull Request #40 · codeitlikemiley/panday · GitHub
Skip to content

M19.3/M19.5/M19.7: the three training gates become functions - #40

Merged
codeitlikemiley merged 1 commit into
mainfrom
m19-executable-gates
Aug 24, 2026
Merged

M19.3/M19.5/M19.7: the three training gates become functions#40
codeitlikemiley merged 1 commit into
mainfrom
m19-executable-gates

Conversation

@codeitlikemiley

Copy link
Copy Markdown
Owner

docs/19 states three acceptance gates in prose. None existed as code, so none could be failed — which is exactly what M19.1 was written against. The models are still absent and that does not change; what changes is that when one arrives, there is something to judge it with.

No fabricated scores. The gates ship, not results.

M19.3 — Score::beats(&incumbent, 10.0)

The milestone's bar is relative ("beats heuristic by ≥10pt"); the existing meets_the_gate is an absolute floor answering a different question.

  • The margin is percentage points, not a fraction. Passing 0.10 and meaning ten points is the obvious way to get this wrong, both readings compile and both look right, so a test pins the unit.
  • A confidently-wrong challenger is refused regardless of margin. The dangerous quadrant is what route-bench exists to measure. A caller checking only the margin would ship a model that is better on average and catastrophic on the cases where the router acts on the answer.

M19.5 — Report::cost_saving / cheaper_than(0.25, …)

reduce-bench had no cost dimension at all. It scored retention and token ratio, which are not money — so "≥25% cheaper" was unmeasurable even with a model in hand.

Cost rather than tokens, because the two move independently: a reducer that keeps 40% of the tokens and hands them to a model at 3× the price is more expensive, and the old ratio-only scorecard would have called that a 60% win.

Prices are per token and supplied by the caller — panday-harness has no price table and must not grow one, or the gate starts drifting with a catalog. A saving against an incumbent that spends nothing returns None rather than a number. Both halves required: cheap-with-regressions is a reducer that saves money by losing the answer.

M19.7 — Scorecard::beats(&base)

On the artifact type every suite already emits, so it isn't agent-bench's alone.

It returns Option<bool>, and the None is the substance. Comparing a tuned run over 10 tasks against a base run over 41 is how this measurement gets faked without anyone lying: both numbers are real, the ratio is meaningless, and a plain bool would have hidden it. Refused on suite mismatch, on an empty run, and on differing case counts. .unwrap_or(false) — what a caller will actually write — fails closed.

Falsified

Injected defectResult
M19.3 margin treated as a fractionred — unit test
M19.5 prices both sides identicallyred — cost test
M19.7 case-count guard removedred — two tests, incl. fail-closed

Green after restoring: 34/34 router, 42/42 harness, 13/13 types.

Numbering

No new milestone numbers — these are the gates of M19.3/5/7. Each bullet in docs/19 now records that its gate is executable while its model is not.

Verification

fmt · clippy --workspace --all-targets -D warnings · 1114 workspace tests, 0 failed · schemas · ts-sdk · sbom (609 components) · deny.

https://claude.ai/code/session_017kFpYDqvz6sKGSkM4YKaRf

docs/19 states three acceptance gates in prose. None existed as code, so none
could be failed — which is precisely what M19.1 was written against. The models
are still absent and that is unchanged; what changes is that when one arrives
there is something to judge it with.
**M19.3 — `Score::beats(&incumbent, 10.0)`.** The milestone's bar is relative
("beats heuristic by ≥10pt"); `meets_the_gate` is an absolute floor and answers
a different question. The margin is percentage points, not a fraction: passing
`0.10` and meaning ten points is the obvious way to get this wrong, both
readings compile, so a test pins the unit. A confidently-wrong challenger is
refused regardless of margin — the dangerous quadrant is what route-bench exists
to measure, and a caller checking only the margin would ship a model that is
better on average and catastrophic where the router acts on it.
**M19.5 — `Report::cost_saving` and `cheaper_than(0.25, ..)`.** reduce-bench had
**no cost dimension at all**: it scored retention and token ratio, which are not
money, so "≥25% cheaper" was unmeasurable with a model in hand. Cost rather than
tokens because the two move independently — a reducer keeping 40% of the tokens
and handing them to a model at 3x the price is *more* expensive, and the old
scorecard would have called that a 60% win. Prices are per token and supplied by
the caller; this crate has no price table and must not grow one, or the gate
drifts with a catalog. A saving against an incumbent that spends nothing is
`None`, not a number. Both halves are required: cheap-with-regressions is a
reducer that saves money by losing the answer.
**M19.7 — `Scorecard::beats(&base)`**, on the artifact type every suite emits.
It returns `Option<bool>` and the `None` is the substance: comparing a tuned run
over ten tasks against a base over forty-one is how this gets faked without
anyone lying — both numbers real, ratio meaningless — and a plain `bool` would
have hidden it. Refused on suite mismatch, on an empty run, and on differing
case counts. `.unwrap_or(false)` fails closed.
No new milestone numbers: these are the gates *of* M19.3/5/7, so each bullet
records that its gate is now executable while its model is not.
All three falsified. Making the margin a fraction reddens the unit test; pricing
both sides of reduce-bench identically reddens the cost test; removing the
case-count guard reddens two, including the fail-closed one. Green after
restoring.
Verified: fmt, clippy -D warnings, 1114 workspace tests, schemas, ts-sdk, sbom,
deny.
Claude-Session: https://claude.ai/code/session_017kFpYDqvz6sKGSkM4YKaRf
@codeitlikemiley
codeitlikemiley marked this pull request as ready for review August 24, 2026 00:15
@codeitlikemiley
codeitlikemiley merged commit 124029f into mainAug 24, 2026
6 checks passed
@codeitlikemiley
codeitlikemiley deleted the m19-executable-gates branch August 24, 2026 00:15
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@codeitlikemiley
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' M19.3/M19.5/M19.7: the three training gates become functions by codeitlikemiley · Pull Request #40 · codeitlikemiley/panday · GitHub
Skip to content

M19.3/M19.5/M19.7: the three training gates become functions - #40

Merged
codeitlikemiley merged 1 commit into
mainfrom
m19-executable-gates
Aug 24, 2026
Merged

M19.3/M19.5/M19.7: the three training gates become functions#40
codeitlikemiley merged 1 commit into
mainfrom
m19-executable-gates

Conversation

@codeitlikemiley

Copy link
Copy Markdown
Owner

docs/19 states three acceptance gates in prose. None existed as code, so none could be failed — which is exactly what M19.1 was written against. The models are still absent and that does not change; what changes is that when one arrives, there is something to judge it with.

No fabricated scores. The gates ship, not results.

M19.3 — Score::beats(&incumbent, 10.0)

The milestone's bar is relative ("beats heuristic by ≥10pt"); the existing meets_the_gate is an absolute floor answering a different question.

  • The margin is percentage points, not a fraction. Passing 0.10 and meaning ten points is the obvious way to get this wrong, both readings compile and both look right, so a test pins the unit.
  • A confidently-wrong challenger is refused regardless of margin. The dangerous quadrant is what route-bench exists to measure. A caller checking only the margin would ship a model that is better on average and catastrophic on the cases where the router acts on the answer.

M19.5 — Report::cost_saving / cheaper_than(0.25, …)

reduce-bench had no cost dimension at all. It scored retention and token ratio, which are not money — so "≥25% cheaper" was unmeasurable even with a model in hand.

Cost rather than tokens, because the two move independently: a reducer that keeps 40% of the tokens and hands them to a model at 3× the price is more expensive, and the old ratio-only scorecard would have called that a 60% win.

Prices are per token and supplied by the caller — panday-harness has no price table and must not grow one, or the gate starts drifting with a catalog. A saving against an incumbent that spends nothing returns None rather than a number. Both halves required: cheap-with-regressions is a reducer that saves money by losing the answer.

M19.7 — Scorecard::beats(&base)

On the artifact type every suite already emits, so it isn't agent-bench's alone.

It returns Option<bool>, and the None is the substance. Comparing a tuned run over 10 tasks against a base run over 41 is how this measurement gets faked without anyone lying: both numbers are real, the ratio is meaningless, and a plain bool would have hidden it. Refused on suite mismatch, on an empty run, and on differing case counts. .unwrap_or(false) — what a caller will actually write — fails closed.

Falsified

Injected defectResult
M19.3 margin treated as a fractionred — unit test
M19.5 prices both sides identicallyred — cost test
M19.7 case-count guard removedred — two tests, incl. fail-closed

Green after restoring: 34/34 router, 42/42 harness, 13/13 types.

Numbering

No new milestone numbers — these are the gates of M19.3/5/7. Each bullet in docs/19 now records that its gate is executable while its model is not.

Verification

fmt · clippy --workspace --all-targets -D warnings · 1114 workspace tests, 0 failed · schemas · ts-sdk · sbom (609 components) · deny.

https://claude.ai/code/session_017kFpYDqvz6sKGSkM4YKaRf

docs/19 states three acceptance gates in prose. None existed as code, so none
could be failed — which is precisely what M19.1 was written against. The models
are still absent and that is unchanged; what changes is that when one arrives
there is something to judge it with.
**M19.3 — `Score::beats(&incumbent, 10.0)`.** The milestone's bar is relative
("beats heuristic by ≥10pt"); `meets_the_gate` is an absolute floor and answers
a different question. The margin is percentage points, not a fraction: passing
`0.10` and meaning ten points is the obvious way to get this wrong, both
readings compile, so a test pins the unit. A confidently-wrong challenger is
refused regardless of margin — the dangerous quadrant is what route-bench exists
to measure, and a caller checking only the margin would ship a model that is
better on average and catastrophic where the router acts on it.
**M19.5 — `Report::cost_saving` and `cheaper_than(0.25, ..)`.** reduce-bench had
**no cost dimension at all**: it scored retention and token ratio, which are not
money, so "≥25% cheaper" was unmeasurable with a model in hand. Cost rather than
tokens because the two move independently — a reducer keeping 40% of the tokens
and handing them to a model at 3x the price is *more* expensive, and the old
scorecard would have called that a 60% win. Prices are per token and supplied by
the caller; this crate has no price table and must not grow one, or the gate
drifts with a catalog. A saving against an incumbent that spends nothing is
`None`, not a number. Both halves are required: cheap-with-regressions is a
reducer that saves money by losing the answer.
**M19.7 — `Scorecard::beats(&base)`**, on the artifact type every suite emits.
It returns `Option<bool>` and the `None` is the substance: comparing a tuned run
over ten tasks against a base over forty-one is how this gets faked without
anyone lying — both numbers real, ratio meaningless — and a plain `bool` would
have hidden it. Refused on suite mismatch, on an empty run, and on differing
case counts. `.unwrap_or(false)` fails closed.
No new milestone numbers: these are the gates *of* M19.3/5/7, so each bullet
records that its gate is now executable while its model is not.
All three falsified. Making the margin a fraction reddens the unit test; pricing
both sides of reduce-bench identically reddens the cost test; removing the
case-count guard reddens two, including the fail-closed one. Green after
restoring.
Verified: fmt, clippy -D warnings, 1114 workspace tests, schemas, ts-sdk, sbom,
deny.
Claude-Session: https://claude.ai/code/session_017kFpYDqvz6sKGSkM4YKaRf
@codeitlikemiley
codeitlikemiley marked this pull request as ready for review August 24, 2026 00:15
@codeitlikemiley
codeitlikemiley merged commit 124029f into mainAug 24, 2026
6 checks passed
@codeitlikemiley
codeitlikemiley deleted the m19-executable-gates branch August 24, 2026 00:15
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@codeitlikemiley
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' M19.3/M19.5/M19.7: the three training gates become functions by codeitlikemiley · Pull Request #40 · codeitlikemiley/panday · GitHub
Skip to content

M19.3/M19.5/M19.7: the three training gates become functions - #40

Merged
codeitlikemiley merged 1 commit into
mainfrom
m19-executable-gates
Aug 24, 2026
Merged

M19.3/M19.5/M19.7: the three training gates become functions#40
codeitlikemiley merged 1 commit into
mainfrom
m19-executable-gates

Conversation

@codeitlikemiley

Copy link
Copy Markdown
Owner

docs/19 states three acceptance gates in prose. None existed as code, so none could be failed — which is exactly what M19.1 was written against. The models are still absent and that does not change; what changes is that when one arrives, there is something to judge it with.

No fabricated scores. The gates ship, not results.

M19.3 — Score::beats(&incumbent, 10.0)

The milestone's bar is relative ("beats heuristic by ≥10pt"); the existing meets_the_gate is an absolute floor answering a different question.

  • The margin is percentage points, not a fraction. Passing 0.10 and meaning ten points is the obvious way to get this wrong, both readings compile and both look right, so a test pins the unit.
  • A confidently-wrong challenger is refused regardless of margin. The dangerous quadrant is what route-bench exists to measure. A caller checking only the margin would ship a model that is better on average and catastrophic on the cases where the router acts on the answer.

M19.5 — Report::cost_saving / cheaper_than(0.25, …)

reduce-bench had no cost dimension at all. It scored retention and token ratio, which are not money — so "≥25% cheaper" was unmeasurable even with a model in hand.

Cost rather than tokens, because the two move independently: a reducer that keeps 40% of the tokens and hands them to a model at 3× the price is more expensive, and the old ratio-only scorecard would have called that a 60% win.

Prices are per token and supplied by the caller — panday-harness has no price table and must not grow one, or the gate starts drifting with a catalog. A saving against an incumbent that spends nothing returns None rather than a number. Both halves required: cheap-with-regressions is a reducer that saves money by losing the answer.

M19.7 — Scorecard::beats(&base)

On the artifact type every suite already emits, so it isn't agent-bench's alone.

It returns Option<bool>, and the None is the substance. Comparing a tuned run over 10 tasks against a base run over 41 is how this measurement gets faked without anyone lying: both numbers are real, the ratio is meaningless, and a plain bool would have hidden it. Refused on suite mismatch, on an empty run, and on differing case counts. .unwrap_or(false) — what a caller will actually write — fails closed.

Falsified

Injected defectResult
M19.3 margin treated as a fractionred — unit test
M19.5 prices both sides identicallyred — cost test
M19.7 case-count guard removedred — two tests, incl. fail-closed

Green after restoring: 34/34 router, 42/42 harness, 13/13 types.

Numbering

No new milestone numbers — these are the gates of M19.3/5/7. Each bullet in docs/19 now records that its gate is executable while its model is not.

Verification

fmt · clippy --workspace --all-targets -D warnings · 1114 workspace tests, 0 failed · schemas · ts-sdk · sbom (609 components) · deny.

https://claude.ai/code/session_017kFpYDqvz6sKGSkM4YKaRf

docs/19 states three acceptance gates in prose. None existed as code, so none
could be failed — which is precisely what M19.1 was written against. The models
are still absent and that is unchanged; what changes is that when one arrives
there is something to judge it with.
**M19.3 — `Score::beats(&incumbent, 10.0)`.** The milestone's bar is relative
("beats heuristic by ≥10pt"); `meets_the_gate` is an absolute floor and answers
a different question. The margin is percentage points, not a fraction: passing
`0.10` and meaning ten points is the obvious way to get this wrong, both
readings compile, so a test pins the unit. A confidently-wrong challenger is
refused regardless of margin — the dangerous quadrant is what route-bench exists
to measure, and a caller checking only the margin would ship a model that is
better on average and catastrophic where the router acts on it.
**M19.5 — `Report::cost_saving` and `cheaper_than(0.25, ..)`.** reduce-bench had
**no cost dimension at all**: it scored retention and token ratio, which are not
money, so "≥25% cheaper" was unmeasurable with a model in hand. Cost rather than
tokens because the two move independently — a reducer keeping 40% of the tokens
and handing them to a model at 3x the price is *more* expensive, and the old
scorecard would have called that a 60% win. Prices are per token and supplied by
the caller; this crate has no price table and must not grow one, or the gate
drifts with a catalog. A saving against an incumbent that spends nothing is
`None`, not a number. Both halves are required: cheap-with-regressions is a
reducer that saves money by losing the answer.
**M19.7 — `Scorecard::beats(&base)`**, on the artifact type every suite emits.
It returns `Option<bool>` and the `None` is the substance: comparing a tuned run
over ten tasks against a base over forty-one is how this gets faked without
anyone lying — both numbers real, ratio meaningless — and a plain `bool` would
have hidden it. Refused on suite mismatch, on an empty run, and on differing
case counts. `.unwrap_or(false)` fails closed.
No new milestone numbers: these are the gates *of* M19.3/5/7, so each bullet
records that its gate is now executable while its model is not.
All three falsified. Making the margin a fraction reddens the unit test; pricing
both sides of reduce-bench identically reddens the cost test; removing the
case-count guard reddens two, including the fail-closed one. Green after
restoring.
Verified: fmt, clippy -D warnings, 1114 workspace tests, schemas, ts-sdk, sbom,
deny.
Claude-Session: https://claude.ai/code/session_017kFpYDqvz6sKGSkM4YKaRf
@codeitlikemiley
codeitlikemiley marked this pull request as ready for review August 24, 2026 00:15
@codeitlikemiley
codeitlikemiley merged commit 124029f into mainAug 24, 2026
6 checks passed
@codeitlikemiley
codeitlikemiley deleted the m19-executable-gates branch August 24, 2026 00:15
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@codeitlikemiley
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' M19.3/M19.5/M19.7: the three training gates become functions by codeitlikemiley · Pull Request #40 · codeitlikemiley/panday · GitHub
Skip to content

M19.3/M19.5/M19.7: the three training gates become functions - #40

Merged
codeitlikemiley merged 1 commit into
mainfrom
m19-executable-gates
Aug 24, 2026
Merged

M19.3/M19.5/M19.7: the three training gates become functions#40
codeitlikemiley merged 1 commit into
mainfrom
m19-executable-gates

Conversation

@codeitlikemiley

Copy link
Copy Markdown
Owner

docs/19 states three acceptance gates in prose. None existed as code, so none could be failed — which is exactly what M19.1 was written against. The models are still absent and that does not change; what changes is that when one arrives, there is something to judge it with.

No fabricated scores. The gates ship, not results.

M19.3 — Score::beats(&incumbent, 10.0)

The milestone's bar is relative ("beats heuristic by ≥10pt"); the existing meets_the_gate is an absolute floor answering a different question.

  • The margin is percentage points, not a fraction. Passing 0.10 and meaning ten points is the obvious way to get this wrong, both readings compile and both look right, so a test pins the unit.
  • A confidently-wrong challenger is refused regardless of margin. The dangerous quadrant is what route-bench exists to measure. A caller checking only the margin would ship a model that is better on average and catastrophic on the cases where the router acts on the answer.

M19.5 — Report::cost_saving / cheaper_than(0.25, …)

reduce-bench had no cost dimension at all. It scored retention and token ratio, which are not money — so "≥25% cheaper" was unmeasurable even with a model in hand.

Cost rather than tokens, because the two move independently: a reducer that keeps 40% of the tokens and hands them to a model at 3× the price is more expensive, and the old ratio-only scorecard would have called that a 60% win.

Prices are per token and supplied by the caller — panday-harness has no price table and must not grow one, or the gate starts drifting with a catalog. A saving against an incumbent that spends nothing returns None rather than a number. Both halves required: cheap-with-regressions is a reducer that saves money by losing the answer.

M19.7 — Scorecard::beats(&base)

On the artifact type every suite already emits, so it isn't agent-bench's alone.

It returns Option<bool>, and the None is the substance. Comparing a tuned run over 10 tasks against a base run over 41 is how this measurement gets faked without anyone lying: both numbers are real, the ratio is meaningless, and a plain bool would have hidden it. Refused on suite mismatch, on an empty run, and on differing case counts. .unwrap_or(false) — what a caller will actually write — fails closed.

Falsified

Injected defectResult
M19.3 margin treated as a fractionred — unit test
M19.5 prices both sides identicallyred — cost test
M19.7 case-count guard removedred — two tests, incl. fail-closed

Green after restoring: 34/34 router, 42/42 harness, 13/13 types.

Numbering

No new milestone numbers — these are the gates of M19.3/5/7. Each bullet in docs/19 now records that its gate is executable while its model is not.

Verification

fmt · clippy --workspace --all-targets -D warnings · 1114 workspace tests, 0 failed · schemas · ts-sdk · sbom (609 components) · deny.

https://claude.ai/code/session_017kFpYDqvz6sKGSkM4YKaRf

docs/19 states three acceptance gates in prose. None existed as code, so none
could be failed — which is precisely what M19.1 was written against. The models
are still absent and that is unchanged; what changes is that when one arrives
there is something to judge it with.
**M19.3 — `Score::beats(&incumbent, 10.0)`.** The milestone's bar is relative
("beats heuristic by ≥10pt"); `meets_the_gate` is an absolute floor and answers
a different question. The margin is percentage points, not a fraction: passing
`0.10` and meaning ten points is the obvious way to get this wrong, both
readings compile, so a test pins the unit. A confidently-wrong challenger is
refused regardless of margin — the dangerous quadrant is what route-bench exists
to measure, and a caller checking only the margin would ship a model that is
better on average and catastrophic where the router acts on it.
**M19.5 — `Report::cost_saving` and `cheaper_than(0.25, ..)`.** reduce-bench had
**no cost dimension at all**: it scored retention and token ratio, which are not
money, so "≥25% cheaper" was unmeasurable with a model in hand. Cost rather than
tokens because the two move independently — a reducer keeping 40% of the tokens
and handing them to a model at 3x the price is *more* expensive, and the old
scorecard would have called that a 60% win. Prices are per token and supplied by
the caller; this crate has no price table and must not grow one, or the gate
drifts with a catalog. A saving against an incumbent that spends nothing is
`None`, not a number. Both halves are required: cheap-with-regressions is a
reducer that saves money by losing the answer.
**M19.7 — `Scorecard::beats(&base)`**, on the artifact type every suite emits.
It returns `Option<bool>` and the `None` is the substance: comparing a tuned run
over ten tasks against a base over forty-one is how this gets faked without
anyone lying — both numbers real, ratio meaningless — and a plain `bool` would
have hidden it. Refused on suite mismatch, on an empty run, and on differing
case counts. `.unwrap_or(false)` fails closed.
No new milestone numbers: these are the gates *of* M19.3/5/7, so each bullet
records that its gate is now executable while its model is not.
All three falsified. Making the margin a fraction reddens the unit test; pricing
both sides of reduce-bench identically reddens the cost test; removing the
case-count guard reddens two, including the fail-closed one. Green after
restoring.
Verified: fmt, clippy -D warnings, 1114 workspace tests, schemas, ts-sdk, sbom,
deny.
Claude-Session: https://claude.ai/code/session_017kFpYDqvz6sKGSkM4YKaRf
@codeitlikemiley
codeitlikemiley marked this pull request as ready for review August 24, 2026 00:15
@codeitlikemiley
codeitlikemiley merged commit 124029f into mainAug 24, 2026
6 checks passed
@codeitlikemiley
codeitlikemiley deleted the m19-executable-gates branch August 24, 2026 00:15
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@codeitlikemiley
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' M19.3/M19.5/M19.7: the three training gates become functions by codeitlikemiley · Pull Request #40 · codeitlikemiley/panday · GitHub
Skip to content

M19.3/M19.5/M19.7: the three training gates become functions - #40

Merged
codeitlikemiley merged 1 commit into
mainfrom
m19-executable-gates
Aug 24, 2026
Merged

M19.3/M19.5/M19.7: the three training gates become functions#40
codeitlikemiley merged 1 commit into
mainfrom
m19-executable-gates

Conversation

@codeitlikemiley

Copy link
Copy Markdown
Owner

docs/19 states three acceptance gates in prose. None existed as code, so none could be failed — which is exactly what M19.1 was written against. The models are still absent and that does not change; what changes is that when one arrives, there is something to judge it with.

No fabricated scores. The gates ship, not results.

M19.3 — Score::beats(&incumbent, 10.0)

The milestone's bar is relative ("beats heuristic by ≥10pt"); the existing meets_the_gate is an absolute floor answering a different question.

  • The margin is percentage points, not a fraction. Passing 0.10 and meaning ten points is the obvious way to get this wrong, both readings compile and both look right, so a test pins the unit.
  • A confidently-wrong challenger is refused regardless of margin. The dangerous quadrant is what route-bench exists to measure. A caller checking only the margin would ship a model that is better on average and catastrophic on the cases where the router acts on the answer.

M19.5 — Report::cost_saving / cheaper_than(0.25, …)

reduce-bench had no cost dimension at all. It scored retention and token ratio, which are not money — so "≥25% cheaper" was unmeasurable even with a model in hand.

Cost rather than tokens, because the two move independently: a reducer that keeps 40% of the tokens and hands them to a model at 3× the price is more expensive, and the old ratio-only scorecard would have called that a 60% win.

Prices are per token and supplied by the caller — panday-harness has no price table and must not grow one, or the gate starts drifting with a catalog. A saving against an incumbent that spends nothing returns None rather than a number. Both halves required: cheap-with-regressions is a reducer that saves money by losing the answer.

M19.7 — Scorecard::beats(&base)

On the artifact type every suite already emits, so it isn't agent-bench's alone.

It returns Option<bool>, and the None is the substance. Comparing a tuned run over 10 tasks against a base run over 41 is how this measurement gets faked without anyone lying: both numbers are real, the ratio is meaningless, and a plain bool would have hidden it. Refused on suite mismatch, on an empty run, and on differing case counts. .unwrap_or(false) — what a caller will actually write — fails closed.

Falsified

Injected defectResult
M19.3 margin treated as a fractionred — unit test
M19.5 prices both sides identicallyred — cost test
M19.7 case-count guard removedred — two tests, incl. fail-closed

Green after restoring: 34/34 router, 42/42 harness, 13/13 types.

Numbering

No new milestone numbers — these are the gates of M19.3/5/7. Each bullet in docs/19 now records that its gate is executable while its model is not.

Verification

fmt · clippy --workspace --all-targets -D warnings · 1114 workspace tests, 0 failed · schemas · ts-sdk · sbom (609 components) · deny.

https://claude.ai/code/session_017kFpYDqvz6sKGSkM4YKaRf

docs/19 states three acceptance gates in prose. None existed as code, so none
could be failed — which is precisely what M19.1 was written against. The models
are still absent and that is unchanged; what changes is that when one arrives
there is something to judge it with.
**M19.3 — `Score::beats(&incumbent, 10.0)`.** The milestone's bar is relative
("beats heuristic by ≥10pt"); `meets_the_gate` is an absolute floor and answers
a different question. The margin is percentage points, not a fraction: passing
`0.10` and meaning ten points is the obvious way to get this wrong, both
readings compile, so a test pins the unit. A confidently-wrong challenger is
refused regardless of margin — the dangerous quadrant is what route-bench exists
to measure, and a caller checking only the margin would ship a model that is
better on average and catastrophic where the router acts on it.
**M19.5 — `Report::cost_saving` and `cheaper_than(0.25, ..)`.** reduce-bench had
**no cost dimension at all**: it scored retention and token ratio, which are not
money, so "≥25% cheaper" was unmeasurable with a model in hand. Cost rather than
tokens because the two move independently — a reducer keeping 40% of the tokens
and handing them to a model at 3x the price is *more* expensive, and the old
scorecard would have called that a 60% win. Prices are per token and supplied by
the caller; this crate has no price table and must not grow one, or the gate
drifts with a catalog. A saving against an incumbent that spends nothing is
`None`, not a number. Both halves are required: cheap-with-regressions is a
reducer that saves money by losing the answer.
**M19.7 — `Scorecard::beats(&base)`**, on the artifact type every suite emits.
It returns `Option<bool>` and the `None` is the substance: comparing a tuned run
over ten tasks against a base over forty-one is how this gets faked without
anyone lying — both numbers real, ratio meaningless — and a plain `bool` would
have hidden it. Refused on suite mismatch, on an empty run, and on differing
case counts. `.unwrap_or(false)` fails closed.
No new milestone numbers: these are the gates *of* M19.3/5/7, so each bullet
records that its gate is now executable while its model is not.
All three falsified. Making the margin a fraction reddens the unit test; pricing
both sides of reduce-bench identically reddens the cost test; removing the
case-count guard reddens two, including the fail-closed one. Green after
restoring.
Verified: fmt, clippy -D warnings, 1114 workspace tests, schemas, ts-sdk, sbom,
deny.
Claude-Session: https://claude.ai/code/session_017kFpYDqvz6sKGSkM4YKaRf
@codeitlikemiley
codeitlikemiley marked this pull request as ready for review August 24, 2026 00:15
@codeitlikemiley
codeitlikemiley merged commit 124029f into mainAug 24, 2026
6 checks passed
@codeitlikemiley
codeitlikemiley deleted the m19-executable-gates branch August 24, 2026 00:15
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@codeitlikemiley
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); })(); M19.3/M19.5/M19.7: the three training gates become functions by codeitlikemiley · Pull Request #40 · codeitlikemiley/panday · GitHub
Skip to content

M19.3/M19.5/M19.7: the three training gates become functions - #40

Merged
codeitlikemiley merged 1 commit into
mainfrom
m19-executable-gates
Aug 24, 2026
Merged

M19.3/M19.5/M19.7: the three training gates become functions#40
codeitlikemiley merged 1 commit into
mainfrom
m19-executable-gates

Conversation

@codeitlikemiley

Copy link
Copy Markdown
Owner

docs/19 states three acceptance gates in prose. None existed as code, so none could be failed — which is exactly what M19.1 was written against. The models are still absent and that does not change; what changes is that when one arrives, there is something to judge it with.

No fabricated scores. The gates ship, not results.

M19.3 — Score::beats(&incumbent, 10.0)

The milestone's bar is relative ("beats heuristic by ≥10pt"); the existing meets_the_gate is an absolute floor answering a different question.

  • The margin is percentage points, not a fraction. Passing 0.10 and meaning ten points is the obvious way to get this wrong, both readings compile and both look right, so a test pins the unit.
  • A confidently-wrong challenger is refused regardless of margin. The dangerous quadrant is what route-bench exists to measure. A caller checking only the margin would ship a model that is better on average and catastrophic on the cases where the router acts on the answer.

M19.5 — Report::cost_saving / cheaper_than(0.25, …)

reduce-bench had no cost dimension at all. It scored retention and token ratio, which are not money — so "≥25% cheaper" was unmeasurable even with a model in hand.

Cost rather than tokens, because the two move independently: a reducer that keeps 40% of the tokens and hands them to a model at 3× the price is more expensive, and the old ratio-only scorecard would have called that a 60% win.

Prices are per token and supplied by the caller — panday-harness has no price table and must not grow one, or the gate starts drifting with a catalog. A saving against an incumbent that spends nothing returns None rather than a number. Both halves required: cheap-with-regressions is a reducer that saves money by losing the answer.

M19.7 — Scorecard::beats(&base)

On the artifact type every suite already emits, so it isn't agent-bench's alone.

It returns Option<bool>, and the None is the substance. Comparing a tuned run over 10 tasks against a base run over 41 is how this measurement gets faked without anyone lying: both numbers are real, the ratio is meaningless, and a plain bool would have hidden it. Refused on suite mismatch, on an empty run, and on differing case counts. .unwrap_or(false) — what a caller will actually write — fails closed.

Falsified

Injected defectResult
M19.3 margin treated as a fractionred — unit test
M19.5 prices both sides identicallyred — cost test
M19.7 case-count guard removedred — two tests, incl. fail-closed

Green after restoring: 34/34 router, 42/42 harness, 13/13 types.

Numbering

No new milestone numbers — these are the gates of M19.3/5/7. Each bullet in docs/19 now records that its gate is executable while its model is not.

Verification

fmt · clippy --workspace --all-targets -D warnings · 1114 workspace tests, 0 failed · schemas · ts-sdk · sbom (609 components) · deny.

https://claude.ai/code/session_017kFpYDqvz6sKGSkM4YKaRf

docs/19 states three acceptance gates in prose. None existed as code, so none
could be failed — which is precisely what M19.1 was written against. The models
are still absent and that is unchanged; what changes is that when one arrives
there is something to judge it with.
**M19.3 — `Score::beats(&incumbent, 10.0)`.** The milestone's bar is relative
("beats heuristic by ≥10pt"); `meets_the_gate` is an absolute floor and answers
a different question. The margin is percentage points, not a fraction: passing
`0.10` and meaning ten points is the obvious way to get this wrong, both
readings compile, so a test pins the unit. A confidently-wrong challenger is
refused regardless of margin — the dangerous quadrant is what route-bench exists
to measure, and a caller checking only the margin would ship a model that is
better on average and catastrophic where the router acts on it.
**M19.5 — `Report::cost_saving` and `cheaper_than(0.25, ..)`.** reduce-bench had
**no cost dimension at all**: it scored retention and token ratio, which are not
money, so "≥25% cheaper" was unmeasurable with a model in hand. Cost rather than
tokens because the two move independently — a reducer keeping 40% of the tokens
and handing them to a model at 3x the price is *more* expensive, and the old
scorecard would have called that a 60% win. Prices are per token and supplied by
the caller; this crate has no price table and must not grow one, or the gate
drifts with a catalog. A saving against an incumbent that spends nothing is
`None`, not a number. Both halves are required: cheap-with-regressions is a
reducer that saves money by losing the answer.
**M19.7 — `Scorecard::beats(&base)`**, on the artifact type every suite emits.
It returns `Option<bool>` and the `None` is the substance: comparing a tuned run
over ten tasks against a base over forty-one is how this gets faked without
anyone lying — both numbers real, ratio meaningless — and a plain `bool` would
have hidden it. Refused on suite mismatch, on an empty run, and on differing
case counts. `.unwrap_or(false)` fails closed.
No new milestone numbers: these are the gates *of* M19.3/5/7, so each bullet
records that its gate is now executable while its model is not.
All three falsified. Making the margin a fraction reddens the unit test; pricing
both sides of reduce-bench identically reddens the cost test; removing the
case-count guard reddens two, including the fail-closed one. Green after
restoring.
Verified: fmt, clippy -D warnings, 1114 workspace tests, schemas, ts-sdk, sbom,
deny.
Claude-Session: https://claude.ai/code/session_017kFpYDqvz6sKGSkM4YKaRf
@codeitlikemiley
codeitlikemiley marked this pull request as ready for review August 24, 2026 00:15
@codeitlikemiley
codeitlikemiley merged commit 124029f into mainAug 24, 2026
6 checks passed
@codeitlikemiley
codeitlikemiley deleted the m19-executable-gates branch August 24, 2026 00:15
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@codeitlikemiley