Skip to content

[finding] The type-check debt ledger banked 118 raw errors of surplus across 4 of 31 entries — the measurement #6376 lacked when it decided a surplus must not go red #12799

Description

@claude

Restart-when: pnpm check:type-check-debt --re-measure prints a non-empty surplus line (surplus: N raw error(s) across M entr(ies) sit BELOW their recorded ceiling) instead of surplus: none
Restart-touch: scripts/check-type-check-coverage.mjs

PM status block, added 2026-08-28 by the domain:devx @ objectstack seat (session session_01CPrUz21stTFhJRUirdc4yw). The two lines above are the machine-readable half of pm:on-hold; the reasoning is in the ACCEPT comment below. The maintainer's ruling (option 3) STANDS — only its precondition is unmet. R24 dispatched this card, took the second reading the ruling requires, and stopped as the ruling itself directs. This is the only change to this body; everything under the rule is the filer's original text, unedited.

Why held rather than queued: returning it to pm:queue would have it picked again and stopped again for the same reason. Both readings now exist — interval start PR #12798 / 96dc446c2026-08-27T20:44:25Z, second reading origin/mainaef1b7e62026-08-28T09:07:59Z, elapsed 12.4 hours — and the second contributes zero surplus entries, so the per-entry size distribution a threshold must be argued from still has exactly one sample (103 / 10 / 3 / 2). The only observed rate (0.97 raw/day) predicts 0.5 raw errors over an interval this short, so the observed zero discriminates nothing.

⚠️Nothing is at risk while this waits. Every ledger entry currently sits exactly at its measurement, so any new raw error is red immediately — the gap this card exists to close is zero today.

⭐⭐ MANDATORY for whoever takes it next, and it is the reason a second attempt was possible at all: record the PER-ENTRY surplus sizes on this card BEFORE running --lower.--lower destroys them the moment it closes the gap. 103/10/3/2 exists only because #12798 wrote it down first.

⭐ A design note that is data, ⛔ not a threshold: the 118 was 103 in one entry with the rest at 10/3/2, so a total-based threshold and a per-entry one behave very differently on the only sample there is — a total of ~50 would have fired at that moment while three of its four entries were tiny.

Timing guidance (⛔ not a restart condition): the shortest interval the observed rate would fill with enough per-entry surpluses to have a size distribution at all is roughly 8–12 weeks from 2026-08-27. The Restart-when: above fires on the signal rather than the date, because nothing in the repo schedules a periodic re-measure — that scheduling gap is the standing cost recorded against option 1, and closing it (a recurring re-measure in CI) is larger than this card. ⚠️#12511 and #12856 also sit on scripts/check-type-check-coverage.mjs undispatched.


Filed unassigned and ungraded by the #12723 dev (session session_01PfaSTikked61BkcsB5Rn69), as the explicit follow-up that card's triage asked for: "If several entries carry slack, file the follow-up with that count as its evidence." Severity not judged, routing not decided — observation-class.

The measurement that makes this fileable

#12723 / PR #12798 re-measured all 31 entries of DEBT + TEST_DEBT in scripts/check-type-check-coverage.mjs, on a closure the gate built itself, at ead731756:

entries carrying slack4 of 31
raw errors of slack118
largest single entry@objectstack/objectql, 354 recorded / 251 measured
othersruntime 227/217, plugin-auth 97/94, plugin-approvals 347/345
DEBT entries carrying slack0 of 13 — every one measured its recorded number to the unit

The gate's own summary line was green for the whole of it, correctly: its contract is "none above its recorded number", and shrinkage is informational by a standing ruling (#6376, PR #6510 — report the surplus and offer --lower, never fail on a surplus alone).

The question this is asking to re-open, and only this one

Should a ledger entry sitting above its measurement make the gate RED — i.e. should the ratchet be self-tightening?

#6376 decided no, and the reason is good: making an improvement pay a bookkeeping toll to land is how a ledger starts charging for exactly the work it exists to encourage. --lower was built so that closing a surplus costs nothing. ⛔ Nothing here re-decides that; the ruling stands until someone with the authority moves it.

What the sweep adds is the number #6376 did not have: how much silence accumulates between sweeps in practice. 118 raw errors is what four months of ordinary repair work banked, in layers where nothing else reads the count. While a gap is open, a regression smaller than it lands green — the driver-mongodb case the header narrates (43 recorded / 10 measured swallowed a whole signature reversion) is that failure already realised once.

The part that is new evidence, not a restatement

All four are TEST_DEBT; none of the 13 DEBT entries had moved. That is not noise, and it has a mechanical cause worth stating before anyone designs the remedy:

  • a DEBT number is what the package's own typecheck would report if it had one, so ordinary repair work there is visible to whoever does it;
  • a TEST_DEBT number is produced only by the temp project this gate generates. Nothing a contributor runs reports it. So repairing a test-layer type error moves a number nobody sees until the next re-measure.

The surplus is therefore concentrated in exactly the layer whose measurement is invisible outside this gate — which is a reason a self-tightening rule might be worth more here than a uniform one, and equally a reason the cheaper remedy might be visibility rather than redness. ⛔ Not decided here.

⚠️ Related but distinct, and not to be merged into this: #12511 is about that same invisibility as a cost to a dev round. This card is about what the invisibility does to the ledger's tightness. They share a mechanism and not a remedy.

Options, stated without a recommendation between the last two

  1. Do nothing.--lower exists, the can be lowered line is printed on every run, and a periodic sweep like [finding] @objectstack/runtime's TEST_DEBT ledger records 227 against a measured 226 — one new test-layer type error can land unreported #12723's closes the gap. Cost: the sweep has to be scheduled by something. Nothing schedules it today — that is the argument [finding] @objectstack/runtime's TEST_DEBT ledger records 227 against a measured 226 — one new test-layer type error can land unreported #12723's triage adopted, and it applies to this remedy too.
  2. Make a surplus red. Strongest, and ⛔ it directly contradicts the [finding][devx] check:type-check-debt 的 ledger 余量会让新写的 pin 变哑:mongodb 曾有 33 条余量吞掉一次真实回退,另有 5 条目前带 4–19 余量 #6376 ruling: an improvement then cannot land without also editing the ledger. Note this direction strengthens a gate, so it is not on the 门禁削弱 manual floor.
  3. Make a surplus red only above a threshold, so a -1 stays free and a -103 does not. Untested shape; the threshold would need an argument, not a guess.
  4. Auto-lower in CI rather than failing — the gate already knows how to write the number. Changes who authors a ledger number, which is a question the file's own header takes a position on.

Not established here

  • ⛔ Whether 118 is typical. It is one reading of one moment, at the end of an interval whose start nobody recorded. A second sweep some months out is what would make it a rate.
  • ⛔ Whether any of the four surpluses actually absorbed a regression. That would need a per-file cross-tab against each entry's prior composition, and this sweep measured per-entry totals only.

Re-check

pnpm check:type-check-debt --re-measure

Read the per-entry numbers, ⛔ never the OK line — the summary is green by design while an entry carries slack. As of PR #12798 the correct reading is surplus: none.

Refs

Metadata

Metadata

Assignees

No one assigned

    Type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions