The record
We caught our own instrument wearing a bad statistic. Here is the whole of it.
One Wednesday — 2026-07-15 — is accounted below, beat by beat. Nothing in this story is about your data or anyone else’s. It is about ours: what we published, what our own benches did to it, and why the discipline that survived is the one this site now runs on. Every sentence is dated and has a file behind it.
Where the bad number came from
Early that morning, a validation session graded its results against a floor statistic called the bootstrap-max — the largest value seen across resamples — and recommended it as the primary floor. The code that computed it lived in a working directory our build gate cannot see, so no gate ever voted on it. By mid-morning the statistic had reached the copy of three live pages. LENDING_receipt-validation §4.4 · FALSE_FLOOR_sweep §1 · 2026-07-15
The conviction — by our own instrument
The same day, another session turned habitat’s own core check on that floor: hold the corpus byte-identical, vary only the number of resamples, and watch. A lawful floor is a property of the data and must converge. The median held (−7%). The 99th percentile held (+5.6%). The bootstrap-max moved +40.6% and kept climbing — it measures sampling effort, not the data. And it was not harmless: at 36,694 resamples, one of our published verdicts inverted, on an unchanged corpus. Measured, not argued; the floor was disqualified the same afternoon. LOCALIZATION_recon §13 · RESULTS_L4_FLIP_POINT · 2026-07-15
The gate that caught its author
The session that ran the conviction then wrote a build gate to make the bad floor structurally impossible — and within the hour, the gate failed its own author’s first “lawful” replacement, which had kept the max around “for reference only.” The author fixed the code rather than the gate. That gate now runs against this site. FALSE_FLOOR_sweep §5.2 · scripts/floor_law_gate.py
The wrong ruling that did not come from carelessness
One careful session, reasoning correctly from a true premise, ruled that a real finding was an artifact of the bad floor. A later session with wider access recomputed it on three corpora and found the finding real — the refutation was sitting in the first session’s own results table. Not every defect enters through sloppiness. A careful chain of reasoning is still not a bench. DISCOVERY2_reconciliation §1 · 2026-07-15
The failure that recurred nine times
Across the day, the dominant failure was never a wrong number. It was a conclusion drawn from a check that could not have seen the thing: a text search that ruled on prose it could not match; a directory listing that ruled on files inside an archive it could not open; a verification that reported success when its command had never run. Nine instances or more, depending on how you count — and in nearly every case the finding was right and its claimed scope was not. The law that came out of it is now printed on this site: a scope claim must state its scope. WRECKAGE_2026-07-15 §4.3
What actually moved the floor: someone else
One session audited its own report five times, adversarially, and found five defects — every one surviving until an independent pass swept the whole artifact. Later, two work orders drafted blind were cross-vetted blind: five defects found in one, three in the other, then five more in round two. Different route, same result. Self-audit has a floor, and the only thing that moved it, all day, was an independent pass. That is why the checkers here are checked. LOCALIZATION_recon §14 · WRECKAGE_2026-07-15 §3
The checkers were AI, and they were not trusted either
The measurements this site now states as its baseline were produced by AI sessions run blind against each other: parallel drafts, blind cross-audits, then two independent executions whose authors could not read each other’s results. Every load-bearing number on this site is one both runs reached separately, byte for byte — including the failures. One run’s own probe was invalid; its report records it. DETECTION_PUBLISH_final · both runs · 2026-07-16
The contradiction we shipped, and the correction
Both runs independently found that four downloadable kit files still asserted a claim the pages had withdrawn the day before — one of them on a data axis that provably cannot vary. The kits were rebuilt, the integrity chain re-anchored, and the correction is stated where the claim used to be. both runs §P5 · kit re-flag 2026-07-16
The lock claim, and the day after
Every session that day carried the same guardrail — a second repository locked, zero writes — and every session truthfully reported it held. It had not held. The first session of the day had written a working kit inside the locked tree at 4 a.m.; every later session scanned only from its own start time, so no scan could see it. A day later, an accounting session ran the scan across the whole day and found it. The engine’s source was untouched — the violation was narrower than an alarm and larger than the claim — but the sentence “the guardrail held” was verified all day by checks that could not have seen the violation. Even the accounting was checked. It lost a clause. WRECKAGE_2026-07-15 §0 · 2026-07-16
What survived
Through all of it — nine benches, twenty-one adversarial batteries, two blind runs — not one published digest moved. Twelve of twelve reproduce from their pinned records, including a 1,258-step dated chain. What died was language: a floor, a rank, a register, a doctrine sentence. What survived was the arithmetic and the discipline — and the discipline is the part we almost failed. The method page states the baseline this left us; the receipts carry their blind spots, with exhibits.
The instrument was never the thing on trial that Wednesday. We were. The record is what acquitted it — and what did not entirely acquit us is printed above, dated, with the files named.
The documents named on this page are our internal working record — dated files, quoted exactly. They are not yet published; whether to publish them is a decision not yet made, and until it is, this page is the summary and the names are the pointers. Nothing here adjudicates any person, dataset, or institution. Detection, not adjudication — including of ourselves.