Seven rule tests raised 64 flags and 48 of them landed on a planted issue. A statistical detector raised 200 and landed on none, and the slow price drift at the centre of the design was missed by every layer.
The company is invented. Saarnitukku Oy is a Vantaa wholesaler of fasteners, seals and maintenance supplies, turning over about €90 million a year, its purchase ledger posted by four people in accounts payable and running to some 50,000 rows.
Nobody reads a ledger that size. Flags are cheap, in that a few rule tests or a statistical model will raise thousands without much effort, and what is scarce is the time of someone competent to decide whether a flag means a duplicate submission, a coding error, a habit of stopping under an approval tier, or nothing at all. The subject throughout is controls and process quality, not accusation.
Three layers screened the same ledger. Seven deterministic rule tests drawn from sourced audit practice ran first, then an Isolation Forest over eight engineered features per transaction, then three AI analyst runs working from one pinned briefing, blinded to each other and to the answer key. Before any of them saw the ledger, 87 anomalies were planted in it, alongside 18 transactions built to be benign and certain to attract a flag.
The rule tests raised 64 flags, 48 of which land on a planted issue. That is the 75% first-pass precision the design set out to hit, and recall on every rule-shaped class is 100%. Most of the remaining 16 are twelve months of a landlord’s rent at a contractual €15,000.
An Isolation Forest raised 200 flags and not one of them touches the answer key. Before any of it ran, the design had committed to freezing and reporting a miss rather than retuning until it passed, so the result is published as it happened, with a diagnosis and no patch. The vendor engineered to drift is measured against its own full-year average, which the drift lifts along with everything else, so a rise of about 19% across the year reads as 2.1 standard deviations against 6.5 from a clean baseline.
Between them the three analyst runs called 37, 35 and 38 of the 39 verdict units correctly, and all 78 detector-side stand-downs were right, each backed by recomputed evidence.
Nobody found the story the case was designed around. From July one vendor folds freight into the price of the goods, the memos saying so in Finnish, and a further rise then compounds inside the bundled line. The rule tests were never built for slow drift, the detector’s best-ranked affected invoice sat 818th of 50,000, and the analysts spent their free look elsewhere. Without the key, the drift would not have been noted as a miss.
The explorer below steps through the catch matrix and all 39 verdict units, run by run, with what each layer flagged, what the key says and the evidence behind each verdict.
The ledger is synthetic. We wrote the generator and hold an answer key recording every plant, kept out of the rule tests, the detector’s features and every file the analysts were allowed to open, and the numbers above come from marking against it. The same held-out-key method applies across our cases, as the data and privacy page describes.
What transfers is the marking. A real ledger has no key, so the substitute is a reconciled control total, a known incident or a manual sample, checked before any layer is trusted unattended. The numbers do not transfer. This extract carries no purchase orders or goods receipts, by design, and a real engagement adds that match. The one method here that would have surfaced the drift is a standing review of each vendor against its own history, which is a report and not a model.
That standing review was then built and run. For each vendor and each month it takes a volume-weighted mean unit price, compares it against a baseline made from that vendor’s own prior months, and accumulates the standardised deviations until they pass a fixed threshold. The three-invoice monthly floor, the three-month minimum baseline and the threshold itself were all set in the design document before the second ledger existed, along with six expectations of what it should do, and none was changed afterwards.
It ran on a second synthetic year of the same company’s ledger, 50,000 fresh rows and 200 vendors carrying three planted price drifts and two benign price rises, and flagged all five. Four of them sit in the top five of the 195 scoreable vendors, the August repricing first, the fast drift second, the bundled-freight drift third and the slow 1%-a-month drift fifth, while the January indexation is further down at 34. The slow drift crossed in July, a month earlier than the pre-registered window. The two benign rises look exactly like the three real ones on the trace; they separate only on the memos behind them.
Two of the six expectations failed, and both are published as they ran. The flag count floods. Forty-two unplanted vendors crossed the threshold against an expected five, because a baseline of only three months is volatile enough to make an ordinary next month read as a large deviation; a 20,000-trial simulation under the same fixed parameters puts that at 18.7% per vendor and predicts 36 to 37 of them. The ranking survives what the count does not, so the instruction is to read the top ten.
The second miss matters more. Run backwards over the first year’s ledger, the report cannot see Teräskontio at all, the vendor whose drift prompted it. It bills twice a month, the method needs three invoices in a month before that month counts towards anything, and it clears that floor in none of the twelve, so it has no rank and no trajectory anywhere in the output. Drift is caught wherever there is volume, a vendor that bills too rarely to build a baseline is invisible, and covering those vendors would need a separate build.
The code is published in the order it ran: the seeded generator, a validator enforcing thirteen coherence rules, the seven rule tests, the detector, the clustering of every flag into the 39 units, the three analyst runs, and a script that marks all three against the key mechanically. The adjudicated scorecard was written by hand from that worksheet, with every miss left in.
The code, the ledger, the three analyst runs and the long-form writeup are in the public repository, together with the drift report of section v, its generator, its second ledger and its marked results.
If your purchase ledger is screened by rules nobody has marked against a known answer, tell us how many flags a month they raise and how many people read them.