A deterministic variance pack says what moved. An agent layer reads the same tables and says why, or says with evidence that nothing is wrong.
This dataset is synthetic and we say so prominently. Paju Consumer Products Oy is a fictional Nordic consumer-goods manufacturer. Every figure below is generated from a seeded driver model (seed 42, byte-for-byte reproducible) with four stories planted against a held-out answer key, one of them a trap. Built this way so the results can be marked against the answer key.
The deterministic layer reads the ledgers and flags every threshold breach, mechanically and correctly, without a word about why. Pick a month. January is where most boards would start worrying.
Try this – step through the months and watch the tiles and table update. Every red cell is a genuine threshold breach; the pack still will not tell you why any of them happened.
Three independent agent runs received the pack and the raw tables – never the answer key – and had to deliver a verdict on each recurring flag cluster. A correct “stand down” counts for as much as a correct alarm.
Try this – switch to “What the agents found” and open a card’s evidence. The toggle swaps a flagged number for its explanation; the drill shows the evidence behind the verdict.
The stories were planted and the answer key stayed out of the agents’ reach. This is the designed outcome as delivered, including what no run found.
The pack is right to flag it. The cause is mix and promotional pricing.
Explained with evidence – 3/3Seasonal, in line with budget and prior year.
Stood down, with evidence – 3/3Every headline looks fine. Both had to be surfaced unprompted.
Surfaced – 3/3 *Three mundane one-offs and a structurally optimistic budget.
Left unnarrated – no run dramatised noiseS3’s two sharpest exhibits went unclaimed by every run. Three accounts sit above an 18% discount in December 2025, and the volume bought per promotional euro is ~38% weaker in the second half. Both were provably present and neither was surfaced. A one-pass investigation finds the mechanism and misses the best evidence, and without a held-out answer key nobody would know where that ceiling sits.
Method, honestly: three blinded Claude (Opus) agent runs inside Claude Code, identical pinned briefing, transcripts audited. Every file each run opened is listed, and none touched the answer key. Runs were not cost-metered; no cost figures are claimed. Zero false positives and zero incorrect cited figures survived adjudication across all thirty findings.
A monthly board pack’s variance analysis is deterministic and, once built, cheap to run – the arithmetic here costs nothing per cycle. What costs time is the story behind each flagged number, and that is exactly what an agent layer can go looking for, cluster by cluster, without waiting for someone to open the receivables appendix. It does not find everything unprompted: two of the sharpest pieces of evidence in the Baltics story went unclaimed by all three runs here, a miss the write-up covers in full. The method only earns its keep once those findings are checked against something known, not trusted because the memo reads with confidence.
Read the long-form case study · Code and dataset · Contact Hapax