Each one runs a method end to end on data we engineered ourselves, with the answer key held out of the pipeline, so every figure on these pages has been marked against that key. One of the ten works from real contracts and an answer key someone else published. The answer key sets out that sequence in full, and data and privacy covers why none of the ten runs on client material.
Month-end at Jalavakoski Konepaja in Tampere arrives as raw text exports from five SAP transactions. The reports have fixed-width columns, German number formats and totals rows distinguishable from line items only by their indentation. One person has read them competently for two decades, and retires this year. An overseer watched a model parse the first few of each report type and wrote ordinary Python for the rest. 184 of 240 reports, 76.7%, finished on that code at an estimated two fifths of the all-model cost, with 13 wrong fields out of 34,470.
Pyökkipaja is a contract-furniture manufacturer in Lahti. Its large domestic suppliers invoice electronically, but the German fittings supplier, the Swedish fabric house, the Estonian steel shop and eighteen one-off vendors all send PDFs that get keyed by hand. A model parsed each PDF, and from those parses the overseer built deterministic code one supplier at a time. 117 of 198 invoices, 59.1%, now finish on parsers, with two wrong fields in 13,650. The nine parsers cost roughly what they saved in their first year; from year two the stream runs at an estimated seventh of the all-model cost.
The monthly board pack at Paju Consumer Products, a Nordic personal-care company, is usually correct and rarely carries the cause. The reason behind a 12% rise in material cost sits three drill-downs below the line item, in a table nobody opens before the meeting, so the obvious explanation goes into the commentary instead. Three blinded agent runs produced thirty findings on the ledger behind that pack, with no false positives at watch severity or above. Four planted stories were found by every run. The two sharpest pieces of evidence in the data were claimed by none of them.
Saarnitukku’s purchase ledger runs to 50,000 rows, posted by four people in accounts payable at a Vantaa wholesaler turning over about €90 million a year. Nobody reads it. Seven rule tests raised 64 flags, 48 of them on planted issues, matching the 75% first-pass precision the design set out to hit. An Isolation Forest raised 200 flags and touched none of the plants, a result frozen and reported because the design committed in advance to publishing it either way. The slow price drift the case was built around was found by no layer at all. A second pass built the standing drift report that result pointed to, and on a fresh year of ledger it put all three planted drifts in the top five of 195 vendors, but could not see the original drifting vendor, which bills too rarely to clear the report’s own minimum.
Sixteen tender documents sit on the desk at Visakoivu, a 180-person maintenance and installation firm in Jyväskylä, with 510 requirements between them. Someone has to answer every one without promising anything the firm cannot deliver. They were answered twice, by 64 attempts from fourteen simulated individuals with a chatbot and by one workflow run once per tender. The workflow answered every requirement on 11 of the 14 tenders it bid on, refused both tenders the firm cannot win, and let no fabricated claim through where 22 of the 64 attempts carried one. The individuals’ per-tender medians ran from 64.3% to 93.3%. The workflow’s three misses are all hidden forms mentioned once in passing inside an annex, the same class the individuals went 0 of 12 on.
Kataja Analytics sells a dashboards-and-reporting platform from Helsinki, 500 mid-market accounts. Bergvik Systems AB, a 31-seat Swedish account at €12,426 a year, never showed a usage dip. What changed was the tone of four support tickets over thirteen months, each closing with some version of "no need to treat this as urgent." A usage model reaches 0.867 AUC on that book; an agent reading the 3,777 tickets takes the combination to 0.890 and lifts recall in a top-100 review list from 59.4% to 66.2%. Thirty accounts that churn for reasons outside the product stay invisible to both layers.
Tammilehto Oy rents construction equipment out of Oulu, about 120 staff, with a 30-clause policy handbook governing deposits, damage charges, late returns, cancellation windows and discount authority. 150 customer emails went through a pipeline that classifies, retrieves the policy, drafts a reply with clause citations and runs 3 deterministic gates before routing. 122 were auto-sent with zero gate violations. 7 of 7 safety and data-request emails were escalated, zero misses in either direction. The consistency gate never fired; no draft attempted to contradict a planted promise, so that safeguard is untested on this corpus. 3 of 6 pre-registered expectations were met.
Every contract in the stack gets the same dozen questions, from who signed it to whose courts govern and whether liability is capped. A junior associate works through them one by one, and most of the answers look the same. These are the real documents in the portfolio, 150 contracts filed with the SEC, marked against the expert key published with them. A model read the first fifty, and from those readings the overseer built pattern matchers for the repeating categories. 173 of 176 review questions, 98.3%, were answered by that code on the hundred held-out contracts. Twelve predictions about which categories would automate went on paper before the run and four were right; none of the three called fully automatable got there, and four expected to stay with the model grew narrow matchers instead.
The evidence offered for corporate training is normally an attendance list and a satisfaction score. Neither says whether anyone in the room can do anything differently. At Kuusiharju, a 200-person technical trade company, forty simulated learners sat a two-day programme with their true learning planted in advance, and three instruments measured it. The pre/post assessment recovered the cohort effect to within 0.0100 of the planted 0.1338 and caught six of the seven people whose confidence rose while their competence did not. The industry’s feedback sheet correlated with actual learning at r = 0.101. Two of the six pre-registered expectations failed and are published as they ran.
A hiring team at Kaikuvaara Oy in Oulu has 240 applications for one signal-processing role, more than it can read carefully against the 12 requirements. The ATS put the keyword stuffer at rank 1 and the hidden gem at 144. A workflow read all 240 and returned evidence dossiers with verbatim citations but no ranking or verdict, catching 14 of 15 planted credential inflations and identifying 19 of 22 keyword stuffers at 96.71% evidence recall. Four of ten pre-registered expectations were met.
The code, the datasets and the long-form writeup behind each case are published at github.com/evammun/hapax-case-studies, one folder per case.