Every case study we publish runs the same sequence, and most of it is fixed before any code is written.
A design document records what will be measured and what counts as a pass before any code runs. Deterministic Python generates the data and a held-out answer key. A language model writes the prose from briefs that already contain the planted signals, adding none and removing none. Every figure on the case pages comes out of a marking script that reads the key. Ten cases are published. Nine run on data we engineered, and the tenth on public contracts with an expert key published by others.
The design documents are in the repository, one per case. They record the data structure, the planted signals, the measurement criteria and the pass conditions. In the anomaly case, the drift report’s threshold, minimum invoice count and six expectations were fixed before the second ledger existed.
Deterministic Python fixes every signal in the data: the archetype, the timing, the severity. A language model writes the prose from a brief that already contains those signals, adding none and removing none. The churn case’s 3,777 support tickets were made this way. The companies are invented, and each case page says so.
Where a case tests a belief of ours, the belief goes on paper first. The contracts case recorded twelve predictions about which review questions would automate. Four were right, and the eight misses run in both directions.
The key is written when the data is made, then kept out of the feature build, the models and every file an agent may open, with one script per project permitted to read it. In the contracts case the key was not ours but the expert annotation published with the corpus, read once after the run closed.
The figures on the case pages come out of that script. 173 of 176 contract questions were answered correctly. 48 of 64 ledger flags landed on a planted issue.
A result that comes out badly is published as it ran, diagnosed but not retuned. The statistical detector in the anomaly case raised 200 flags on a 50,000-row purchase ledger and touched none of the 87 planted issues. The drift report built afterwards to cover that gap flags three planted drifts in the top five of 195 vendors, but cannot see the vendor whose drift prompted it, because it bills too rarely to clear the report’s minimum.
The code, the data, the keys and the decision logs go out together.
Making the quietly unhappy accounts in the churn case invisible to the usage metadata took four iterations. The model kept finding a route back in through unresolved-ticket counts, then through ticket volume itself, since those accounts filed 6.65 tickets in a trailing six months against 1.04 for a healthy one. Each attempt meant regenerating the book and re-running everything downstream.
Four of the contracts case’s twelve predictions were right, and the eight misses run in both directions. The tender case prints four of its six pre-registrations as not met, and the rest carry their negatives in the body.
All of it is public at github.com/evammun/hapax-case-studies, one folder per case. Each holds the pipeline in run order, the generated data, the answer key, the design document and the decision log kept while the work was under way, the reversed decisions included. Where a case reads real documents, the repository records the source, the licence and a checksum to verify it against.
Two things are absent. The superseded churn datasets from the earlier attempts at that archetype were kept but not published, and the drift report from the anomaly case follows shortly.
A real engagement has no answer key, so the first question is what stands in for one. It might be a reconciled control total, a closed incident, or a sample somebody has already checked by hand. Tell us what exists and we will tell you what can be measured against it.