Hapax · Case study · Contract review
A language model read fifty real commercial contracts and answered twelve review questions about each. An overseer watched its accepted answers, wrote ordinary code for the questions that looked formulaic, and retired the model from them – under a gate that could refuse a matcher and shadow checks that could kill one in production. Everything was then marked against annotations we didn’t write.
I · The two ends of contract language
The whole case turns on this distinction. A governing-law clause is written nearly the same way in thousands of contracts; a negotiated term is written once. Code can learn the first kind. Only judgment covers the second.
II · The stream
Each batch the model was asked only the questions no matcher had earned. Gold markers are matcher births; wine markers are demotions – a matcher killed in flight by a shadow check. The narrowing was won, and then partly lost, and both movements are in the data.
III · The centrepiece
A automatable P partial – code answers the formulaic instances, the model keeps the rest R model-resident. The misses run in both directions: every category we promised would fully automate died by demotion, and four categories we wrote off grew narrow, accurate matchers. Wrong predictions stay in print.
Try this – filter to the eight we got wrong, then click into one. Each row opens the real contract language it was learned from, the matcher’s actual code, and a live disagreement against the key on a named contract.
IV · The examination
Active matchers ran alone over the holdout, at zero marginal cost. Where a matcher chose to answer it was near-perfect on presence – but read the bars honestly: matchers answer only the instances they trust, so their accuracy is over an easier subset than the model’s. Audit Rights activated and then matched nothing at all; Expiration Date got two of its four answers wrong – a silently wrong deterministic date, exactly the failure mode our prediction table warned about.
V · Trust, then verify
VI · The adjudication
The key was never edited. As-marked scores stand; the adjudication is published beside them. Against everything marked, the arguable key errors amount to roughly 2% – the ordinary margin of expert annotation.
VII · Cost
VIII · Beyond this dataset
An overseer watches which review questions keep getting answered by the same handful of phrasings and writes ordinary code to take them over; the model keeps everything genuinely negotiated, permanently. What doesn’t transfer is the count – which categories automate, and how many, depends on your documents. We pre-registered a split and let the data mark eight of our twelve guesses wrong.
Read the long-form case study · Code and dataset · Contact Hapax