Hapax · Case study · Contract review

The clauses that retire the model

A language model read fifty real commercial contracts and answered twelve review questions about each. An overseer watched its accepted answers, wrote ordinary code for the questions that looked formulaic, and retired the model from them – under a gate that could refuse a matcher and shadow checks that could kill one in production. Everything was then marked against annotations we didn’t write.

Real data · marked against an external key This case runs on real contracts. The corpus is CUAD – 510 real contracts filed with the SEC, annotated by law students and attorneys under The Atticus Project (CC BY 4.0). The external expert answer key was held out of the pipeline throughout and read exactly once, after the stream closed, by the one script that scores against it – then never edited. We pre-registered a prediction for every category before running anything; the data marked us 4 of 12, wrong in both directions. Of the 80 pipeline-versus-key disagreements, read by hand rather than sampled, about 15 turned out to be the key’s own errors or arguably so. Nothing on this page is legal advice.

I · The two ends of contract language

Some clauses are liturgy. Some are negotiation.

The whole case turns on this distinction. A governing-law clause is written nearly the same way in thousands of contracts; a negotiated term is written once. Code can learn the first kind. Only judgment covers the second.


II · The stream

Fifty contracts, eight batches, a shrinking job

Each batch the model was asked only the questions no matcher had earned. Gold markers are matcher births; wine markers are demotions – a matcher killed in flight by a shadow check. The narrowing was won, and then partly lost, and both movements are in the data.


III · The centrepiece

We predicted each category’s fate before any run. The data marked us: 4 of 12.

A automatable   P partial – code answers the formulaic instances, the model keeps the rest   R model-resident. The misses run in both directions: every category we promised would fully automate died by demotion, and four categories we wrote off grew narrow, accurate matchers. Wrong predictions stay in print.

Try this – filter to the eight we got wrong, then click into one. Each row opens the real contract language it was learned from, the matcher’s actual code, and a live disagreement against the key on a named contract.


IV · The examination

One hundred contracts the model never read

Active matchers ran alone over the holdout, at zero marginal cost. Where a matcher chose to answer it was near-perfect on presence – but read the bars honestly: matchers answer only the instances they trust, so their accuracy is over an easier subset than the model’s. Audit Rights activated and then matched nothing at all; Expiration Date got two of its four answers wrong – a silently wrong deterministic date, exactly the failure mode our prediction table warned about.


V · Trust, then verify

Three demotions, three different failure modes


VI · The adjudication

When we disagreed with the attorneys, we read every case by hand

The key was never edited. As-marked scores stand; the adjudication is published beside them. Against everything marked, the arguable key errors amount to roughly 2% – the ordinary margin of expert annotation.


VII · Cost


VIII · Beyond this dataset

What this would look like on your contracts

An overseer watches which review questions keep getting answered by the same handful of phrasings and writes ordinary code to take them over; the model keeps everything genuinely negotiated, permanently. What doesn’t transfer is the count – which categories automate, and how many, depends on your documents. We pre-registered a split and let the data mark eight of our twelve guesses wrong.

Read the long-form case study · Code and dataset · Contact Hapax