173 of 176 review questions, 98.3%, were answered on a hundred contracts no model had read, by pattern-matching code an overseer wrote, at no cost per contract.
A due-diligence or legal-operations team asks the same dozen questions of nearly every contract in the stack, such as who signed, when it starts, whose courts govern it, whether it can be assigned, whether liability is capped, and whether competition is restricted. The work is expensive and most of it repeats. The question is how much of it repeats closely enough for ordinary code to take over, and what has to stay with a model that reads.
The contracts here are real, which makes this the one study in our portfolio not working from documents we wrote ourselves.
A model read 50 contracts, answering all twelve questions of each against one schema. An overseer watched the answers it accepted and, where a category kept turning on the same handful of phrasings, wrote a deterministic matcher for it. A matcher is pattern-matching code with no model behind it at run time. Each matcher had to reproduce the model’s own earlier answers before it went live, and every contract read afterwards doubled as a shadow check alongside it.
Twelve matchers were written. One failed that activation gate and never answered a question. Three went live and were demoted later, each on a single disagreement. One was a contract naming Texas, Singapore and Belgium law where the matcher knew only US states, one a preamble scan that overran and took two defined terms for signatories, and one a schedule the key itself records under two titles. The eight survivors answered 176 categories across 100 held-out contracts, 173 of them correctly on presence, 98.3%, at zero marginal cost. A matcher answers only the instances it trusts and refers the rest back to the model, so that figure is measured over an easier subset by construction.
Twelve predictions about which categories would automate were written down before anything ran. Four were right. None of the three we called fully automatable got there. Two were demoted, and the third answered 21% of the holdout. Four we expected to stay with the model grew matchers that answer a narrow formulaic minority and refer the rest back, and two of those four did badly. Expiration date, where the answer is a computed date instead of a quoted phrase, was wrong on two of the four holdout answers it gave; audit rights activated on ten induction patterns and then matched nothing across the hundred held-out contracts.
The explorer below holds the run: the twelve categories, the matchers as the overseer wrote them, the stream that activated and demoted them, and the marking against the key.
The documents are real. They are 150 commercial contracts filed with the SEC, selected from CUAD, a corpus of 510 that law students and attorneys annotated under The Atticus Project.1 The answer key is theirs rather than ours. It stayed out of the pipeline throughout and was read once, after the stream had closed and the twelve predictions were on paper. The data and privacy page covers how the case studies handle data generally.
Marking against someone else’s key cuts both ways. All 80 disagreements between the pipeline and the key were read by hand, and about 15 of them turned out to be the key in error or arguably so. The examples include a misspelled party name, a date in the wrong format for CUAD’s own convention, and an agreement date contradicting the span the key quotes as its own evidence. That is roughly a fifth of the contested rows and about 2% of everything marked. The annotators did expert legal work, valued by the dataset’s authors at over two million dollars, and we are marking against it for free.
The method transfers to another stack of contracts, from learning what the model accepted through gating on replay and demoting on disagreement to marking against something the pipeline cannot see. The counts do not. Which questions matter and what share of them automates depends on the contracts in front of you, and reading a sample of those is where an engagement starts.
The steps stand in the order they ran: corpus preparation, key derivation, a validator that checks the key stayed out of reach, the model path, the overseer, the router that replays the stream, and one evaluation script, which is the only file permitted to open the answer key. All twelve matchers are kept as the overseer wrote them, the frozen and demoted ones included, each with a header naming the contracts it learnt from.
The code, the derived data and the long-form writeup are in the public repository. CUAD’s licence permits us to publish the derived answer-key files,2 and the 150 contracts this case read are there with them. The full CUAD archive is not ours to redistribute, so the repository records where to download it and the checksum to verify it against.
Nothing on this page is legal advice.
Sources
If the same questions get asked of every contract that crosses your desk, tell us which questions they are and roughly how many contracts a month.