In most commercial AI work, model confidence is taken at face value. Quantifying that confidence, and designing systems that act on it, is a research discipline before it is a deployment pattern. This page describes the deployment shape and collects what the ten case studies showed about it.
A model produces an answer. In most deployments, how far that answer should be trusted is either estimated from a benchmark or assumed from the model’s general reputation. Neither says anything about a specific output on a specific input.
Uncertainty quantification asks the question explicitly. Confidence can be measured by probing a model repeatedly on the same input and observing the spread of its answers, by calibrating its stated probabilities against observed outcomes, or by testing it on held-out data where the right answers exist. Each approach produces a confidence estimate for a given output, not for the model in general, and that estimate can be acted on.
One of the founders has spent his research career on measuring how far a model’s output should be trusted and on the calibration methods that make the measurement practical. The founders are also writing a paper on calibrating uncertainty in language-model systems, in progress.
A query arrives and a small model handles it. If the model is confident, the answer is sent. If it is not, the query moves to a larger model. What policy reserves for a person goes there regardless of confidence.
Setting the threshold is measurement work, done on examples from the client’s own processes. A threshold set too high sends routine work to the expensive model. Set too low, it lets errors through.
Four of the ten cases tested different parts of this design.
The inbox case built a customer-email pipeline with three deterministic gates and an escalation lane for safety and personal-data emails. 7 of 7 emails in those categories were escalated, with zero misses in either direction. The straight-through rate came in at 81.3%, above its pre-registered 55–70% band, because 6 needs-info emails were misclassified as routine and auto-sent. None of those 6 invented an answer, and each deferred to the policy handbook, so the overshoot was safe.
In the talent-screening case, the workflow read 240 applications against a twelve-requirement profile and returned evidence with verbatim citations. Where an application did not contain evidence for a requirement, the system recorded it as not stated and proposed a question for the screening call. It never inferred a qualification the application did not contain. 19 citations failed verification, 17 from transcription slippage and 2 fabricated; both fabrications were surfaced and published.
The training-outcomes case tested the pre/post assessment instrument attached to every training engagement. Its scorer was 94.23% reliable across 9,600 graded items. When every verdict was recomputed under a perfect scorer, two of the six pre-registered expectations flipped, in opposite directions. A 6% error rate in the scoring changed the conclusion.
In the anomaly-detection case, a statistical detector raised 200 flags on a 50,000-row purchase ledger and none landed on the 87 planted issues. The slow price drift at the centre of the design was found by nobody. Both results are published as they ran.
A frontier model on every call is the obvious route and the expensive one. Most calls in most workflows are routine, and a small model handles them. What varies is where the boundary sits, and getting it right requires measurement on the client’s own tasks.
A confidence-aware system routes the majority of calls to the small model and reserves the large one for the minority that need it. Monitoring the confidence distribution over time also surfaces drift. If the share of uncertain calls rises, something in the data or the task has changed.
Hapax earns nothing when a client chooses one model over another, so the routing advice carries no revenue interest.
If you are running a model and want to know how much of its output to trust, tell us what it handles and what the volume looks like.