This is writing that comes out of the work. We publish when something is finished, so the gaps are long.
Large language models produce answers without a reliable signal of how much trust those answers deserve. We are writing a practical Bayesian framework for calibrating uncertainty in LLM systems: when to act on an output, when to route it for human review, and how to measure whether the system is as well calibrated as it claims. It is written for practitioners.
Acerbi & Martin · in progress
Ten worked cases are now published, each with the run behind it open to inspection. They have a section of their own.
When something is finished it will appear here. There is nothing to subscribe to, but if you would like to hear when a note is published, write to enquiries@hapax.fi and we will email you when one is. We use the address for that and nothing else.