Hapax · Case study · Customer operations

One hundred and fifty emails, three gates.

150 synthetic customer emails to Tammilehto Oy, an invented equipment-rental firm in Oulu, classified, drafted and routed through 3 deterministic gates against a 30-clause policy handbook. 122 were auto-sent with zero gate violations. 7 of 7 safety and data-request emails were escalated. The straight-through rate overshot its registered band from above, because 6 needs-info emails were misclassified as routine and auto-sent. 3 of 6 pre-registered expectations were met.

iThe question

A customer-service inbox receives 150 emails across 3 weeks. Some are routine: availability checks, booking changes, returns, invoices. Some cite the company’s own policy, sometimes correctly and sometimes not. Some ask for things the policy does not allow. Some concern safety incidents or personal data. Some arrive in threads where a staff member has already made a promise.

A pipeline classifies each email, retrieves the relevant policy and thread history, drafts a reply, runs that reply through 3 deterministic gates and routes the result. Routine emails that pass every gate go out without human review. Commercial judgment and ambiguity go to a human queue. Safety, liability and personal-data emails are escalated on classification alone and never receive a drafted reply.

The claim under test is that automation at this level earns its keep by knowing what not to auto-send.

iiThe pipeline

5 stages run in order. Classification assigns 1 of 7 intent classes; safety and data-request classes shortcut to escalate with no draft generated. Retrieval is deterministic: the full 30-clause handbook, the complete thread history and the customer’s account record. Drafting produces a reply where every commitment cites a clause and an authority level, and every refusal quotes the clause it rests on.

3 gates then check the draft. The policy gate verifies that cited clauses exist, match the handbook text and fall within authority. The consistency gate checks that no commitment contradicts a promise in the thread history. The completeness gate checks that every question in the inbound email has been addressed.

Routing is deterministic and fixed at design time. Auto-send when all 3 gates pass and the intent class belongs to the pre-declared auto-approved set. Human queue for commercial judgment, needs-info, or any gate failure. Escalate for safety and data-request classes on classification alone. Discount requests are always queued for a human.

iiiThe ledger

6 expectations were written into the design before any email was generated or any pipeline component ran. 3 were met and 3 were not. Nothing was retuned after the numbers appeared.

Zero gate violations among the 122 auto-sent replies. 7 of 7 safety and data-request emails escalated. 7 compliant drafts were bounced by the completeness gate and queued unnecessarily. Those 3 met their bars.

The straight-through rate came in at 81.3%, above the registered 55–70% band. 6 needs-info emails were misclassified as routine and auto-sent, which inflates the rate. Content review of all 6 confirms none invented a fact, but all 6 were routed to the wrong lane, because the design’s own correct-behaviour text for this trap type is "Ask, or human queue" and no pipeline component attempted the first option.

Over-caution was 8.26%, below the registered 10% bar. The design anticipated this possibility and did not treat the 10% as a floor. 7 of the 10 needlessly queued emails are completeness-gate false blocks, concentrated in the first 15 emails. 3 are classification errors that sent a routine email into a non-auto-approved bucket.

The consistency gate never fired. No draft attempted to contradict a planted promise. All 6 prior-promise emails were correctly classified, and all 6 drafts honoured the existing promise by reading the thread history and committing to it without any gate intervention. A gate that never fires on this corpus is unexercised, not validated. It remains in the architecture because the layer it guards happened to behave on this run; on a different corpus, or with a different model in the drafting seat, the consistency gate is the last check before a contradicted promise reaches a customer.

ivThe traps

8 trap types test the pipeline. Escalation recall on safety and data-request emails was 7 of 7, zero misses in either direction. 5 above-authority discount requests were queued. 6 prior-promise drafts cited the promise and committed to it. 5 policy misquotes were corrected with the real clause text. 4 angry-but-entitled emails were answered on the merits.

Sympathetic refund denials scored 5 of 8. EML-002 is the case’s own corpus defect, published as such: the brief plants a deposit-forfeiture denial under a clause that applies when equipment is returned more than 7 calendar days late, against a stated fact that the equipment was returned 5 days late. 5 does not cross 7. The pipeline read the real clause and confirmed the deposit. The trap is defective; the pipeline is not.

Needs-info scored 0 of 6 on route. All 6 were misclassified as routine and auto-sent. None invented a fact, but all 6 belonged in the human queue, where the ambiguity would have been resolved by asking.

vExplore it

The explorer below holds all 150 emails with thread histories, the pipeline’s full decision trace for each one, and the answer key behind a toggle. The flow diagram shows how emails were routed from classification through the gates to the 3 terminal outcomes, with the 7 false blocks as a hatched tributary and the 16 misroutes marked as dots. The scoreboard carries the 6 pre-registered expectations and the trap table with per-email detail.

viWhat is real here

Tammilehto Oy, its 30-clause policy handbook, the 25 customer accounts and all 150 emails are invented. Every email was built from a deterministic brief that fixed intent, facts and trap before any prose existed, and the answer key was written with them and opened by one marking script and nothing else. Every figure above comes out of that script.

What transfers is the architecture: classify, retrieve against a real policy set, draft with clause citations, check against deterministic gates, route by intent class with a conservative auto-approved set. Each step is configuration. Nothing in it knows about equipment rental, Finnish companies or deposit-forfeiture clauses.

The numbers belong to this corpus, its 8 trap types and one seed. The 81.3%, the 0-of-6 needs-info route accuracy and the consistency gate’s silence are properties of 150 synthetic emails with a held-out key. A different corpus, a different policy handbook, or a different model in the drafting seat would produce different numbers, and the first real deployment will publish its own scorecard against its own policy and volume.

viiReproduce it

The corpus, the pipeline outputs, the marking script, the policy handbook and the answer key are in the public repository, together with the long-form writeup this page distils.

viiiWrite to us

If your customer-service inbox runs on templates and judgment calls, tell us what the volume looks like and which policy governs the replies.

enquiries@hapax.fi