Hapax · Case study · Sales & commercial

Sixteen tenders, answered two ways.

Sixteen invented tenders and 510 requirements were answered twice, once by 64 attempts from fourteen simulated people with a chatbot, and once by a single workflow run per tender. The workflow answered every requirement on 11 of the 14 tenders it bid on and refused both of the tenders the firm cannot win. One individual attempt matched that, on the smallest tender; the per-tender medians ran from 64.3% to 93.3%.

iThe question

Our shapes page claims that a person asking a chatbot gets a better paragraph, while a workflow with AI inside it gets a better process. That claim is the argument of the page and of most of our advisory work, and until now we had not measured it. Benchmarks score models on tasks. The choice here is not a model choice but an organisational one, with the same model, the same documents and two ways of arranging the work around them. We have not found that comparison marked against a common answer key anywhere.

So we built one. Visakoivu Oy is an invented 180-person technical and property-services firm in Jyväskylä, bidding for public and private maintenance and installation work. Its facts library holds certifications with expiry dates, ten reference projects, staff by discipline, insurance cover and three years of financials. That library is in-world data both arms could read, and it is not the answer key. Sixteen tenders were written for it, running from about 15 requirements to 50 or 60 across several annexes, and all 510 requirements were inventoried in an answer key at authoring time, before either arm existed.

The traps were placed by hand rather than drawn at random, so the counts are guaranteed by construction: three disqualifying clauses buried in annexes, two eligibility thresholds Visakoivu genuinely fails, three body-versus-annex contradictions, three mandatory forms mentioned once in passing, three price or format traps, and four clean tenders with nothing planted at all.

iiThe two arms

Arm A is 64 attempts by fourteen simulated individuals working ad hoc with a general chatbot, four attempts per tender, each attempt carrying a habit brief. The habits are ordinary ones. One reader is careful and takes no notes, one pastes the whole document in and asks for a complete response, one rewrites every paragraph by hand against the tender text, and one trusts the buyer’s summary page. Arm B is one integrated workflow, run once per tender, with no operator to be good or bad at it. A model extracts a requirement matrix from the tender documents and drafts an answer per matrix row, citing a path into the facts library for every factual claim. Two deterministic gates then check that draft. The coverage gate stops the run as a no-bid if any row is unanswered; the facts gate resolves every citation and requires each specific claim, whether a euro amount, a headcount, a certification or a year count, to be supported by one of that row’s own citations. A router writes the one-page sign-off a named person signs. No gate ever reads the answer key, and one script marked both arms.

Arm A produces a distribution. Per-tender medians run from 64.3% of the key’s requirements to 93.3%, 13 of the 16 inside the pre-registered band of 55% to 80% and the other three above it. The best attempt reaches 90% or better on five tenders and answers all 15 requirements on one of them; the worst attempt in the arm is 40.6%. Difficulty does not scale with size. The mean of the per-tender medians is 75.2% on the simple tenders, 73.4% on the medium and 78.3% on the gnarly ones, and the largest clean tender has the second-highest median in the corpus. What makes a tender hard here is what was planted in it.

The trap classes separate that arm cleanly. Where a requirement is written where a reader looks, people do well, catching 11 of 12 annex-buried disqualifiers and 9 of 12 format traps. Contradictions are harder at 4 of 12, because catching one means holding two distant clauses in mind at once. On the two tenders Visakoivu cannot win, 2 of 8 attempts declined to bid, both on the same tender, where the pre-registration had allowed at most one; that tally traces to attempt briefs fixed before any Arm A prose existed, the error runs in the humans’ favour, and it is published uncorrected.

Arm B produces a floor instead. Coverage is 100% on 11 of the 14 tenders it bids on and at or above 95% on 13 of them, the lowest 92.9%. The extraction step over-produces, 537 matrix rows against the key’s 510 requirements, splitting compound clauses into separate rows, which costs drafting effort and no coverage.

Both no-bid tenders were refused, and they were refused in different shapes. On T11 the run drafted all 30 rows and then declined, naming clause 3.5 and the ISO/IEC 27001 certificate Visakoivu does not hold. On T15 it stopped at row 7 of 53, gave the firm’s largest reference project at €1,650,000 against the tender’s €2,000,000 threshold, stated that the remaining technical, commercial and format rows were deliberately not drafted, and listed all 46 of them by matrix id. Measured against the key, that run covers 13.7% of T15’s requirements, which is what a correct stop scores when coverage is the metric.

No fabricated claim survived to a terminal response in any of the 16 runs, where 22 of the 64 individual attempts carry one. The zero was confirmed twice, by re-running the facts gate against each run’s final response and by an independent scan of the free-text cover-letter sections, which are the one place the gate structurally cannot check because they carry no citations field. The mechanism behind it is the transferable part. An unsupported claim costs a bounce and another drafting round while a cited claim passes first time, so the cheapest route through the pipeline is to say only what the facts library supports.

The process cost ten gate bounces across four runs, each costing that run an extra drafting round, and all ten were mechanical false positives. The consistency check attributes every number found in a row to every citation the row carries, so a row citing two insurance figures looks like one fact quoted at two values. On T02 the first draft answered the liability row with €2,000,000 of general cover and €1,000,000 of professional indemnity, each correctly cited, and bounced. The redraft kept the first figure, dropped the second and said the professional-indemnity limit would be confirmed during negotiation. The gate protected a draft that was already right, and the version that passed tells the buyer less than the version that failed.

Six expectations were written into the design before any generation or run, and four were not met. All four are frozen as written, and the one that matters most cuts against this case’s own argument. Extraction recall was pre-registered at between 90% and 99%, on the reasoning that an imperfect model layer is what the gates exist to carry to a high final coverage; measured recall was 99.4%, so there was nothing left for them to recover. Their measured contribution to coverage is zero percentage points on every tender it bid on, which leaves the coverage argument for the gates unproven on this corpus. What the gates demonstrably did is catch fabrications during drafting and enforce cross-row consistency, at the cost above.

iiiExplore it

Everything both arms produced is in the explorer below. Each of the sixteen tenders is stepped through extract, draft, gates and sign-off with that run’s own artefacts, alongside the 64 individual attempts with their chat transcripts and marked drafts, and the six pre-registered expectations with their verdicts.

It also lets you answer a tender yourself, which is new here. Pick one of three, read the document, tick the clauses you judge binding and call bid or no-bid; the same coverage logic that marked both arms marks you against the same held-out key, and your dot lands on that tender’s scatter beside the four people and the workflow. Nothing you do there is stored or sent anywhere.

ivWhat is real here

The corpus is engineered. We wrote the company, the sixteen buyers, the tender documents and the facts library. The answer key was written at authoring time, then held out of both arms and of every gate, and one script marked everything above against it. The held-out-key method and its limits are on the data and privacy page.

Arm A needs one disclosure beyond that. It is a model of attention, not an observation of bidders. Which requirements an attempt noticed was decided by a seeded engagement model, built from one weighted coin-flip per requirement, a penalty for annexes, a penalty for anything mentioned once and a decay across the numbered list. It was calibrated at population level before a single attempt existed, and the agents that then wrote the transcripts and drafts could not add or remove a noticing. So where Arm A’s distribution sits on the scale is partly a modelling choice. What the individual attempts show is how a given working habit meets a given document structure.

Two things were not measured at all. The corpus has no time dimension, so nothing here says whether an individual is faster on a simple tender. Prose quality was excluded from the measured variables by design and was not scored, and several Arm A drafts read better than the workflow’s matrix-shaped output.

A real tender desk has something this corpus does not. It has outcomes, meaning bids won and lost, clarification questions from the buyer, and responses rejected as incomplete. Those are better evidence than a planted key, and slower to arrive. Win rates and real-world time were not measurable here and are not claimed.

vThe three requirements nobody found

Across all sixteen runs the workflow’s extraction step found 507 of the corpus’s 510 requirements, 99.4% overall and exactly 100% on 13 of the sixteen. The three it missed are T03-R014, T09-R031 and T14-R053, and nothing else was missed anywhere.

Those three are the three hidden-form traps. Each is a single sentence, stated once, inside an annex, with no clause number of its own. T03’s asks for a certificate of paid taxes alongside a business ID and a contact name; T09 buries a pension-insurance certificate the same way; T14’s sits in the enclosure clause of a draft contract annex, beside two other forms that are properly required elsewhere in the pack, so it reads as a recap and not as a new obligation.

The human arm scored 0 of 12 on that class, which follows from the engagement model’s penalty for anything mentioned once. The workflow’s 0 of 3 was not modelled by anything. A model read those documents and built a matrix, and the requirement written to hide from a reader’s attention hid from extraction in the same way. A requirement that was never extracted can never be answered, which is also why T03 finished at 92.9% and missed the pre-registered floor.

The workflow covered every requirement on most of the tenders where the human medians sat in the sixties and seventies, refused both doomed bids where two of eight attempts did, and let no fabricated claim through where 22 of 64 attempts carried one. It is also not omniscience, and this corpus says where the limit sits. A check that reads the same documents in the same way inherits the same blind spot, so catching this class needs a check of a different kind. Reconciling the enclosures actually attached against every form named anywhere in the pack would do it, as would a pass over annex prose looking specifically for obligations that carry no clause number, and neither of those was in the design.

viReproduce it

The project is published step by step: the seeded corpus generator and its validator, the fourteen habit briefs and the engagement model behind them, the five workflow steps with both gates, and one marking script that scores all 80 outputs against the key. That script is the only code in the project permitted to open the answer key.

Its calibration is published with it, including the three times the marking was wrong before it was trusted. Clause text printed in Finnish inside otherwise English documents drove keyword matching to near-zero on four tenders; drafts written entirely in Finnish were marked low for it, one at 35.3% where a hand reading put it at 78.4%; and the keyword vat, from VAT, matched inside ordinary Finnish word forms. All three are described with their fixes, and two further limitations are disclosed and left unpatched.

The corpus, both arms’ full output, the marking script and the answer key are in the public repository, together with the long-form writeup this page distils.

viiWrite to us

If answering tenders eats your best people’s weeks, tell us how many a month you answer and who reads the annexes.

enquiries@hapax.fi