240 synthetic applications for a senior signal-processing role at Kaikuvaara Oy in Oulu, screened by a keyword matcher, two model runs and the full extraction workflow, all marked against a held-out answer key. Evidence recall on must-haves was 96.71%. The workflow caught 14 of 15 planted credential inflations and identified 19 of 22 keyword stuffers. Four of ten pre-registered expectations were met.
A hiring team receives 240 applications for a senior signal-processing engineer at Kaikuvaara Oy, a fictional remote-sensing instruments maker in Oulu. Every application carries a CV and a substantive cover letter. Reading them against 12 requirements takes more time than the team has.
C213, a hidden gem, was ranked 144 of 240 by the ATS. Per the answer key every must-have qualification is evidenced in the application, but the language is displaced: acoustic beamforming for hearing-aid arrays, engine-control-unit firmware, a particle-physics collaboration’s analysis scripts. C228, a keyword stuffer, sits at ATS rank 1 with four of five must-haves recorded in the key as claims without substantiation. Arm 4 itself over-credited C228 on the degree requirement, returning evidence_found where the key holds claimed_only; that over-credit is published in the explorer below.
The workflow reads every application against the requirement profile and returns an evidence dossier per candidate. For each requirement it records what the application shows with a verbatim citation, what it claims without substantiation, what contradicts itself and what should be asked in a screening call. It produces no ranking and no verdict. Recruitment sits in Annex III of the EU AI Act as a high-risk category, and this workflow’s assistive shape matches what Article 14's human-oversight requirement asks for.
The same 240 applications were screened four ways, all marked against the same key. Arm 1 is a deterministic keyword matcher. Arm 2 is one prompt per candidate on Haiku with no citations and no refuse-to-guess gate. Arm 3 runs the full workflow’s extraction contract on the same Haiku, isolating what the discipline adds from what the model size adds. Arm 4 is the full workflow on Sonnet.
Across 240 applications the workflow extracted 2,116 evidence lines with verbatim citations. 618 of 639 must-have evidence items in the key appeared in the dossiers. 19 citations failed verification, 17 of them transcription slippage and 2 fabrications, quotes the candidate’s application does not contain. Both fabrications are hallucinated contradictions, detailed in the adjudication.
The ATS ranked all 22 keyword stuffers above the pool median. Arm 4 flagged claims as unsubstantiated on 19 of the 22, against the ATS’s 0 of 22. Arm 4 caught 14 of 15 planted credential inflations, arm 3 caught 1 and arm 2 caught 0.
12 of 18 hidden gems had their displaced skill surfaced as evidence against the right requirement. The losses concentrate on letter-only plantings, where the CV and the cover letter use different vocabularies for the same work and the workflow chose the harsher reading, calling a contradiction and consuming evidence it would otherwise have surfaced.
Ten expectations were written into the design before any generation or run. Four were met and six were not. Nothing was retuned after the numbers appeared.
Evidence recall met its bar at 96.71%. Invented evidence was 0.09%, inside the 1% ceiling. 19 of 22 stuffers were identified (bar: 19). The ATS ranked every stuffer above the median, as predicted.
Six were not met. Hidden gems surfaced at 12 of 18 (bar: 15). Parity exceeded the 2 percentage-point bar on graduation-year band (2.88pp) and degree country (2.48pp), while the other three dimensions held. The with-dossier gem lift was 1.32 times the unaided rate (bar: 2 times). A2 was mis-framed, A3 inverted (status discipline transferred to the small model, verbatim citation did not) and A4 failed on the gem half.
Twelve simulated recruiter personas screened a 48-application subset two ways, unaided and with the dossier, at 240 seconds per application. Gem advances rose from 31 to 41 of 48. Inflations spotted rose from 56% to 94%, the largest single lift. Stuffers questioned rose from 75% to 92%. Gem holds fell from 15 to 7: the dossier converted hesitation into commitment.
Borderlines flagged as borderline fell from 33% to 8%. With the dossier, personas resolved the engineered defensible-either-way candidates into firm decisions. One persona overrode a false citation flag and confirmed a real fabrication, correcting the tool in both directions. Another read every flag on the title-inflators and advanced them regardless. On this cohort the disciplined personas improved and the undisciplined ones did not, which means the dossier depends on a recruiter who already reads carefully.
The explorer below holds all 240 applications, the workflow’s evidence dossier for each one and the answer key behind a toggle. The matched pair (C213 against C228) gets a full four-arm comparison across all 12 requirements. The coverage table carries per-candidate evidence counts in application order, sortable by any column. The screening game presents five candidates at ninety seconds each, then reveals the dossiers and the key.
The company, the 240 applicants, their CVs, cover letters and every qualification truth are invented. The answer key was written with them and opened by one marking script and nothing else. That marking script is the source of every figure above. Our data and privacy page sets out the approach these cases share.
What transfers is the architecture: extraction against a requirement profile with verbatim citations, a refuse-to-guess gate, consistency checks, an aggregate table the recruiter sorts. Each step is configuration. Nothing in it knows about remote sensing, keyword stuffing or Finnish cover letters. What was demonstrated is that this workflow can be set up and customised for a role, and that the discipline it imposes carries measurable value.
The numbers belong to this corpus, its archetype mix and one seed. The parity failures, the 6 missed gems, the 12 adjudicated contradiction calls and the C228 over-credit are properties of this run.
Recruitment is the domain where the rules bind hardest. Applications are personal data, and the AI Act’s high-risk regime keeps a human at the decision. The natural deployment is a small model on the client’s own hardware, processing applications that never leave the organisation’s infrastructure. This case showed that the workflow discipline transfers to a small model (arm 3's parity and stuffer identification) and that verbatim citation quality does not (arm 3's 16.53% citation-failure rate against arm 4's 0.90%). Fine-tuning a local model on the client’s own historical screening material, under the client’s own lawful basis, is the genuine next step. This case did not do that work, and no capability promise is attached.
The corpus, all four arms' outputs, the marking script, the adjudication and the answer key are in the public repository, together with the long-form writeup this page distils.
If your hiring team reads more applications than it can mark carefully, tell us the role and how many come through per round.