A simulated cohort of forty sat a two-day programme with their true learning planted in each of them beforehand. The honest instrument recovered the cohort effect to within 0.0100 of the planted 0.1338, and caught six of the seven people whose confidence rose while their competence did not. The feedback sheet the industry runs instead correlated with actual learning at r = 0.101.
The evidence normally offered for corporate training is an attendance list and a satisfaction score. Neither measures what anyone can do afterwards, so a company buying training cannot tell from the paperwork whether the room learned anything.
We have trained no commercial cohort, so the case about how much a real cohort learned is not available to write yet. The instrument is. This case builds the assessment that gets attached to every training engagement we run, which is pre and post testing on held-out task sets, and tests it on forty simulated learners whose true learning was planted in advance and written into an answer key that only the marking script opens.
The claim it supports is a narrow one. It is not that our training works, but that when we train your people the effect will be measured with an instrument whose accuracy, and whose failure modes, are published here.
Kuusiharju Oy is an invented technical trade company of about 200 people. Forty of its staff in finance, operations, sales and service administration sit the two-day AI-Augmented Workflows programme from our own training page. Each learner carries a deterministic profile made of a baseline ability, a planted true shift from the training, a confidence trajectory, an assessment-noise level and a propensity to volunteer. The confidence trajectory is drawn from its own distribution and never as a function of the shift, which is the separation the whole case rests on, so it was given a pre-declared band. The cohort-wide correlation between planted shift and confidence gain had to land between 0.20 and 0.65; at this seed it is 0.322.
Six archetypes make up the forty. Eight strong improvers, twelve modest ones, seven confident non-learners, seven non-responders, three ceiling cases who start with nowhere to go, and three fast forgetters. The confident non-learners are deliberately the largest problem group, because they are the case a feedback form cannot distinguish from someone who genuinely improved. Their mean planted shift is about 0.03 against a mean confidence gain of 1.63 points on a seven-point scale.
Three instruments then run on the same cohort. The honest one, which is the product, is two ten-task forms drawn from the programme’s own skill claims and counterbalanced so that no learner answers the same questions twice, with twenty sitting A first and twenty sitting B. Both forms are published in full, because a client should be able to read exactly what would be measured, and a real engagement generates fresh forms by the same published method. The tasks are ordinary work. A9 reads “You now generate the weekly ops report with AI assistance. Describe the specific check you would build into that process – not ‘I’d read it over’ but a named, repeatable step – and say exactly where in the weekly workflow it sits.”
The feedback sheet is the industry standard, run honestly as a comparator. It is post-only self-reported satisfaction and self-assessed improvement, on the same forty people. The naive evaluation is how training is usually measured when anyone tries, and it carries three documented bad practices by construction. It takes volunteers only, gives them the same form before and after, and blends self-report into the score at a weight of 0.5. The practice effect from sitting a form twice is planted as a flat bonus to effective ability.
All three read the same corpus of 240 response files, one per learner per form per occasion, with occasions at pre, post and eight weeks after. Deterministic Python fixed every learner’s ability, shift, noise and rubric hits; prose agents wrote the answers to those briefs and could not add or remove a rubric item. Scoring then reads the answer text and never the briefs, which keeps generation and marking apart.
Six expectations were written into the design before any data existed. Four were met and two were not, and nothing was retuned after the numbers appeared. Scores are fractions of the rubric points available on a form, so a shift of 0.13 means the cohort gained about a seventh of the available marks.
The instrument measured a mean gain of 0.1437 against a planted 0.1338, a gap of 0.0100 inside a tolerance of 0.0124, which is a tenth of the planted shifts’ own standard deviation. That margin is 0.0024 wide, which matters once the same figure is recomputed under a perfect rater below.
Six of the seven confident non-learners were flagged, by a single frozen statistic. Each learner’s confidence gain and measured gain are z-scored across the cohort’s own forty values, and anyone whose confidence z exceeds their performance z by more than one standard deviation is flagged. The threshold was set before the run and never adjusted. The rule returns eight names, the six intended ones plus two genuine modest improvers who combined small real gains with large confidence gains. It missed one, L30, whose confidence gain of 1.19 points is the smallest in that archetype.
The feedback sheet correlates with actual learning at r = 0.101, and satisfaction does slightly worse at 0.056, both computed against the planted shift for all forty learners. On this cohort the standard evidence for training carries almost no information about whether anyone learned anything, and the expectation that it would come in under 0.30 was written before the numbers existed.
Two expectations failed and both are frozen as they ran. Form equivalence is the first. Mean pre-score was 0.4400 on Form A and 0.3688 on Form B, a difference of 0.0713 against a tolerance of ±0.05, so Form B is the harder paper by about seven percentage points of rubric coverage. The forms’ structural item difficulties were matched at design time and passed their check; what diverged is the realised difficulty once the same rubric items were written out as prose and scored from text. It is not a scoring artefact either, since the gap widens to 0.0900 under the perfect-scoring pass. Equivalent forms are hard to write, and a real engagement pilots both on a non-client group before either is relied on. Counterbalancing limits the damage. With twenty learners starting on each form, the cohort mean is protected even when an individual’s pre/post pair is not.
The naive arm is the second. It overstated the cohort effect by 41.3%, measuring 0.1890 where the cohort truth is 0.1338, and the pre-registration had asked for at least 50%. The prediction was right about the direction of distortion but wrong about its size.
The decomposition underneath it is the more useful half, and it is messier than the design anticipated. Switching off one flaw at a time and holding the other two at the naive arm’s actual settings, selection is worth 24.1% of the overstatement and the repeated form 42.8%, but self-report weighting is worth minus 57.9%. Blending confidence into the score pulls the estimate down and not up, because the repeated-form performance delta is already inflated, since the flat practice bonus pushes several rubric-band thresholds at once, while the confidence gain rescaled to the same units is comparatively modest. The interaction term is 90.9% of the total, which means the three flaws do not add. An evaluation carrying two of them cannot have its bias estimated by adding up what each one is worth alone.
Every number above was produced by scoring response text with a rubric detector whose item-level agreement with the briefs’ intended scores is 94.23%, or 9,046 of 9,600 item checks, against a design target of 97%. That is the closest analogue this case has to a human marker’s reliability. Someone reading free text against a rubric disagrees with a co-marker, or with themselves on a second pass, some fraction of the time, and this scorer’s fraction is measured rather than assumed.
So every verdict was computed a second time on the briefs’ own intended scores, which is what a perfectly reliable rater would produce. Two of the six flip, in opposite directions. The recovery of the cohort effect goes from met to not met, because the perfect-scoring gap between planted and measured effect is 0.0212 against the same ±0.0124 tolerance. The naive arm’s overstatement goes from not met to met, because under perfect scoring it reaches 57.5% and clears the 50% bar the scored pass missed at 41.3%. Under the noisy scorer the instrument looks more accurate than it is, and the naive arm looks less distorted than it is.
The confident-non-learner rule changes composition without changing its verdict. Under perfect scoring it catches all seven, with two false positives either way, one of them a different name.
The feedback sheet is identical across both passes. A form that never reads an answer cannot disagree with one, so its correlation with planted learning is immune to rater error by default.
Neither pass is the true answer. The scored pass is what a real marking programme produces, the perfect pass is what it is trying to approximate, and the distance between them is what this case has to report. For a real engagement that settles three things. Forms are piloted before they are relied on. Marking uses two raters, or one calibrated rubric with its agreement rate measured and published. And a conclusion about a cohort is stated with the reliability of the scoring layer attached to it, since two of the six verdicts here turned on exactly that.
The design named this failure in advance. A pre/post instrument measures the day the course ends, and three of the forty learners were built to lose most of what they gained within two months.
At post they are invisible. The three fast forgetters average a measured gain of 0.2667 against 0.2150 for the twenty genuine improvers, whose own standard deviation is 0.1204. They sit inside a single standard deviation, and if anything they look slightly better than the people who actually kept the skill. The eight-week occasion exists in the corpus and no named instrument reads it. Against it, the fast forgetters retain 21.9% of their post-training gain and the genuine improvers retain 87.8%.
One learner’s answers show the decay in plain text. At post, asked for an instruction that returns the same three-line summary every time, Markus Hirvonen specified the field order, the exact output template and a rule for missing data, then explained that the format and the missing-data rule together are what make the output independent of the email’s wording. Eight weeks later, on the same task, he wrote “Give it a fixed structure – date, item and quantity, then the change, in that order. Keeps the three lines the same each time.” The answer keeps the topic and loses the mechanism that earned the marks.
This is the published argument for the follow-up a month later that our training page promises. A pre/post pair measures whether the room learned. A third measurement some weeks out measures whether the workplace kept it, and on this cohort the two questions have different answers for three of the forty people.
The explorer below holds the cohort and the marking. All forty learners can be filtered by role and archetype and opened individually, with their planted shift, the instrument’s estimate and their own answers side by side. The gallery of the seven confident non-learners is the part to read first, because it shows what a satisfaction score would have reported about each of them.
The third view is the robustness flip, which puts the six verdicts as scored against the six under a perfect rater and marks the two rows that change.
The cohort is simulated. We wrote the company, the forty learners, their planted shifts and the 240 answers, and the answer key was written with them and opened by one marking script and nothing else. That script produces every figure above. The method is common to all ten cases and described on the data and privacy page.
That has one consequence. What this page tests is the ruler, not the training. Nothing here is evidence that our training works, or that any training works.
The numbers do not transfer either. They belong to this cohort, its archetype mix and one seed. The seven confident non-learners and the three forgetters are there because we put them there, and a real cohort has whatever mix it has. What transfers is the instrument and the discipline around it, both published in full.
What a real cohort learns is unknown before it is measured, which is the reason for measuring it. The first real engagement’s measured outcomes will be published with the client’s permission or described without identification, whatever they show.
The files run in the order the work did: the cohort generator and its seeded profiles, both task forms with their rubric, the 240 responses, the scoring layer, and one marking script that is the only code permitted to open the answer key.
The scorer’s calibration is published with it. It began at 74.80% item-level agreement and reached 94.23% across eight rounds of reading disagreements by hand and broadening patterns only where the detector had demonstrably missed text that matched its brief. One round was a net regression. Widening a pattern to catch a visible false negative made it fire on a neighbouring item as well, collapsing two items into one trigger and costing about 14 percentage points of agreement on two tasks, and it was measured, reverted and left documented in the code. A stratified sample of 60 of the 554 residual disagreements is published with the response text, so the classification is checkable.
The forms, the cohort, all 240 responses, the marking script, the scorer calibration and the answer key are in the public repository, together with the long-form writeup this page distils.
If your training evidence is a feedback form, tell us what the programme was meant to change and we will tell you how we would measure whether it did.