Field note
How to measure whether AI training worked
A happy-sheet score at the end of a session tells you almost nothing about Monday.
The only real measure of AI training effectiveness is what people do afterward, not how the session felt. Track three things in order: whether staff use the tool on a genuine task within a week without being told to, whether that use continues after 30 days, and whether it visibly cuts time or improves output on the task it targeted.
Day 1: reaction
Did people follow along and see the point
Week 1: first use
Did anyone use it on a real task unprompted
Day 30: usage holds
Habit survives once the novelty wears off
Day 30: the result
Time or quality gained on that task
Why a happy-sheet score doesn't tell you anything
Most training gets judged on the wrong level. A post-session survey asking whether people enjoyed it or found it clear measures reaction, and reaction correlates weakly with whether anyone changes what they do at their desk on Monday. It is easy to run and easy to report, which is exactly why it is overused: it produces a number without producing evidence.
If a session scores well and three months later nobody on the team has touched the tool on a real task, the training did not work, whatever the survey said. Reaction is worth collecting because a session people actively disliked is unlikely to change anything, but treat it as a screen for failure, not proof of success.
What to check in the first week
Before you train anyone, name the one task the session is meant to change: first-draft replies, meeting summaries, a specific report. That single task is what you measure against, not general AI usage.
A week after training, check how many people used the tool on that exact task without being reminded. Not logged in, not asked a question in the session, actually applied it to real work unprompted. That is the first honest signal, and it is usually lower than the room's energy on the day suggested it would be.
What to check a month later, because usage fades fast
The harder and more important check comes at day 30. Novelty carries a lot of week-one usage on its own, and it drops off once the task gets busy again and the old habit is still faster to fall back on. Whether usage survives past that point is the real test of whether training changed behaviour or just produced a good afternoon.
Microsoft's 2026 Work Trend Index found that organisational factors such as culture, manager support and how the work itself is structured explain more than twice the variance in reported AI results that an individual's own skill level does. In practice that means a manager who checks in on the task and expects the new habit to hold does more for day-30 usage than another hour of training content. If usage is fading, look at what happens after the session before you blame the session itself.
How to tie it to a business result
Do not try to prove company-wide productivity gains from one training session. That number is too noisy to attribute to a single intervention and chasing it wastes the credibility of the exercise. Measure the one task you named at the start: time a handful of real examples before training and the same task after 30 days of use, or track a simpler proxy like turnaround time or error rate on that task specifically.
A result on one well-chosen task, reported honestly, is more convincing to the rest of the business than a vague claim about overall AI adoption, and it gives you a template you can repeat on the next task.
Organisational factors like culture, manager support and how work is structured explain far more of whether AI training pays off than an individual's own skill level, which is why measuring the session alone misses most of the story.
The takeaway
Before you run another AI training session, pick the one task it should change and check three things in order: who used it unprompted in week one, who is still using it at day 30, and whether that specific task is measurably faster or better now.
