How to get the baseline a pilot needs by measuring tasks in aggregate rather than timing named individuals, so staff do not start performing for the metric.
We wrote last time about the four things a pilot needs before it starts, and one of them was a real, measured baseline. That word, measured, is doing more work than it looks like. Measure a process carelessly and you end up measuring the people doing it instead, and the moment staff feel watched rather than assessed, the numbers you get back stop being the truth and start being a performance for the metric.
This piece is about the line between the two, and how to stay on the right side of it without losing the evidence the pilot needs.
Measurement targets the task. Surveillance targets the person. The distinction sounds abstract until you see how easily the same data collection slides from one to the other without anyone deciding to cross a line. Timing how long a task type takes, averaged across a team, is measurement. Timing how long each named individual takes on each task, visible to a manager as a running list, is surveillance, even if nobody called it that and even if the intention was entirely benign.
Aggregate cycle time for a task type, across the whole team, tells you whether the process got faster. It does not need a name attached to each data point to do that. A sample of outputs, reviewed for quality against a rubric agreed in advance, tells you whether accuracy held up. It does not need every single output attributed to whoever produced it. Total volume processed across the pilot period tells you about capacity. It does not need an individual leaderboard to be useful, and a leaderboard is usually where the actual damage starts.
We have written before about psychological safety as one of the ten readiness dimensions, and about how the fastest adopters are often the most afraid of being seen to struggle. Named, individual-level metrics make that fear concrete: people learn exactly who is watching which number, and they start managing the number instead of the work. Someone slows down to avoid a visible error. Someone hides a mistake rather than reports it. None of this shows up as dishonesty, it shows up as a suspiciously clean result, which is its own kind of warning sign. Aggregate, task-level measurement removes the personal incentive to distort the data, because there is no individual number to protect.
You can measure the change without measuring the person who made it.
Aggregate data is the right default, and it has a real cost: it cannot tell you which specific person needs more support, only that the team-level number moved or did not. Sometimes an organisation genuinely needs individual-level information, to help someone who is struggling rather than to judge them. The honest answer is that this is a different conversation entirely, a consented, developmental one between a manager and a named person, not something folded quietly into a pilot's evidence-gathering. Keep the two apart, or the pilot's aggregate data earns none of the trust it was designed to protect.
Measure the task, not the person: aggregate, anonymised, and announced in advance; anything narrower than that needs its own explicit, consented reason, agreed separately from the pilot.
No external statistic cited; this article extends the organisation's own pilot-design and readiness-dimension work rather than presenting third-party evidence.
A pilot produces evidence rather than anecdote only if four things are fixed in writing beforehand, namely a success metric, a measured baseline, an end date and a named honest reporter.
AI Readiness Assessment · 3 minGuide · 13 September 2026What low scores on workforce impact, internal politics, psychological safety and future-readiness look like, why they stall adoption, and why they move only when leadership changes its behaviour.
AI Readiness Assessment · 4 minIf this is the question on your desk, a thirty-minute call tells you whether the service fits, or that you do not need us yet.