M.A.I. Consulting
GuidePractitioner15 September 20264 min read

Reading a vendor benchmark without being taken in

Why a vendor's headline benchmark is a hypothesis rather than evidence, and the four questions to ask about who chose the task and the success measure before it influences a purchase.

Every AI vendor now arrives with a number: forty per cent faster, ninety per cent accurate, six hours saved a week per user. The number is usually real, in the narrow sense that somebody genuinely measured it. It is also, almost without exception, the best number the vendor's own team could produce, from a task they chose, run under conditions they controlled.

This piece is not an argument against vendor numbers. It is a short list of what a benchmark actually tells you, what it quietly does not, and the questions worth asking before a number is allowed to move a procurement decision.

The benchmark is real. The generalisation is the vendor's guess.

A demo that completes a task forty per cent faster genuinely did complete it forty per cent faster, on that task, that day. The part nobody tests on stage is what happens once the tool meets your actual work: your document formats, your edge cases, your staff who have never used it before. Independent research on this exact pattern found professionals using an AI tool were twenty-five per cent faster and completed twelve per cent more tasks, and also nineteen percentage points less likely to reach a correct answer once the task fell outside the tool's core competence. The speed gain was real. So was the accuracy cliff, and the vendor's own demo would never have shown you where it starts.

Who chose the task, and who decided what counted as success

A benchmark run by the vendor, on a task the vendor selected, judged against criteria the vendor wrote, is not evidence about your work. It is evidence about the vendor's ability to choose a favourable task. None of this makes the vendor dishonest. It makes the benchmark a marketing artefact rather than an evaluation, and the two need different levels of trust.

Four questions worth asking before a number moves a decision

A benchmark tells you what a tool can do on its best day, with its best task. It does not tell you what happens on your worst day, with your actual work.

The trade-off in insisting on your own pilot

Refusing to act on a vendor number until you have run your own pilot, on your own task, with a real baseline, is the right instinct, and it has a genuine cost: weeks that a signed contract and a vendor's slide deck would not have required. The honest answer is that this cost buys something the slide deck cannot: evidence about what actually happens on your work, with your staff, rather than a demo built to succeed. A pilot done properly, the way we described for measuring change without surveillance, is slower than trusting the number. It is also the only way to find out whether the number was ever going to be yours.

The practical takeaway

Treat every vendor benchmark as a hypothesis worth testing on your own task, not a result worth acting on directly; ask who chose the task and who decided what counted as success before the test began.

Sources

Series · Choosing where to start · part 4 of 5
Keep reading
05 ยท Custom Agents & Tools

Can the tool we already pay for do this task for us, every time?

If this is the question on your desk, a thirty-minute call tells you whether the service fits, or that you do not need us yet.