Note

How to Judge an Enterprise AI Vendor's Claims

An accuracy number without a denominator is not a claim. The questions that separate a product from a demo, from someone who runs vendor selections.

·Published ·8 min read·#enterprise-ai#selection#controls

Method 4 of 5 see the reading order →

Reviewed by Tan Gravam

On this page

Judge an enterprise AI claim by its denominator, not its percentage — on whose data, against what ground truth, over which population, and what happens to the cases it gets wrong. A number without those four is not a measurement, and treating it as one is how a selection becomes a negotiation about slideware.

The textbook approach to evaluating a vendor is the demo, the reference call and the accuracy figure. All three still matter. All three are also, for a probabilistic product, much easier to stage than for deterministic software — and the usual selection process was designed for software that behaves the same way twice.

I have run and reviewed vendor selections in finance systems for eighteen years, and published a scorecard and an RFP structure for doing it. What follows is that discipline pointed at products whose central claim is about accuracy.

The four questions behind any accuracy number

A vendor says the model is 95% accurate. That figure is meaningless until you can answer:

On whose data? Theirs, a public benchmark, or a customer's real production mix. A number produced on curated data tells you the ceiling under ideal conditions, which is not the number you will operate at.

Against what ground truth? Somebody decided what "correct" meant. If that somebody was the vendor, on cases the vendor selected, the exercise measured agreement with themselves.

Over which population? The one that matters includes the hard cases. Ask explicitly whether difficult or ambiguous items were in the sample or excluded from it — the answer is frequently "excluded", and it is frequently not volunteered.

And what happens to the other 5%? This is the question that actually decides the economics. Five per cent of a large volume is a queue, and somebody works it. The cost of the product is the licence plus that queue, and vendors rarely price the second half because it is not theirs.

A vendor comfortable being shown their own error rate is a different proposition from one who steers you back to the script.

Make the demo yours

The single highest-yield change to a selection is refusing the prepared demo, or rather accepting it and then insisting on a second one you control.

Bring your own data. Unrehearsed, including the awkward cases you already know about — the multi-entity one, the one with the broken reference, the one everybody in your team mentions when this process comes up. Then watch the failures rather than the successes. Where does a wrong answer surface? Who sees it? How much work is correcting it? What does the system record about the correction, and can you get that record out?

This is the same instinct behind asking demo questions that cannot be prepared for. It just matters more here, because a deterministic system's demo is a fair sample of its behaviour and a probabilistic one's is not.

Weight before you watch

Decide what matters, and how much, before the first meeting.

That sounds procedural and it is the mechanism that makes a selection survive contact with a good presenter. Weighting after the demos means weighting in the presence of persuasion, and the criteria drift toward whatever the most impressive vendor happened to be strong at. Fix the weights while nobody is in the room, score each vendor on the same evidence, and let the arithmetic be uncomfortable if it wants to be. That is the whole argument of running a selection as a process rather than a series of impressions.

The four unglamorous questions

These separate a product from a demonstration, and none of them is about the model.

What happens when you are unavailable? Does our process stop, queue, or fall back to something? A dependency whose failure mode is "the work stops" needs to be priced as such.

What does the audit trail record? At what grain — a summary, or the individual action with its inputs? Can we export it, retain it on our schedule, and hand a single record to someone who has never met the system? That is what audit evidence has to survive.

Who is the action attributed to? If the product acts inside our systems, whose identity does it use, and what appears on the record afterwards? The answer is a design decision with consequences that cannot be fixed after go-live.

How do we learn that quality has drifted? Not that the service is down — that answers have quietly got worse. If the answer is "customers tell us", the product has no observability and you will be the monitoring.

What a good answer sounds like

Worth saying plainly, because scepticism is easy and calibration is the useful part.

A good vendor answer contains a number with its conditions attached, a named limitation, and an unprompted description of where the product does badly. "Around 90% on invoice formats similar to the samples you sent, materially worse on handwritten annotations, and here is what the queue looks like at your volume" is a strong answer. It is more useful than 98% with no context, and a vendor who talks that way is describing something they have run in production.

The reference call matters too, and the question that works is not "are you happy". It is: what surprised you after go-live, what does your team do now that they did not expect to, and what would you scope differently. People answer that honestly far more often than they answer a satisfaction question.

What I would decide

Fix the weights before the demos. Run the product on your own difficult data and watch the failures. Price the residue — the exception queue is part of the total cost and it is on your side of the line. Ask the four unglamorous questions early, because they are the ones whose answers cannot be improved later. And treat a vendor's willingness to state a limitation as evidence, because it is the closest thing to a controlled variable you will get.

None of that is AI-specific procurement wisdom. It is ordinary selection discipline, applied to a product category where the demo is unusually easy to stage and the interesting number is the one nobody puts on the slide.

See also how to find workflows worth redesigning with AI and where to start with enterprise AI.

Frequently asked questions

4

What does an AI vendor's accuracy percentage actually mean?

On its own, almost nothing, because the number has no denominator you can see. Ask three things and the figure becomes interpretable or it collapses: measured on whose data — theirs, a public benchmark, or a customer's real production mix; measured against what ground truth, and who decided what correct meant; and measured on which population, including whether the hard cases were in the sample or filtered out of it. A vendor who can answer all three is describing a measurement. One who cannot is quoting a marketing figure.

What should you ask an AI vendor in a demo?

Ask to run your own data, unrehearsed, including the cases you know are difficult — and watch what happens to the ones it gets wrong rather than the ones it gets right. The demo is designed to show the clean path; the product is what happens to the residue. Ask who sees a failure, how they know it happened, what it costs to correct, and what the system records about the correction. A vendor comfortable being shown their own error rate is a materially different proposition from one who steers you back to the script.

How do you compare AI vendors fairly?

Fix your weighting before the demos, not after. The mechanism that makes a selection defensible is deciding what matters and how much while nobody is in the room persuading you, then scoring each vendor against those weights on the same evidence. Without that, a polished presentation quietly re-weights the criteria — which is the same failure as any software selection, and the reason a scorecard is worth building before the first meeting rather than after the third.

What questions expose a thin enterprise AI product?

Four, and they are all about the unglamorous half. What happens when your system is unavailable — does the process stop, queue, or fall back? What does the audit trail record, at what grain, and can a customer export it? Who is the action attributed to when the product acts on our behalf? And how do we find out that quality has drifted, before a person notices something wrong? A product built for demonstrations has thin answers to all four, and the answers do not improve after signature.