Note

How to Measure ROI From an AI Workflow

Licences times adoption times a guess measures deployment, not outcome. The ten metrics that move when work is redesigned, and how to baseline them credibly.

·Published ·10 min read·#enterprise-ai#process#governance

Method 2 of 2 see the reading order →

Reviewed by Tan Gravam

On this page

The number in most AI business cases is built the same way: seats, multiplied by an assumed hours-saved-per-person-per-week, multiplied by a loaded hourly rate, discounted by an adoption percentage. Versions of that arithmetic get presented in steering meetings with a straight face, and the striking thing is that only one input in it — the seat count — was ever observed. Everything else was asserted at the point the slide was built.

What that calculation measures is deployment. How many people have the tool, and how many opened it. It says nothing about whether any piece of work got faster, cheaper, better or more reliable, and it cannot be checked afterwards, which is why these numbers are usually retired quietly rather than revisited.

It survives because it is available early. It can be produced before anything has changed, in the week the budget submission is due, which is exactly when a business case is needed and exactly when no evidence exists. That is a real constraint, not a moral failing. But it means the number arrives with no way of ever being right or wrong, and sceptical executives can smell that from across the table.

The alternative is not more sophisticated modelling. It is measuring the work.

Measure the work, not the tool

A redesigned workflow changes observable things about how cases move: how long they take, how many hands touch them, how often they come back. Those are the numbers a finance director will accept, because they are the same numbers used to justify every process change since long before anyone deployed a model.

Ten of them matter. None is a model metric, and none requires a data science function to produce.

The four that are about time

Cycle time. Wall-clock elapsed from the trigger event to the case being done and usable by whoever was waiting. Instrument it from two timestamps that already exist — record created and record closed — rather than from a stopwatch or a survey. The trap is where you start the clock: teams measure from "when we picked it up", which conveniently excludes the queue in front of the work, and the queue is usually where most of the elapsed time sits.

Touch time. The minutes a human actually spends working a case. It is the hardest of the ten to obtain honestly; sample it by observation across a normal week rather than asking people to self-report, or derive it from gaps between system events. The trap is its seductiveness. Touch time is the metric AI improves most easily and the one the business feels least — a step that got ten times faster inside a process that still takes eleven days is a rounding error, which is the whole difference between automating a task and redesigning the work.

Queue and wait time. Cycle time minus touch time, broken down by which queue. Instrument it from status-change timestamps; if the statuses do not exist, that absence is your first finding and probably your first fix. The trap is ownership — waiting is always attributed to the other team, so nobody owns it. Report it per handoff, with a name against each, and it stops being weather.

Decision latency. The gap between the moment a decision could have been made, because all the evidence was available, and the moment it was made. Approximate it as time from a case reaching "ready for review" to approval. The trap is starting the clock when the approver opened the item rather than when the evidence was complete; measured that way, every approval SLA in the enterprise looks excellent while the business waits days.

Flow, and what a case actually costs

Throughput. Cases completed per period, in total and per person. Count closures in the system across enough weeks to see the period-end shape, not a fortnight chosen because it was quiet. The trap is mix: throughput rises on its own when the work gets easier. Always report it next to the case mix — type, size band, complexity — or you will eventually claim a redesign for a soft quarter and be found out by someone who has the volume report.

Cost per case. Fully loaded cost of running the workflow divided by cases completed: people time, system cost, licences, inference, and the exception team everyone forgets to include. The trap is a cost the manual version did not carry as a system cost: inference and per-case run charges that scale with volume rather than sitting in a fixed licence. Leave it out and your business case gets worse as the workflow succeeds, which is a discovery you want to make now rather than in year two.

An hour saved is not money until somebody decides what to do with the hour.

That sentence is where most AI business cases quietly break. Saved minutes spread thinly across forty people are capacity, not cash. They become cash only when a role is not backfilled, a contract is not renewed, or a volume increase is absorbed without hiring. Say which of those you are claiming, and say who has agreed to it. If none applies, call it capacity and state what the capacity will be used for — that is a defensible claim, and the number survives contact with a finance business partner.

Quality, and the metrics that are easiest to fake

Rework. Cases touched more than once because something was wrong, missing or misrouted. Instrument it by counting reopens, returns, second approvals and versions. The trap is that rework frequently lives outside the system altogether, as an email asking for a corrected file, so a redesign that formalises it can look as though it created rework. Establish where rework currently hides before you claim to have reduced it.

Exception rate. The share of cases leaving the happy path, and the share of total effort those cases consume. Define the exception types before counting anything, and count effort by type rather than incidence alone. The trap is direction: a redesign that handles the clean cases straight through will push the exception rate up while genuinely improving the workflow. Report exception effort as a share of total effort alongside the rate, or the honest improvement reads as a regression.

Human interventions per case. How many times a person had to touch a case for it to complete. Instrument it by counting actions attributed to a human identity in the audit log. This is the honest test of autonomy and it is unforgiving: a mandatory review step is an intervention, whatever the deck calls it. Watching this number is also how you notice that "automation" has quietly relocated the work to a checking queue rather than removed it.

Error and reversal rate. Outputs that turned out to be wrong, measured by what had to be undone — reversals, corrections, credit notes, restatements, complaints. Instrument it from the correction mechanisms you already have, because finance and audit have been tracking those for years. The trap is the worst one on this list: this metric can improve by looking less hard. A redesign changes detection as well as performance, so if you find more errors after go-live, say so and report detection separately. A programme that reports a falling error rate and a falling detection rate has measured nothing.

Baseline first, or you have an anecdote

None of these ten means anything without a before. And the before has to be taken before the change, which sounds obvious and is skipped on most programmes I have seen, because by the time anyone wants the number the workflow has already moved.

Take the baseline over at least one full cycle including a period end — a workflow measured only in a quiet fortnight is not the workflow. Prefer history that already exists in your systems over a study conducted for the occasion, because the moment a team knows their process is being examined, the queue gets cleaned and the behaviour changes. A baseline taken after the kick-off meeting flatters the old world and understates whatever you eventually build.

Record more than the numbers. The case definition. The start and end events. The mix. Who is in scope and who is not. What is excluded and why. Six months later, that record is what lets somebody reconstruct the comparison instead of taking your word for it.

Freeze the basis

The other half of credibility is that before and after have to be the same measurement. This is the point the cash-forecasting teardown makes about accuracy — a forecast that was edited after cut-off cannot be graded, because you would be marking a version that had already seen the answer.

The same logic applies to a business case. Freeze the definitions on the day you baseline and put them in writing: what counts as a case, what counts as closed, which case types are in scope, how exceptions are categorised. If the definition of "closed" moves during the project, the before-and-after comparison is fiction, and it will be exposed by the first person who asks how a case was counted.

Then keep a log of everything that changed and was not the AI. A team reorganisation, a volume drop, a new system release, a policy change, three experienced people leaving. Not to excuse the result — to make it attributable. A comparison period or a parallel region running the old way is better still, and where the workflow runs in several entities that is often available at no cost beyond the discipline of leaving one alone for a quarter.

What a defensible case looks like on the day it is challenged

Assume the reader's job is to find the hole in it. Six things make that difficult in the right way:

  • One named workflow, with the case defined in a sentence, rather than "AI in finance operations".
  • An observed before number, with the dates, the method and the mix stated, taken before the project was announced wherever the history allowed it.
  • The metric the business feels first — cycle time or cost per case — leading, with touch time supporting rather than starring.
  • The full cost, including the run cost that scales with volume, the exception team, and the control and audit work that the redesign creates.
  • What did not change, stated explicitly. A case that claims improvement everywhere is not credible; one that says "rework was flat and here is why" usually is.
  • A named person accountable for the number after go-live, and a date on which it is taken again.

Add the way the claim could be wrong — mix shift, seasonality, a parallel change — before anyone else does. Naming your own confounders is the fastest route to being believed, and it costs a paragraph.

The number is a design constraint, not a slide

The useful thing about deciding these measures early is that they change what gets built. If cycle time is the claim, then the design has to attack waiting rather than typing. If cost per case is the claim, inference cost belongs in the architecture conversation from the first week. If interventions per case is the claim, the review step needs a genuine reason to exist rather than being the default comfort.

That is also what separates work that becomes the way things run from work that stays a capability demonstration: a pilot proves something is possible, and an operating system for the work has to keep proving it. Choose the workflow with an assessment that can say no, state the outcome as something measurable before you start, and take the baseline while the old process is still running. Everything after that is arithmetic anyone can check — which is the only kind worth putting in front of someone who intends to challenge it.


See also how to find workflows worth redesigning with AI and enterprise AI transformation.

Frequently asked questions

3

Why do AI ROI calculations overstate the benefit?

Because most of them are built from seats, an assumed number of hours saved per person per week, a loaded hourly rate and an adoption percentage. Only the seat count is observed; everything else is asserted, and the result measures how widely a tool was deployed rather than whether any work got better. The calculation also converts saved minutes straight into money, which only holds if someone actually removes the cost. And because nothing in it was ever measured, nothing in it can be checked afterwards — which is why these numbers are quietly dropped rather than revisited.

Which metrics show that an AI workflow actually improved?

Ten, grouped in three families. Time: cycle time from trigger to done, touch time, queue and wait time, and decision latency. Flow and cost: throughput and fully loaded cost per case. Quality: rework, exception rate, human interventions per case, and the error or reversal rate. Cycle time and cost per case are what the business feels; touch time is the one that improves most easily and matters least on its own. Human interventions per case is the honest test of autonomy, because a review step is an intervention even when the deck calls it straight-through.

When should the baseline for an AI project be taken?

Before the project is announced, ideally from timestamps that already exist in your systems rather than from a study run for the occasion. The moment a team knows their process is being examined, the queue gets cleaned and the behaviour changes, so a baseline taken after the kick-off flatters the old world and understates the gain. Cover at least one full cycle including a period end, record the case mix alongside the numbers, and write down the definitions the same day. A baseline nobody can reconstruct six months later is an anecdote with decimal places.