How to Find Workflows Worth Redesigning With AI
Eight dimensions you can observe rather than debate — frequency, effort, input, judgement, exceptions, error cost, state, control — and the verdict they give.
Method 1 of 2 see the reading order →
Reviewed by Tan Gravam
On this page
The textbook way to find AI opportunities is a discovery workshop: gather the process owners, brainstorm use cases, plot them on a value-versus-effort grid, and take the top-right quadrant into a pilot. That workshop is older than AI, and I have sat in enough of them to know what the grid actually holds. It holds the opinions of the people who were free that afternoon, converted into coordinates, and then treated for the rest of the programme as evidence.
What comes out is a long list of plausible things ranked by enthusiasm. Something near the top gets piloted, it works, and it never becomes the way the work runs — because whatever made it a good demonstration was never the thing that would have made it worth doing.
So here is the assessment I would run instead. Eight dimensions, each of which can be observed rather than debated, and a verdict at the end that is allowed to be "leave this one alone". It should take days rather than months, and the people who can answer it already work there.
The unit is a workflow, not a use case
"Invoice processing with AI" is a use case. It has no beginning, no end and no owner, so it cannot be measured and it cannot be finished.
A workflow has a trigger, a defined case, a sequence of steps that crosses people and systems, an end state, and someone who wants it to be over. Score use cases and you are scoring a technology. Score workflows and you are scoring work.
Define the case before anything else, and write it down: one supplier invoice, or one payment run; one customer dispute, or one month of disputes. Every number that follows depends on that choice, and teams that skip it end up comparing a before figure and an after figure that counted different things.
Three dimensions that size the prize
Frequency. How often the work actually runs. High means daily or per-transaction, thousands of cases a year, and it is what pays for a build that has to survive a control review. Observe it by counting cases in the system over a full quarter — and look at the distribution, not just the total. A workflow that runs four hundred times a month with three hundred of those in the last three days is a queueing problem wearing a volume problem's clothes.
Human effort per case. The minutes a person spends, and whether that time scales with volume. Do not ask people to estimate it; they answer with the worst case they remember. Sit with someone for a morning, or take the gap between two system timestamps and ask what happened in between. The answer is often "I was waiting", which is a different finding and a more important one.
Exception load. What share of cases leave the happy path, and — the part that matters — what share of the total effort they consume. High is the norm in enterprise work: a minority of cases eating most of the week. Observe it by asking the team to name the last five things that made them stay late, then look at how those are tracked. If exceptions are not counted anywhere, that absence is itself the finding.
Exception load is the dimension people misread most often. A heavy exception load is not a reason to walk away. Paired with repeatable judgement it is the best opportunity in the building, because the assembly work in front of each exception is exactly what a model is good at. Paired with unrepeatable judgement it is the worst, because you would be automating the manufacture of confident wrong answers.
Two that decide whether AI is the right tool at all
Unstructured input. How much of what arrives is prose, PDFs, email bodies, attachments, and spreadsheets that no two entities fill in the same way. High means the first thing anyone does each morning is read something to work out what it is. Observe it by opening the shared mailbox and the attachments folder, not by reading the process design. If the input is already a clean table in a system you own, AI is not the lever — integration is.
Repeatability of the judgement. Would two experienced people, given the same case, reach the same answer? Note that this is not "is the rule written down". Plenty of good judgement has never been written down and is still highly repeatable; plenty of documented policy is applied four different ways in four regions.
Observe it with a blind pass. Take twenty closed cases from last quarter, including several awkward ones, hand them separately to two experienced people, and compare both the answers and the reasons. It costs an afternoon and it is the single most informative thing in this whole assessment.
The dimension that kills most candidates is not difficulty. It is that two experienced people, given the same case, do not agree on the answer.
Three that set the cost and the ceiling
Cost of error. What happens when the output is wrong, how quickly anyone notices, and how hard it is to undo. High means irreversible, external, regulated or monetary: a payment released, a ledger posted, a commitment made to a customer. Observe it by asking what the last mistake cost and how it was caught. Two answers should worry you. "We have never had one" means either the work is benign or errors are invisible, and invisible is worse. "The reviewer catches it" means the reviewer is the control, and removing the step without replacing the control is not a decision the project gets to make on its own.
State required. What the work must remember between steps and across people: status, owner, versions, evidence, what it is waiting for. Observe it by asking where the truth lives between Monday and Thursday. If the answer is a mailbox and a spreadsheet on a shared drive, state is high and currently unheld — which is not a disqualification, but it tells you honestly that you are building an application rather than writing a prompt. That distinction is most of the difference between a pilot and something that runs the work.
Control burden. How much of this process exists because it is a control, and what evidence has to be produced. Observe it by reading the risk and control matrix and asking internal audit which controls in this process they actually test. Control burden is not a veto either. It is a multiplier on build cost and a hard ceiling on how far autonomy can go, and knowing that early saves a design from being dismantled at its first control review.
Reading the eight together
Do not add them up. A weighted total gives false precision, and it lets a high frequency score outvote unrepeatable judgement, which it must never do. Two of the eight are gates and six are magnitudes. Read them as a shape.
| Verdict | The reading that points here |
|---|---|
| Don't use AI | Input already structured and judgement already codified — an integration or rules problem in AI clothing. Or judgement unrepeatable with a high cost of error. Or frequency times effort too small to fund a build. |
| Automate | High frequency, high unstructured input, repeatable judgement, error cheap and quickly visible, light control burden — and an output that deterministic code can check before it moves. |
| Augment | High effort and high unstructured input, judgement repeatable enough, but cost of error or control burden high. The model reads, assembles and drafts; a named person decides. |
| Autonomy | All of the augment conditions plus a genuinely multi-step sequence across systems, real state, and actions that are bounded and reversible. The sequence itself has to be the cost, not one step in it — autonomy in the three-mode sense, running across an orchestrator placement. |
A few combinations are worth stating plainly. High cost of error plus repeatable judgement is not a reason to stop; it is the definition of augmentation, and it is where most enterprise work lands — which is why augmentation is the right default far more often than programme narratives suggest. High control burden plus high state means the expensive part is the application and the evidence, not the model, and the plan should be costed that way. And a single-step workflow never justifies an agent: if there is no sequence, there is nothing for an agent to hold.
A worked pass: the daily cash position
Take a workflow most finance leaders recognise. Every morning, treasury establishes how much cash the group has, where it is, and what it is going to do today.
Frequency is as high as it gets: every business day, every entity, before the funding window closes. Human effort per case is real and almost entirely retrieval — logging into portals, pulling statements, keying balances, chasing the accounts whose file did not arrive.
Then the two gates. Unstructured input is low. Bank statements arrive in structured formats, and what looks like unstructured work is format heterogeneity and missing connectivity, which is a different problem with a different fix. Repeatability of judgement is so high it is barely judgement: which account, which currency, which value date, which balance — rules, all of them.
The rest confirms it. Exception load is real but narrow: the file that did not arrive, the item that does not map to a known account, the transaction code nobody has seen before. Cost of error is high, because a wrong position leads to a funding decision. State required within the day is modest. Control burden is heavy on anything that moves money and light on the reporting itself.
Verdict: don't use AI. The honest answer is bank connectivity and mapping, and a programme that adopted this as its AI flagship would spend a year and deliver a conversational interface in front of a data problem. There is exactly one AI-shaped seam in it — proposing a classification for the statement lines that never map, for a person to accept, with the accepted mapping written back so the next one lands automatically. That is a useful small augmentation inside a workflow whose real problem is plumbing.
Move one desk along and the reading inverts. The forecast cycle has high unstructured input, judgement that is repeatable enough, high effort, an error cost that is recoverable because nothing moves money, high state and moderate control burden. That reads as augment, and the full pass over cash forecasting shows what falls out of it.
"Not this one" is the most useful answer you will get
Run this over a workshop's long list and most candidates fail on one of three readings: the input is already structured, the judgement is not repeatable, or the prize will not fund a build that has to pass a control review. That is the expected result, not a disappointing one.
It is also the finding worth carrying into the steering meeting. "We assessed forty workflows and are doing two properly" is a stronger position than forty pilots. The costly mistake here is not rejecting too much; it is choosing a workflow that demonstrates well and spending a year proving a capability nobody needed — the gap between adopting AI and redesigning the work.
What to do with the two that survive
Take them apart properly before committing anything: what stays deterministic, what a model genuinely does, what a named human decides, and what the workflow has to hold between steps. That teardown is its own discipline and the rest of the AI workflow design notes cover it.
And before a single thing changes, measure the workflow as it runs today. The assessment tells you which work is worth redesigning; it says nothing about whether the redesign worked, and the only version of that number anyone will believe is the one taken before the project was announced. That is the measurement problem, and it is easier to do first than to reconstruct afterwards.
See also how to measure ROI from an AI workflow and the anatomy of an AI-native enterprise workflow.
Frequently asked questions
3
How do you choose which workflow to redesign with AI first?
Assess candidates on eight things you can observe rather than argue about: how often the work runs, how much human effort each case takes, how much of the input arrives as prose or documents, whether experienced people agree on the judgement, how much of the effort sits in exceptions, what a wrong answer costs, what the work has to remember between steps, and how much of the process is a control. Three of those size the prize, two decide whether AI is the right tool at all, and three set the cost and the ceiling. Pick the one or two that survive all eight, not the one that demonstrates best.
What makes a workflow a poor candidate for AI?
Three readings account for most rejections. The input is already structured and the rules are already written down, in which case the bottleneck is integration or configuration and a model adds nothing but a new failure mode. The judgement is not repeatable, meaning two experienced people looking at the same case reach different answers, so there is no stable target for a system to hit. Or the prize is simply too small: frequency times effort will not fund a build that has to survive an audit. None of those is a technology limitation, which is why a better model does not change the verdict.
How can you tell whether the judgement in a workflow is repeatable?
Run a blind pass. Take twenty closed cases from the last quarter, including several awkward ones, and give them separately to two experienced people with no discussion between them. Compare the answers and, more importantly, the reasons. If they agree on most cases and can explain the ones they got right, the judgement is repeatable enough to be supported by a system. If they disagree materially, you have a policy problem rather than an AI opportunity, and deploying a model would only make one person's interpretation the house standard without anyone deciding that it should be.