Where to Start With Enterprise AI
Not with a pilot. A sequence borrowed from enterprise systems delivery — and the three steps in it that change when the component is probabilistic.
Method 3 of 5 see the reading order →
Reviewed by Tan Gravam
On this page
Start with an inventory of what you already run, then pick one workflow using an assessment that is allowed to say no, and build the decision rights before you build the model. The pilot is not step one. It is a way of answering a written question somewhere in the middle, and running one first is the commonest reason a programme spends a year proving something everybody already believed.
The textbook advice is the opposite: find a low-risk quick win, pilot it, build momentum. That advice is not wrong so much as incomplete — it optimises for showing something soon, and the thing it shows is almost never the thing that was hard.
I have sequenced enterprise systems programmes for eighteen years. The sequence below is the one that works there, with the three steps that genuinely change when the component in the middle is probabilistic marked as such. It is deliberately unglamorous. Most of the value in sequencing is refusing to skip the boring parts.
The sequence
1. Inventory what you already have. Not what the programme plans — what is already running. AI arrived in most organisations as features of software already owned, switched on by people who were not asked to file a project. Until that list exists, every subsequent step is being applied to an unknown population. This takes about a week and gets worse every quarter you leave it.
2. Pick one workflow, with an assessment that can return "don't". Not the one that demos well and not the one whose owner is most enthusiastic. An assessment that can say no is the only kind worth running — if the method has never produced a "leave this alone", it is a ranking exercise wearing the clothes of an evaluation.
3. Decide the measure before you build. Which number you are claiming, baselined on the current process, with the confounders named in advance. This is the step most often done afterwards, at which point the baseline is a reconstruction and nobody believes it.
4. Write the decision rights and the control path. Who may approve what, on what evidence, and what enforces it. This is a sequencing change, not a governance formality — it determines the architecture, so discovering it at a control review means rebuilding. Most of the answer already exists in your delegation of authority; the work is mapping onto it rather than inventing a parallel regime.
5. Design the exception path before the clean case. (Genuinely different.) In a deterministic programme the exception is a defect to be driven down. Here it is a permanent population with a size you can measure, and the path it takes is most of the operational design. Building the happy path first produces a demo and a backlog.
6. Run it in parallel, on real volume. Not a sample, not last quarter's file. The parallel run is where the exception rate becomes a number instead of an assumption, and where you find out whether reviewers can actually work at the pace the design assumes. It is also where the gates that let a pilot into production either close or do not.
7. Go back and check. Six months later, against the measure from step 3. This is benefits realisation, it is skipped almost universally, and skipping it is how an organisation accumulates capabilities nobody can defend.
If your method has never returned "leave this one alone", it is not an assessment. It is a ranking of enthusiasm.
The three steps that are genuinely different
Most of that sequence is ordinary delivery discipline, and I would be suspicious of anyone claiming otherwise. Three steps do change, and they change in ways that break habits rather than tools.
Fit-gap does not work. Fit-gap analysis compares a requirement against what a system does and returns fit, gap, or workaround. You cannot do that against a component whose answer varies. The honest substitute is empirical: run real cases, measure how often the output is usable, and treat the residue as the requirement. That is a different activity with a different cost, and programmes that budget for a workshop and need a data exercise discover it late.
Readiness means something else. In a systems programme, readiness means environments, cutover rehearsals, trained users, a run team. Those still apply. What is added is data readiness of a specific kind: not "is our data clean" in the abstract, but is the record this answer depends on reachable from where the capability runs, current enough to be right, and scoped so it does not also expose things it should not. That question has a long lead time and no shortcut.
The pilot moves. In a normal programme a proof of concept sits early, to de-risk technology choice. Here the technology is not the risk — the exception rate, the control path and the reviewer's capacity are. So the useful experiment happens after steps 1 to 4, and its purpose is to produce numbers for a design decision, not confidence for a steering committee. This is the same reason pilots do not become operating systems: one that was never asked a question cannot fail, and something that cannot fail teaches nothing.
What this looks like when it goes wrong
Two patterns, both recognisable from a distance.
The enthusiasm-led programme. Started at step 2 with whoever volunteered, skipped 3 and 4, produced a working demo in six weeks and then spent nine months in review cycles that keep discovering constraints the design cannot accommodate. Every one of those constraints existed on day one and was findable.
The governance-led programme. Started at step 4, wrote a policy, built a council, produced a framework — and never reached step 2, because there was no workflow to test it against. A control regime designed against no concrete case is invariably calibrated for the worst thing anyone imagined, and it makes the first real proposal look non-compliant.
The sequence exists to avoid both. Steps 1 to 3 are cheap and produce facts. Step 4 is where most of the design is decided. Steps 5 to 7 are where the work actually is.
What I would decide
Spend the first month on the inventory and the assessment, and expect the assessment to eliminate most candidates — that is it working. Write the measure and the decision rights before anything is built, because both change the architecture. Design the exception path first. Run in parallel on real volume, not a curated sample. And put the six-month check in the plan at the start, because nobody schedules it afterwards.
The uncomfortable part is that almost none of this is about AI. That is the point: the transformation is in the work, not the tool, and the sequence that gets you there is the one enterprise delivery already knows — with three steps swapped for ones that survive a component you cannot fully predict.
See also what AI transformation actually means and how to find workflows worth redesigning.
Frequently asked questions
4
Should the first enterprise AI project be a pilot?
A pilot is a good way to answer a question and a poor way to start a programme, and the difference is whether you wrote the question down first. Run one to find out something specific — whether the input is machine-readable enough, what the real exception rate is, whether reviewers can work at the required pace. Do not run one to demonstrate that the technology works, because that answer is already known and the demonstration teaches you nothing you can build on. A pilot with no written question becomes a slide.
What is the first step in an enterprise AI programme?
An inventory, and it is duller and more useful than it sounds. Most organisations cannot list the AI-touching processes they already run, because a good number arrived as features of software they already owned rather than as projects anyone approved. You cannot sequence, govern or classify what you have not enumerated, and the enumeration takes about a week and gets harder every quarter you leave it.
How is sequencing an AI programme different from a normal systems programme?
Three steps change. Fit-gap analysis does not work the same way, because you cannot compare a requirement against a probabilistic component's feature list and get a yes or no — the honest equivalent is measuring the exception rate on real data. Readiness stops meaning environments and training and starts meaning whether the data an answer depends on is reachable and current. And the exception path has to be designed before the clean case rather than after it, because with a deterministic system the exception is a defect, while here it is a permanent, sized population.
What should you build before the model?
The decision rights and the path an action takes to become real. Who may approve what, on what evidence, and which component enforces it — because that determines the design, and retrofitting it means rebuilding. Teams routinely build the capability first and discover at the control review that the workflow they designed cannot be permitted, which is not a governance problem; it is a sequencing one.