AI Workflow Teardown: Bank Reconciliation
Bank reconciliation taken apart: why the matching engine was never the bottleneck, what AI does with the residue, and which breaks need a named human.
Teardown 2 of 2 see the reading order →
Reviewed by Tan Gravam
On this page
The textbook version of bank reconciliation is a two-column exercise: statement lines on one side, ledger entries on the other, tick off what agrees, investigate what is left. Every manual describes it that way, which is the problem: the description makes the ticking sound like the work. The ticking has been automated for decades.
This is the second in a series of teardowns taking one real workflow apart with the method; the first took cash forecasting through the same headings. Reconciliation is the better demonstration: the split between rule, model and named human is unusually clean.
The workflow as it runs today
Statements arrive overnight, one file per bank account, in whichever format that bank sends — MT940, camt.053, BAI2. An import job loads them and matching rules run: amount against open item, reference against document number, date inside a window, tolerance where policy allows. Recognised transaction types post automatically.
Everything the rules could not decide lands in a queue, and the human day starts. An accountant works down it — the bank portal, the remittance mailbox, the open-item list, an email to the customer. Some items resolve in a minute; some sit for weeks. At period end the rest are aged, written off, provided for or explained in a note, and the reconciliation has to tie out regardless.
Where the time actually goes
Map the cycle honestly, step by step with the reason each takes as long as it does, and the same shape tends to appear. The matching engine is not in it: it runs unattended and finishes before anyone logs in. The elapsed time sits in the residue, in four forms.
Identification: who a payment is from, when the only clue is a truncated company name and a reference from the payer's system. Assembly: the open items, the payment history, the credit note, the dispute, the remittance advice that arrived as a PDF two days early. Waiting: a query to a customer does not return the same day, and the item ages while the queue grows. Repetition: the same counterparty deducts the same charge every month and someone derives it again each time, because the last person to work it out wrote it in a comment field nobody searches.
Reconciliations are rarely late because the matching engine was slow. They are late because a handful of items have no explanation and one person owns all of them.
An unexplained break is not a clerical nuisance, either: it is an unverified statement about cash, ageing while the pressure at period end is to shrink the queue rather than understand it.
The deterministic core
Now the split. The first pile is everything that must be identical every time it is asked, or that a control depends on.
- Amount matching, including currency and tolerance — a stated band in policy, not judged item by item.
- Dates. Value date against booking date, cut-off, which period a line belongs to.
- Reference and identifier matching. End-to-end identifiers, structured remittance references, document and payment-run numbers. Where a structured reference exists, matching on it is a rule and a model adds nothing.
- Balancing. Opening plus movements equals closing, on both sides, and the reconciliation still ties out. Non-negotiable.
- Postings and clearing, generated by code from validated data, with a reversal path.
- Ageing, routing and segregation of duties, with four-eyes above a threshold, checked rather than requested.
None of it is a candidate for a model, ever. A reconciliation whose totals could vary between two runs on the same statement is not a reconciliation; it is an opinion.
Which is what most AI proposals for this workflow miss. The matching engine is not the interesting part and never was — rules already clear the routine lines cheaply, so any gain there is measured against a step that already costs nothing.
What AI can genuinely do here
The second pile is what an experienced reconciliation clerk does all day: read things and build a case.
- Interpreting free text. Narrative and remittance fields, structured or unstructured; the advice emailed as a PDF; the deduction schedule sent as a spreadsheet.
- Categorising the exception. Timing difference, charge deducted at source, FX difference, partial payment, duplicate, misdirected receipt.
- Proposing root-cause hypotheses, ranked, with the evidence each rests on: this credit looks like two invoices less a settlement discount this customer has taken all year.
- Synthesising the evidence. One pack per break — statement line, candidate open items, the payer's history, the correspondence, and how the last ones like it were closed.
- Drafting the remediation. The note to the counterparty, the clearing proposal, the write-off memo, the period-end explanation.
What these share is that being wrong is visible and cheap: a rejected hypothesis costs a click from someone who would otherwise have spent twenty minutes on it. Nothing clears because a model was confident. Where being wrong is neither visible nor cheap — anything that posts or moves money — the work belongs in the first pile or the third. That is the automation, augmentation and autonomy split, applied item by item.
What stays with a named human
The third pile is decisions, and the test is not capability. It is who answers for it when it is wrong.
Write-offs, because money leaves permanently and does not come back: a band, an approver, four eyes. Accounting treatment: whether a difference is a bank charge, an FX movement, a discount taken, a loss or a receivable that survives, and which period it belongs to. Material exceptions: anything above a threshold, and anything touching a counterparty relationship, where the cheapest answer and the right one differ.
And the unresolved break at period end — carry it, provide for it, escalate it, park it in suspense with an owner and a date, or decide that none of those is acceptable and the close waits. That needs a person, because it is the decision an automated queue would rather not make.
The redesign does not remove that person; it turns their work from investigation into decision.
The state this workflow has to hold
Here is where reconciliation redesigns quietly fail, and it has nothing to do with models. The statement has a system; the ledger has a system; the break has neither, and lives in a spreadsheet, a comment field and one person's memory.
- The break itself, with an identity, a status, an age and an owner who is a person rather than a team.
- Evidence gathered — what was looked at, what was asked, who answered, when — so the third person to touch it does not start again from the statement line.
- Resolution and reason code, in a vocabulary you can count rather than free-text prose.
- Historical patterns. This counterparty always pays net of their bank's charges; this entity's payroll arrives as one debit with no detail. That knowledge is the asset here, and it leaves when the person holding it does.
- An audit trail. Which action, by which identity, at which time, on which item.
That list is why this becomes an application with a model inside it rather than a chat window with a good prompt — and one reason pilots stall short of running the work.
The exception path
Here the exception path is not part of the design; it is the design. Six recurring shapes; the redesign is judged on them.
- A statement line with no ledger counterpart. An unexpected receipt or charge, needing identification before a posting — and identification is the free-text problem.
- A ledger entry with no statement line. A payment that never landed, an item in transit: a timing difference until proven otherwise, and proving otherwise is the work.
- A difference inside tolerance. Cleared by rule, but recorded with a reason rather than swallowed — otherwise the mechanism that makes small differences painless destroys the signal that one counterparty deducts a fee every month.
- A difference outside tolerance. A short payment or deduction, needing the evidence pack: which invoices, which deduction, whether it recurs or is disputed.
- Matched to the wrong thing. The expensive one, because it looks resolved: two invoices of the same amount, the wrong customer cleared, found when the right one gets a dunning letter. Detection must be designed.
- The break nobody can close. It ages, stays visible, stays owned, and is never closed automatically for tidiness.
Each needs the same four things: detection, a route, context that travels with it, and a write-back so the next costs less. The write-back is what compounds — it is how a pattern stops being folklore and becomes state.
What it has to integrate with
The reads are broad: statements from every bank in every format, open items from receivables and payables, the ledger with its clearing and suspense accounts, payment runs and their references, remittance advice arriving by whatever channel. Where a group runs a treasury management system, much of that already sits in one place and the integration is narrower than it looks.
The writes are where this differs from the forecasting teardown. That workflow published a document; this one posts to a ledger — clearing, reclassification, write-off — so the control surface is real. Which forces one design rule: the model proposes, code posts. A posting is generated deterministically from an approved decision, under a traceable identity, with a reversal path, never emitted by a model. Get that boundary wrong and nothing downstream is defensible, which is why enterprise AI is an architecture problem before it is a model problem.
The redesigned workflow
Statements land, the rules run, the routine lines clear exactly as they always did. What changes is the residue. Every unmatched item becomes a break record the moment matching finishes, owned by the account, entity and type rather than by whoever opens the queue first. Before a person looks at it the evidence is assembled: narrative parsed, candidate open items retrieved, the payer's history summarised, correspondence found, hypotheses ranked. Where a recorded pattern applies — this counterparty always deducts their bank's charges — the proposal says so and cites the breaks behind it.
The accountant's session becomes review, not investigation: accept, reject, or ask for the one fact the assembly could not find. Accepting generates a deterministic posting under their identity, four-eyes above a band. Rejecting is recorded, because a rejected proposal is both signal and measure.
What is left at day's end is the residue of the residue: items needing a decision, aged and visible and owned. And the reconciliation still ties out, which nothing here may make easier by hiding anything.
What the business case is made of
Four measures, named before the build and taken before and after. None is a model metric.
- Auto-match rate, stated as a baseline precisely so nobody claims credit for moving it. A better rate beside a queue the same size on the twentieth means something already free got optimised.
- Time to resolve per break, by category — the one that should move, because assembly is what was removed.
- Open breaks at period end, and their age. The control measure, and the one a finance director already watches.
- Manual touches per break: portals opened, emails written, spreadsheets built.
Stated as a shape, illustratively rather than measured: the change is not fewer breaks. It is breaks that arrive with their evidence already gathered, so an afternoon of assembly becomes a morning of decisions. What justifies the work is not the clerk hours. It is the cost of a period closed over items nobody understood — and of an unexplained credit sitting in suspense for six weeks, a control failure with a plausible story attached.
What this generalises to
The piles fall out unusually cleanly here, but what made the redesign work is not specific to banks.
The engine everyone wants to improve is usually already the cheap part. In most mature processes the deterministic core is solved, fast and dull, and the work sits in what it could not decide. Ask what share of the elapsed time is inside the step being automated, and be ready for almost none.
Propose, never post. Wherever a workflow writes consequentially to a system of record, the model's output is a proposal and the write is code generated from an approved decision — approved by a person, or by a remit narrow enough to have been written down in advance. That one boundary is what makes a probabilistic component acceptable to an auditor.
The pattern in someone's head is the asset. Every exception-heavy process has someone who knows why this counterparty is always short. Capturing that as state, cited in the next proposal, is worth more than any model choice.
An unresolved item must be allowed to stay unresolved. Owned, aged, visible, never closed for tidiness. A design that cannot hold an open item indefinitely will eventually hide one, which turns an efficiency project into a control finding.
The next teardown takes another workflow through the same headings. They do not change, which is rather the point: a method that works only on its author's favourite process is not a method.
See also the anatomy of an AI-native enterprise workflow and bank statement formats: MT940, camt.053 and BAI2.
Frequently asked questions
3
Why does a faster matching engine not speed up bank reconciliation?
Because the matching engine was never the slow part. Rules that compare amount, reference and date within a tolerance run unattended overnight, and where remittance references are structured they clear the routine lines without anyone watching. The calendar is consumed by the residue — the items the rules could not decide — and each of those costs a person a search through a bank portal, an open-item list, a remittance mailbox and a colleague's memory. A redesign that makes matching faster improves a step that already finishes before anyone arrives at work, and leaves the queue exactly the size it was.
What can AI safely do with an unmatched bank statement line?
It can read it and build the case, but it must not clear it. Interpreting a narrative or remittance field, identifying the likely payer, retrieving candidate open items, recognising that a shortfall matches a recurring deduction, and drafting both the clearing proposal and the note to the counterparty are all reading and synthesis, where being wrong is visible and cheap. Turning that proposal into a posting stays deterministic: code generates the entry from an approved decision, under a traceable identity, with a reversal path. Propose, never post, is the boundary that keeps the whole design auditable.
Who should own an unresolved reconciliation break at period end?
A named person, not a team and not a queue. Someone has to decide whether an item nobody could explain is carried forward, provided for, escalated, or parked in a suspense account with an owner and a date attached — and sometimes the correct answer is that the close waits. Automation may assemble the evidence and age the item, but it must never close it for tidiness. An unexplained credit that quietly disappears from a queue is not an efficiency gain; it is a control that stopped reporting, which is precisely what a reconciliation exists to prevent.