How to measure AI ROI is the wrong question if you only count the upside. The right question is whether the value a system produces, over a real time window, clears the full cost of building it, running it, and having a human approve what it writes. Most AI ROI models fail because they price the promise and skip the meter, the maintenance, and the approval time. This is a framework a CFO will accept, written by people who build these systems and run them inside their own operating group.

Key takeaways

  • Return has four honest forms: hard dollars recovered, hours removed from payroll-bearing work, error rate reduction with a dollar value attached, and risk avoided with a documented exposure.
  • Ongoing cost is not zero. Metered cloud usage, human approval time, monitoring, and vendor drift all belong on the cost side of the ledger.
  • A system that writes to money-moving records should never be judged on speed alone. The value often lives in what it prevents.
  • Ninety days is usually too short to judge an AI build. Six to twelve months of steady operation is a fairer window.
  • If the candidate project cannot show a concrete recovered or avoided dollar tied to a real document or record, the honest answer is often not to build.

Start with what actually counts as return

Return has to be something a finance team can audit. That rules out "productivity" as a headline number. It also rules out vendor case studies that quote percentage improvements without a denominator.

Four categories hold up under scrutiny. Hard dollars recovered: rebates captured, invoices corrected, refunds issued that would not have been. Hours removed: work that used to sit on a payroll line and now does not, priced at fully loaded cost. Error rate reduction: mistakes prevented, valued at the average cost of the mistake (a returned shipment, a re-keyed order, a customer credit). Risk avoided: exposure that had a real number attached before the system existed.

If a proposed AI build cannot map cleanly to one of those four, be suspicious. The value probably lives somewhere else, or nowhere.

The cost side vendors do not quote

Every AI ROI pitch shows a clean number on the upside. The cost side is where the model gets honest or dishonest.

Build cost is the easy part. It is a fixed fee, quoted in writing, and it either fits the budget or it does not. What people forget is what happens after go-live.

Cloud usage is metered like any cloud service, quoted in writing before a build. Volume changes month to month. A tool that reads ten thousand documents costs more than one that reads a thousand, and that meter runs as long as the system runs.

Human approval time is real payroll. If a person on your team reviews every entry the system proposes before it posts to a live record, that review has a fully loaded hourly cost. Multiply it out. A tool that saves two hours a day but requires ninety minutes of approval saves thirty minutes, not two hours.

Monitoring is a job. Someone watches for input formats that change, vendors that revise their document layouts, and outputs that start drifting. That person is not free.

Vendor drift is the quiet one. Underlying models change. APIs deprecate. A system that worked cleanly for eight months can need attention in month nine. Budget for it.

For a fuller treatment of the ongoing cost side, our capability page on wiring generative AI into your systems walks through what actually gets touched in production.

How to model ROI on a system that touches money

Systems that read and write to live records deserve a stricter model than internal chat tools. The stakes are different, and so is the math.

Start with the record. Pick one workflow: a supplier rebate posting, a customer credit memo, an invoice reconciliation. Count the volume of documents or events per month. Estimate the current miss rate, meaning the percentage of value that leaks because nobody has time to catch it. Multiply. That is the ceiling of what a system could recover if it caught everything, which it will not.

Now discount. A reasonable intelligent document automation build will not surface every entry, and not every surfaced entry will be approved. Assume the system catches sixty to eighty percent of what a diligent human would catch, and that the approver accepts most but not all of what it surfaces.

Subtract the cost side: build fee amortized over your judgment window, plus metered cloud, plus approval time, plus monitoring. What is left is your honest ROI number.

If it is negative, do not build. If it is positive but thin, look at whether the process itself is worth automating or whether a policy change would do the same work.

A grounded example from our own group

We built a rebate-and-margin recovery system inside Kelsan, a multi-state distributor running Epicor P21. The tool reads the supplier rebate documents Kelsan already receives, matches them against invoices in P21, and surfaces entries for a person on the Kelsan team to approve before anything posts. It has recovered over $100,000 in margin that would otherwise have been lost.

Here is what that ROI conversation actually looked like. On the return side, the number is a real dollar figure tied to real invoices, not a productivity estimate. On the cost side, we counted the build fee, the metered cloud usage for reading the documents, the approver's time reviewing surfaced entries, and the monitoring overhead when a supplier changes a document format. The margin recovered had to clear all of that.

It did, and it does. But we would not have built it if the volume of rebate documents and the historical miss rate had not shown a real ceiling worth chasing. That was the audit question before it was a build question. If you run a distribution business on P21, our page on AI use cases for wholesale distributors covers similar patterns.

When to judge the build

Ninety days is too short. AI systems that touch documents and records need a full seasonal cycle of the underlying business before the number stabilizes. In distribution, that means a quarter of billing cycles, a quarter of supplier statement runs, and a quarter of price changes.

Six months is a fair first judgment. Twelve months is when the number is honest. Cost the build across that window, not against a single month of recovery, and the ROI conversation becomes something a board will actually engage with.

When the honest answer is not to build

Some projects should be killed at the audit stage. If the documents involved are inconsistent to the point where a human cannot reliably read them either, a system will not save you. If the process changes every quarter because of internal policy shifts, the maintenance cost will eat the return. If the total recoverable value is small relative to the build and run cost, buy a spreadsheet template instead.

A fixed-fee AI Capability Audit exists to answer this question before anyone spends build money. Sometimes the deliverable is a plan. Sometimes the deliverable is a memo that says, plainly, do not do this. Both are worth the fee.

About the author

Throughline is a small team of builders inside Keller Group. We build AI systems into our own operating companies first, then into yours.