Why Most AI ROI Numbers Collapse Under Questioning
Almost every AI programme produces a benefit number within a few months of going live, and almost none of those numbers survive a serious finance review. The pattern repeats: a team measures how long a task used to take, multiplies an estimated time saving by the number of people using the tool, converts the result to money at an average salary, and presents a figure with three significant digits. That figure is not dishonest, but it is built from a chain where every link looks reasonable and the product is fiction. Measuring honestly means building the number from the other end, starting with a unit of account that can actually be audited.
The gap shows up in what teams say is blocking them. An NVIDIA 2026 survey reported 30% of respondents citing a lack of ROI clarity as a blocker to adoption, ranking alongside data problems and a shortage of expertise. The cause is rarely a missing spreadsheet; it is that nobody agreed, before launch, on what would count as evidence. Once the system is live every measurement becomes contested, because each person in the room now has a position to defend and a story that explains whichever numbers they prefer. That is why measurement design belongs in the same document as scope, alongside the broader case for adopting AI in a business.
What follows is a measurement design rather than a business case template. It moves in order: establish a baseline before launch, define cost per workflow execution as the unit of account, separate deflected cost from realised cost, discount throughput gains by the error rate they carry, solve the counterfactual with a holdout or a staggered rollout, split leading from lagging indicators, and finish with a payback calculation and the criteria for killing a project. Every currency figure in the worked examples is illustrative, expressed in Turkish lira, and drawn from ranges we consider plausible rather than from any measured industry average.
Hours Saved Are Not Money Until Someone Converts Them
The conversion from saved time to saved money runs through six steps, and each one has to be defended separately. First, time saved per item, measured rather than estimated. Second, that figure multiplied by actual volume, giving gross hours. Third, gross hours multiplied by a realisation rate, giving convertible hours. Fourth, convertible hours multiplied by a fully loaded hourly cost, giving gross money. Fifth, subtract the run cost of the system itself. Sixth, subtract the cost of any quality degradation the system introduced. What reaches the bottom of that chain is usually somewhere between a quarter and two thirds of what was claimed at the top, and the gap is almost never in the arithmetic.
Step three is where most business cases die, and it is almost always missing. Suppose a tool saves 30 minutes per person per day across 40 people. That is 20 hours a day, roughly 2.5 full-time equivalents at 160 hours a month. If the team is still 40 people at the end of the quarter, output is unchanged and no contractor invoice fell, the company saved zero lira. It bought slack. Slack has value, but it is not a line a finance director can point to. Saved hours become money through exactly three mechanisms: a role that is not filled, an external spend that is reduced, or capacity redeployed onto work that carries its own revenue or cost line.
Two smaller errors compound the main one. Teams use gross salary instead of a fully loaded rate, which understates the true hourly cost by roughly 35% to 60% once employer contributions, tooling, workspace and management overhead are counted. And they hold volume constant even when cheaper analysis induces more of it: we have seen review volumes rise 20% to 40% in the two quarters after a workflow became fast enough to use casually. Before a line of code is written, name the conversion mechanism, name the person accountable for executing it, and put a date on it. If nobody will sign for any mechanism, the honest forecast for the labour component is zero, and the project has to justify itself on cycle time, quality or capacity instead. That conversation belongs in the scoping phase, where changing the answer is still cheap.
Build the Baseline Before Launch, Not After
The single most common measurement failure in this category is the missing baseline. Two to four weeks before go-live, instrument the current process and record six things: volume per week, cycle time from request to completion at the median and the ninetieth percentile, human touch time separated from waiting time, error and rework rate under a written definition, the cost of an average rework event, and who performs each step. Mean cycle time is the least useful of these and the one most often reported. The distribution is where the money hides: a process with a median of two hours and a ninetieth percentile of nine days has a queueing problem, not a speed problem. Hand-timing 150 to 400 items across a representative mix of easy and hard cases is usually enough to establish a defensible median.
If the process already produces timestamps, use them, but audit what they mean before trusting them: in the projects we run, the recorded start time is often the moment a ticket was assigned rather than the moment work began, which inflates the baseline by hours. Retroactive baselines are worthless for three reasons and should be treated as unavailable rather than approximate. First, memory anchors on the worst case: ask anyone how long a task takes and you will get the painful example, not the median. Second, the definition drifts to fit the result, because the person constructing the baseline already knows what the new system produced. Third, the people who could attest to the old process have already changed how they work. A baseline reconstructed after launch is a negotiation, not a measurement, and everyone in the room knows it.
Baselines also fail for a mundane reason: the events needed to compute them do not exist in any system. Cycle time cannot be measured if the process lives in email and a shared drive. Error rate cannot be measured if corrections are made silently in place. In practice, roughly a third of the measurement work on a first project is plumbing, which is why baseline instrumentation belongs in the same workstream as preparing the data the system will run on. Freeze the finished baseline in a one-page document with definitions, dates, sample sizes and a signature, and circulate it before anyone has seen a single result.
Cost per Workflow Execution as the Unit of Account
Programme-level cost figures are useless for decisions because they cannot be compared to anything. Cost per workflow execution can be compared: it is the total cost of running one complete unit of work through the system, and it has four components. Take an illustrative contract triage workflow running 1,200 executions a month. Inference: with retrieval-heavy prompts of roughly 10,000 to 14,000 input tokens and around 900 output tokens, assume 1.20 TL per execution, and take that number from your provider’s current price list rather than from an article. Retrieval and infrastructure: vector database, embedding refresh, object storage, orchestration and logging at 9,000 TL a month, which is 7.50 TL per execution. Human review: 30% of executions are routed to a reviewer for four minutes and the rest get a 30 second spot check, giving 1.55 minutes on average at a fully loaded 450 TL per hour, so 11.63 TL.
Amortised build cost is the component teams leave out, and at low volume it dominates everything else. A build of 1,200,000 TL amortised over 24 months is 50,000 TL a month, which at 1,200 executions is 41.67 TL per execution. Total fully loaded cost is therefore 1.20 plus 7.50 plus 11.63 plus 41.67, or 62.00 TL per execution. Run-only cost, the number that matters for the marginal decision, is 20.33 TL. The manual baseline for the same unit of work is 22 minutes of the same reviewer, or 165 TL. Two very different comparisons follow from those two figures, and confusing them is how boards get misled.
Volume changes the picture more than any engineering optimisation available in the first year. At 3,000 executions a month the same build amortises to 16.67 TL and infrastructure to 3.00 TL, taking fully loaded unit cost to 32.50 TL. At 600 executions those two rise to 83.33 and 15.00, giving 111.16 TL, which sits uncomfortably close to the manual cost and leaves no margin for error. Before optimising tokens, check whether the workflow has enough volume to deserve a custom build at all. That question sits at the heart of what an AI project actually costs, and it is usually asked far too late.
Deflected Cost Versus Realised Cost
Deflected cost is work the system absorbed that a person would otherwise have done, valued at a rate. Realised cost is money that left the profit and loss account differently: a role not filled, an outsourcing invoice reduced, overtime not paid, a licence not renewed, revenue recognised earlier. Deflection is real and worth measuring, but it is capacity, not cash. It converts to cash only through a decision that a named person has to make and defend. Boards confuse the two constantly, and the confusion is not innocent, because deflected numbers are larger, arrive sooner and require nobody to do anything uncomfortable.
The error compounds when deflected figures are added across projects. Three separate initiatives each claim to save 20% of a ten-person team, and the portfolio review reports a saving of six people. Ask who the six people are and the answer is that they are all still employed. The sanity check is arithmetic: the sum of all claimed savings for a function must not exceed that function’s total cost. In the projects we run, applying that ceiling removes between a third and a half of the claimed portfolio benefit before anyone argues about methodology.
Report both numbers, in two columns, and never blend them into one headline. The realised column shows what finance can book, and every line in it carries a named owner and a target month. A project that shows 400,000 TL of deflection and 0 TL of realisation is not a failure; it is a project whose conversion decision has not been made yet. Saying that plainly is far better than inventing a blended figure, because the blended figure is the one that gets quoted back a year later when nobody can find the savings.
Quality-Adjusted Throughput: Speed That Costs Accuracy Is Not Speed
Throughput gains that arrive with a higher error rate are not gains, and the arithmetic is unforgiving. Errors do not simply reduce the count of good units; they consume capacity through rework. Any throughput claim therefore needs two adjustments: subtract the defective output, then subtract the time spent fixing it. Teams that report only the first adjustment still overstate the gain substantially, and teams that report neither can produce a number four times larger than the truth. Measure the error rate on the same rubric, before and after, with a blind re-review of a sample.
An illustrative example. A review team clears 100 items a day at a 3.0% error rate. With an assistant in the loop the same team clears 160 items a day, but the error rate rises to 5.2%. The headline gain is 60%. Adjusting for defects alone gives 97 good items against 151.7, a gain of 56%. Now price the rework: each escaped error costs 45 minutes to find and fix. The manual process generates 3 errors and 2.25 hours of rework, so 100 items consume 10.25 hours. The assisted process generates 8.3 errors and 6.2 hours of rework, so 160 items consume 14.2 hours in total.
Effective throughput is therefore 9.8 items per hour before and 11.2 after, a gain of 15%, not 60%. And this is the optimistic version, because it assumes every error is caught internally. If a share of them reach a customer, a regulator or a counterparty, the multiplier on rework cost is not one but five to twenty, and a small error-rate increase can turn a positive project negative. Error rate belongs on the same dashboard as volume, sampled continuously rather than audited annually, which is the practical argument for running evaluations and observability in production rather than only before launch.
The Counterfactual Problem, and How Holdouts Fix It
Comparing after to before assumes nothing else changed, which is never true. Volume moves with the season, the team gets better at the task independently, a process change lands in the same quarter. Any of these can produce the entire measured effect on its own. The only reliable answer is a concurrent control: a portion of the same work, in the same period, handled the old way. Without it you are not measuring the system, you are measuring the quarter, and the finance team will eventually ask a question you cannot answer.
A workable holdout is 20% to 30% of eligible volume, run for at least six to eight weeks, with the randomisation unit chosen to match the decision. Randomising by item is cleanest statistically but leaks badly when the same person handles both arms, because whatever they learn from the assisted items changes how they do the manual ones. Randomising by person removes that leak and introduces another, since people talk to each other. Randomising by team is the most robust and the least sensitive, and usually needs eight weeks rather than four to produce a signal you would act on.
Where a holdout is politically impossible, a staggered rollout does most of the same work. Bring four teams live at two-week intervals: each team acts as its own control before it switches, and every team that has not yet switched is a concurrent control for the ones that have. The comparison is a difference in differences, which is straightforward to build once the baseline is properly instrumented. Staged rollout is usually the right answer anyway for reasons that have nothing to do with measurement, which is why it features so heavily in moving from pilot to production.
Guard against five specific confounds. Seasonality: month-end and quarter-end distort both volume and error rate, so never run a four-week test that straddles only one of them. Learning curve: reviewer performance with a new tool typically keeps improving for three to four weeks, so the first fortnight understates the steady state. Selection: if users choose what to route through the system they will route the easy cases, and the measured gain is mostly case mix. Spillover: contamination between the two arms. And concurrent change: freeze other process changes for the duration, or accept that you will never separate the effects afterwards.
Leading and Lagging Indicators Belong on Different Clocks
Weekly dashboards should carry leading indicators only, and there should be at most seven to nine of them, each with a threshold and a named owner. The useful set is adoption as a percentage of eligible volume, escalation or fallback rate, human override rate, evaluation pass rate on a fixed regression set, retrieval hit rate, latency at the median and the ninetieth percentile, cost per execution, and the volume routed to low-confidence handling. These move within days, respond to engineering work, and tell you whether the system is on a trajectory. They are the things that must be true for a benefit to appear later.
Quarterly reviews carry the lagging set: cost per item against baseline, cycle time at the ninetieth percentile, error and rework rate, realised conversion against the plan with owners and dates, and payback progress against the original forecast. Putting cost per item on a weekly dashboard guarantees that somebody reacts to a two-week fluctuation, ships a change, and destroys the comparison. In the projects we run, a sensible rhythm is a weekly fifteen-minute review of the leading set, a monthly written note on drift, and a quarterly session where the payback model is rebuilt with actual figures rather than quietly adjusted.
The two failure modes are symmetric. Teams that track only leading indicators report high adoption and good latency for four quarters and never discover that unit cost never fell below the manual baseline. Teams that track only lagging indicators discover a problem two quarters after it started, when the cause is untraceable and all that is left is an argument. Both sets are cheap to instrument if the logging is designed at the start and expensive to retrofit afterwards, because the events required for the lagging set have to be emitted by the production system from the first day it handles real work.
A Payback Calculation, Start to Finish
Return to the illustrative workflow: 1,200 executions a month, a build cost of 1,200,000 TL, manual human cost of 165 TL per execution and post-launch human cost of 11.63 TL per execution. Human time saved is therefore worth 153.37 TL per execution, or 184,044 TL a month. Machine run cost is inference plus infrastructure, 8.70 TL per execution or 10,440 TL a month. The naive monthly benefit is 173,604 TL, and the naive payback is 1,200,000 divided by 173,604, which is 6.9 months. This is the number that reaches the steering committee pack, and it is wrong by roughly a factor of two.
Now apply the realisation rate. The freed time is 409 hours a month. Of that, 160 hours corresponds to a reviewer role that will genuinely not be backfilled, worth 72,000 TL a month at 450 TL an hour. Another 130 hours displaces work currently outsourced for 36,000 TL a month. The remaining 119 hours are absorbed into slack, meetings and unpaid overtime that nobody was tracking, and are worth zero. Realised labour value is 108,000 TL a month, a realisation rate of 59%. Net of the 10,440 TL run cost, the honest monthly benefit is 97,560 TL and payback moves out to 12.3 months.
Then apply the quality adjustment. Suppose the error rate on the same rubric moves from 3.0% to 4.8%, which is 21.6 additional errors a month at 1,200 executions, each costing 45 minutes of rework at 450 TL an hour, or 337.50 TL. That is 7,290 TL a month. Net monthly benefit falls to 90,270 TL and payback to 13.3 months, so the system breaks even during month 14 rather than month 7. It involved counting the run cost, the conversion decision and the errors, all three of which were present in reality from the first day.
Publish the sensitivity, because one assumption moves the answer more than all the others combined. At a realisation rate of 40% the monthly benefit is 55,888 TL and payback is 21.5 months. At 59% it is 90,270 TL and 13.3 months. At 80% it is 129,505 TL and 9.3 months. That is the difference between a project a board funds and one it does not, and it turns entirely on whether somebody will commit to not filling a role. Put those three columns on one page, name the assumption, and let the decision be made about the assumption rather than about the model.
Knowing When to Stop, and How We Run Measurement
Some projects have no measurable return, and the discipline is to say so on a schedule rather than when patience runs out. Gartner’s June 2025 press release predicted that over 40% of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls, based on a poll of over 3,400 respondents. Unclear business value is the interesting item on that list, because it is the one a measurement plan actually prevents. A project that cannot say what it is worth will eventually be cancelled by someone who also cannot say, and the organisation learns nothing from either decision.
Set the kill criteria before launch and review them one quarter after the pilot converts. Adoption under 30% of eligible volume with no diagnosed and fixable cause. Quality-adjusted throughput gain under 10% and not trending upward across two consecutive months. Cost per execution not falling and still above the manual cost per item. A realisation rate under 25% with no owner willing to sign for the conversion. A payback forecast that has slipped by more than half on two separate occasions. Two of those five, with no credible plan attached, is a stop. Stopping need not waste the work: the labelled data, the evaluation set, the baseline instrumentation and the retrieval pipeline usually outlive the application, so write a two-page post-mortem answering whether it was the use case, the data, the workflow design or the adoption.
At HatsonTech we put the measurement one-pager into the scope document before the first prompt is written, and we ship a cost meter with the first release that logs tokens, retrieval calls, latency and human touch time per execution, so cost per execution exists from day one instead of being reconstructed later. We default to staggered rollouts so there is always a concurrent control. Most of the measurement failures we see are data engineering failures, which is why the baseline work usually starts as a data pipeline and warehouse job rather than a modelling one. When the numbers do not support continuing, we say so, on the terms we set at the beginning.