By Mert Dönmezler11 min read

Modeling Automation ROI Under Uncertainty

How to model automation ROI as a distribution: break the benefit into its terms, take priors from a reference class of past projects, and decide on percentiles instead of the mean.

  • roi
  • automation
  • decision-analysis
  • research
  • monte-carlo
  • forecasting
  • risk

Automation proposals are almost always defended with one number: a payback period, an annual saving, a return multiple. That number is imprecise, and it is also biased upward for structural reasons. Every term in the standard return expression is a random variable, most of those variables are skewed in the same direction, and the return is a nonlinear function of them. So the return you compute from typical inputs is not the typical return. The sections below show how to replace the point estimate with a decomposed model, a reference-class prior, a small Monte Carlo run, and decision rules stated on percentiles instead of means.

Anatomy of an automation ROI point estimate

Begin with the coarse form. Over a horizon of T years, return on investment is ROI = (B − C) / C, where B is cumulative benefit and C is cumulative cost. At that level there is nothing to act on, so break both sides down.

Annual benefit is B_a = F × H × R × A, where F is the number of times the process runs per year, H the human hours one run takes today, R the fully loaded hourly cost of those hours, and A the fraction of the work the automation actually absorbs. Cost is C = D + M × D × T + X, where D is delivery cost, M an annual maintenance rate applied to D, and X the integration, training and change-management cost that proposals tend to leave off the page.

Each term is uncertain in its own way. F is usually the best-known term and often the only one measured properly. An operational log settles it, which is one reason process mining belongs before automation and not after it. Our own RPA workflow for a card printing company, which pulls ERP order data, formats shipping labels and pushes them to fulfillment, runs 800 or more times a day. We took that figure from the execution logs, and nobody had to estimate it in an interview. H is harder. Asked how long a task takes, people describe the attentive case and leave out the interruptions, the rework and the waiting. R is harder again, because saved hours only become saved money if the freed capacity is redeployed or headcount actually changes. A is where the optimism collects. The residual exception path, the five percent the system hands back to a human, often takes more supervisory attention per unit than the ninety-five percent it replaced.

Two structural points follow. First, the theory of constraints (Goldratt, 1984) implies that time recovered at a non-bottleneck step turns into throughput only by accident, so R should be near zero there. Little's Law (Little, 1961) says the same thing formally. For a stable system L = λW, so cutting the time in system W raises throughput only if the arrival rate λ is free to rise. Second, ROI is a ratio and net present value is a discounted sum with a delay term, and both are nonlinear. Putting mean inputs into a nonlinear function does not give you the mean output. This is the flaw of averages (Savage, 2009), and it is why a plausible spreadsheet can be wrong in a predictable direction.

The planning fallacy in delivery cost and schedule estimates

The second source of bias is how the delivery term gets produced. Teams build D and the go-live date by imagining the plan: they break the work down, add up the parts and put a margin on top. Kahneman and Tversky (1979a) called this the inside view. They showed that the fix for it, distributional information about similar past cases, is usually available and routinely ignored. Kahneman (2011) describes the same distinction as inside view versus outside view.

The resulting interval is lopsided as well as wide. Schedule and cost overruns cannot go below zero and have no practical upper limit, so their distribution is right-skewed and its mode sits below its mean. A project whose most likely timeline is three months may well have a mean of five. The long tail is thin but heavy: an integration that turns out not to exist, a data owner who leaves, a compliance review nobody scoped. Benefit only starts after go-live and is discounted, so expected value computed at the modal timeline overstates expected value computed across the whole delay distribution. The gap grows with the discount rate, and the finite horizon makes it worse, because any benefit pushed past T drops out of the calculation entirely.

Prospect theory (Kahneman & Tversky, 1979b) adds a second-order effect on the human side of the decision. Losses loom larger than equivalent gains, and outcomes are judged against a reference point. A sponsor who has publicly committed to a point estimate is therefore reluctant to revise it downward, so the estimate tends to stick even after it has been shown to be wrong.

Reference-class forecasting as a corrective

The practical fix is to forecast the class of projects this one belongs to, instead of the project itself (Flyvbjerg, 2006). The procedure has three steps and needs no new methodology.

Define the class narrowly enough to tell you something and broadly enough to have members. "Automation projects" is too broad. Something like "internal process automations touching one ERP and one external system, delivered by this team" works. For each past member, record the planned and actual duration, the planned and actual delivery cost, and the fraction of forecast benefit actually realized twelve months after go-live. Teams almost never keep that last one. Then apply the empirical distribution of ratios to the current inside-view estimate as an uplift, and skip the argument about whether this project is different.

Hubbard (2007) makes this workable at small scale. In his framing a measurement reduces uncertainty and does not have to remove it, and the only measurements worth making are those that could change the decision. Five observations from your own history are enough to put useful bounds on a median, and that tells you more than a confident number with no history behind it.

A Monte Carlo model of automation ROI

With a decomposed model and reference-class priors, the simulation is short. The pseudocode below is the whole method.

for i in 1..N:
    F ~ Triangular(low, mode, high)        # runs per year
    H ~ Triangular(low, mode, high)        # human hours per run, today
    R ~ Truncated Normal(mu, sigma, >0)    # realizable loaded hourly cost
    A ~ Beta(a, b)                         # share of work absorbed
    D ~ Lognormal(...)                     # delivery cost
    L ~ Lognormal(...)                     # go-live delay, months
    M ~ Triangular(0.10, 0.175, 0.30)      # annual maintenance rate
    benefit[i] = discounted sum of (F/12)*H*R*A for months m > L, m <= 12T
    cost[i]    = D + discounted sum of M*D/12 for months m > L, m <= 12T
    npv[i]     = benefit[i] - cost[i]
report P10, P50, P90 of npv; report Pr(npv < 0); report Pr(payback <= 12 months)

The table below is an illustrative example. It describes a hypothetical mid-sized internal automation, and every figure in it is an assumption we chose to show the method. None of it is observed data.

TermIllustrative distributionStated assumption
F, runs per yearTriangular(600, 900, 1,100)Frequency taken from a stable operational log
H, hours per runTriangular(0.4, 0.8, 1.6)Self-reported mode, right tail for rework
R, realizable €/hourTruncated Normal(28, 9)Loaded cost, discounted for non-redeployable time
A, share absorbedBeta(8, 2)Mean near 0.8, exception path stays human
D, delivery costLognormal, median €12,000Assumed median for a mid-sized build, right tail for scope growth
L, delay in monthsLognormal, median 2, P90 near 6Right-skewed per reference class
M, maintenance rateTriangular(0.10, 0.175, 0.30)Assumed band, deliberately wide

Reading the output matters more than the arithmetic. Use the median as the headline figure, not the mean. Give P10 as the number the sponsor should be prepared to live with. Report Pr(NPV < 0) explicitly as well, because a project with an attractive median and a twenty-five percent chance of loss is a very different proposal from one with the same median and a three percent chance. If the mean and median of the output are far apart, that gap is itself the finding. It tells you the case rests on the upper tail. A tornado-style sensitivity check then shows which input, varied across its own range, moves the output most, and so where the next measurement should go. That is exactly Hubbard's criterion.

One practical note on the D term. Whatever floor a team uses for the smallest engagement it will take on is useful in this model only as a lower bound. It tells you what a narrow workflow cannot cost less than, and nothing about what a given project will cost. Using a commercial floor as a delivery estimate is a common way to smuggle an unjustified point estimate back into a distributional model.

The option value of a staged rollout

A staged rollout is better thought of as a sequence of options than as a slower version of one project. Stage one buys an observation of the terms the model is least sure about, usually A and L, and it buys the right, but not the obligation, to fund stage two. When the output distribution is wide, that right is valuable, because it cuts off the left tail. The bad draws get abandoned at pilot cost instead of full cost.

Optionality has a price, and the price belongs in the model. Staging adds throwaway scaffolding, duplicated integration work and a second round of stakeholder attention. The most expensive item is deferred benefit, which the discount factor charges for directly. A workable decision rule is to stage when the pilot really does measure the uncertain term, when the pilot costs much less than the full build, and when the full build's downside is large enough that being able to walk away is worth something. Our own QC work followed this pattern. We proved the validation pipeline first, and only after that built a follow-on pre-press tool on top of it that converts approved designs into press-ready montage files.

The tail matters for the abandon decision itself (Taleb, 2007). If the plausible failure mode is unbounded, for example a compliance exposure or a silent data corruption that spreads for months, then no median-based ROI justifies going ahead without a staged, reversible path, because the quantity being averaged has no stable mean.

Maintenance and drift as a recurring liability

Maintenance usually ends up as a footnote. It should be a line item with its own distribution. The usual form is a planning band expressed as a percentage of build cost per year, and the band should be wide, because several different mechanisms feed it. Rule-based automation degrades when the systems underneath it change: a selector moves, an export format gains a column, an API deprecates a field. These are the ordinary entries in a failure taxonomy for robotic process automation, and each one turns directly into maintenance hours. Model-based components degrade differently, through drift in the input distribution. That is worse, because it is silent and shows up as a loss of quality instead of as an exception.

Capitalizing the liability changes the rankings. Treated as a perpetuity at rate M on build cost D with discount rate r, the present value of maintenance is M × D / r. At M = 0.175 and r = 0.10 that is 1.75 × D, so the ongoing obligation is larger than the build. The assumptions are explicit and strong: constant rate, no growth, infinite life. Even cut back to a five-year horizon, the term is large enough to reverse the order of two projects that looked equivalent on delivery cost alone.

Automation earns its maintenance through lower variance as well as lower mean time. In our automated QC pipeline for a collectible card producer, 90 or more validation checkpoints turned an eight-hour manual pass into roughly fifteen minutes, with about 95% of the QC process automated and production running around 300% faster. The time saving is what gets quoted. Over the life of the pipeline, what matters just as much is that the checkpoint set does not get tired at hour seven.

What not to automate on ROI grounds

Four categories fail the test on their own. Low-frequency, low-duration work, where F × H is so small that no plausible A recovers D. Work at a non-bottleneck, where recovered time does not convert (Goldratt, 1984). Processes that are already scheduled to change, since automation locks in the current process and makes it more expensive to alter. And broken processes, where automation just produces the same errors at higher throughput.

There is one case that runs the other way. A process with a modest mean saving but highly variable manual output, such as inspection, reconciliation or compliance checks, can be a strong candidate even when the hours look unimpressive. The benefit is in the tail of defects that no longer happen, more than in the hours recovered.

Limitations

We have not shown that distributional forecasting leads to better automation outcomes. There is no controlled comparison here, and none of the numbers in the worked example are measurements. The first-party figures (the QC results and the RPA run frequency) come from our own engagements and are not a sample you can draw general parameters from. The Monte Carlo model also assumes independent inputs, which is false in practice. Delivery cost and delay move together, and both move with the share of work actually absorbed, so a naive simulation understates the joint downside. Reference-class forecasting needs a class, and small teams may not have one. The method then becomes a structured guess with its assumptions in plain view, which is more transparent but not more accurate. Finally, this model leaves out benefits that are hard to put a price on, such as lower key-person risk, faster iteration and staff time freed for more important work. A decision rule built only on the quantifiable terms will keep under-investing in those.

References

  • Flyvbjerg, B. (2006). From Nobel Prize to Project Management: Getting Risks Right.
  • Goldratt, E. M. (1984). The Goal.
  • Hubbard, D. W. (2007). How to Measure Anything.
  • Kahneman, D. (2011). Thinking, Fast and Slow.
  • Kahneman, D., & Tversky, A. (1979a). Intuitive Prediction: Biases and Corrective Procedures.
  • Kahneman, D., & Tversky, A. (1979b). Prospect Theory: An Analysis of Decision Under Risk.
  • Little, J. D. C. (1961). A Proof for the Queuing Formula: L = λW.
  • Savage, S. L. (2009). The Flaw of Averages.
  • Taleb, N. N. (2007). The Black Swan.