When people call a business process slow, the diagnosis is almost always about execution speed: a step takes too long, so make the step faster. But in most processes that have reached the point of complaints, cases spend most of their lead time waiting, and speeding up the tasks themselves barely registers. This article states Little's Law precisely and uses Kingman-style variability reasoning to show why lead time becomes unstable above roughly 80 percent utilisation. It then argues that reducing variability, removing handoffs and limiting work in progress do much more than removing keystrokes. The same arithmetic lets you split a lead time into touch time and waiting time, estimate how sensitive the wait is to utilisation, and tell in advance whether a proposed automation will change anything measurable.
Little's Law Stated Precisely
Little's Law (Little, 1961) relates three quantities of a queueing system observed over a long interval: the average number of items in the system, the average arrival rate, and the average time an item spends in the system.
L = λ · W
L = average work in progress (items in the system)
λ = average arrival rate (items per unit time)
W = average time in system per item (lead time)
The law is a conservation identity and does not model any particular queue. It assumes nothing about the arrival or service distributions, the number of servers, the queue discipline, or whether items are handled singly or in batches. That generality is what makes it usable on business processes, where none of those things are known.
Its assumptions still matter, because most misuse comes from ignoring them. The system has to be observed over a period in which it is broadly stationary: the arrival rate is not trending, and the queue is neither growing without bound nor draining to nothing. The boundary of the system has to be defined, and every item that enters must eventually leave and be counted. Cases that are silently cancelled, absorbed into another workstream or reclassified break that count. All three quantities must be averages over the same interval and the same population.
Rearranged, the law becomes a diagnostic. W = L / λ says lead time depends on how much work is in progress relative to how fast work leaves. That is a statement about inventory, and effort does not appear in it. If the number of open cases doubles while the completion rate stays the same, each case takes twice as long, even though nobody is working any slower.
Touch Time and Lead Time Are Different Measurements
Lead time is the elapsed interval from the moment a case enters the defined boundary to the moment it leaves. Touch time (also called work content or process time) is the sum of the intervals during which someone or something is actively advancing that specific case. Their ratio is flow efficiency, and in office and knowledge work it is routinely in the single digits.
The two are also measured differently, and this is where process documentation usually goes wrong. Ask people how long a step takes and they report touch time, because that is the part they experience. Lead time can only come from timestamps on state transitions, with separate marks for when a case became available to a step and when someone started working on it. Without that separation, queueing time gets booked as service time, and the model can't tell a slow worker from a long line. Those transitions can be recovered from system event logs instead of interviews, which is the argument for treating process mining as a prerequisite for automation.
Our automated quality-control pipeline for a collectible card producer shows the gap up close. The manual process it replaced was described as an eight-hour pass, and very little of that time was inspection. Most of it was files sitting in an inbox, a designer waiting until a reviewer was free, a batch held back until there were enough items to justify a session, and one artefact passing between three people who each needed context the others had. The automated version runs more than 90 validation checkpoints and finishes in roughly 15 minutes, with about 95 percent of the process automated. Most of the gain comes from removing the waits and handoffs between checks, and much less from faster checking. The checkpoint design itself is a separate question, covered in what automated quality control catches that human inspection cannot.
The Utilisation Cliff: Why Queueing Time Explodes Above 80 Percent Load
Waiting dominates because queueing time grows non-linearly with load. For a single-server queue in heavy traffic, Kingman's approximation (Kingman, 1961) gives the expected wait before service as
Wq ≈ ( ρ / (1 − ρ) ) · ( (ca² + cs²) / 2 ) · τ
ρ = utilisation (arrival rate ÷ service capacity)
ca = coefficient of variation of inter-arrival times
cs = coefficient of variation of service times
τ = mean service time (touch time)
The middle factor is variability and the last is work content. The first factor is the one that wrecks lead times, because ρ / (1 − ρ) grows without bound as utilisation approaches one. With variability and touch time held fixed, the multiplier looks like this.
| Utilisation ρ | Queue multiplier ρ / (1 − ρ) | Relative wait vs. 50% load |
|---|---|---|
| 0.50 | 1.0 | 1× |
| 0.70 | 2.3 | 2.3× |
| 0.80 | 4.0 | 4× |
| 0.90 | 9.0 | 9× |
| 0.95 | 19.0 | 19× |
| 0.98 | 49.0 | 49× |
Between 50 and 80 percent load the wait quadruples, and between 80 and 95 percent it nearly quintuples again. This is the arithmetic behind the claim that keeping every resource busy ruins flow. If you staff a review function to 95 percent utilisation, you accept a nineteen-fold waiting penalty to get rid of idle time that never showed up on the balance sheet anyway. Reinertsen (2009) treats this trade-off as the central economic error in managing product development. Ohno (1988) insisted on stopping the line rather than keeping it loaded, which is the same point applied on the shop floor: above the cliff, keeping each station busy works against the throughput of the whole system.
That makes utilisation a poor dashboard metric. It lags, it is non-linear, and on a bar chart a step at 82 percent looks much like one at 94 percent, although the two behave completely differently. Queue length and the age of the oldest item in the queue start moving earlier, before the cliff is reached.
Arrival and Service Variability as Hidden Drivers of Waiting Time
The middle factor of Kingman's formula is the one most organisations never measure, and it scales the whole waiting time. Halving both arrival and service variability, in coefficient-of-variation terms, cuts waiting by roughly a factor of four at the same utilisation and the same touch time, without anyone working faster.
Arrival variability
Organisations mostly create their own arrival variability. Upstream batching, month-end cycles, campaigns that release work in bursts, and priority overrides that reorder the queue all raise the arrival coefficient of variation. Deming's (1986) distinction between common-cause and special-cause variation applies directly here. Expediting individual late cases raises variability. Higher variability raises the average lead time, and that produces more late cases. What helps is levelling the release of work so that it arrives at an even rate.
Service variability
Service variability comes from mixed case types sharing one undifferentiated queue, from missing inputs that force a case to be set down and picked up again, and from rework loops. A rare class of case that takes ten times the mean can dominate the variance term on its own, and moving that tail into its own stream is usually cheaper than shortening it. This is the operational version of tail-latency reasoning in distributed systems, where Kleppmann (2017) shows that high percentiles, more than means, govern what users actually experience. Nielsen's (1993) response-time limits (about 0.1 second for direct manipulation, 1 second for uninterrupted thought, 10 seconds for holding attention) apply to human steps too. What people notice there is the slow tail, much more than the average.
Where Automation Actually Helps a Queue-Bound Process
If waiting dominates lead time, and utilisation and variability dominate waiting, then automating τ has limited value. Cutting touch time does lower utilisation, which helps a process sitting near the cliff. But take a process at 60 percent utilisation with 4 percent flow efficiency. An initiative there that is justified only by faster data entry will deliver almost nothing anyone notices. That is one reason business cases built on hours saved tend to overstate their returns, which I take up in modelling automation returns under uncertainty.
In a queue-bound process, automation pays off through three mechanisms, listed here from most to least valuable. The first is removing handoffs. Every transfer between people or systems creates a queue, so removing a handoff removes a whole waiting stage rather than shortening one. The second is smaller batches. Batching is the cheapest way to create both work in progress and arrival variability, and software has no economic reason to batch. Third, a deterministic automated step has a service coefficient of variation near zero, which shrinks the variability factor for everything downstream of it.
An RPA workflow we built for a card printing company is a clear case of the second mechanism. It pulls order data from the ERP system, formats shipping labels and pushes them to fulfilment, and it runs more than 800 times a day. It does format a label faster than a person would, but the bigger effect is that the economic batch size dropped to one. Nobody holds labels back any more until a run is worth someone's afternoon, so that queue never forms, and work reaches fulfilment at an even pace instead of in lumps.
Work-in-Progress Limits as the Cheapest Lead-Time Intervention
Little's Law also gives you the cheapest lever, and the only one you can pull with a policy instead of engineering work. If W = L / λ and demand sets λ, then capping L caps W directly. A work-in-progress limit is a management decision. You can put one in place in an afternoon, and it needs no software and no extra headcount.
The side effects are worth even more than the arithmetic. A binding limit makes the constraint visible, because work stops piling up in front of it and gets refused at the entry point instead. That is how Goldratt's (1984) theory of constraints finds a bottleneck without a study. A limit also swaps a cost nobody sees, cases sitting in a full queue, for one that everybody notices and finds uncomfortable: a person with nothing to do. In practice, set the cap near the current average WIP, lower it in steps, and stop when queues stop getting shorter or the constraint starts running out of work.
A Worked Illustrative Example: Diagnosing an Approval Process
The figures below are constructed to show the arithmetic and do not come from a real engagement. I state the assumptions so you can put in your own.
Assume an approval process with six sequential steps, arrivals of 20 cases per working day, and a measured average of 60 cases open at any time. Little's Law gives W = 60 / 20 = 3 working days. Suppose reported touch times add up to 55 minutes. Flow efficiency is then 55 minutes against three eight-hour days, about 3.8 percent.
Compare two interventions. Automating the two largest steps to remove half of all touch time takes roughly 27 minutes off a lead time of 1,440 minutes. That is about 1.9 percent, and nobody will see it against week-to-week noise. Capping WIP at 30 cases, with the completion rate unchanged, gives W = 30 / 20 = 1.5 days, a 50 percent reduction from a policy change that costs nothing to develop. There is also a third option. If the fourth step runs at 0.93 utilisation, shifting a modest amount of capacity to bring it down to 0.80 cuts its queue multiplier from about 13 to 4, and that typically outweighs both other effects.
The decision rule is mechanical. Measure flow efficiency first. Below roughly 15 percent the process is queue-bound, and touch-time automation has to be justified by handoff or batch-size reduction rather than by hours saved. Above 40 percent, work content really does dominate and task-level automation is the right target. In between, measure utilisation per step and fix anything above 0.85 before writing code.
Limitations
The analysis is a way to reason about lead time, and it makes no forecast. Little's Law holds wherever its boundary and stationarity conditions hold. Kingman's expression is an approximation that fits heavy traffic and single-server queues best. Multi-server steps, priority rules, blocking and workers shared across processes all deviate from it, sometimes by a lot. The table is arithmetic on the utilisation factor alone and predicts nothing about a real process's waiting time, and the worked example is constructed, with every assumption stated. The first-party figures are our own results in their specific settings. When we attribute the drop from eight hours to fifteen minutes mainly to removing waits and handoffs, that is our engineering judgement about that project, and we did not run a controlled measurement. Touch-time reduction still has value, but flow efficiency caps it, so a project should know the flow efficiency of the process before it estimates its own return.
References
- Deming, W. E. (1986). Out of the Crisis.
- Goldratt, E. M. (1984). The Goal.
- Kingman, J. F. C. (1961). The Single Server Queue in Heavy Traffic.
- Kleppmann, M. (2017). Designing Data-Intensive Applications.
- Little, J. D. C. (1961). A Proof for the Queuing Formula: L = λW.
- Nielsen, J. (1993). Usability Engineering.
- Ohno, T. (1988). Toyota Production System.
- Reinertsen, D. G. (2009). The Principles of Product Development Flow.