Classical cost-of-quality reasoning assumes a production run long enough for prevention spending to amortise across many near-identical units. Short-run production breaks that assumption. When the effective run length gets close to one and every unit has its own specification, it no longer follows automatically that prevention beats detection. This article sets out an explicit defect-escape cost model, shows why the jump from internal to external failure cost gets larger when runs are short, and derives a few rules a bespoke producer can use to decide where its first quality budget goes.
Prevention, appraisal and failure costs when the run length is one
The split of quality cost into prevention, appraisal, internal failure and external failure comes from the mid-century quality literature (Juran, 1951; Feigenbaum, 1961). Its force rests on one observation. The categories act as substitutes for one another, and the cheap ones sit upstream. Crosby's claim that quality is free (Crosby, 1979) is the strong form of the same argument. ISO 9001 builds the structural version into a standard, requiring that process controls be planned in advance instead of discovered after the fact (International Organization for Standardization, 2015).
Every step of that argument depends on run length, because each category has a different ratio of fixed to variable cost.
| Category | What the spending buys | Behaviour as run length approaches one |
|---|---|---|
| Prevention | Specifications, tooling, jigs, training, design rules | Largely fixed per specification, so with one unit per specification it cannot amortise |
| Appraisal | Inspection, measurement, review of the unit itself | Largely variable per unit, and unchanged per unit as run length falls |
| Internal failure | Scrap, rework, re-setup before shipment | Rises, because a bespoke unit cannot be replaced from stock |
| External failure | Returns, reprints, expedited freight, lost future orders | Rises sharply, because the escaped unit is the entire order line |
The problem sits in the right-hand column. Classical reasoning wants to put the money into prevention, and prevention is the category whose economics collapse when the denominator is one. Appraisal is per-unit by nature, so run length does not change it. That makes it the only upstream lever whose cost per unit does not climb as runs get shorter.
A defect-escape probability model for unique units
An escape is a non-conformance that leaves the plant undetected. To reason about escapes you have to separate three things that informal discussion tends to lump together: how often a defect is generated, how likely the control system is to catch it, and what it costs when it is missed.
The model
Let defect classes be indexed by j: wrong bleed, missing overprint, transposed card back, out-of-tolerance colour, mislabelled carton. For a single unique unit, write the expected escape cost as
E[C_esc] = Σ_j λ_j · (1 − D_j) · L_j
L_j = C_rework + C_material + C_delay + C_reputation
D_j = 1 − Π_i (1 − d_ij)
Here λ_j is the incidence of class j per unit, d_ij is the detection probability of checkpoint i against class j, D_j is the aggregate detection probability of the checkpoint chain, and L_j is the loss if the defect escapes, broken down into rework, wasted material, delay and reputational components.
Two properties drive everything that follows. First, escape probability multiplies across checkpoint misses, so adding an independent checkpoint of modest sensitivity cuts escapes faster than improving an existing one. That is the reasoning behind poka-yoke devices at each operation instead of a single final gate (Shingo, 1986). Second, the sum runs over classes. A system that is excellent against three classes and blind to a fourth gets almost all of its escape rate from the fourth, so while any class is still uncovered, covering more classes matters more than making one checkpoint more sensitive.
Estimating the parameters
You can estimate λ_j from job history if non-conformances are logged by class and not as one undifferentiated defect count. d_ij is harder, and most estimates go wrong here. The reliable method is seeded-defect testing. Put known non-conformances of each class into files or units, route them through the live control chain, and record what fraction the chain rejects. Without that test, a checkpoint's sensitivity is a guess. Human visual checks in particular get worse the longer someone has been on the task, and a single fixed number does not capture that. Classes defined by a continuous tolerance instead of a yes-or-no rule, colour being the obvious one, also need a fixed acceptance threshold before a seeded test means anything. Setting that threshold is a question of perceptual colour tolerancing, separate from detection.
The last term of L_j is a policy decision more than a measurement. Rework, material and expedited-freight costs can be pulled from accounting records, but nobody observes the reputational component directly. In practice, state it as an explicit assumption and check how sensitive the conclusion is to it. Setting it to zero is the modelling choice that quietly reproduces the behaviour the quality literature has criticised for decades (Deming, 1986).
Why short production runs amplify failure-cost escalation
A rule of thumb that circulates widely in the quality literature says a defect costs on the order of one unit of effort to fix at design, ten at production and a hundred once it reaches the customer. It is best read as a rough ordering, not a measured ratio. What matters here is that both ends of the ordering move against the producer when runs are short.
At the upstream end, prevention cost per unit is F/N for a fixed prevention investment F over a run of N units. As N falls to one, prevention per unit equals the whole of F. Write a design rule for a single card face, build a jig for one carton size or brief an operator on one job, and the full fixed cost has to be recovered from one unit of output. Long-run manufacturing hides this because N is large enough for F/N to disappear into overhead.
At the downstream end, L_j grows. In volume production an escaped defect is one unit out of thousands. It is replaced from inventory, the order is still filled, and the cost is scrap plus handling. In bespoke production the escaped unit is the order. There is nothing in stock to replace it with, the replacement needs a new setup at full setup cost, and the delay pushes back a delivery date that may be the reason the customer chose this supplier. In volume work the delay and reputational terms are rounding errors. Here they become as large as the material term, or larger.
Both movements point the same way. The only category whose economics hold up is appraisal, and only if its per-unit cost can be pushed close to zero. That is a question of automation, because more diligence will not get the cost there. A check whose marginal cost is mostly human time cannot be run on every unique unit at an acceptable cost, but one whose marginal cost is compute can. This is the economic reason automated quality control catches what human inspection cannot in high-variety work: it can afford to check every unit.
Worked example: escape cost under manual versus automated appraisal
The numbers below are constructed for illustration from stated assumptions, and none of them is an observed result. Take a bespoke print job of 250 distinct card faces produced once. Suppose a spec-conformance defect class has an incidence of λ = 0.02 per face, so five defective faces are expected per job. A manual review pass detects 70% of them and an automated rule chain detects 98%. The loss per escape is €200 rework, €450 wasted material, €650 expedited freight and schedule recovery, plus a stated reputational allowance of €500, giving L = €1,800.
Expected escape cost under manual review is 5 × 0.30 × €1,800 = €2,700 per job. Under the automated chain it is 5 × 0.02 × €1,800 = €180 per job. The avoided escape cost is €2,520. The appraisal cost differs as well. An eight-hour manual pass at €40 per hour costs €320, against roughly €10 of compute and supervision for the automated pass, which saves a further €310. Total avoided cost is €2,830 per job.
If the automated chain costs €12,000 to build, breakeven is F / Δ = 12,000 / 2,830 ≈ 4.2 jobs. A single breakeven point computed from point estimates is optimistic in the usual way, and every parameter above is uncertain, so the figure is better read as a distribution. Modelling automation ROI under uncertainty covers how to do that. The fixed cost that could not amortise over the units in one job amortises over the jobs instead, because the control logic can be reused and the product cannot. That is how the short-run problem gets resolved: move fixed spending out of anything unit-specific and into a control system whose denominator is the number of jobs.
Our automated QC pipeline for a collectible card producer follows the same pattern. It carries 90+ validation checkpoints and automates roughly 95% of the QC process. An eight-hour manual pass became about 15 minutes, and production ran about 300% faster. Escapes mattered more to the economics than throughput did: wasted materials and delayed shipments on jobs where every SKU is bespoke. A companion pre-press tool converts approved designs into press-ready montage files, which removes a manual re-handling step at the point where late-stage escapes were most expensive. A separate RPA workflow for a card printing company pulls ERP order data, formats shipping labels and pushes them to fulfilment, running 800+ times a day. Its value is of the same kind: it removes a class of per-order transcription errors that only surface at the customer.
Taguchi's loss function: defect cost as a curve rather than a step
Cost-of-quality arithmetic tends to treat conformance as binary: inside tolerance costs nothing, outside tolerance costs L. Taguchi's loss function (Taguchi, 1986) drops that framing and models quality loss as growing continuously with deviation from target, approximately L(y) = k(y − T)² for a characteristic y with target T. A unit at the edge of tolerance already carries some loss. Where a specification limit sits is a commercial decision about how much accumulated loss is acceptable, and nothing physical changes at that line.
For short runs this matters in two ways. Loss is quadratic, so barely acceptable units cost disproportionately when they cluster, and in bespoke work they cluster by job, because a job shares one setup, operator and substrate. Also, with no long run to average over, the unit shipped is the sample. Process capability computed over volume production has no direct equivalent when the run is one. This is where the usual statistical process control apparatus meets machine vision and has to be rebuilt around evidence from each unit.
Where to spend the first quality budget
The model gives a handful of decision rules.
Spend first on the class nobody checks for, and only then on the weak checkpoint. Escape probability multiplies within a class and adds across classes, so whichever class has no checkpoint at all dominates total escape cost. Listing the classes and confirming that each has at least one control is worth more than tightening a control that already works.
Automate appraisal before adding prevention. Prevention written for one specification cannot amortise over a run of one, while appraisal automation amortises over jobs. That reverses the usual order, but only because of how the costs are structured. Prevention is still worth having.
Put controls at the constraint. Inspection uses capacity, and capacity spent at a bottleneck is lost output (Goldratt & Cox, 1984). A defective unit that passes through the constraint has used capacity that cannot be recovered, so detection belongs immediately upstream of it.
Prefer many cheap, independent checks to one authoritative gate. This is source inspection restated in economic terms (Shingo, 1986; Ohno, 1988).
Log by class from the start. None of the parameters above can be estimated unless non-conformances are recorded as typed events, and a single defect count supports none of the decisions the model can express.
Limitations and threats to validity
The model describes a cost structure and does not measure an effect. Its central claim, that automating appraisal beats adding prevention as run length approaches one, rests on assumptions that will not always hold. The multiplicative form assumes checkpoints are independent, and real ones are correlated. A bad input file can get past several rule-based checks at once, so the model overstates detection. The reputational term in L_j is an assumption, and conclusions near breakeven depend heavily on it. Seeded-defect testing only measures performance against the classes someone thought to seed. It tells you nothing about classes nobody anticipated, and those are the expensive ones.
The worked example was built to make the arithmetic easy to follow. It is not a benchmark, and none of its rates were observed. Our own figures come from specific engagements with their own defect mix, substrate and customer tolerance, and they do not carry over to other producers. Finally, the framework puts a price on escapes and leaves out the design and negotiation work that reduces λ_j at source. While the specification is still open, that work is the cheapest intervention available.
References
- Crosby, P. B. (1979). Quality Is Free: The Art of Making Quality Certain.
- Deming, W. E. (1986). Out of the Crisis.
- Feigenbaum, A. V. (1961). Total Quality Control.
- Goldratt, E. M., & Cox, J. (1984). The Goal.
- International Organization for Standardization. (2015). ISO 9001: Quality management systems — Requirements.
- Juran, J. M. (1951). Quality Control Handbook.
- Ohno, T. (1988). Toyota Production System: Beyond Large-Scale Production.
- Shingo, S. (1986). Zero Quality Control: Source Inspection and the Poka-Yoke System.
- Taguchi, G. (1986). Introduction to Quality Engineering.