Statistical process control (control charts, capability indices, acceptance sampling) is taught as a set of statistical tools, but the assumption underneath it is economic. Inspection costs money, so only a fraction of output can be examined. Machine vision drives the cost of inspecting one more unit toward zero, and that removes the premise the sampling compromise was built on. Below we restate the compromise formally, look at what changes when every unit is measured instead of only judged, and argue that the binding constraint then moves from detection to specification. Whether that move works out depends on two things: measurement system analysis, and the economics of false alarms.
The economics of acceptance sampling and inspection cost
Shewhart framed the control chart as an economic instrument from the start (Shewhart, 1931). The aim was to get the most usable information out of each unit of inspection effort. Acceptance sampling comes from the same line of thinking. The published attribute plans (ISO 2859-1) are compromise tables indexed by lot size and acceptable quality level, and they exist because examining every item was assumed to be uneconomic.
The trade-off fits in one line. Let N be lot size, f the inspected fraction, c_i the cost of inspecting one unit, p the fraction of units carrying a defect, c_e the expected consequence of one escaped defect, and m the probability that inspection misses a defect that is present. Uninspected units escape with probability one, and inspected units escape with probability m:
C(f) = f · N · c_i + N · p · c_e · [1 − f · (1 − m)]
The derivative with respect to f is the constant N · c_i − N · p · c_e · (1 − m), so the model has no interior optimum. Inspect everything when c_i < p · c_e · (1 − m), otherwise inspect nothing. Classical sampling plans get their interior optima from two features the model leaves out, and both come from human inspection. Inspection capacity is a limited shared resource, so c_i rises steeply once the station becomes the bottleneck. And accept/reject is decided per lot rather than per unit, so a sample informs a decision about items nobody examined. A system that evaluates every unit in milliseconds has neither feature. As c_i approaches zero the inequality holds for any positive defect rate and escape cost, and the optimum becomes full inspection in every case. In this model, cost was the only argument for sampling.
Assumptions behind the Shewhart control chart
A Shewhart chart is a hypothesis test run over and over. Limits sit roughly three standard deviations from the centre line, with the standard deviation estimated from within-subgroup variation, and the chart asks whether between-subgroup variation is larger than within-subgroup variation predicts. Four assumptions follow. The process is stationary, so the parameters being estimated exist and stay put. Observations are effectively independent, because positive autocorrelation deflates the within-subgroup estimate and floods the chart with false alarms. Subgroups are rational, meaning common causes act inside a subgroup and special causes act between subgroups. Finally, the limits need a baseline, conventionally twenty to twenty-five subgroups collected while the process is behaving.
High-variety, short-run production breaks the first and the last of these by its nature. If every job is a different design on different stock, there is no single process to characterise, only a series of short-lived ones that are retired before anyone could estimate their limits. Wheeler (1993) treats the discipline as interpreting variation rather than computing statistics, and that reading is what keeps it useful in short runs. Limits computed for one job say little about the next one, but the habit of separating signal from noise carries over. Deming (1986) warned against tampering, which means adjusting the process in response to common-cause variation and so raising its variance. Short runs invite exactly that, since without a baseline every job looks like a special cause.
From pass/fail verdicts to measured distributions
The naive reading of universal inspection is that it catches every defect. The more useful point is that each unit now yields a vector of measurements instead of a verdict. A unit produced at the centre of its tolerance and one produced 0.01 mm inside the boundary get the same verdict, but as data they are very different.
Taguchi's quadratic loss function was formulated to recover this information (Taguchi, 1986). Instead of a step at the specification limit, loss is modelled as L(y) = k · (y − τ)², which grows with the squared deviation of characteristic y from target τ. Sampled data supports little beyond the fraction out of specification, so the step function is the only model it can carry. With full measurement, E[L] = k · [(μ − τ)² + σ²] can be computed per job, and its two terms line up with the two things you can do about it: re-centre the process, or reduce its variance.
What gets monitored changes too. The question on the line becomes whether the distribution has moved, which means watching the mean of a characteristic across a run, its dispersion, the share of units close to a tolerance boundary as an early warning, and capability as C_p = (USL − LSL) / (6σ̂), with C_pk accounting for off-centring. A run whose C_pk is falling while its reject count is still zero is a signal that sampling cannot give you.
The automated quality pipeline we built for a collectible card producer is one example. It runs 90+ validation checkpoints per design and automates roughly 95% of what used to be an eight-hour manual pass, which now takes about fifteen minutes. Throughput went up by roughly 300%. For this article the more interesting change is that checkpoints which used to give subjective verdicts now give numbers on every design. Drift in incoming artwork shows up gradually, as a shifting distribution. With verdicts alone it would only have appeared once it produced a cluster of rejections. A follow-on pre-press tool converts approved designs into press-ready montage files. That closes a second way for defects to escape: a design that is correct on its own but imposed wrongly on the sheet.
One caveat belongs here. Automated inspection errors are systematic. A human inspector's misses are roughly independent across units, while a misconfigured detector misses the same feature on every unit. The upside is that a systematic error only has to be found once, and the fix then covers every later unit.
Specification as the binding constraint on automated inspection
When detection is nearly free, deciding whether a unit is acceptable comes down to having a definition of acceptable. Full inspection moves the bottleneck to specification, and most organisations find out on the day the cameras go live that nobody ever agreed what "good" means.
A usable specification has three properties that informal quality language lacks. First, it is numeric. "Colours match" becomes a ΔE tolerance, "registration is clean" becomes a maximum misregistration in millimetres, and "elements are placed properly" becomes a minimum distance from the trim edge. It also states its measurement condition, since the same sample gives different colour numbers under different illuminants and backings. And it names an owner who is allowed to change it, because a tolerance nobody owns is either never tightened or quietly relaxed under schedule pressure.
Tolerances like these can be written down, and the industry already does it. The offset process control standards (ISO 12647-2) specify aim values for solid ink colorimetry and tone value increase, and ISO 9001 carries the general requirement that inspection criteria be defined. Only the numbers are local to a plant. An enforceable specification looks like configuration:
registration_tolerance_mm: 0.20
delta_e_2000_max: 3.0
trim_safe_margin_min_mm: 2.0
min_effective_ppi: 300
on_failure: hard_stop # hard_stop | route_to_review | auto_repair
The last line is the one people most often leave out. If the specification says nothing about what happens when it is violated, the decision falls to whoever is standing at the machine, and the variation the specification was written to remove comes back.
Measurement system analysis for machine vision
Observed variance is a sum: σ²_obs = σ²_process + σ²_measurement. A chart fed by a measurement system nobody has characterised cannot tell a process shift from gauge drift, and it will blame the process for the drift. The usual screen expresses measurement variation as a proportion of tolerance width. When that proportion is large, fix the gauge instead of charting the process.
For vision the components are unfamiliar, but you can list them. Repeatability is assessed by presenting the same unit again and again with nothing changed. Reproducibility is assessed across whatever the system treats as interchangeable, such as stations, shifts, lighting states, cameras and lenses. Bias and linearity are checked against calibrated reference targets, and those targets age too. Some contributors have no counterpart in mechanical gauging at all: focus and depth of field, lens distortion and its correction, resampling, lossy compression applied upstream, and the colour transform chain from sensor response to reported value. Each of them adds to σ²_measurement. None of them stays constant on its own, so someone has to decide to hold each one fixed.
False-alarm economics and operator trust in inspection systems
What decides whether operators actually use an inspection system is mostly its alarm rate. If a unit passes k independent checks each with per-check false-alarm probability α, a conforming unit raises at least one alarm with probability 1 − (1 − α)^k. The figures below are a worked example built on stated assumptions, and none of them come from observed data. With k = 90 and α = 0.001, a rate most engineers would call negligible, the unit-level false-alarm probability is about 8.6%. Holding unit-level false alarms to 1% across 90 checks takes an α of roughly 1.1 × 10⁻⁴.
Base rates make it worse. Assume a true defect prevalence of 2%, detection sensitivity of 0.95, and the 8.6% false-alarm rate above. The share of alarms that correspond to real defects is (0.02 × 0.95) / (0.02 × 0.95 + 0.98 × 0.086), about 18%. Four alarms in five are spurious, and operators will predictably stop reading them.
Human-in-the-loop design under automation bias has to plan for this failure. Parasuraman and Riley distinguish disuse, where people reject or disable automation, typically because of false alarms, from misuse, the complacent over-reliance that follows unwarranted trust (Parasuraman & Riley, 1997). An unmanaged alarm budget pushes operators toward disuse, and alarms that nobody audits invite misuse. The fixes are structural. Set the false-alarm budget at unit level and divide it across the checks. Separate hard-stop rules from advisory ones, and let only the hard stops halt production. Log every override with its reason, so that a rule that fires too often gets retuned instead of tolerated. Juran insisted that quality control be organised as a managed function with defined responsibilities (Juran, 1951), and the rule set needs the same treatment.
Decision rules for inspection design
- Estimate
c_i, the cost of inspecting one additional unit, before designing any sampling plan. If it is near zero, the plan solves a problem you no longer have. - Do not chart a characteristic you cannot specify numerically. It needs a specification decision first.
- Complete measurement system analysis before the first chart, and repeat it after any change to lighting, optics, or the image processing chain.
- Chart the measured variable. Charts of pass/fail verdicts throw information away and react late.
- Budget false alarms at unit level and allocate that budget explicitly across checks.
- Keep sampling only where measurement is genuinely destructive, slow, or costly.
- Treat the specification as a controlled document. A tolerance edit is a process change and belongs on the timeline beside the data it explains.
Limitations of this analysis
This is an argument about cost structure, and it goes no further than that. It does not show that machine vision can detect any particular defect class. Detectability has to be tested characteristic by characteristic, and characteristics that are contextual, semantic or aesthetic may not be specifiable in numbers at all. For those, the specification bottleneck described here is a permanent limit, and more work will not clear it. The model assumes that defect opportunities are independent and that consequence costs add up, and both are simplifications. The numbers in the worked examples are assumptions picked to show the arithmetic. Our own figures come from one engagement in one industry. Last, nothing here shows that full inspection improves a process. It changes what you can see. A process improves when someone acts on its causes, and more thorough measurement of output does not do that by itself (Deming, 1986).
References
- Shewhart, W. A. (1931). Economic Control of Quality of Manufactured Product.
- Deming, W. E. (1986). Out of the Crisis.
- Juran, J. M. (1951). Quality Control Handbook.
- Taguchi, G. (1986). Introduction to Quality Engineering.
- Wheeler, D. J. (1993). Understanding Variation: The Key to Managing Chaos.
- Parasuraman, R., & Riley, V. (1997). Humans and Automation: Use, Misuse, Disuse, Abuse.
- International Organization for Standardization. (1999). ISO 2859-1: Sampling procedures for inspection by attributes.
- International Organization for Standardization. (2013). ISO 12647-2: Process control for the production of half-tone colour separations, proof and production prints — Offset lithographic processes.
- International Organization for Standardization. (2015). ISO 9001: Quality management systems — Requirements.