On a design document, a human review step is usually recorded as a control: a person checks the machine, so the machine counts as supervised. In practice the step creates a second job, supervisory monitoring, and that job has its own error rates and its own way of getting worse over time. I argue that automation bias and skill atrophy make naive "the human checks everything" designs worse than either well-scoped automation or plain manual work. Three rules follow from that: escalate on calibrated uncertainty, give the reviewer a decision to make instead of a verdict to confirm, and keep practising the manual skills the escalation path depends on. The later sections show how to compute an escalation threshold and specify an audit slice, and how to tell when a review step has turned into a formality.
Supervisory Control Roles: Operator, Reviewer, Adjudicator
The phrase "human in the loop" covers three different jobs. An operator does the work with help from the machine. A reviewer looks at machine output and accepts or rejects it. An adjudicator decides the cases the machine has explicitly passed on. When a design doesn't say which of these it means, it has usually ended up with the weakest one, a reviewer with no authority to decide and no context from the operation.
The standard starting point for splitting functions between people and machines is Fitts (1951). The habit it established is still the default way of thinking about the problem: list what humans do well, list what machines do well, and assign tasks accordingly. Its weakness is that it treats allocation as dividing up a fixed task list, when the allocation itself changes the list. That is one reason why understanding the process before automating it has to come first. Once a machine takes over the routine part of a process, what remains for the person should be analysed as a new job, because it no longer resembles the old one with a few steps removed.
Parasuraman and Riley (1997) give us the vocabulary for the failure modes. Automation can be used, misused through overreliance, disused through unwarranted distrust, or abused through a deployment that ignores what the allocation does to the people involved. The naive review step falls in the last category: nobody ever decided that a person should classify two thousand borderline items a day, three seconds each.
Automation Bias and Complacency
Automation bias is the tendency to accept automated output as correct and to search too little for evidence against it. It shows up in two ways. In omission errors, the supervisor misses a condition the automation did not flag. In commission errors, the supervisor acts on an incorrect recommendation although contradictory information was available. Commission errors are the more worrying kind, because the evidence that would have caught the mistake was there and nobody used it.
The mechanism is what Kahneman (2011) calls substitution, where a hard question is replaced by an easier one, and it has little to do with carelessness. Answering "is this classification correct?" means reconstructing the case. Answering "does this look like what the system usually gets right?" does not. Under time pressure, with a reliable system, the substitution pays off on nearly every individual item, which is why people carry it over to the few items where it is wrong.
Complacency is the same effect stretched over time: attention drifts away from a channel that has been reliable. So a designer cannot assume a fixed level of supervisory attention. It depends on how reliable the system has looked so far, and it drops as the system gets better.
The Ironies of Automation and the Residual Task
Bainbridge (1983) described the structural problem. Automation removes the easy, well-specified cases that make up most of the volume and leaves the operator the remainder. Those remaining cases tend to be ambiguous or rare, and they carry a disproportionate share of the consequences.
Two things follow. First, estimate the difficulty of the human task from the remainder, never from the average case. If 95 percent of decisions are automated, reviewers get the hardest 5 percent, and that is a very different thing from 5 percent of the old cognitive load. Second, people are poor at monitoring for rare failures, so a programme that gives them the monitoring role has handed them the one function the allocation tradition would keep away from them.
The Vigilance Decrement in Rare-Signal Review
Mackworth's work on prolonged visual search (Mackworth, 1948) established the vigilance decrement: detection performance declines over a watch period, and it declines fastest when signals are rare. A high-accuracy system puts its reviewer in exactly that position.
For a batch of N items with per-item system error rate p and reviewer detection rate d:
escaped defects = N · p · (1 − d)
reviews per defect = N / (N · p · d) = 1 / (p · d)
The second quantity is the cost. Note that d depends on p: as p falls, the signals the reviewer is watching for get rarer, and d falls with it.
The table is a worked example with assumed inputs. It takes 10,000 items, all reviewed, with a detection rate of 0.80 at a 10 percent error rate and 0.30 at 1 percent. I picked these figures to make the structure visible, and none of them were measured.
| System error rate | Defects in 10,000 | Assumed detection rate | Escaped defects | Reviews per defect caught |
|---|---|---|---|---|
| 10% | 1,000 | 0.80 | 200 | 12.5 |
| 1% | 100 | 0.30 | 70 | 333 |
Escaped defects still go down with the better model, as they have to. The productivity of review, though, collapses: the same effort now catches one defect per 333 inspections, against one per 12.5 before. At that yield the review step mostly produces the belief that the work has been checked, and this is the sense in which 99 percent accuracy is harder to supervise than 90 percent.
Escalation Design for Automated Decisions
The alternative to reviewing everything is routing, and the routing rule should come from an explicit cost comparison rather than a percentage that felt responsible.
In our automated quality-control pipeline for a collectible card producer, roughly 95 percent of the process is automated across 90-plus validation checkpoints. The remaining 5 percent are cases where the checkpoint evidence is ambiguous. The pipeline sends those to an experienced inspector on purpose: the machines flag them and the inspector decides. The manual pass used to take eight hours and now takes about fifteen minutes, because inspectors only spend time where their judgment changes the outcome.
Calibrated confidence thresholds
A confidence score is usable for routing only if it is calibrated: among items scored 0.90, close to 90 percent should be correct. An uncalibrated score routes items by some arbitrary transform of the truth, and the escalation sets it produces are large and tell you little. Measure calibration per segment on held-out data, and measure it again on a schedule, because it drifts when the inputs drift. That puts it under the same production evaluation discipline any deployed model needs.
Once the scores are calibrated, the threshold follows from costs. Let p(x) be the estimated probability that autonomous handling of item x is wrong, C_miss the cost of that error, C_review the cost of a human decision:
escalate if p(x) · C_miss > C_review
i.e. p(x) > C_review / C_miss
That ratio is the design parameter. To state it honestly you have to cost prevention, appraisal and failure separately, the way the cost-of-quality accounting for short production runs does. If a team can't state the ratio in some unit (money, hours, reworked units, complaints), it can't defend its threshold either, and it will end up with one picked to fill whatever review capacity happens to be available.
Disagreement sampling
Confidence is only one of the signals you can route on. Suppose you have two independent checks, such as a rule engine and a model, two models with different inductive biases, or the current and previous version of a system. The items they disagree on contain hard cases at a higher rate than a confidence band does, because disagreement is evidence about the item, while a confidence score is one estimator reporting on itself. Watch for correlated error, though. Two systems trained on the same data make the same mistakes, and then their agreement says nothing about whether they are right.
Mandatory audit slices
Confidence routing has a built-in blind spot. Nobody looks at the high-confidence region, so there is no data that would show it degrading. The only way to keep an unbiased estimate of the automated path's error rate is a small random sample of auto-approved items, drawn without regard to confidence and reviewed as carefully as the escalations. Budget this slice as a measurement cost, separate from rework. It also keeps reviewers seeing ordinary cases, so their sense of the base rate stays accurate, and with it their judgment on escalations.
Review Interface Design and Reviewer Skill Maintenance
Preserving situation awareness
Endsley (1995) breaks situation awareness down into perception of the relevant elements, comprehension of what they mean together, and projection of what follows. A review screen that shows a verdict and two buttons supports none of these. It gives the reviewer a conclusion with the evidence taken away, and that is when the substitution described earlier comes easiest.
In practice this means a few specific things. Put the evidence that drove the verdict on screen next to the verdict. Make the alternatives the system considered visible. Include the item's history so a repeat failure is recognisable, and where the workflow allows, present the evidence before the recommendation. Nielsen's usability heuristics cover the rest: visibility of system status, match to the real world, recognition rather than recall. A common anti-pattern is a layout where the cheapest action is also the least informed one, such as a large green Approve button next to a small "see details" link. That layout invites rubber-stamping.
Exercising the manual path
An escalation path depends on the competence of the people it routes to, and that competence fades when it isn't used. If inspectors spend a year seeing only machine-adjudicated cases with machine-written summaries, the skills the design counts on will have worn down by the time they are needed: reading a raw case, noticing an unfamiliar failure mode, overruling the system.
The countermeasures are ordinary: periodic fully manual runs on a small case set, seeded cases with known answers mixed into the queue, rotation between adjudication and audit work, and a documented re-validation interval for reviewer competence. ISO 9001 already requires organisations to determine and maintain the competence of people whose work affects quality performance. In a heavily automated process the competence that matters is handling the hard remaining cases, and that needs regular practice. A one-time qualification does not keep it up.
Design Rules for Human-in-the-Loop Review
- Work out the cost ratio before you pick a threshold. If C_review / C_miss can't be expressed in any unit, review capacity ends up setting the escalation rate instead of risk.
- Never ship a step where the only thing the person can do is approve. The reviewer has to be able to modify the item, ask for information, or reject it with a reason that changes what happens downstream. Otherwise the step records consent, and no judgment goes into it.
- Where per-item error rates are low, prefer full automation with an audited sample to full review. At low p, reviewing everything catches little and uses up the attention the audit slice needs.
- Measure the reviewers as well as the model: overturn rate by reviewer and segment, time on task, detection rate on seeded cases. An overturn rate near zero means either the system is close to perfect or the review has become a ritual, and only seeded cases tell you which.
- Treat the escalation rate as a commitment of staff capacity. Review quality drops quietly when the queue backs up, so a threshold that sends more escalations than the team can decide carefully gives you a weaker control, even though it looks stricter.
Limitations of This Analysis
This is a design argument built on established human-factors findings. I have not run a study. The detection rates in the numeric example are assumptions picked to show a structural relationship, and I am not claiming that any real system has those values. Three things can only be measured in place: how strong automation bias is in a given domain, how detection rate changes as signals get rarer, and how fast a professional skill fades without practice. The first-party figures come from one pipeline in one production setting, so they are not a benchmark. The threshold rule also treats C_miss as a single expected value. That is not good enough when consequences are heavy-tailed or fall on people other than the operator. In those cases the threshold has to be set from the tail, which is a different analysis from this one.
References
- Bainbridge, L. (1983). Ironies of automation.
- Endsley, M. R. (1995). Toward a theory of situation awareness in dynamic systems.
- Fitts, P. M. (1951). Human engineering for an effective air-navigation and traffic-control system.
- Kahneman, D. (2011). Thinking, Fast and Slow.
- Mackworth, N. H. (1948). The breakdown of vigilance during prolonged visual search.
- Nielsen, J. Usability Engineering.
- Parasuraman, R., & Riley, V. (1997). Humans and automation: Use, misuse, disuse, abuse.
- International Organization for Standardization. (2015). ISO 9001: Quality management systems — Requirements.