Writing
What Manufacturing Taught Me About Agent Failure Modes
The shop floor has been running autonomous systems against physical reality for decades
Manufacturing offers useful examples for anyone designing an agent that can act without a person confirming each step. On a shop floor, designers have long had to decide when a system should continue, when it should stop, and who can intervene.
I would examine those practices when designing an agent for an enterprise system, while paying attention to where the analogy stops working.
A useful starting assumption is that records can drift from what is happening on the floor. The design then needs a way to find and resolve the difference.
Compare the record with the work
An inventory system says there are four hundred units. There are three hundred and ninety one. There are several possible explanations. Scrap happened, a count was off, something was consumed against the wrong order, something is physically present and logically committed elsewhere.
Cycle counting provides a recurring check: a scheduled, resourced, permanent ritual of comparing the record to the thing, with a tolerance, an investigation threshold, and an owner. The divergence is not anecdotal; it has been measured across a whole store network and it is pervasive (DeHoratius & Raman, 2008).
Reconciliation belongs in the operating process even when the system is working as designed. That is a useful assumption for agents too. An agent's understanding of the system's state is a model of a model, one degree further from the world, and it will drift for reasons nobody did anything wrong to cause.
The practical consequence is that an agent touching operational data needs a scheduled comparison against ground truth, with somebody who owns the discrepancies. Set aside time for reviewing and resolving the differences. Governance frameworks put continuous measurement in the same structural position (National Institute of Standards and Technology, 2023).
Stopping is a designed outcome
Andon is the practice of stopping the line and signaling for help (Ohno, 1988). What makes it interesting is not the cord, it is that the authority to stop was deliberately given to the person nearest the problem, and that using it is treated as correct behavior rather than as failure.
If an agent is evaluated mainly on completed tasks, a halt can look like a failure to remove. The evaluation should also recognize cases where stopping prevented a bad action.
Continuing through an unresolved error has a cost. On a line, a system that continues while wrong produces defective units at speed, and the cost of the run is far higher than the cost of the stop. Automating the straightforward cases and leaving only the hard residue to a human is the oldest known trap in this design space (Bainbridge, 1983). The same arithmetic holds in an ERP: an agent that proceeds through ambiguity produces a batch of consistent, confidently wrong records, and consistency makes them harder to find later than a single obvious error would have been.
Design the stop as carefully as the successful path: what it signals, who receives it, what context it carries, and how quickly a human can resolve it and resume. A page without enough context leaves the recipient to reconstruct the problem before they can help. A capable system that people route around has failed as completely as one that is wrong (Parasuraman & Riley, 1997), and the fix is designing for efficient correction rather than for fewer interruptions (Amershi et al., 2019).
Make the wrong action impossible rather than unlikely
Poka yoke is mistake proofing: designing the fixture so the part cannot be inserted backward, rather than training the operator to be careful. The physical world rewards this because parts get inserted backward regardless of training, at three in the morning, in the last hour of a long shift.
The software equivalent is constraining at the interface, not at the agent. If an agent must never post to a closed period, the enforcement belongs in the permission set and the interface, not in the prompt. Enforce that boundary through access controls, so a prompt cannot grant the missing permission.
This one matters more for agents than it did for people, because an agent will attempt the same wrong thing at machine speed and without the hesitation that makes a human pause and ask. Instructions alone leave the boundary dependent on the agent following them correctly. Defences work when they are layered and structural, not when they depend on everyone remembering (Reason, 2000).
Keep the inputs needed to investigate
In a regulated plant you can take a finished unit and reconstruct it: which lot each component came from, which machine, which operator, which shift, which process parameters. Genealogy exists so that when a defect surfaces months later, you can determine the blast radius precisely rather than recalling everything.
Application logging usually does not do this. It records that an action occurred and often what the result was. It rarely records what the system believed at the time.
For an agent, the equivalent of genealogy is the reasoning inputs: the data it read, the version of the rules it applied, the model and prompt in use, and the alternative it rejected. Regulated manufacturing already requires this shape of record, and has for years (U.S. Food and Drug Administration, 2025). Store that with the output. When something is discovered to be wrong six weeks later, the investigation needs to establish what happened and which other outputs depended on the same inputs or rules.
Those records help bound a recall or a correction when the underlying problem becomes clear. Undeclared consumers and entangled inputs are what make the blast radius hard to bound after the fact (Sculley et al., 2015).
Where the analogy stops
Manufacturing's physical constraints impose a natural rate limit. A machine can only run so fast, a line only produces so many units an hour, and that ceiling has quietly bounded the damage of every bad decision the system has ever made.
Software can operate on a much shorter timescale. An agent can make ten thousand wrong writes in the time a line makes one bad part, and it will do so without any of the sensory signals, the noise, the smell, the pile of scrap, that tell a floor something has gone wrong before the report does.
So the floor's disciplines transfer, but the margin does not. Tight coupling without a natural rate limit is close to the definition of a system whose accidents are normal rather than exceptional (Perrow, 1999). Which argues for the rate limits being explicit and designed rather than inherited from physics: batch sizes, throttles, daily caps, staged rollouts by entity or site. Those limits give someone time to notice a problem before it spreads further.
Make intervention part of the work
The practices belong together: compare the records, give people authority to stop, enforce the boundaries, and retain enough history to investigate. Each needs an owner and time in the operating process.
For an agent, I would make those responsibilities visible in the interface and support process. They provide a practical way to build reliance around demonstrated reliability (Lee & See, 2004).
References
- Ohno (1988). Toyota Production System: Beyond Large-Scale Production. Productivity Press. www.routledge.com/Toyota-Production-System-Beyond-Large-Scale-Production/Ohno/p/book/9780915299140 Jidoka, andon, and the authority to stop the line, described by the system's principal architect.
- DeHoratius & Raman (2008). Inventory record inaccuracy: An empirical analysis. Management Science, 54(4), 627-641. doi.org/10.1287/mnsc.1070.0789 Empirical evidence that inventory records and physical reality diverge continuously.
- Bainbridge (1983). Ironies of automation. Automatica, 19(6), 775-779. doi.org/10.1016/0005-1098(83)90046-8
- Reason (2000). Human error: models and management. BMJ, 320(7237), 768-770. pmc.ncbi.nlm.nih.gov/articles/PMC1117770
- Perrow (1999). Normal Accidents: Living with High-Risk Technologies. Princeton University Press. press.princeton.edu/books/paperback/9780691004129/normal-accidents
- U.S. Food and Drug Administration (2025). 21 CFR Part 11: Electronic Records; Electronic Signatures. Code of Federal Regulations. www.ecfr.gov/current/title-21/chapter-I/subchapter-A/part-11 The regulated analogue of genealogy: records, audit trails and traceability requirements.
- Parasuraman & Riley (1997). Humans and automation: Use, misuse, disuse, abuse. Human Factors, 39(2), 230-253. doi.org/10.1518/001872097778543886
- Lee & See (2004). Trust in automation: Designing for appropriate reliance. Human Factors, 46(1), 50-80. doi.org/10.1518/hfes.46.1.50_30392
- Sculley et al. (2015). Hidden Technical Debt in Machine Learning Systems. Advances in Neural Information Processing Systems 28. proceedings.neurips.cc/paper_files/paper/2015/hash/86df7dcfd896fcaf2674f757a2463eba-Abstract.html
- National Institute of Standards and Technology (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1. U.S. Department of Commerce. nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf
- Amershi et al. (2019). Guidelines for Human-AI Interaction. CHI Conference on Human Factors in Computing Systems. dl.acm.org/doi/10.1145/3290605.3300233