Writing
The Reversibility Ladder
Where agents may write in a system of record
When considering an agent for an enterprise system, I would ask what it can do and what happens when it gets something wrong. Reading an invoice, drafting a requisition, and posting an entry leave different kinds of mistakes to correct.
A capability score alone does not tell the system owner which of those operations to permit.
My starting rule would be:
An agent should be allowed to write in inverse proportion to how hard that write is to reverse.
The team needs evidence that the agent can do the task. The person responsible for the system also needs to understand the consequences of granting it access. Reversibility gives them a useful way to examine each proposed write.
What a wrong result would require
A ninety eight percent accurate agent sounds excellent until you ask what happens during the other two percent.
If the answer is "a draft requisition is wrong and somebody edits it," you have a productivity tool with a small correction cost. If the answer is "a journal entry posted to a closed period," you have an accounting problem, an audit finding, and a conversation with your external auditors. The same accuracy rate can describe both cases while leaving their correction costs unexplained.
Accuracy needs to be measured on the task, and reversibility needs to be examined in the receiving system. Deployment research describes the importance of constraints beyond model performance (Paleyes, Urma & Lawrence, 2022). Automating the tractable part and leaving the residue to a human is the oldest known irony in this field (Bainbridge, 1983).
The ladder
Rank operations by what it costs to undo them, then let that ranking set agent authority. Governance frameworks say the same thing in the abstract: authority is a thing to be mapped and managed rather than assumed (National Institute of Standards and Technology, 2023), while Article 14 of the EU AI Act provides for proportionate human oversight and the ability to interrupt high-risk systems within its scope (European Parliament & Council, 2024).
Tier 1, free to undo. Drafting a requisition, summarizing a variance, proposing a journal, drafting a customer email, explaining why a report moved. Nothing has happened yet in the system of record. The cost of being wrong is somebody's attention. This is where I would start, with review proportionate to what the draft will be used for.
Tier 2, cheap to undo. Updating a descriptive field, tagging a record, attaching a document, setting a non-controlling flag. A wrong value is corrected by writing the right one, and nothing downstream has already consumed it. Let agents write, and log it as agent-originated so a human can find every one later.
Tier 3, costly to undo. Inventory adjustments, price changes, credit limits, master data creation, releasing a production order. These are technically reversible. The problem is that by the time you notice, the value has already been read by something else and acted upon. I would have the agent propose these changes for a person to review and commit.
Tier 4, effectively irreversible. Posting to the general ledger. Executing a payment run. Closing a period. Tax determination on a filed return. Anything that leaves the building or enters the audit record. My policy would keep direct agent writes out of this tier. A better model or a confirmation dialog at the end of a batch would not, by itself, justify changing that boundary.
An operation can be easy to execute and costly to correct. That is why a technically simple action may belong in Tier 4.
Why the boundary sits exactly there
The surrounding controls help explain the boundaries between the tiers.
Reversal is not deletion. Where posted ledger entries are retained, a correction requires a reversing entry and leaves both in the history. That distinction matters to the design of financial reporting controls (Public Company Accounting Oversight Board, 2024). The correction leaves a record of both entries. The error and its correction both become part of the financial history that someone will later have to explain. The reversal therefore needs to be considered as an accounting process, including who authorizes and explains it.
The account needs an accountable owner. Every consequential action in a controlled system answers the question "who did this, and were they allowed to." A service account can identify the software that acted while leaving unclear which person authorized it. Until an agent can be a party to segregation of duties rather than a hole in it, the control does not survive its involvement (Committee of Sponsoring Organizations of the Treadway Commission, 2013). What an agent may reach is a least-privilege question, and has been since 1975 (Saltzer & Schroeder, 1975).
Blast radius is not proportional to the size of the write. One wrong inventory number is a small write. Planning consumes it, generates demand, purchasing acts on the demand, and cash moves. The original error may be difficult to trace once several downstream processes have acted on it. This is why Tier 3 exists as its own rung: those writes look survivable in isolation and are not, because downstream processes may have already acted on the value. Tight coupling is precisely what converts a local error into a system-wide one faster than anyone can intervene (Perrow, 1999), and the undeclared downstream consumer is a documented failure mode in its own right (Sculley et al., 2015).
Some processes require reproducible results and an explanation. Tax determination, statutory reporting, and period close have to produce the same answer twice and defend it to somebody external. A strong accuracy score does not establish that the result can be reproduced and explained under review. Those requirements belong in the design alongside model performance.
Access also depends on timing. Close periods, freeze windows, statutory cutoffs, batch schedules: a system that is available to write to is not necessarily open to write to.
How to use the ladder
Use the tiers to review a list of proposed operations.
Start by listing what the agent would actually write, not what it would understand. Some proposals turn out to need only a draft or recommendation. If that meets the need, there may be no reason to grant write access.
Then place each write on a rung, and make somebody name the reversal path out loud. This is defence in depth rather than exhortation: the useful question is what conditions make the error likely, not who to blame afterward (Reason, 2000). Not "we would catch it." The literal sequence: who notices, how, how fast, and what they do. If the team cannot explain the reversal path, leave the permission unresolved until it can.
Then check propagation before authority. Ask what reads this value within an hour. Include those consumers when estimating the cost of a wrong value.
Use that analysis to set the performance and review requirements for the model.
What this changes about pilots
A demonstration that completes a consequential action can be persuasive. Before making it the first release, I would check whether the required access is compatible with the controls already in place.
I would begin with useful Tier 1 and Tier 2 work, measure the results, and establish the audit and monitoring needed for later decisions. Any move toward Tier 3 should include evidence and a correction process that people can use efficiently (Amershi et al., 2019). I would keep Tier 4 outside that plan and make the boundary explicit from the beginning.
The first release may be modest, but it gives the team operating experience to use in the next decision.
Leave room to learn from use
A serious error can also cause people to stop using an otherwise capable system. Disuse belongs alongside over-reliance in the review of automation risks (Parasuraman & Riley, 1997). There is useful work available without granting the most consequential permissions.
I would begin with work a person can check and correct, then use what we learn to decide whether the agent should take on more. Some permissions may stay with people. That needs to be agreed before the team builds around them.
References
- Bainbridge (1983). Ironies of automation. Automatica, 19(6), 775-779. doi.org/10.1016/0005-1098(83)90046-8 The classic result that automating the easy parts leaves humans with the harder residue.
- Parasuraman & Riley (1997). Humans and automation: Use, misuse, disuse, abuse. Human Factors, 39(2), 230-253. doi.org/10.1518/001872097778543886 Distinguishes misuse (over-reliance) from disuse (rejecting a capable system).
- National Institute of Standards and Technology (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1. U.S. Department of Commerce. nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf Govern, Map, Measure, Manage. The Govern function is where write authority belongs.
- European Parliament & Council (2024). Article 14: Human oversight, Regulation (EU) 2024/1689 (AI Act). Official Journal of the European Union. eur-lex.europa.eu/eli/reg/2024/1689/oj/eng Article 14 concerns human oversight of high-risk AI systems within the regulation's scope, including intervention and safe interruption.
- Public Company Accounting Oversight Board (2024). AS 2201: An Audit of Internal Control Over Financial Reporting. PCAOB Auditing Standards. pcaobus.org/oversight/standards/auditing-standards/details/AS2201 An auditing standard for internal control over financial reporting. The essay uses it as control context, not as a product-specific ledger specification.
- Committee of Sponsoring Organizations of the Treadway Commission (2013). Internal Control, Integrated Framework. COSO. www.coso.org/guidance-on-ic Source of the control environment and segregation of duties expectations cited here.
- Saltzer & Schroeder (1975). The protection of information in computer systems. Proceedings of the IEEE, 63(9). web.mit.edu/Saltzer/www/publications/protection Sets out the principle of least privilege used here when discussing access boundaries.
- Perrow (1999). Normal Accidents: Living with High-Risk Technologies. Princeton University Press. press.princeton.edu/books/paperback/9780691004129/normal-accidents Discusses interactions and tight coupling in complex systems. The enterprise examples here are applications of that framework.
- Reason (2000). Human error: models and management. BMJ, 320(7237), 768-770. pmc.ncbi.nlm.nih.gov/articles/PMC1117770 Describes system conditions and multiple layers of defence in the study of human error.
- Amershi et al. (2019). Guidelines for Human-AI Interaction. CHI Conference on Human Factors in Computing Systems. dl.acm.org/doi/10.1145/3290605.3300233 Guideline 9, Support efficient correction: the guidance the correction-cost argument here rests on.
- Sculley et al. (2015). Hidden Technical Debt in Machine Learning Systems. Advances in Neural Information Processing Systems 28. proceedings.neurips.cc/paper_files/paper/2015/hash/86df7dcfd896fcaf2674f757a2463eba-Abstract.html Entanglement and undeclared consumers: the mechanism behind downstream propagation.
- Paleyes, Urma & Lawrence (2022). Challenges in Deploying Machine Learning: A Survey of Case Studies. ACM Computing Surveys, 55(6). arxiv.org/abs/2011.09926 Survey evidence that deployment failures cluster outside model quality.