Writing

The Months After Go-Live

Following adoption long enough to see whether the work changes

The unused terminal A working system can sit unused while the familiar process continues beside it.

A program ships. The system works, the integration holds, the accuracy is good. Six months later the promised benefit has not appeared, and an investigation finds that roughly the same number of people are doing roughly the same work, now with an additional system alongside it.

I have seen systems deliver less than promised because the work around them barely changed. The productivity literature also gives organizational change a substantial role: returns to information technology show up only alongside complementary organizational change (Brynjolfsson & Hitt, 2000).

I find it useful to distinguish five stages: delivered, available, used, trusted, and sustained. A delivery report tells us little about the later stages unless someone has measured them too.

Why this hits AI programs harder

Agents add a few demands that deserve attention during rollout.

The output requires a judgment the user did not previously have to make. A report is either run or not run. An agent's proposal has to be assessed: is this right, do I trust it, what happens to me if I approve it and it is wrong. That is a new cognitive task added to somebody's day, and if it takes longer than doing the work, a rational person stops using it. They may never report why they stopped. Research on technology acceptance examines perceived usefulness and ease of use (Davis, 1989), with social influence and facilitating conditions added later (Venkatesh et al., 2003).

Where the value leaks From delivery through sustained use, each stage needs its own measure.

Trust is asymmetric and slow to rebuild. After a bad output, a person may spend more time checking subsequent results. Usage can remain high even while the review burden removes much of the benefit. The asymmetry is experimentally documented: confidence drops sharply after seeing an algorithm err, even one that outperforms the human (Dietvorst, Simmons & Massey, 2015). It is not uniform, though, and the same literature finds people often prefer algorithmic judgment when they have not seen it fail (Logg, Minson & Moore, 2019). Click counts alone will not tell you how much confidence the reviewer has lost.

Accountability also needs attention. When an agent proposes and a human commits, the human is answerable for the commit. They have taken on review work while remaining accountable for the result. Without a clear policy for handling an approved mistake, they may spend so long checking each proposal that little time is saved. Reliance tracks perceived reliability and perceived risk together, not capability alone (Lee & See, 2004) (Hoff & Bashir, 2015).

Following the work after delivery

Between delivered and available you lose access, permissions, and licensing: practical requirements that can delay use if no one is responsible for completing them.

Between available and used you lose to workflow placement. If the agent lives in a system the user does not have open, in a tab they must remember, it will not be used, regardless of quality. Each extra step makes returning to the familiar process easier (Amershi et al., 2019).

Between used and trusted you lose to the review burden described above, and to the absence of an obvious way to report a bad output. A user with no channel to say "this one was wrong" concludes that nobody wants to know, and may stop trying to report problems.

Between trusted and sustained you lose to the old process never being retired. Sustained adoption is enacted in daily practice over time rather than achieved at go-live (Orlikowski, 1996). When both paths remain available, people under pressure may return to the familiar one. Keeping the change in use requires attention after the initial rollout. Follow the exceptions over time to see whether the new process is taking hold or whether people are mainly reporting work they still do the old way.

Decisions to make during rollout

Retire the old path deliberately, on a date, with a named owner. Not as a hard cutover on day one, which is reckless, but as an explicit decision with a schedule. Leaving the old process available indefinitely makes it easier for people to return to it under pressure. Short-term wins, consolidated deliberately, are what carry a change past the point of reversal (Kotter, 1995).

Placement is the second lever. The agent belongs inside the system they already have open, in the queue they already work through, in the document they already produce. Watch how much extra navigation a person has to do each time.

The accountability question needs an answer in writing, before rollout, at a level of seniority that makes it real. Something on the order of: you are accountable for reviewing to a stated standard, and if you did that and the output was still wrong, that is the program's failure and not yours. Vague reassurance in a town hall does not do this. People need to know what happens on a bad day, and they will assume the harshest answer until told otherwise.

Make reporting a bad output take one action. And then close the loop visibly, so the person who took the trouble to report it can see what happened. The reports are also your best eval cases, so this pays twice.

Every rollout has people who are already trying to make it work. Find them and give them the time. They are usually not the ones nominated, they are already doing this on top of their real job, and giving them protected hours lets them contribute without adding an unpaid second job.

Measure trusted and sustained, not delivered. Ask what share of proposals are accepted without being redone by hand, and whether the manual path has actually stopped. Compare those measures with usage counts. More clicks can accompany more manual correction.

Give someone authority to change the incentives

Responsibility can fall between the technical program and the sponsor. A communications workstream may receive the assignment without the authority to change what managers measure. Training and a launch email can explain the system, but they cannot resolve that conflict.

Consider the person who is expected to adopt it. If a person's manager measures them on throughput, and the agent makes them slower for six weeks while they learn to trust it, they will not use it, and they are behaving correctly. No amount of enablement fixes an incentive that points the other way. It is also the predictable end state of automating the easy portion and leaving a person the hard remainder plus the review (Bainbridge, 1983). A capable system that people rationally decline to use is a documented outcome, not an anomaly (Parasuraman & Riley, 1997). Someone with organizational authority has to change what is measured, temporarily, and absorb the dip.

The retired process Decide when and how to retire the old process, with someone responsible for the transition.

References

  1. Kotter (1995). Leading change: Why transformation efforts fail. Harvard Business Review, 73(2). hbr.org/1995/05/leading-change-why-transformation-efforts-fail-2 Discusses organizational change, including short-term wins and the work needed to sustain them.
  2. Davis (1989). Perceived usefulness, perceived ease of use, and user acceptance of information technology. MIS Quarterly, 13(3), 319-340. doi.org/10.2307/249008 Examines perceived usefulness and ease of use in technology acceptance.
  3. Venkatesh et al. (2003). User acceptance of information technology: Toward a unified view. MIS Quarterly, 27(3), 425-478. doi.org/10.2307/30036540 The unified model, adding social influence and facilitating conditions to the acceptance picture.
  4. Brynjolfsson & Hitt (2000). Beyond computation: Information technology, organizational transformation and business performance. Journal of Economic Perspectives, 14(4), 23-48. doi.org/10.1257/jep.14.4.23
  5. Orlikowski (1996). Improvising organizational transformation over time: A situated change perspective. Information Systems Research, 7(1), 63-92. doi.org/10.1287/isre.7.1.63 Change as something enacted and adjusted in practice rather than specified once and executed.
  6. Parasuraman & Riley (1997). Humans and automation: Use, misuse, disuse, abuse. Human Factors, 39(2), 230-253. doi.org/10.1518/001872097778543886
  7. Lee & See (2004). Trust in automation: Designing for appropriate reliance. Human Factors, 46(1), 50-80. doi.org/10.1518/hfes.46.1.50_30392
  8. Dietvorst, Simmons & Massey (2015). Algorithm aversion: People erroneously avoid algorithms after seeing them err. Journal of Experimental Psychology: General, 144(1). doi.org/10.1037/xge0000033
  9. Logg, Minson & Moore (2019). Algorithm appreciation: People prefer algorithmic to human judgment. Organizational Behavior and Human Decision Processes, 151. doi.org/10.1016/j.obhdp.2018.12.005 The other side of the finding: people often prefer algorithmic judgment, which makes context decisive.
  10. Hoff & Bashir (2015). Trust in automation: Integrating empirical evidence on factors that influence trust. Human Factors, 57(3), 407-434. doi.org/10.1177/0018720814547570 Reviews influences on trust in automation, including experience and system transparency.
  11. Bainbridge (1983). Ironies of automation. Automatica, 19(6), 775-779. doi.org/10.1016/0005-1098(83)90046-8
  12. Amershi et al. (2019). Guidelines for Human-AI Interaction. CHI Conference on Human Factors in Computing Systems. dl.acm.org/doi/10.1145/3290605.3300233