Writing
Migration Archaeology
Finding the decisions and workarounds a replacement needs to account for
A long-running enterprise system can become difficult to explain even to the people who depend on it. Replacing it means finding out which behavior they still need.
A replacement estimate can cover the visible features while missing years of exceptions and informal workarounds. That systems change continuously and grow in complexity unless effort is spent reducing it was formalized decades ago (Lehman, 1980), as was the mechanism by which they age: modification by people who no longer hold the original design (Parnas, 1994).
I would investigate those exceptions before committing to an estimate. The behavior of the system, the corrections people make, and the history of its code can help recover decisions that the documentation no longer explains.
Four sources to read together
Each of these sources can tell you something the others leave out. Where they disagree, there is more to investigate.
Observed behavior shows what the current system does. What the system does in production, under real volume, with real data, including the paths that only fire at quarter end and the one that only fires for a single country. Compare that behavior with the intended requirements. A discrepancy can mean the documentation is stale, the implementation is wrong, or the requirements changed.
The corrections people make around it reveal work the interface may not show. Ask about the spreadsheets, manual checks, and reminders people use before posting. Some compensate for defects. Others preserve a business decision the replacement will need to support.
The code and its history can help explain how the present implementation came about. It also contains intentions that were abandoned midway, which is why systems accumulate structures that are referenced nowhere and deleted by nobody. Commit history and change records, where they survive, are the closest thing to an interview with people who have left.
The written documentation can preserve requirements and decisions that are hard to infer from behavior. Check when it was maintained and whether its examples still match production. Documentation decay is a predictable consequence of continuing change rather than a local failure of discipline (Parnas, 1994).
Reading the documents alongside the running system gives the team a better chance of finding missing requirements before acceptance testing.
Find out why the exception exists
The classic rule says do not remove a fence until you know why it was put there (Chesterton, 1930). That question is useful when the people who added a rule have left and its purpose is no longer clear.
I would investigate four possible explanations for an unexpected behavior.
It is a regulatory or contractual requirement, which must be understood before the behavior changes. Retention periods, immutable audit trails and signature rules are written down somewhere, frequently in a regulation rather than in your repository (U.S. Food and Drug Administration, 2025). These are the most dangerous to clean up because they usually look arbitrary. A field that must be retained for a specific number of years, a sequence that must never have gaps, a rounding rule that seems wrong and is legally specified.
It is a workaround for a defect in another system, which may or may not still exist. Worth checking, because if the upstream defect was fixed six years ago the workaround may no longer be needed, and if it was not, you are about to reintroduce a problem somebody already solved.
It is an accommodation for a customer, a plant, or a country that is still operating on the assumption it exists. A team outside the change program may still rely on that accommodation. The undeclared consumer is the same hazard that machine learning systems rediscovered under a newer name (Sculley et al., 2015).
Or it is genuinely obsolete, a remnant of a process that ended. Confirm that nothing still depends on it before removing it.
Code can suggest an explanation. Following the behavior to the people who depend on it helps establish whether that explanation is still valid.
The method
Start from the outputs, not the code. Every report, file, feed, and document the system produces has a consumer, and the consumer can tell you which parts they actually use. This is the fastest way to find both the critical path and the large amount of the system that produces things nobody has read in years.
Follow the corrections. Ask each team what they fix by hand and what they check before trusting it. Those conversations can reveal requirements missing from the specification, and you will find the workarounds that need to become features rather than being migrated as-is. Those manual corrections are also where the real data quality cost has been sitting, unbilled, for years (Redman, 1998).
Interview the leavers while they are still reachable. Make time with people who are about to leave. Their knowledge may be tacit and difficult to recover after they have gone (Nonaka, 1994). Book the time. Ask how the process works and why the exceptions were added. The latter is often harder to recover from the implementation.
Instrument before you rewrite. Where you can, log which paths actually execute. Dead code is much easier to retire with evidence than with argument, and the evidence also protects you when somebody insists a path is critical that has not run in two years.
Write the fence register. A living list: the strangeness, what you currently believe explains it, the confidence, and who could confirm. A later team member should be able to see what remains uncertain and whom to ask.
Where this goes wrong
Treating all existing behavior as a requirement can carry unnecessary workarounds into the replacement. Worse, the new structure will tend to reproduce the communication structure of whoever rebuilt it rather than the problem it was meant to solve (Conway, 1968), and the difficulty that remains is essential rather than accidental (Brooks, 1987).
The opposite failure is treating it as noise, declaring a clean-slate design, and rediscovering each fence in production, one incident at a time, in the order of how badly each one hurts. In a tightly coupled system that discovery process is not gradual, it is an outage (Perrow, 1999).
Record what should remain, what can change, and the evidence behind each decision. Keep unresolved cases visible so someone can investigate them.
Allow time for the investigation
An investigation can be difficult to fund when a competing proposal offers an earlier delivery date. Its value lies partly in the problems the team finds before making a commitment.
Some of that learning will continue during the replacement. Organizations that succeed here tend to treat the change as something that unfolds and is adjusted in situ rather than something specified once and executed (Orlikowski, 1996). I would leave room in the plan for it, with the people who know the old system still available. The first estimate should make those unknowns visible.
References
- Lehman (1980). Programs, life cycles, and laws of software evolution. Proceedings of the IEEE, 68(9), 1060-1076. doi.org/10.1109/PROC.1980.11805 The laws of software evolution: continuing change, and growing complexity unless work is done to reduce it.
- Parnas (1994). Software aging. Proceedings of the 16th International Conference on Software Engineering. doi.org/10.1109/ICSE.1994.296790 Why systems age: accumulated modification and the loss of the people who understood the design.
- Brooks (1987). No Silver Bullet: Essence and accidents of software engineering. Computer, 20(4), 10-19. doi.org/10.1109/MC.1987.1663532
- Chesterton (1930). The Thing: Why I Am a Catholic. Dodd, Mead & Company. archive.org/details/thingwhyiamcatho0000ches The fence parable, from the chapter 'The Drift from Domesticity'.
- Nonaka (1994). A dynamic theory of organizational knowledge creation. Organization Science, 5(1), 14-37. doi.org/10.1287/orsc.5.1.14 Discusses tacit knowledge and its relationship to organizational knowledge creation.
- Conway (1968). How do committees invent?. Datamation, April 1968. www.melconway.com/Home/Committees_Paper.html
- Sculley et al. (2015). Hidden Technical Debt in Machine Learning Systems. Advances in Neural Information Processing Systems 28. proceedings.neurips.cc/paper_files/paper/2015/hash/86df7dcfd896fcaf2674f757a2463eba-Abstract.html
- U.S. Food and Drug Administration (2025). 21 CFR Part 11: Electronic Records; Electronic Signatures. Code of Federal Regulations. www.ecfr.gov/current/title-21/chapter-I/subchapter-A/part-11
- Redman (1998). The impact of poor data quality on the typical enterprise. Communications of the ACM, 41(2), 79-82. doi.org/10.1145/269012.269025
- Perrow (1999). Normal Accidents: Living with High-Risk Technologies. Princeton University Press. press.princeton.edu/books/paperback/9780691004129/normal-accidents
- Orlikowski (1996). Improvising organizational transformation over time: A situated change perspective. Information Systems Research, 7(1), 63-92. doi.org/10.1287/isre.7.1.63