Legacy modernization rarely fails because the new technology is hard. It fails because the old system knows things nobody wrote down: the discount that only applies on the last business day of a quarter, the field that means something different for customers onboarded before a merger, the manual check an operator does because the system once got it wrong.
Those rules are requirements. A rewrite that misses them is a regression, however modern the stack. AI does not remove that problem, but it makes recovering the rules far cheaper than it used to be. That changes how a modernization programme should be run.
Why big rewrites go wrong
The classic approach is to specify the new system from interviews and documentation, build it, and switch over. The specification is incomplete because the knowledge is in the code and in people's habits, not in documents. The gaps surface after cutover, when they are most expensive to fix.
Step 1: recover the rules
Language models are good at reading code that nobody wants to read. We use them to summarize modules, trace how data moves through the system and draft a catalogue of the business rules they find. Your domain experts then confirm, correct or reject each rule.
The confirmed catalogue becomes the specification for the new system. It is also the first durable documentation the old system has had, which is useful whether or not you replace it.
A hallucinated business rule is worse than a missing one, because it looks authoritative. Every rule the model drafts needs a person to confirm it.
Step 2: pin today's behaviour with tests
Before changing anything, capture what the system does now. Characterization tests record current behaviour from production samples and logs, so any change in behaviour is visible. AI generates candidate tests quickly; engineers curate them and decide which behaviours are intended and which are bugs to fix on purpose.
Step 3: digitize the inputs
Many legacy workflows start on paper: scanned forms, emailed PDFs, faxed documents. Document extraction models can turn these into structured data, but accuracy varies by document type and by field. Measure it on a sample of your real documents before setting targets, and design a review queue so that low-confidence fields go to a person while the rest flow straight through.
Step 4: replace one capability at a time
Put a facade in front of the old system and move capabilities behind it one by one. Run old and new in parallel and reconcile their outputs automatically. Retire each legacy component once the new one has matched it in production for long enough to trust. The business keeps running throughout.
Where AI does not help yet
- Deciding what the business should do. When two rules contradict each other, a person has to choose.
- Accountability for regulated decisions. The model can draft; a named person signs off.
- Verifying its own output. Recovered rules and generated tests are hypotheses until checked.
What to measure
Track the share of business rules confirmed by experts, test coverage of the critical paths, extraction accuracy per field, manual touches per case, and the number of legacy components retired. These tell you whether the programme is reducing risk, not just producing code.