Skip to main content
onext technology
AI 1 October 2026 - 11 min read

Before an agent touches your legacy code: specify what it does, not what it should do

AI is already very good at explaining what a legacy system does. The risk begins when that explanation is treated as a requirement and an agent “fixes” something someone needs exactly as it is. Legacy needs two documents, not one: what the code does today and what's going to change.

Jordi García
Tech Lead at onext
A developer compares an old code listing printed on continuous fanfold paper with the same module on screen, at dusk in an office: specifying the behaviour of legacy code before an AI agent changes it

A made-up but plausible example: an agent is asked to add a field to a fifteen-year-old billing module. It reads the code, notices that rounding happens in an unusual order, “fixes” it along the way and leaves the tests green. Three weeks later, the report a client reconciles with its accounts every month is out by a few cents per line. The odd rounding wasn't a bug: it was the behaviour that client depended on.

On the face of it, nobody did anything wrong. The agent found something that looked like a defect and corrected it. The problem is that nobody had told it which behaviour it couldn't change, because nobody had ever written it down. In a legacy system, the most reliable documentation is the code itself, accidents included. And an agent working without knowing which of those accidents are contract is a risk, however good the model.

What AI already does well with legacy code

It's worth starting with what works, because it works well. In November 2025, the Thoughtworks Technology Radar placed using generative AI to understand legacy codebases in “Adopt”, its highest ring: based on its experience across multiple clients, it has become “a practical default rather than an experiment”. Tools such as Claude Code, Cursor or Copilot surface business rules, summarise logic and identify dependencies in systems nobody on the current team wrote. We cover this in more detail in GenAI for understanding legacy code.

The most cited case is Morgan Stanley's. According to The Wall Street Journal in June 2025, as reported by Entrepreneur, its in-house tool DevGen.AI had reviewed nine million lines of old code in five months, saving an estimated 280,000 developer hours. The interesting part is what it actually does: it translates code in old languages into plain-English specifications, which developers then use as a reference to rewrite it. It doesn't write the new code; that's still people's work.

In other words, the value isn't AI rewriting the system, but AI producing a text that says what it does. That text is the right starting point. What you need to understand is what kind of text it is.

A description isn't a specification

When AI reads legacy code and explains what it does, it produces a description. It says what the code does, not what it should do. It includes the odd rounding, the field padded with zeros when it arrives empty and the edge case handled unexpectedly because someone was in a hurry in 2011. A specification is something else: it says what the system must do, and someone signs it off.

The difference matters because in a system with real users, much of what looks like a bug is, in practice, a contract. Hyrum Wright put it into words, and it's now known as Hyrum's Law: with a sufficient number of users of an API, it doesn't matter what you promise in the contract, all observable behaviours of your system will be depended on by somebody. A legacy system has spent years accumulating users of behaviours nobody promised.

The Thoughtworks Radar points to this in another November 2025 entry, which proposes using AI-generated descriptions as an intermediate step for rewriting systems: the goal isn't to hide implementation details, but to introduce “a temporary abstraction” that lets you reason about what the system does before deciding how to rebuild it. Temporary is the key word. The description is a thinking tool, not a requirement.

Two documents, not one

That's where the central idea of this piece comes from. Before an agent modifies a legacy module, you need two separate documents, with different rules:

Record of current behaviour Change specification
What it states What the code does today, bugs included What will be different after the change, and nothing else
Who produces it AI reads the code; a person reviews it and flags anything suspicious A person, together with whoever knows the business
How it's verified With characterization tests that pin each behaviour With new tests that fail before the change and pass after it
Who signs it off Nobody: it's a record, not a requirement Whoever is accountable for the change
What the agent may do Nothing that contradicts it, except what the change specification says Implement exactly that
How long it lives As long as the module exists; updated with every approved change Until the change is merged; then it moves into the record

onext's proposal. The names matter least; what matters is not mixing the two documents

The rule that makes the pair work is simple: the agent may only change a behaviour in the record if the change specification names it. Everything else stays as it is, even if it looks like a bug. If the agent finds something suspicious, it logs it; it doesn't fix it. Fixing it is a business decision, and it goes into the next change specification with its own sign-off and test.

In the rounding example, the record would have said “the amount is rounded per line before summing, not on the total”, with a test pinning it and a “suspicious, ask finance” flag. The agent would have added the field without touching the rounding. And someone would have asked, instead of finding out from the client.

It's the same logic we argue for with new code in the specification is where you sign off, with one difference: in legacy there's a prior document nobody signs, because nobody decided it. It's simply recorded.

The safety net: characterization tests

A record in prose is a hypothesis. The AI may have misread a branch, and the text won't say so. What turns the record into something reliable is pinning each statement with a test that runs.

There's been a technique with its own name for that for twenty years. Michael Feathers coined the term characterization test in 2004, in Working Effectively with Legacy Code: a test that describes the actual behaviour of existing code to protect it from unintended changes. It doesn't say whether the code is correct; it says whether it has changed. When one fails, someone decides whether the change was intended.

There's a twist worth pointing out here, because it contradicts something we've written. In AI-generated tests we explained that a test written from the code doesn't test it: it only describes it, which is why it doesn't catch bugs. In new code that's a problem. In legacy it's exactly what you want. A characterization test is, by definition, a description of the code. AI generating them quickly and in volume is a real advantage, on one condition: someone reviews which behaviours have been pinned and flags the known bugs, so that a year from now nobody reads them as requirements.

The question to ask before handing a change to an agent: if this module starts behaving differently in something the change specification doesn't mention, which test would fail? If the answer is “none”, the agent isn't working over a safety net, but on trust that it won't touch anything else.

Specify the seam, not the system

The obvious objection is cost. Specifying an entire legacy system before touching it is a months-long project that goes stale before it's done. And Spec-Driven Development tools don't help as much as you might expect: in her October 2025 analysis of Kiro, Spec Kit and Tessl, Birgitta Böckeler noted that introducing two of them into an existing codebase seemed to be even more work, and when she tried Kiro on a small bug the full workflow felt like using a sledgehammer to crack a nut.

The way out is the one gradual modernisation has used for twenty years. Martin Fowler calls it the strangler fig: instead of rewriting in one go, you identify seams along which to split the system and move functionality across bit by bit, so that investment and returns also arrive gradually. Applied to specifications, it means specifying only the seam where the change comes in: the module you're about to touch, its inputs and outputs, and the behaviours its neighbours expect from it. The record grows at the pace of the changes, not ahead of them.

In practice, the order is this. First, AI reads the module and drafts the record. Second, the characterization tests are generated and a person flags anything suspicious. Third, the change specification is written, short, naming each behaviour in the record that will change. Only then does the agent act, with an explicit instruction to change nothing else. Reviewing the result no longer asks “is this right?”, which hardly anyone can answer in legacy, but “has anything changed that wasn't in the specification?”, which the tests answer on their own. It's the question we proposed in the review that still expects an author.

What you gain, beyond the change

There's a side effect that's usually worth more than the change itself: after a few iterations, the team has, for the first time, a reliable document of what the system does, backed by tests that prove it. It's what almost no legacy system has, and what makes everything else possible, from the next modification to a full migration. And it's exactly the kind of artefact Spec-Driven Development puts at the centre: not the code, but the verifiable description of what it must do.

What you don't gain is speed on the first change. The first one costs more than asking the agent to do it directly. The second, in the same module, costs less. And the third is the one you can delegate with confidence, because the safety net is already there. It's also a good way to see what a developer is capable of: someone who finds odd behaviour and logs it instead of fixing it is showing the judgement we describe in what to measure when hiring developers.

Frequently asked questions

Can Spec-Driven Development be applied to a legacy system?

Yes, but not by starting with a specification of the whole system. SDD tools were designed mainly for new code, and Thoughtworks' Birgitta Böckeler noted in October 2025 that introducing two of the three she analysed into an existing codebase seemed to be even more work. What works is specifying only the part you're about to touch, with two separate documents: one describing what the code does today and another stating what will change.

What's the difference between a description of the code and a specification?

A description says what the code does, including its errors and accidents. A specification says what it must do, and someone signs it off. When AI analyses legacy code it produces descriptions, and they're very useful. The problem starts when they're treated as specifications and an agent “fixes” a behaviour that someone in production needs exactly as it is.

What is a characterization test?

It's a test that pins down the actual behaviour of existing code, not the behaviour it should have. Michael Feathers coined the term in 2004, in “Working Effectively with Legacy Code”. It doesn't tell you whether the code is correct, but whether it has changed: if a characterization test fails after a modification, someone has to decide whether that change was intended or a side effect.

Can AI generate the characterization tests?

Yes, and it's one of the few cases where a test generated from the code is exactly what you need. In new code, a test written from the implementation merely describes it. In legacy, describing current behaviour is the goal. The condition is that a person reviews which behaviours have been pinned and flags the ones that are known bugs, so nobody mistakes them for requirements.

What should I do with the bugs AI finds while analysing legacy code?

Log them, don't fix them in the same pass. Odd behaviour may be a bug or something a client, a report or an integration nobody remembers depends on. You pin it with a characterization test flagged as suspicious, ask whoever knows the business and, if you decide to fix it, it goes into the change specification as an explicit, signed-off modification with its own test.

How much of the system needs specifying before you start?

The minimum for the change you're about to make: the module or seam where the modification comes in, its inputs and outputs, and the behaviours that must not change. Specifying an entire legacy system before touching it is a months-long project that goes stale before it's finished. You specify piece by piece, at the pace of the changes, as in a gradual migration.

Sources

Jordi García
Written by
Jordi García
Tech Lead at onext

Jordi García is Tech Lead at onext. He works on bringing AI into governed production across development and product teams —with Spec-Driven Development, context engineering and human verification at every step— and authors onext's technical insights on the method, quality and cost of applied AI.

LinkedIn →

The first legacy module you hand to an agent, with a safety net

We pick a seam with your team, draft the record of current behaviour, pin it with characterization tests and make the first specified change. The method stays with your team for the next one.

See how we work

No system rewrite. No stopping delivery.