Technology
Software

Safety Net First: How AI Generates Characterization Tests for Legacy Code With No Tests

Miłosz Cupiał
Head of Delivery
September 18, 2026
9
min read

Executive summary The biggest obstacle to modernising legacy software is rarely the technology itself. It is fear: nobody knows exactly what the system does, there are no tests to confirm it, and every change risks breaking something that customers depend on. Characterization tests solve this by recording how the code behaves today, before anyone changes it. Writing them by hand has always been slow and tedious, which is why many teams skipped them. AI changes that equation. Used correctly, it can analyse unfamiliar code, propose test cases and generate much of the scaffolding in days rather than months. Used carelessly, it produces tests that look convincing but protect nothing. This article explains how AI-generated characterization tests work, where they add real value, where they fall short and how to build a safety net you can actually trust.

Why Legacy Code Without Tests Is So Hard to Change

Most organisations have at least one system that everyone is afraid to touch. It works, it earns money, and it has grown for ten or fifteen years through the hands of developers who have long since moved on. The problem is not that the code is old; it is that its behaviour is undocumented and unverified.

Without automated tests, every change becomes a gamble:

  • developers cannot tell whether a modification breaks an edge case nobody remembers,
  • manual regression testing grows with every release and still misses things,
  • refactoring is postponed indefinitely because the risk seems higher than the benefit,
  • knowledge stays in the heads of a few senior people, who become bottlenecks.

The result is a vicious circle. Code without tests is risky to change, so nobody improves it, so it becomes even harder to change. Breaking that circle requires a safety net before any structural work begins.

What Characterization Tests Are and Why They Come First

The term was popularised by Michael Feathers in Working Effectively with Legacy Code. A characterization test does not check what the system should do. It records what the system actually does today and fails if that behaviour changes.

That distinction matters. In legacy systems, the specification is often missing, outdated or contradicted by years of patches. Customers may even rely on behaviour that was originally a bug. Characterization tests accept the current behaviour as the reference point, so that any change during refactoring becomes visible and deliberate rather than accidental.

In practice, characterization tests take several forms:

  • Unit-level tests that call individual functions or classes with representative inputs and assert the current outputs.
  • Golden master tests that feed a large set of inputs through a module or the entire system and compare the complete output with a stored reference.
  • API or integration-level tests that capture how services respond to requests, including error handling and edge cases.

Once this net is in place, refactoring stops being a leap of faith. If behaviour changes, a test fails, and the team can decide whether the change was intended.

Where AI Changes the Economics

Characterization testing has always been a sound idea with a practical problem: it takes a lot of time. A developer must read unfamiliar code, trace the execution paths, find meaningful inputs, isolate dependencies and write the assertions. For a large codebase, that could mean months of effort before any visible improvement.

AI coding tools shift this balance in several concrete ways.

Code comprehension at scale. AI can read a module, summarise what it appears to do, list its branches and point out hidden dependencies such as database calls, file access or global state. This shortens the most time-consuming part of the work: understanding the code.

Systematic input discovery. AI is good at proposing inputs that exercise different paths, including boundary values, empty collections, invalid formats and unusual combinations that a developer might not think of under time pressure.

Scaffolding and isolation. Much of the effort in legacy testing goes into setting up test harnesses, creating test doubles and breaking dependencies. AI can generate this boilerplate quickly and consistently across many modules.

Documentation as a by-product. The same analysis that produces tests also produces readable descriptions of how the code behaves. For a system with no documentation, this alone can be valuable.

In our experience, teams using AI-assisted test generation within a structured delivery process can increase automated test coverage significantly in a matter of weeks. In one engagement described on our AI-assisted software delivery page, automated test coverage rose by around 15 to 20 percentage points. The gain comes from combining AI speed with engineering judgement, not from AI alone.

Where AI Falls Short

AI-generated tests carry specific risks that teams must manage deliberately. Ignoring them produces a false sense of security, which is worse than having no tests at all.

Assertions based on assumptions, not observations. An AI tool may write an assertion based on what it thinks the code returns. For characterization tests, that is a fundamental error. Expected values must come from actually running the current code, not from the model's interpretation of it.

Weak tests that always pass. Tests can execute code without checking anything meaningful. They raise coverage figures while protecting nothing. High coverage is not the same as a strong safety net.

Non-deterministic behaviour. Legacy code often depends on the current time, random values, external services or shared state. Tests that ignore this become flaky, and flaky tests are quickly ignored by the team.

Sensitive data. Golden master approaches sometimes use production-like data. Copying real customer data into test fixtures or AI tools raises confidentiality and GDPR concerns. Data should be anonymised or synthesised.

Code confidentiality. Feeding proprietary code into AI tools requires approved tools with appropriate enterprise terms and clear rules on what may be shared.

A Practical Process for AI-Generated Characterization Tests

A reliable approach combines AI generation with verification steps that keep the results honest.

1. Choose where to start. Do not try to cover the whole system at once. Focus on modules that combine high business importance, frequent changes and high complexity. These are the places where a safety net pays off first and where refactoring is usually most urgent.

2. Let AI map the behaviour. Ask AI to analyse the selected module, describe its responsibilities, list execution paths and dependencies, and propose a set of test scenarios. A developer reviews this map and corrects or extends it.

3. Generate tests and pin real outputs. AI generates the test code and inputs, but the expected values are captured by running the existing code. Every characterization test must pass against the current system before it is accepted.

4. Stabilise the tests. Control time, randomness and external calls through test doubles or fixed configurations. Run the suite repeatedly to catch flaky tests early.

5. Check that the tests actually protect. Coverage shows which lines were executed, not whether the tests would catch a change. Mutation testing, which introduces small deliberate changes into the code and checks whether tests fail, is a far better measure of how strong the safety net really is.

6. Integrate into the pipeline. Characterization tests only help if they run automatically on every change. Adding them to the CI/CD pipeline, supported by solid DevOps and cloud practices, turns them into a permanent guard rather than a one-off exercise.

7. Refactor, then evolve the tests. With the net in place, the team can refactor in small, verifiable steps. Over time, characterization tests are gradually replaced or complemented by tests that describe intended behaviour, especially where the team decides that existing behaviour was actually a bug.

How to Tell Whether the Safety Net Is Good Enough

Before starting larger structural changes, it is worth checking the net against a few simple criteria:

QuestionWhat good looks like
Do all tests pass reliably on the current code?Yes, across repeated runs, without manual intervention.
Are expected values based on real execution?Every assertion was captured from the running system, not inferred.
Would the tests catch meaningful changes?Mutation testing shows that most introduced changes cause failures.
Are critical business paths covered?The most valuable and most frequently changed flows are protected first.
Do the tests run automatically?They are part of the CI/CD pipeline and block faulty changes.
Is sensitive data excluded?Fixtures use anonymised or synthetic data.

If the answer to most of these questions is yes, the team can refactor with confidence. If not, it is better to strengthen the net than to start the modernisation too early.

Characterization Tests in a Broader Modernisation Strategy

Characterization tests are not a goal in themselves. They are the foundation for everything that follows: refactoring, gradual replacement of components, migration to new architectures or AI-assisted code transformation.

They also make modernisation plans more credible. A team that can show a stable safety net around its critical modules can estimate refactoring effort more accurately, take on larger changes with less risk and demonstrate progress to management with hard data. The same evidence is valuable in a technology due diligence, where test coverage and the ability to change code safely are key signals of how much a product is really worth.

For organisations that are unsure where to begin, an AI refactoring assessment helps identify the modules where a safety net matters most, estimate the effort and decide which parts of the system should be refactored, replaced or left as they are.

Modernise With a Net, Not a Leap of Faith

Legacy systems do not become safer by waiting. Every year without tests adds more undocumented behaviour, more workarounds and more reliance on a shrinking group of people who understand the code. AI does not remove the need for engineering discipline, but it makes the first and most important step far cheaper: capturing what the system does today.

Teams that invest a few weeks in a trustworthy safety net gain something that is hard to achieve any other way: the freedom to change their core systems without fear. If you want to modernise a critical system without risking what already works, let's talk about how our product and application engineering teams can build that net and take the modernisation forward.

‍

FAQ

FAQ - Safety Net First: How AI Generates Characterization Tests for Legacy Code With No Tests

What is the difference between characterization tests and regular unit tests?

Regular unit tests check whether code behaves as specified. Characterization tests record how code behaves today, regardless of whether that behaviour is correct. They are designed for legacy systems where the specification is missing or unreliable, and their main purpose is to make any change in behaviour visible during refactoring.

Can AI write characterization tests on its own?

Not reliably. AI can analyse code, propose scenarios and generate most of the test code, but expected values must be captured by running the existing system, and a developer must review the scenarios and results. Without these steps, AI may produce tests that reflect assumptions rather than real behaviour.

What if the tests capture existing bugs?

That is intended. Characterization tests document current behaviour, including defects, so that nothing changes unnoticed. When the team decides to fix a bug, it updates the relevant test deliberately. This makes every behavioural change a conscious decision rather than a side effect.

How much test coverage do we need before refactoring?

There is no universal number. It is more important to protect the critical business paths and the modules you plan to change than to reach a specific percentage across the whole codebase. Mutation testing is a better indicator of readiness than line coverage alone.

Is it safe to use AI tools on proprietary legacy code?

It can be, provided the organisation uses approved AI tools with enterprise terms, has clear rules on what code and data may be shared, and keeps sensitive production data out of prompts and test fixtures. These safeguards should be in place before the work starts.

How long does it take to build a safety net for a legacy module?

It depends on the size and complexity of the module, but with AI-assisted generation and a structured process, a first reliable safety net for a critical module can often be built in days or a few weeks rather than months. Covering an entire large system is a gradual process that typically follows the refactoring roadmap.

Articles you might be interested in

Measuring the ROI of AI in the SDLC: DORA Metrics Instead of Productivity Claims

September 30, 2026
Minutes

AI in Software Delivery for Regulated Industries: How to Keep Auditability and Compliance Intact

September 30, 2026
Minutes

Pre-Exit Readiness: How to Prepare a Portfolio Company for Buyer Tech Due Diligence in 90 Days

September 30, 2026
Minutes