Technology

Measuring the ROI of AI in the SDLC: DORA Metrics Instead of Productivity Claims

Miłosz Cupiał
Head of Delivery
September 30, 2026
9
min read

Executive summary Almost every engineering organisation now uses AI tools, and almost every board now asks the same question: what are we getting for it? The answers usually come from vendor statistics, suggestion acceptance rates or developers' own impressions. None of these proves business value. Independent research shows that developers can feel faster with AI while actually working more slowly, and that faster coding does not automatically translate into faster, more stable releases. A credible ROI case needs a baseline, a small set of delivery and quality metrics, and an honest comparison of benefits and costs. This article explains why the usual numbers mislead, which metrics to use instead, how to set up measurement and what realistic results look like.

The Problem With "Our Developers Feel 30% Faster"

Ask a team using AI coding assistants whether they are more productive, and most will say yes. The trouble is that perceived productivity and measured productivity are not the same thing.

In mid-2025, the research organisation METR published a randomised controlled trial with experienced open-source developers working on repositories they knew well. Before the study, the developers expected AI to make them 24% faster. When they used AI tools, tasks actually took 19% longer on average. Even after the study, participants still believed they had been around 20% faster.

This does not mean AI tools are useless. The study covered a specific setting, and other research has found gains in different contexts. What it does show is that developer impressions are an unreliable basis for investment decisions. The same applies to the numbers most often quoted in vendor materials:

  • Lines of code generated say nothing about whether the code was needed, correct or maintainable.
  • Suggestion acceptance rates measure how often developers click "accept", not whether the accepted code created value.
  • The share of code written by AI can rise while quality and delivery speed fall.

Why Faster Coding Does Not Automatically Mean Faster Delivery

Writing code is only one step in getting a change to production. Requirements, review, testing, security checks, deployment and incident handling all take time. If AI speeds up coding but nothing else changes, the bottleneck simply moves. Reviewers face more and larger pull requests, test suites struggle to keep up and more changes wait in the queue.

Google's DORA research programme observed a similar pattern. Its 2024 report found that higher AI adoption was associated with small improvements in documentation quality, code quality and code review speed, but also with an estimated 1.5% decrease in delivery throughput and a 7.2% decrease in delivery stability for every 25% increase in AI adoption. One explanation the researchers offer is that AI makes it easier to produce larger changes, and large changes are riskier to release.

The lesson is not to avoid AI, but to measure outcomes at the level of the whole delivery system, not at the level of individual keystrokes.

What to Measure Instead

A useful measurement framework works on three levels: delivery performance, quality and flow, and business impact.

Delivery performance: the DORA metrics

The four DORA metrics have become the most widely used indicators of software delivery performance:

  • Deployment frequency: how often the team releases changes to production.
  • Lead time for changes: how long it takes for a committed change to reach production.
  • Change failure rate: the share of deployments that cause a failure requiring remediation.
  • Time to restore service: how quickly the team recovers when something goes wrong.

Their strength is balance. Two metrics measure speed and two measure stability, so a team cannot improve one side by quietly sacrificing the other.

Quality and flow

DORA metrics show the outcome but not always the cause. A few supporting metrics help explain what is happening:

MetricWhat it reveals
Pull request review timeWhether AI-generated changes are creating a review bottleneck
Pull request sizeWhether changes are growing larger and riskier
Automated test coverage on critical pathsWhether the safety of changes keeps pace with their volume
Rework rateHow much recently written code has to be changed again soon after
Defects found in productionWhether quality problems are reaching customers

Business impact

Finally, delivery metrics must be connected to what the business cares about: time to market for new features, engineering capacity freed for roadmap work, and the cost of incidents and rework. This is where the ROI calculation happens.

What not to measure

Just as important is what to leave out. Avoid measuring individual developers' output, lines of code or AI usage per person. Individual metrics encourage gaming, damage trust and say little about team outcomes. Once a measure becomes a target, people optimise for the measure rather than the result.

How to Set Up Credible Measurement

1. Establish a baseline before rollout. Collect at least four to eight weeks of data before introducing or expanding AI tools. Without a baseline, any later improvement is an assumption.

2. Start with a pilot and a comparison. Roll out AI to one or two teams first, while similar teams continue as before. This separates the effect of AI from seasonal changes, reorganisations or roadmap shifts.

3. Collect data automatically. The version control system, CI/CD pipeline and issue tracker already contain most of the data needed. Automated collection is more reliable and less burdensome than manual reporting. Well-instrumented CI/CD pipelines and cloud operations make this considerably easier.

4. Allow for a learning curve. Teams often become slower before they become faster, as they learn which tasks AI handles well and adjust their workflows. Evaluate results over at least one to three months, not after the first week.

5. Combine data with developer feedback. Short, regular surveys on satisfaction, cognitive load and trust in AI output explain the numbers and reveal problems before they show up in delivery metrics.

6. Adjust the process, not just the tools. If review time grows, introduce smaller pull requests or AI-assisted review. If stability drops, strengthen automated testing. Measurement is only useful if it leads to changes in how the team works.

From Metrics to ROI: A Simple Calculation

A credible ROI calculation compares realistic benefits with full costs. The following example uses illustrative figures only to show the logic.

A product organisation has 20 engineers with an average fully loaded annual cost of €90,000, giving a total engineering cost of €1.8 million per year. After a structured rollout, the measured net efficiency gain across the delivery process is 10%. That corresponds to roughly €180,000 worth of engineering capacity per year that can be redirected to roadmap work.

On the cost side, the organisation needs to include tool licences, the enablement programme and training, the time spent setting up workflows and measurement, and any additional review effort. If these total €40,000 in the first year, the net benefit is around €140,000 and the return is roughly 3.5 times the investment.

Two points matter more than the exact figures. First, the gain must be net: time saved in coding but lost in review or incident handling does not count. Second, freed capacity only creates value if it is used for work that matters, such as shipping roadmap features earlier or reducing technical debt.

What Realistic Results Look Like

Structured AI adoption in software teams rarely doubles productivity, but it can deliver consistent, measurable improvements. In engagements described on our AI-assisted software delivery page, results have included:

  • around 20 to 25% shorter average pull request review time,
  • an increase of around 15 to 20 percentage points in automated test coverage,
  • an improvement in deployment frequency of about one tier on the DevOps maturity scale,
  • roughly 15% of engineering capacity freed for new feature development,
  • in a regulated fintech environment, a net efficiency gain of around 10 to 15% with a lower change failure rate and full audit traceability.

These are modest numbers compared with marketing claims, and that is exactly why they are credible. A sustained double-digit improvement, backed by delivery data, is a strong business case.

Results also depend on the starting point. In large legacy codebases with little test coverage, AI tools struggle to deliver value safely. In such cases, an AI refactoring assessment can help identify where modernisation should come first.

Common Pitfalls

Measuring too early. Judging AI after two weeks captures the learning curve rather than the long-term effect.

Tracking tool usage as success. High adoption shows that people use the tools, not that the business benefits.

Ignoring stability. Faster throughput at the cost of more incidents is not a gain.

Measuring individuals. Individual metrics distort behaviour and undermine the trust that honest measurement depends on.

Skipping the baseline. Without a starting point, even real improvements cannot be demonstrated convincingly.

If You Can't Measure It, You Can't Scale It

AI in software development is moving from experimentation to budget line. Organisations that can show measurable improvements in delivery speed and stability will find it far easier to justify further investment, while those relying on impressions will face growing scepticism from finance teams and boards.

Investors are asking the same questions. In a technology due diligence, a team that can present DORA metrics over time and explain the effect of AI on its delivery performance stands out clearly from one that can only describe its tool stack. If you want to introduce or scale AI in your engineering teams with a clear baseline and measurable results, let's talk about how our AI-assisted software delivery programme starts with a focused pilot and a measurement framework from day one.

FAQ

FAQ - Measuring the ROI of AI in the SDLC: DORA Metrics Instead of Productivity Claims

What are DORA metrics?

DORA metrics are four indicators of software delivery performance developed by the DevOps Research and Assessment programme, now part of Google Cloud: deployment frequency, lead time for changes, change failure rate and time to restore service. Together they measure both the speed and the stability of delivery.

How long does it take to see the effect of AI tools in the metrics?

Teams usually need several weeks to adapt their workflows, and some initially become slower. A meaningful assessment typically requires one to three months of data after rollout, compared with a baseline collected before it.

Should we measure the productivity of individual developers?

No. Individual metrics tend to encourage gaming, damage trust and reveal little about team outcomes. Measuring at team and system level gives a more accurate picture and avoids these side effects.

Is the share of AI-generated code a useful metric?

Not on its own. A higher share of AI-generated code says nothing about whether delivery became faster, more stable or more valuable. It can be useful as context, but it should never be treated as a success measure.

What ROI is realistic for AI in software development?

It depends on the starting point, but structured adoption often delivers a net efficiency gain in the range of around 10 to 15%, alongside improvements in review time and test coverage. Whether that translates into a strong ROI depends on keeping costs under control and using the freed capacity for valuable work.

Which data sources do we need to measure DORA metrics?

In most cases, the version control system, the CI/CD pipeline and the issue tracker provide everything needed. Incident management data is also required to measure change failure rate and time to restore service accurately.

Articles you might be interested in

AI in Software Delivery for Regulated Industries: How to Keep Auditability and Compliance Intact

September 30, 2026
Minutes

Safety Net First: How AI Generates Characterization Tests for Legacy Code With No Tests

September 30, 2026
Minutes

Pre-Exit Readiness: How to Prepare a Portfolio Company for Buyer Tech Due Diligence in 90 Days

September 30, 2026
Minutes