Most teams that have adopted AI in their development still measure their performance exactly as they did three years ago. Velocity. Story points. Number of closed tickets. Logged hours. The problem isn't AI. The problem is that the measurement system hasn't evolved.
And when you change how you produce software but don't change how you measure it, the metrics start to lie.
The problem: we're still measuring teams with pre-AI metrics
Before generative AI, the cost of producing code was relatively proportional to human effort. Not anymore.
A developer with AI can generate more code in less time, prototype faster and solve repetitive tasks in seconds. But that doesn't necessarily mean the team delivers more value.
Traditional metrics start to distort:
If AI reduces implementation effort, story points stop representing real complexity.
A rise in velocity can hide a rise in rework.
More lines can mean more technical debt.
Closing more tasks doesn't mean the product advances strategically.
What really changes when AI enters the team
1. The marginal cost of code trends to zero
Code stops being the scarce resource. What becomes scarce is:
- Clarity in decisions
- Shared context
- Architectural quality
When code is cheap, bad decisions are far more expensive.
2. The bottleneck shifts
The bottleneck used to be technical. Now it's usually in:
- Problem definition
- Prioritization
- Architecture
- Alignment between product and technology
AI accelerates. It amplifies good practices, solid architecture and clear decisions.
AI multiplies the chaos. More inconsistent code, more variability, more debt accumulated in less time.
3. Architectural risk increases
AI can generate "seemingly correct" solutions that don't respect system patterns, introduce unnecessary complexity and make future maintenance harder.
If there's no solid system (SDD, structured context, clear criteria), the impact shows up months later.
The 5 metrics that do matter in teams working with AI
If code is no longer the bottleneck, what should we measure? These are the metrics that really indicate sustainable productivity.
Not from "In Progress". From the strategic decision to the moment the feature generates impact in production.
If AI is working, this number should drop consistently.
How much of the generated code needs to be rewritten, refactored or fixed within 30 days?
If this ratio rises, AI isn't accelerating: it's generating debt faster.
Each feature should be analyzed in terms of added complexity, new couplings and future maintenance cost.
If debt grows faster than the capacity to simplify, the system degrades.
How long does it take the team to correctly understand what it should build? When implementation is fast, interpretation errors become more expensive.
This is where productivity is really gained.
Instead of measuring technical output, measure the impact generated per decision taken and value delivered per strategic cycle.
A mature team with AI doesn't produce more code. It produces better decisions per unit of time.
Metrics you should drop (or reinterpret)
It doesn't mean you should eliminate them, but you should stop using them as primary metrics:
Activity metrics, not impact metrics
In the AI era, measuring activity is dangerous: it can give a sense of progress while the system's fragility increases.
A mini diagnostic for CTOs
If you're adopting AI in your team, ask yourself these questions:
Self-assessment: is AI amplifying structural problems?
If you answer "yes" to three or more, AI is probably amplifying structural problems.
Conclusion: measuring badly is more dangerous than implementing badly
Implementing AI without changing the system is risky. But measuring the impact badly is even worse.
Because you may believe you're accelerating when you're actually accumulating complexity.
In the AI era, code stops being the bottleneck.
The system becomes it.
And if you want to know whether your team is really improving its productivity, start by measuring what has changed.
What SDD solves: At onext we implement Spec-Driven Development precisely so teams measure what matters: impact per decision, not output per hour. Structured specifications that align the team before writing code. Teams working with SDD cut time per feature by 75% because they eliminate rework at the source: the definition.
Further reading: Code quality in the AI Coding era | You measure incidents, but ignore the metrics that matter
Methodology: At onext we help CTOs redesign their measurement systems as part of our AI Centers of Excellence. From activity metrics to sustainable-impact metrics.

Jordi García is Tech Lead at onext. He works on bringing AI into governed production across development and product teams —with Spec-Driven Development, context engineering and human verification at every step— and authors onext's technical insights on the method, quality and cost of applied AI.
LinkedIn →