Skip to main content
onext technology
AI May 17, 2026 - 8 min read

The 70% gap: your AI adoption in the development team isn't a technical problem — it's an organizational problem

If you've spent months paying for premium Copilot, Cursor or Windsurf licenses for your whole team and the internal conversation is still "we're not seeing the return we expected", you're very likely attacking 30% of the problem.

Jordi García
Tech Lead at onext
CTO in a Spanish tech-startup meeting room reviewing AI adoption metrics in their development team

If you've spent months paying for premium Copilot, Cursor or Windsurf licenses for your whole development team and the internal conversation is still "we're not seeing the return we expected", you're very likely attacking 30% of the problem. The other 70% — the one no vendor will sell you because it doesn't come on an invoice — is organizational. This piece explains what's in that 70%, why you mistake it for a technical problem every quarter, and the four concrete questions your next technology committee can use to tell the two apart.

The most expensive misunderstanding of the quarter

I've spent several months seeing the same pattern in CTOs and VPs of Engineering at mid-sized companies. The conversation with their leadership team plays out on an almost identical script: "we've invested tens of thousands of euros in Copilot licenses for the developers; the internal survey says the team is happy with the tool; but release velocity hasn't changed, production bugs keep rising, and the PR review ratio doesn't improve".

It's exactly the pattern we see at mid-sized companies that move from per-developer licenses to the operating model of an enterprise AI platform governed by the client: the problem isn't solved by adding more Copilot seats, but by redefining how you operate with AI at the organizational level.

When the conversation reaches the CFO or the board, the question that lands on the table is this: "what do we do? Do we buy another, more powerful tool? Do we move up to Cursor Enterprise? Do we contract Copilot Workspace?". The natural intuition is to scale the technical investment: if one tool doesn't return, a better one will.

That intuition is what at onext we call the 70% misunderstanding. And it's the reason why more than 40% of agent-based projects will be cancelled before the end of 2027, according to Gartner.

The figure few vendors will tell you

After working with twelve development teams that have completed our AI transformation program — teams of between twenty and two hundred developers — the pattern we consistently see is this: roughly 70% of AI success in a team isn't technical. It's organizational.

The remaining 30% is technical: the tool you contract, the model, the stack, the IDE integration. It's important. But it's bounded, relatively easy to buy, and every vendor in the market competes for that 30%. The other 70% — the one that really separates productive teams from stuck teams — lives in five organizational dimensions that no tool solves on its own. They're the five that an enterprise platform like onext Enterprise AI is designed to operationalize from day one — not as an add-on, as a primary layer.

1. Defined human-agent lanes

Who decides what, who reviews what, at which point in the flow the human signs off. Without this, each developer builds their own informal arrangement with the AI and the organization loses consistency.

2. Operational metrics framework

Not "we go faster with Copilot" but average PR review cycle, production bugs, onboarding time. Concrete, comparable, monthly figures.

3. Quality gate governance

Who defines when an agent's output can be merged. What gets validated automatically. What gets validated by a human. And, critically: who updates those rules when the team learns. This is exactly the logic that at onext we turn into operational quality-gate governance at the platform level, not a Confluence page nobody reviews.

4. Dual training

10-20% creators (they design workflows, project constitutions, SDD rules) and 80-90% power users (they operate with that in their day-to-day). Confusing the two roles is the recipe for a transformation that doesn't scale.

5. Sustaining plan

What happens after six months when the novelty fades and workflows start going stale. Who updates them. How often. Who measures whether the team keeps accumulating advantage or the curve flattens. Without an enterprise operating model with measurable sustaining, the transformation erodes past the second quarter.

All five are designed with method. They aren't bought. And that explains why the CTO who only buys better tools, cycle after cycle, keeps seeing the same result: the tool works, the problem persists.

The METR study as evidence of the gap

The study by the Model Evaluation & Threat Research Institute documented something worth having on any technical committee's table: in its experiment on experienced developers, participants using AI reported +20% subjective productivity while being −19% less productive measured in useful commits, PR review cycle and production defects.

The paradox isn't a flaw in the technology — the technology in that experiment was first-class. The paradox is the direct consequence of the organizational gap: developers accelerated by the tool without method lose time reviewing PRs the agent signed off badly, rewriting tests the agent didn't understand, debugging latent bugs the generated code doesn't exhibit immediately. Input acceleration doesn't translate into useful-output acceleration. The team feels it's going faster. The team goes slower.

That's the operational signature of the missing 70%.

The four questions for your next committee

If you're taking the AI adoption topic to your organization's next technical committee, these are the four questions that distinguish a technical problem from an organizational problem.

Question 1: Do we have the human-agent lanes for each critical workflow defined in writing, or does each developer improvise them by their own judgment? If the answer is "they improvise them", the problem isn't the tool — it's the absence of operational governance. Buying a better tool doesn't solve it.

Question 2: Do we measure real productivity (useful commits, PR cycle, production defects, onboarding time) or perceived productivity (how fast we feel)? If the answer is "perceived", we're probably in the METR paradox. Buying a better tool amplifies the paradox, it doesn't solve it.

Question 3: Do we have people with judgment on the team to define and update the quality gates, or is the team waiting for the vendor to do it for them? If the answer is "the vendor does it", the organization has ceded governance to the provider. That isn't transformation — it's dependency.

Question 4: Does the organization have a six-month sustaining plan, or are we assuming the tool updates itself and the processes do too? If the answer is "we assume it updates itself", sustaining doesn't exist. The curve will flatten at six-to-nine months and you'll be back in today's conversation.

If all four answers point to an organizational gap (not a technical one), the next step isn't buying another license. It's designing the enterprise operating model that closes that gap through method — exactly what we do at onext Enterprise AI for development teams of 100 to 1,000 people.

If three or four answers are negative: your next step isn't buying a more expensive tool. It's redesigning operational lanes. That doesn't come on an invoice — it's designed with method.

Want to know if your team is trapped in the 70%?

We've turned the four questions above into an interactive 16-question checklist. You complete it in 5–7 minutes and receive a personalized diagnosis by email: well-covered dimensions, specific gaps and priorities for the next 90 days.

Complete the free checklist

The internal benchmark behind the thesis

When we say "70% of AI success is organizational" it isn't an intuition. It's the accumulated pattern from twelve development teams we've supported through AI transformation — a twelve-to-sixteen-week program, with four overlapping phases (diagnosis, redesign, deployment, transfer). The consistent result when the program is executed well:

  • ×7 delivery velocity (specific to each team, depending on the prior baseline).
  • 0 sprints lost during the organizational redesign.
  • −50% time-to-production: half the time between a ticket leaving Discovery and landing in production.
  • +40% code consistency measured in linting, test coverage and post-merge rework.
  • 1-2 senior FTE/year freed up per team on tasks now absorbable by supervised agents.

All five figures come from the organizational redesign. The tool is the same — Claude, mainly, with Cursor or Copilot in the IDE. What changes is how the team relates to the tool.

The honest quote that closes the conversation

"Agents aren't yet reliable enough for complex engineering."

— Andrej Karpathy

The person saying this isn't an anti-AI skeptic. He's one of the most respected AI engineers in the world. The operational conclusion isn't "don't use AI". It's: use it with method.

And the method isn't bought. It's designed, installed in the team, maintained. The five organizational dimensions I described above — lanes, metrics, gate governance, dual training, sustaining — are that method. And they're what the 30% of the technical problem solved by Copilot, Cursor or this quarter's model will never cover.

If in your next technical committee the conversation is "we're not seeing a return with AI, which tool do we buy now?", you have the wrong diagnosis. The return won't come from the tool — it'll come from the method with which your organization operates with it.

We've prepared four questions for your next committee that distinguish a technical problem from an organizational one in under ten minutes. One A4 page, no email required:

Download the 4 questions for your next committee

If the four answers leave you uncomfortable, let's talk: info@onext.es — subject "AI organizational audit". A 30-minute conversation, no commitment.

Sources: METR study on developer productivity with AI tools (the +20%/−19% paradox); Gartner on the cancellation of agent-based projects before 2027; public statements by Andrej Karpathy on agent reliability in complex engineering.

Further reading: Month 6: the sustaining system that prevents it · The 3 canonical claims of the AI Engine

Jordi García
Written by
Jordi García
Tech Lead at onext

Jordi García is Tech Lead at onext. He works on bringing AI into governed production across development and product teams —with Spec-Driven Development, context engineering and human verification at every step— and authors onext's technical insights on the method, quality and cost of applied AI.

LinkedIn →

Does your organization have the organizational 70% covered?

At onext we diagnose the five organizational dimensions of your development team with AI and design the transformation plan without stopping deliveries.

See how we work

A 30-minute conversation. No commitment.