A test for this week: pick the last bug closed by someone junior on your team and ask them to explain, without opening the assistant, why it was failing. Not what changed in the code, but why it failed. If the answer is "the agent found it", the bug is closed, but that person hasn't learned anything from it.
That isn't a criticism. It's the sensible thing to do when the assistant proposes the fix in thirty seconds and the sprint is tight. The problem isn't the junior. It's that the work that used to train them has changed hands without anyone deciding it should.
The argument of this piece is that AI doesn't make juniors unnecessary: it makes unnecessary the work through which a junior became a senior. And since the people who will review and sign off agent-generated code five years from now are today's juniors, that work has to be put back into the process, deliberately, in the place where the work now happens.
Two studies that seem to contradict each other
The first is one of the largest field experiments published on coding assistants. Cui, Demirer, Jaffe, Musolff, Peng and Salz analysed three randomised trials of GitHub Copilot at Microsoft, Accenture and a Fortune 100 company, covering 4,867 developers in total. Those with the assistant completed 26% more tasks. And the gain wasn't evenly spread: developers with shorter tenure and in more junior roles adopted it more and improved more; for those with longer tenure and in senior roles, the effect wasn't significant.
The second is a controlled experiment Anthropic published in January 2026, by Judy Hanwen Shen and Alex Tamkin. 52 developers, mostly junior and all with more than a year of Python, had to learn Trio, an asynchronous programming library they didn't know, and implement two features with it. Half were given an AI assistant. Afterwards, everyone took a quiz on the concepts they had just used. The AI group averaged 50%; the group without AI, 67%. The largest gap was on debugging questions. And the AI group didn't finish significantly faster: about two minutes, within the margin of error.
Read together, they don't contradict each other. They measure different things. The first measures what goes out the door this week. The second measures what stays in the head of the person who did it. A company can improve a lot on the first while getting worse on the second, and no productivity dashboard will show it, because those dashboards measure closed tasks, not judgement gained. It's the same trap we described in the ROI of Copilot and Cursor: what's easy to measure isn't what decides the outcome.
What the junior used to do and no longer does
A junior didn't learn because someone taught them, but because they did a particular kind of work that forced them to understand. That work is precisely what has been automated the most.
| Junior task | What it taught | What happens with an agent |
|---|---|---|
| Writing the first draft | How to break a problem down and which decisions need making | The draft arrives finished; the decisions come already made and nobody sees them |
| Reading the documentation | A mental model of the library, not just the call needed today | The agent uses the library without anyone having read it |
| Debugging their own mistake | Forming hypotheses and ruling them out: the basis of judgement | The error is pasted in and the fix accepted; it's where Anthropic's study saw the largest gap |
| Writing the tests | What the code should do and when it breaks | Tests are generated from the code and describe what it does, not what it should do |
| Defending their pull request | Explaining a change and taking reasoned corrections | The reviewer argues with an agent's code; the author has nothing to defend |
onext's own analysis
We've covered the last two rows from the quality side already: tests generated from the code that prove nothing, and the review that still expects an author. Here we see the other side of the same thing. When the agent takes over those tasks, it's not just a quality control that goes missing: it's the place where people learned how to exercise it.
How the group that did learn used AI
The most useful part of Anthropic's study isn't the average, but what lies underneath it. The authors reviewed the recordings of every session and found six distinct ways of using the assistant. Three were associated with low scores and three with high scores.
- Low scores: handing all the code to the AI from the start; starting alone and gradually handing over more and more; and using AI to debug by pasting errors until something works.
- High scores: generating the code and then working to understand it; asking for code and explanation together; and using AI only for conceptual questions while writing the code themselves.
The groups are small —between two and seven people each— so they're best read as pointers, not laws. But the authors' conclusion is clear: patterns that involve cognitive effort preserve learning even with AI in the loop. What does the damage isn't the assistant, it's stopping thinking while you use it.
Earlier work, in a different setting, points the same way. Bastani and colleagues gave nearly a thousand secondary school students in Turkey access to GPT-4 during their maths practice sessions, in two versions: one that mimicked the standard chat interface and another with instructions designed to protect learning. While they had access, both improved grades. When access was removed, those who had used the standard chat scored 17% lower than students who never had it; for those who used the safeguarded version, that negative effect was largely mitigated. Published in PNAS in 2025, their conclusion reads almost like an instruction for a CTO: the design choices behind the rollout decide whether people learn.
Move the learning to where the work now is
With agents, engineering work doesn't disappear: it moves. It shifts from writing code to saying what the code must do and checking that it does it. That's the idea behind the specification as the place where you sign off. If the work has moved there, the learning has to move with it. In practice, that comes down to three concrete changes.
1. The junior writes the specification and the acceptance criteria
Before the agent touches anything, the junior writes what the change must do, which edge cases exist and how you'll know it works. A senior reviews it in ten minutes. It's the draft that used to be the code: it forces them to break the problem down and make the decisions the agent would otherwise make silently. And it has an advantage the code never had: reasoning errors show up before a single line exists. The acceptance criteria are also the tests the agent can't write on its own without simply describing its own code.
2. In review, the author explains their change without the assistant
A simple team rule: whoever opens the pull request is its author, even if an agent wrote it, and must be able to explain what the change does, what happens if the external call fails and why the obvious alternative was ruled out. Without opening the chat. If they can't, the change isn't ready, even if the tests pass. It takes five minutes per pull request and turns review back into what it used to be: the place where a junior learns from someone with more judgement. It's also how you spot the "hand everything over" pattern the study linked to the lowest scores.
3. For learning, the assistant explains instead of solving
The tools already support this. Claude Code ships an output style called Learning: the assistant explains its choices and, when it reaches a piece with a real design decision, leaves a few lines marked TODO(human) for the person to write, and waits. It can be switched on per user or set in the project configuration for the whole team. There's no need to use it all the time; it makes sense in the first weeks with a new technology, in a module the junior will have to maintain, or when debugging an incident. With a caveat the documentation itself makes: it's an instruction the model follows, not a guaranteed control. It works if the team wants to use it, just like shared instructions.
None of these three changes slows delivery noticeably, and all three reinforce controls the team would need anyway. It's the difference between training as a separate activity, which always loses to the sprint, and training inside the workflow, which is the only kind that survives. We saw this when discussing why transformations break in month 6: what isn't in the process doesn't last.
The uncomfortable lesson: the problem arrives in three years, not now
There's a convenient reading of all this: if the agent does junior work, let's hire fewer juniors. The data suggest some companies are already doing so. Erik Brynjolfsson, Bharat Chandar and Ruyu Chen at Stanford analysed payroll records for millions of workers from ADP, the largest payroll software provider in the United States. By September 2025, employment of software developers aged 22 to 25 had fallen by nearly 20% from its peak in late 2022, while employment of more experienced workers stayed stable or kept growing. Across the occupations most exposed to AI, the relative decline in early-career employment was 16%.
Two honest caveats. This is US data, and the authors present it as early evidence consistent with an AI effect, not as proof of causation. But the argument we care about doesn't depend on that figure. The whole way of working with agents that we advocate —specify, verify, and have a person sign off where there's risk— needs people with the judgement to sign. That judgement can't be bought ready-made on the market indefinitely: someone has to develop it. A company that stops training juniors today is deciding who won't be signing off its code in 2030, and it'll find out when it tries to hire seniors and competes with every other company that made the same decision.
That's why this isn't an HR topic but a question of team architecture. We raised it at the point of entry in what to measure when hiring developers: if the agent can pass the technical test, what you need to assess is judgement. This piece picks up from there: once someone is on the team, that judgement either develops or withers depending on how the work is set up. And that is up to the CTO.
Frequently asked questions
Do junior developers learn less when they use AI?
It depends on how they use it. In the controlled experiment Anthropic published in January 2026, 52 developers, mostly junior, learned a new Python library with and without an assistant. The AI group scored 50% on the follow-up quiz and the group without AI 67%, with the largest gap on debugging questions. But those who used AI to ask conceptual questions or request explanations, rather than to hand over the code, kept their learning.
If AI makes juniors more productive, where's the problem?
Productivity and learning are measured in different places. The field experiments with GitHub Copilot at Microsoft, Accenture and a Fortune 100 company found 26% more completed tasks, with the largest gains among junior staff and recent hires. That measures what ships today. The comprehension quiz measures whether that person will be able to review and sign off code in a few years' time. A company can gain on the first while losing on the second without noticing.
Should juniors be banned from using AI assistants?
No. In Anthropic's data, the usage patterns that preserved learning did use AI: they asked for explanations, asked conceptual questions, or generated code and then worked to understand it. In the study by Bastani and colleagues with nearly a thousand students, a tutor with instructions designed to protect learning largely mitigated the negative effect. The problem is use without design, not the tool.
What is Claude Code's Learning output style?
It's one of the output styles that ship with Claude Code. With it, the assistant explains its choices and, when it reaches a piece with a real design decision, leaves a few lines marked TODO(human) for the person to write, and waits. It can be switched on per user or set in the project configuration for the whole team. The documentation itself warns that it's an instruction the model follows, not a guaranteed control.
How do you check in code review whether the author understands their change?
Ask them to explain it without the assistant open: what the change does, which cases it covers, what happens if the external call fails, and why the obvious alternative was ruled out. If they can't, the pull request isn't ready, even if the tests pass. It's a five-minute question, and it's what turns review into a learning moment rather than just a filter.
Are companies hiring fewer junior developers because of AI?
There are signs of it in the United States. Brynjolfsson, Chandar and Chen at Stanford analysed ADP payroll data and found that employment of software developers aged 22 to 25 had fallen by nearly 20% by September 2025 from its peak in late 2022, while employment of more experienced workers stayed stable or kept growing. The authors present this as early evidence consistent with an AI effect, not as proof of causation.
Sources
- Judy Hanwen Shen and Alex Tamkin (Anthropic), "How AI assistance impacts the formation of coding skills", 29 January 2026, and the full paper, "How AI Impacts Skill Formation", arXiv:2601.20245.
- Kevin Zheyuan Cui, Mert Demirer, Sonia Jaffe, Leon Musolff, Sida Peng and Tobias Salz, "The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers", February 2025 version.
- Hamsa Bastani, Osbert Bastani, Alp Sungu, Haosen Ge, Özge Kabakcı and Rei Mariman, "Generative AI without guardrails can harm learning: Evidence from high school mathematics", PNAS 122 (26), 2025 (full text on PubMed Central).
- Erik Brynjolfsson, Bharat Chandar and Ruyu Chen, "Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence", Stanford Digital Economy Lab, 13 November 2025.
- Claude Code documentation, "Output styles".

Jordi García is Tech Lead at onext. He works on bringing AI into governed production across development and product teams —with Spec-Driven Development, context engineering and human verification at every step— and authors onext's technical insights on the method, quality and cost of applied AI.
LinkedIn →