Skip to main content
onext technology
Leadership 30 September 2026 - 10 min read

If an agent can pass your technical test, you're assessing the agent: what to measure when hiring developers

Canva requires candidates to use AI in its technical interviews; Anthropic asks them not to unless it says otherwise. Both are right, because they're measuring different things. What has become scarce isn't writing code, but knowing when code that looks correct isn't. We explain how to assess that in an interview.

Jordi García
Tech Lead at onext
Two developers review a pull request together on a monitor during a technical interview, one pointing at a line of code, in an office at dusk: hiring developers who can review the code AI generates

Before your next interview, try this: paste your technical test into Claude Code or Cursor and watch the clock. If within a few minutes you have a solution that passes your tests, the test no longer tells a good developer apart from someone who knows how to paste a brief. And if you run it live with AI banned, you're watching how someone works in conditions they'll never have in the job.

You don't have to choose between banning and allowing. You have to decide what you want to measure and build the test for that. And what's worth measuring has changed, because in a team that works with agents, writing code is no longer the hard part. The hard part is spotting the defect in code that looks correct.

Two AI companies, two opposite answers

In June 2025, Canva published a post on its engineering blog with a title that leaves no room for doubt: “Yes, You Can Use AI in Our Interviews. In fact, we insist”. For backend, machine learning and frontend roles, it replaced the computer science fundamentals interview —problems like implementing Conway's Game of Life— with more ambiguous, realistic problems, along the lines of “build a control system for managing aircraft takeoffs and landings at a busy airport”. And it changed what it assesses: how an ambiguous requirement gets broken down, what technical decisions are made with AI's help, whether defects in generated code are spotted and fixed, and whether the result meets production standards.

The most interesting part is what it learned from the first trials. The strongest candidates asked clarifying questions, used AI for specific subtasks without losing control of the whole, and reviewed what it generated with a critical eye. And candidates with no experience of these tools struggled, according to Canva, because they lacked the judgement to guide the AI.

Anthropic keeps a guide for candidates, updated in July 2025, that goes the other way on what matters here. For the CV, AI is fine as long as the first draft is yours. But take-home assessments are done without Claude unless stated otherwise, and in live interviews, in its words, “this is all you”: no AI assistance unless expressly indicated. It wants to see how the person reasons in real time.

Both policies are consistent, because each knows what it's measuring. Anthropic wants to see individual reasoning; Canva, how someone works with the tool they'll use. What's inconsistent is what many companies do: allow AI without changing the test, which measures the model's speed, or ban it in the interview and then expect productivity with agents from day one in the job.

What has become scarce

If writing code is no longer the bottleneck, what is? Two pieces of data make it fairly clear.

The first comes from METR, a research organisation that evaluates AI models. In July 2025 it published an experiment with sixteen experienced open-source developers, who completed 246 real tasks in repositories they regularly contributed to, some with AI and some without. With AI they took 19% longer. Before starting they expected to be 24% faster, and afterwards they still believed they had gained 20%. The authors stress that the study doesn't show AI slows down most developers: it's a specific sample, on projects they knew very well. But the gap between what's perceived and what's measured isn't a detail, because it shows up even in very experienced people.

The second comes from Stack Overflow's 2025 annual survey. 84% of respondents use or plan to use AI tools for development, but only 33% trust their accuracy, against 46% who distrust it. The most cited frustration, for 66%, is solutions that are “almost right, but not quite”. And 45% say debugging AI-generated code is more time-consuming.

Together they describe the skill to look for. It isn't writing fast; the tool does that. It's spotting the “almost right” before it reaches production, and being able to measure honestly how much the tool helped. Neither shows up in an algorithm test, nor in a live coding session where AI is allowed and success means the exercise gets solved.

The question the test has to answer: does this person spot the defect in code that looks correct, can they say what needs building before handing it to an agent, and can they estimate how much it really helped them? If your test has no moment where the candidate has to say “this is wrong”, it isn't measuring what you need.

What each test measures today

By that standard, the usual tests look like this:

Test What it measures today What it's good for
Whiteboard algorithm, no tools Fundamentals and thinking out loud. Says nothing about how they'll work with agents A short complement, not the main test
Unsupervised take-home Whether they can use an agent, or nothing, if you don't then ask how and why they did it Only if reviewed with the candidate
Live coding with AI allowed and the usual brief The model's speed Not much: the brief no longer discriminates
Specifying an ambiguous requirement before implementing it Whether they can say what needs building, what's out of scope and how you'll know it's right Central in a team that works with specs
Reviewing an AI-generated pull request with seeded defects Judgement about someone else's code: the task they'll do every day Central
Debugging a defect in plausible code What takes the most time, according to developers themselves Very useful, and quick to prepare

onext's assessment. The tests aren't mutually exclusive: what changes is which one carries most weight in the decision

The two central rows have something in common: they reproduce the job. In a team that uses Spec-Driven Development, a developer's day is largely about stating precisely what needs building, letting the agent implement it and checking that what was implemented is what was asked for. If the interview doesn't include those three moments, it's assessing a different job.

A session in three parts

This is how we'd set up a technical interview of about ninety minutes for a development role in a team that works with agents. It isn't the only way to do it, but each part has a specific signal to watch for.

Specify (about 20 minutes)

You hand over a deliberately ambiguous requirement, similar to Canva's airport, and ask for a short specification before any code is written: what's in and what's out, what acceptance criteria it would have and what they'd ask the client. Here you look at what questions they ask and whether their criteria can be checked. Someone who writes “it must be fast” hasn't specified anything; someone who writes how many operations per second and under what conditions has. It's the same reasoning we develop in the specification is where you sign off.

Delegate (about 40 minutes)

The candidate implements a part with their own tool —the one they use day to day, not one the company imposes— and talks through what they're doing. You don't look at whether they finish. You look at how they split the work with the agent, what they accept without reading, when they stop it and what they do when the result doesn't match the specification they've just written.

Review and break (about 30 minutes)

They're given an AI-generated pull request for the same problem, with three seeded defects of the kind AI really makes: a test that describes the code instead of testing it, an edge case nobody has handled and a criterion from the specification that was silently skipped. You look at what they find, how they explain why it's a defect and what they'd ask for before approving it. It is, almost literally, the review that still expects an author they'll be doing every week.

And one last question, straight out of the METR study: “how much do you think AI saved you in this session?”. The point isn't to penalise a poor estimate, but to see whether the person noticed. Someone who answers “a lot on the implementation, but I lost time in the review because the test wasn't testing anything” is measuring their own work. Someone who answers “loads” without qualification probably isn't. It's the same problem we see in companies when Copilot's ROI is measured in the wrong place, only in a single person.

Fundamentals don't disappear: you ask about them at the end, without tools, in a short conversation about why they made each decision. This is where Anthropic is right. Without fundamentals nobody spots a subtle error in generated code, because they don't know it's an error.

What AI shouldn't be doing in your hiring process

There's an irony in all this. Many companies that ask candidates not to use AI use it themselves to screen CVs. And there the problem isn't consistency, it's legal.

The EU AI Act classifies as high-risk AI systems intended for the recruitment or selection of people, in particular to analyse and filter job applications and to evaluate candidates (Annex III, point 4). Following the amendment that entered into force on 27 July 2026, those obligations apply from 2 December 2027. The postponement isn't permission to wait: a screening tool bought today will still be in use then, and you'll need to be able to explain how it decides and who oversees what it rejects.

There's a less legal reason too. A process that screens with a model and assesses without one sends an odd message to someone you'll be asking to work with agents and review what they generate. The sensible approach is for decisions about people to be made by people, and to say so in writing. That's what we've done in our own job opening: no AI filters or scores the applications, and our privacy policy sets this out in the section on candidates (in Spanish).

The interview has to look like the job

All of the above comes down to an idea that isn't new: the interview has to look like the job. What's new is that the job has changed faster than interviews have. If your team specifies, delegates to an agent and verifies, that's what you need to see in the interview. If the test can be solved by pasting the brief, it tells you nothing about the candidate. And if at no point does anyone have to find a defect, it leaves out precisely the skill that's most needed.

A bad hire is expensive, and we've already worked out what it costs to leave a technical role unfilled for months. Hiring someone who writes fast with an agent but doesn't see their own errors is expensive too, it just takes longer to show.

By the way, we're hiring: we're looking for a Full Stack Developer to work with AI agents and Spec-Driven Development, the job opening is here (in Spanish).

Frequently asked questions

Should candidates be allowed to use AI in the technical interview?

It depends on what you want to measure, and you have to decide that first. If you want to see how someone reasons on their own, it makes sense to ask them not to use it, as Anthropic does in its assessments and interviews unless it says otherwise. If you want to see how they'll work in the job, it makes sense to require it, as Canva has done since June 2025. What doesn't work is allowing it without changing the test: then you're measuring the model's speed, not the person.

Are algorithm tests still useful?

For checking fundamentals, yes, and fundamentals still count: without them nobody spots a subtle error in generated code. But they no longer work as the main test, because they measure how someone works without the tools they'll use every day. A short conversation without tools about why they made each design decision gives you the same information in less time.

How do you assess whether a candidate can review AI-generated code?

By giving them an AI-generated pull request to review, seeded with the kind of defects AI really makes: a test that describes the code instead of testing it, an edge case nobody handles, a requirement from the specification that was silently skipped. You look at what they find, how they explain why it's a defect and what they'd do to stop it reaching production. It's the task they'll be doing every day.

What does the EU AI Act say about using AI to screen candidates?

That AI systems intended to analyse and filter job applications and to evaluate candidates are high-risk (Annex III, point 4). Following the amendment that entered into force on 27 July 2026, those obligations apply from 2 December 2027. That's no reason to wait: a screening tool bought today will still be running then, and you'll need to be able to explain how it decides.

Are developers who use AI faster?

Not always, and they can't measure it themselves. In METR's July 2025 study, sixteen experienced open-source developers took 19% longer on their tasks with AI, when they had expected to be 24% faster, and afterwards they still believed they had gained 20%. The authors warn that this can't be generalised to all developers, but the gap between perception and measurement is the useful lesson for an interview.

What profile should you look for in a team that works with specifications and agents?

Someone who can say what needs building and how you'll know it's right before writing any code, who hands the agent the parts they can verify and who stops to read what they accept. In the interview it shows in three things: the questions they ask about an ambiguous requirement, the defects they find in code that looks correct, and how well they estimate how much the tool actually helped them.

Sources

Jordi García
Written by
Jordi García
Tech Lead at onext

Jordi García is Tech Lead at onext. He works on bringing AI into governed production across development and product teams —with Spec-Driven Development, context engineering and human verification at every step— and authors onext's technical insights on the method, quality and cost of applied AI.

LinkedIn →

Specifying, delegating and verifying can be learned, in your team too

We build with your team the method for working with agents —specs, review and verification— on your own codebase. And once that method exists, you also know what to ask for in your next hire.

See how we work

Without stopping delivery.