Skip to main content
onext technology
AI 4 October 2026 - 9 min read

What a software development tender should ask for so that AI reaches the code

Writing "with AI" into the subject matter of the contract does not change how software gets built. What changes it is what the tender asks to be verified: checkable specifications, traceability, human verification and delivery metrics. And Spanish public procurement law allows asking for it, subject to legal validation.

Jordi García
Tech Lead at onext
A bound printed document with a pen on top, next to a laptop with blurred code on the screen, on a wooden work desk at dusk

A test before you publish the tender: imagine the winning bidder writes in its offer "we will use AI assistants throughout development". Which clause in your tender obliges it to prove it? And which one lets you check, in month three, that AI has reached the code and did not stay in the offer?

If the answer is "none", nobody is acting in bad faith. Some development and maintenance tenders were written to buy a team's time, and three words have now been added to them. The tender still rewards what it always rewarded, and the supplier does what the tender measures.

The thesis of this piece is that a tender that mentions AI but scores price and hours buys hours, with or without AI. For AI to reach the code, the tender has to ask for four things that can be checked: verifiable specifications, traceability, human verification and delivery metrics. We write for those who draft or review tenders and for those who answer them: they are the same questions from either side of the table. The bidder's side, the offer, we cover in how to write the technical proposal of a public-sector tender; this one is the buyer's.

What we saw reading 21 tenders

On 3 October 2026, onext reviewed 21 public tenders that its radar had flagged for their link to AI or to software, and logged each one in its internal register. It is our own count, not a representative sample or a market study; that is why we give only aggregate figures and no contracting authority by name.

  • In 5 of the 21, AI appears in the subject matter of the contract, as a product, a licence or training. In the other 16 it does not: they are maintenance and enhancement of systems, data services or supplies.
  • In 6 files the register records the weight of price in the scoring: 50%, 50%, 51%, 60%, 80% and 95%. None below 50%. The other 15 do not record it, so we do not know what it weighed there.
  • In one of the files the register notes that spec-driven development (SDD: the specification is the source of truth and the code is derived from it) carried no points.

Those figures do not prove that the administration buys badly. They say something more modest and more useful: in what we reviewed, the software tenders we reviewed mostly describe a team, profiles or hours, and a price, and AI, when it appears, appears as a label. That is understandable: asking properly for a way of working that has changed in two years is not trivial, and in what we have seen there are few examples to copy.

Why "we will use AI" is not evidence

In July 2025, METR published a randomised trial with 16 experienced developers and 246 real tasks in open-source repositories they had contributed to for years (on average, over 22,000 stars and a million lines of code). With AI tools —mostly Cursor Pro with Claude 3.5 and 3.7 Sonnet, the frontier models at the time— they took 19% longer to close the tasks. Before starting they expected to be 24% faster; afterwards they still believed they had been 20% faster.

The result matters for the last figure, not the first. If the people doing the work cannot tell, task in hand, whether they are faster or slower, a sentence in an offer proves nothing, and neither does a satisfaction survey. A measurement is needed.

What this study does not say: the authors themselves make clear it does not show that AI fails to speed up most developers, or that it does not help in other contexts, or that it cannot be used better. It was one specific setting (experts in repositories they knew very well) with early-2025 tools. The reading for a tender is not "AI does not work" but "do not trust perception; ask for the measurement".

Public procurement law already allows asking for it

It is sometimes assumed that asking for a way of working in a tender is risky. Spain's Public Sector Contracts Act (Law 9/2017) leaves more room than is usually used:

  • Article 126.2: technical specifications "may refer to the specific process or method of production or provision" of the services, as long as they are linked to the subject matter of the contract and proportionate to its value and objectives.
  • Article 145.2: qualitative criteria may include quality, "including technical value", and the organisation, qualification and experience of the staff; they must be accompanied by a cost-related criterion.
  • Article 145.5: criteria must be linked to the subject matter, formulated objectively and "accompanied by specifications that make it possible to effectively verify the information provided by bidders".
  • Articles 145.3.g and 145.4: in intellectual services price cannot be the only factor, and quality criteria must add up to at least 51%. Whether software development falls into that category is for the contracting authority to decide: the law does not say so expressly.

This is not legal advice, and each wording has to be validated by the authority's legal service. But the underlying idea is sound: asking how work is done is not unusual, it is something the law contemplates as long as it is linked to the subject matter of the contract and can be checked. Article 145.5 is also what we use as a design criterion: if a requirement cannot be checked, it does not score.

The four asks

There are four because each closes a gap through which AI stays in the title. All of them ask for a trail that exists if the work is done a certain way and does not exist if it is not.

Ask What the winning bidder delivers How the buyer checks it
1. Verifiable specification For each increment, a specification with executable acceptance criteria, versioned in a repository the buyer can access, before any code is written Validates it and can run the criteria on its own
2. Traceability Each change linked to the requirement that originates it and to the verification that covers it; a record of whether AI assisted, with tool and model, as audit data Sampling: pick five changes at random and reconstruct the path requirement, change and test
3. Human verification Each change reviewed by a person other than the one who submitted it, with name and role; tests that are not generated from the code itself Reviews the review records; at acceptance, the author explains the change without the assistant
4. Delivery metrics A baseline at the start and monthly measurement: time from a change being accepted to being in production, failed changes, rework and defects found after delivery The data comes from the repository and the deployment system, not from a hand-written report

onext's own elaboration

1. Verifiable specification: the contract is signed before the code

When an agent writes the code, the place where what gets built is decided becomes the specification. It is the underlying idea of the specification as the place where you sign off and of spec-driven development. For a buyer the translation is direct: the supplier delivers, before programming each increment, a document with acceptance criteria that can be executed, and the buyer validates it. It is the opposite of the pool of hours: the buyer knows what it has agreed to build.

2. Traceability: reconstruct where each change comes from

Asking for each change to link to its requirement and to the test that covers it is not bureaucracy: it is what makes auditing possible. The record of whether AI assisted, and with which tool and model, goes in as audit data, not as a score: it helps to understand a failure, not to punish anyone. With the sampling of five changes at random, the buyer checks with a sample what no report can guarantee.

3. Human verification: a person signs off, and it is not the one who submitted the change

AI speeds up writing code; reviewing it is still the job of a person with judgement. The tender can ask that each change be reviewed by someone other than the author, that tests not come from the very code they verify (tests generated from the code describe what it does, not what it should do) and that, at acceptance, whoever submitted the change can explain it without the assistant in front of them. It is the contractual version of a review that still expects an author. And in public administration it links to something already demanded of AI that serves citizens: human oversight that is designed, not assumed.

4. Delivery metrics: measure before, measure after

The lesson of METR is a lesson about measurement. A baseline at the start and a monthly reading of four or five numbers (time to production, failed changes, rework, defects after delivery) show whether AI has reached the code. The data comes from the repository and the deployment system, not from a hand-written report. What to measure, and what to stop measuring because it lies, we develop in KPIs for AI development teams.

How to score it without becoming subjective

The risk of asking for a method is ending up with criteria that depend on who reads the offer. Article 145.5 gives the way out: objective criteria with specifications that make it possible to check what the bidder says. A ladder with checkable steps works better than "up to 10 points for the quality of the approach". An example wording, as our own proposal that the legal service will have to validate:

  • 0 points: no evidence of how the bidder will work, or it merely states that it will use AI.
  • Intermediate score: describes the procedure (how specifications are written, how changes are recorded, who reviews and what is measured).
  • Top score: besides describing it, provides a real, anonymised example of a change with its specification, its trail to the requirement and its review record.

The third step is the one that separates those who already work this way from those who promise to: the first leaves examples, the second leaves sentences.

Price: if you pay by the hour, you cannot ask for savings

There is an incentive here worth keeping in mind, and it is our own reasoning, not a data point. If the contract pays hours and AI reduces the hours needed, the supplier earns less for working better. It is not an accusation: nobody voluntarily shares a saving that lowers their invoice. There are two usual ways out: tie the price to accepted increments rather than hours, or include a clause that shares the saving measured against the baseline. Both depend on the fourth ask (without metrics there is no saving to share) and both require a design that the authority's legal service validates. For whoever answers the tender, the flip side is the same one we saw in AI for public tenders: buy or build: the advantage lies in having material and method already in order before the tender arrives.

What a tender should not ask for

  • A specific tool or model. It clashes with equal treatment of bidders and ages in months. Ask for the result and the trail, not the brand.
  • A percentage of AI-generated code. It measures usage, not outcome, and is easy to game. When a figure becomes a target, it stops measuring what it used to measure.
  • AI without a data condition. If execution involves handing public-sector data to the contractor, article 202.1 already requires a special data-protection condition. With AI it is worth specifying which third-party services the buyer's code and specifications may reach.

One final distinction, so as not to mix tools. When what is being bought is an AI system, and not a development done with AI, there are the European model clauses for AI public procurement (MCC-AI), from the Commission's community of practice, available since June 2025 in the 24 official languages. This piece deals with the opposite case: AI as the supplier's way of working, not as a product that is delivered.

A ten-minute test before you publish

Take the tender and answer four questions, without looking at anyone's offer:

  1. Which document proves that what would be built was agreed before it was built?
  2. If I pick a change at random, can I reconstruct where it comes from and who reviewed it?
  3. Who signs off that the code is correct, and is it a different person from whoever submitted it?
  4. Which number will show, in month three, that delivery is improving, and where does it come from?

If two answers are "none" or "I don't know", the tender buys hours even if it says AI on the cover. If all four have an answer, AI has somewhere to land. And it is the same test a serious bidder would run on its own team: whoever cannot answer them cannot promise what they promise.

Frequently asked questions

Can a public tender require AI to be used in development?

Spain's Public Sector Contracts Act allows technical specifications to refer to the "specific process or method of production or provision" of the service (art. 126.2), as long as they are linked to the subject matter of the contract and proportionate to its value and objectives. That is why it is better to ask for results and evidence than for a particular tool. Each wording must be validated by the contracting authority's legal service.

Why is scoring the team's AI experience not enough?

Because declared experience does not show how work will be done on this contract. The law allows scoring the qualification and experience of staff (art. 145.2), but the METR trial shows that self-perception fails: experienced developers believed they were 20% faster with AI and had in fact taken 19% longer. Useful evidence is a verifiable trail, not an impression.

How much should price weigh in an AI development tender?

There is no single figure. The law requires that, in intellectual services, price cannot be the only factor (art. 145.3.g) and that quality criteria add up to at least 51% (art. 145.4); whether software development counts as "intellectual in nature" is up to each contracting authority, because the law does not say so expressly. The more price weighs, the more hours are being bought. In the six files of our count that record the weight, it ran from 50% to 95%.

How do you ask for AI traceability without over-monitoring the team?

As audit data, not as a score or a sanction: record whether AI assisted a change, with which tool and which model. It is used to reconstruct the path requirement, change and test when something fails. If the contract involves handing public-sector data to the contractor, article 202.1 already requires a special data-protection condition: with AI it is worth specifying which third-party services the data may reach.

What if the winning bidder does not use AI and meets the metrics?

Then the tender is met. That is the reason for asking for measurable results and not usage: if the team achieves the same lead times, quality and traceability without AI, the buyer gets what it wanted. If it uses AI and does not achieve them, that shows too. The metric protects both sides.

Does the same apply to a private client?

Yes, and it is easier, because it does not have to fit the LCSP. The four asks (verifiable specification, traceability, human verification and delivery metrics) fit any development contract or RFP. The discipline is the same: ask for evidence of how work is done, not for the promise of a tool.

Sources

Jordi García
Written by
Jordi García
Tech Lead at onext

Jordi García is Tech Lead at onext. He works on bringing AI into governed production across development and product teams —with Spec-Driven Development, context engineering and human verification at every step— and authors onext's technical insights on the method, quality and cost of applied AI.

LinkedIn →

Let AI reach the code, in your contract too

We work with development teams on a method —specification, traceability, human verification and metrics— that can be asked for in a tender and checked with a sample. The method stays with your team.

See how we work

No tool promises. Evidence.