In public administration, AI makes sense as a tool that supports public servants, not as a replacement for their judgement. The AI searches and drafts; the person checks, corrects and decides. It's the reasonable position, and it's the one almost every public guideline takes — starting with the criteria for using generative AI that the Government of Catalonia approved in April 2025: AI may not take or replace the decision, only support it.
The trouble with reasonable positions is that they don't tell you where to start or how to know whether it's working. For that there's data. Since late 2023, the French, British and Australian administrations have published results from their generative AI pilots, and read together they tell a more interesting story than “support yes, replacement no”. They show where time is really saved, and why human oversight isn't something you declare: it's something you design.
Two kinds of evidence that don't say the same thing
Let's start with the numbers, including a column that's almost never shown: how each one was measured.
| Initiative | Result | How it was measured |
|---|---|---|
| France · replies to citizen reviews on Services Publics+ (since late 2023) | Average response time from 19 days to 3. 70% of the AI's proposals are reused. 11% higher satisfaction with the reply | Platform data and user ratings. No published methodology or control group |
| UK · Copilot across 12 organisations, 20,000 staff (late 2024) | An average of 26 minutes saved a day | Survey with 7,115 responses. The report itself warns the savings are self-reported |
| UK · Copilot at the Department for Business and Trade, 1,000 licences | No evidence that time savings led to improved productivity. Summarising reports much faster and better; spreadsheet analysis slower and less accurate | Usage diary, interviews, a control group and observed tasks scored blind (small samples) |
| Australia · Copilot at the Treasury, 218 staff (2024) | Lower use than expected, unsuitable for complex tasks, no clear evidence of better work outcomes | Evaluation by the Australian Government's evaluation centre |
| France · “L'Assistant”, an assistant for public servants (2025-2026) | 12% less time on drafting and up to 16% on summarising documents | Experimental focus groups with blind-assessed quality, plus surveys |
Results published by the administrations themselves. The more demanding the method, the more modest the number
The caption under the table is the first lesson. The more demanding the method, the more modest the number. The 26 minutes a day come from asking people; when the same kind of tool is evaluated with a control group and observed tasks, the effect on productivity doesn't clearly show up. It's the same pattern we see in companies, and the one we described in why Copilot's ROI gets measured in the wrong place: the perceived saving is real, but it isn't a measurement.
And one detail from the UK evaluation deserves more attention than the headline: 22% of diary respondents had spotted hallucinations, 11% weren't sure whether they had, and another 22% chose not to answer. When a significant share of staff can't say whether the tool got something wrong, the time saving they report includes work someone will have to redo later.
Why they don't contradict each other: a queue isn't a licence
The French citizen-review case and the Copilot evaluations aren't measuring the same thing, and that difference is the most useful lesson here.
Services Publics+ applied AI to a queue: thousands of messages from citizens that needed an answer, with a response time that could be measured before and after, a repeatable type of text and known sources — the official administrative information sheets. The system proposes a draft; the agent uses it or not, edits it, rewrites it or throws it away, and approves the final reply. When the bottleneck is drafting, and the AI drafts, the response time drops. And because the response time was measured beforehand, you can say by how much.
The Copilot trials applied AI to a licence: a general-purpose tool for each person to use however they like. The saving exists, but it's scattered across loose minutes, on tasks that weren't the bottleneck for anything, and partly eaten up by checking. No response time drops, because no response time was chosen.
Every administration has obvious candidates: replies to enquiries and complaints, requests to correct applications, procedural reports, case summaries for whoever has to decide, answers to repetitive objections. They all share three traits: high volume, a text that looks similar from one case to the next, and sources that exist and can be closed. It's the same criterion we recommend in companies for choosing and measuring the first use case.
Albert: what happens when you start with the hardest case
France also offers the counter-example, and it's worth telling because it's often cited as proof of the opposite. Albert, the generative AI developed in-house by the French administration, was trialled in 48 citizen service offices. In January 2026, according to AFP, the state's digital directorate decided not to roll it out “in its current form”, after staff reported technical failures and wrong answers.
It wasn't abandoned: the infrastructure carried on, and the assistant for public servants was rebuilt on a different technical base. In June 2026 it was extended to all state staff after an evaluation that, this time, measured time and blind-assessed quality. But the sequence leaves a clear lesson. A general-purpose assistant answering open questions at the counter is the hardest case, not the first one: there's no bounded queue, the questions don't resemble each other, the sources aren't closed, and errors go straight to the citizen.
There's also a technical reason. A system that searches official documents before drafting reduces fabrications, but doesn't eliminate them: the retriever can miss a relevant exception or pull in a repealed rule, and the model can combine two correct fragments badly. In public administration that risk has a name: currency. A corpus of regulations that doesn't know which version is in force is worse than no corpus, because it produces well-argued answers about rules that no longer exist. That's why the value is in the documents, not the model, and the first investment is having them organised, versioned and dated.
The 70%: the number to watch
Of all the French data, the most interesting figure isn't the response time. It's that 70% of the AI's proposals are reused. That's good news — the tool is right often enough to be useful — and at the same time it's the number to watch, because if it drifts up towards 98% over time, it can mean two very different things: that the tool has improved, or that staff have stopped looking.
Research adds a nuance worth knowing here. In three experiments in the Netherlands — two with citizens and one with 1,345 civil servants — Saar Alon-Barkat and Madalina Busuioc did not find the blind trust in the machine you might expect. They found something more uncomfortable: in the first two, people followed the recommendation mostly when it confirmed a stereotype. In the third, with civil servants and after the childcare benefits scandal, the pattern disappeared: the scandal had made them more cautious. Oversight doesn't fail only because people trust AI too much; it also fails because they trust it selectively, and the organisation's context matters.
And process design pushes one way or the other. According to documents obtained by Privacy International, in a recommendation tool used by the UK Home Office, caseworkers had to justify in writing when they rejected a recommendation, but not when they accepted it. Nobody has to decide to approve without looking: it's enough that looking costs more than approving.
That gives three design rules that don't depend on the model:
- A reason is always written down, on accepting and on rejecting. Even a single line. If accepting costs nothing and rejecting costs a paragraph, the system is designed to approve.
- The amendment rate is measured: what share of proposals gets changed, by whom and in which kind of case. A rate that falls without any change to the tool is an alarm, not a success.
- The reviewer has authority to change the decision and all the facts of the case in front of them, not just the draft. It's what the EU AI Act requires for high-risk cases, and what we covered in the exception path: a review that approves a hundred per cent of cases in twenty seconds isn't oversight.
The line between drafting and deciding, in law
All of this also has a legal framework, and it's worth being clear about it because it changed in July. The line that matters is the familiar one: AI drafting something a person then reviews isn't the same as AI taking part in a decision about someone's rights.
| Use | What it requires | Since when |
|---|---|---|
| Drafting internal documents, summarising case files | Not high-risk in principle. Staff AI-literacy measures and data protection do apply | Already |
| Publishing generated text to inform the public on matters of public interest | Disclose that it was AI-generated, unless it has gone through human review and someone holds editorial responsibility | Since August 2026 |
| Assessing whether a person is eligible for an essential public benefit or service | High-risk: human oversight by people with competence, training and authority; awareness of automation bias; a fundamental rights impact assessment before deployment; registration. A purely preparatory task may fall outside, unless it involves profiling | 2 December 2027 (postponed in July 2026) |
| Automated administrative action, with no direct involvement of a public employee (Spain) | Law 40/2015, art. 41: designate in advance which body defines, supervises and quality-controls the system, and which is responsible for the purposes of appeals | In force since 2016 |
An indicative summary of the EU AI Act as amended in July 2026 and of Spain's Law 40/2015. Not a substitute for legal analysis of each case
Three observations on the table. First: the postponement to December 2027 isn't permission to wait. The impact assessment has to be done before deployment, and a system that goes live in 2027 without having planned for it will be late. Second: the “preparatory task” exception is narrow and will be argued over a great deal. A case summary that systematically leaves out a certain kind of circumstance is shaping the decision, whatever you call it. Third: transparency already has case law in Spain. In September 2025, the Supreme Court recognised the Civio foundation's right to access the source code of BOSCO, the application that decides who can receive Spain's subsidised electricity rate, with reasoning that applies to any system involved in granting social rights: the more it decides, the more it has to be explainable.
There's also a less legal reason to take the line seriously. The two great failures of automated decision-making in Europe and Australia — childcare benefits in the Netherlands, which ended with the government resigning in January 2021, and Robodebt in Australia — weren't generative AI cases. They were rule-based systems and data matching that flagged families as suspected fraudsters or raised debts that didn't exist. But they left the same lesson: when the system proposes and nobody with authority and time reviews, errors are multiplied by volume. Generative AI doesn't change that lesson; it only makes volume easier to produce.
Where we'd start
With all of the above, this is how we set up a first project with a public body, whether a city council, a consortium or a department:
First, a queue, not a licence. A specific process with volume, a response time measured today and a text a person writes. If there's no measured response time, the first step is to measure it: without it there's no way to say whether the project worked.
Second, closed, dated sources. Which regulations, which information sheets, which reply templates, in which version and who maintains them. The system cites the source of every statement and the reviewer checks it. If the corpus isn't up to date, that's where the project starts.
Third, design the review step before choosing the model. A reason on accepting and on rejecting, the amendment rate on a dashboard, and a person with real authority over the outcome. It's the part that almost never gets budgeted, and the one that decides whether oversight is effective or a mere formality.
Fourth, measure like the Australian Treasury, not like a survey. One group working with the tool and one without, the same kind of cases, and quality reviewed by someone who doesn't know which is which. It takes a few more weeks and it's the only thing you can defend before an oversight body.
Fifth, decide where the system lives. France ended up building its assistant on its own certified infrastructure, and the Spanish government has announced a sovereign AI platform in its cloud. You don't need to wait for either to start, but you do need to know from the outset which data leaves, where it goes, and what happens if you have to switch provider tomorrow. We cover it in the split of control across eight axes.
Human judgement isn't something you declare
The starting position is right: operational efficiency and human judgement aren't opposed, and an administration that combines them well will serve people better. What the data adds is that human judgement doesn't appear because a guideline says so. It appears when someone chooses a specific queue, closes the sources, designs a review step that costs the same in both directions and measures how often the civil servant says no. If that number exists and someone looks at it, there's oversight. If it doesn't, there's a formality with an approve button.
The good news is that all of this fits into a first project of a few weeks, on a single queue and without changing platform. The bad news is that there's no shortcut: the saving you can defend is the one you've measured.
Frequently asked questions
Does generative AI really save time in public administration?
It depends on where it's applied, and the published evidence shows this quite clearly. Where it targets a specific queue with a measured delay — like the replies to citizen reviews in France, where the average response time went from 19 days to 3 — the effect is large and visible. Where it's handed out as licences for all staff, the most rigorous evaluations find no clear improvement in productivity: the self-reported saving is real, but it gets diluted or eaten up by checking what the tool produces.
What's the difference between a self-reported trial and a rigorous evaluation?
The first asks people how much time they think they save; the second measures tasks, compares against a group that doesn't use the tool and assesses quality without knowing who did what. The UK has both: the twelve-organisation trial reports 26 minutes saved a day, based on a survey, while the Department for Business and Trade evaluation, with observed tasks and a control group, found no evidence that those savings turned into higher productivity. They don't contradict each other: they measure different things.
What does the EU AI Act say about these uses in public administration?
It distinguishes by use. Assessing whether a person is eligible for an essential public benefit or service is high-risk, with obligations for human oversight, staff with competence and authority, and a fundamental rights impact assessment for public bodies. Those obligations were postponed: for these systems they apply from 2 December 2027. Drafting replies or summarising documents is not, in principle, high-risk. And the transparency obligations for generated text have applied since August 2026.
Is it enough for a civil servant to approve what the AI proposes?
Not if the approval step isn't designed. A process that requires a justification to reject but not to accept pushes people towards approving. Experiments run in the Netherlands found something subtler than blind trust: people followed the recommendation mostly when it confirmed a stereotype. Oversight works when the reviewer has the authority to change the decision, all the facts in front of them, a reason to write down whether they accept or reject, and someone measuring how often they amend.
What happened to Albert, the French state's AI?
It was trialled in 48 citizen service offices, and in January 2026 the French administration decided not to roll it out in its then-current form, after staff reported technical failures and wrong answers. The project didn't disappear: its infrastructure lives on, and the assistant for public servants, rebuilt on a different technical base, was extended to all state staff in June 2026 after an evaluation that blind-assessed quality. The lesson is that a general-purpose assistant facing citizens is the hardest case, not the first one.
Where should a public administration start?
With a queue, not a licence: a process with volume, a measured delay and a text a person writes today, such as replies to enquiries, requests to correct applications or procedural reports. With closed, current sources, a reviewer who gives a reason both when accepting and when rejecting, and a measurement with a comparison group and blind-reviewed quality. If that works on one queue, you know what it's worth and what it costs to supervise; then you choose the next one.

Jordi García is Tech Lead at onext. He works on bringing AI into governed production across development and product teams —with Spec-Driven Development, context engineering and human verification at every step— and authors onext's technical insights on the method, quality and cost of applied AI.
LinkedIn →