Insights
Analysis, learnings and perspectives on IT transformation, artificial intelligence and modern architectures. Written by consultants who implement these technologies every day.
Latest articles
Reviewing AI-generated code: your review still expects an author
Merges with no review at all, human or agentic, are up 31.3%, and bugs per developer up 54%. This is not carelessness: the review is asking for something that can no longer be given. This piece separates the real bottleneck — which is almost never review, as Steve Fenton's objection shows — from where the risk concentrates, and dismantles the four assumptions that held code review together: that an author exists to ask, that writing costs more than reading, that diff size reflects the decision, and that new code resembles what is already there. With Faros AI, LinearB, GitClear, Veracode, DORA 2025, Stack Overflow and METR's trial on the 39-point error in self-assessment.
Prompt injection: your agent's security lives in the permissions, not in the prompt
Inside a language model there is no separate channel for instructions and another for data, so prompt injection cannot be patched — it has to be bounded. This piece moves the defence from where it does not work — the system prompt and filters — to where it does: the tool permission table. Willison's lethal trifecta and Meta's rule of two as design criteria; the nine fields of a tool record; the six patterns with provable resistance and what they cost in utility; and the anatomy of two real incidents, EchoLeak in Microsoft 365 Copilot and the toxic agent flow through GitHub's MCP server. With OWASP LLM01:2025, the UK NCSC warning, "The Attacker Moves Second", CaMeL, AgentDojo, InjecAgent and spotlighting.
The golden set: how to build the cases that decide whether your AI works
"Show me your golden set" organises a conversation about AI in production faster than any other question. This piece builds the object the rubric and the baseline take for granted: where the cases come from (traffic, incidents and risk hypotheses), why you stratify by capability rather than volume, how many cases you actually need, why your labelling error rate is the ceiling for everything else, and why a held-out set is a consumable resource with a looking budget. With CheckList, Northcutt's work on label errors, Aroyo and Welty's crowd truth, the reusable holdout from Dwork and colleagues in Science, GSM1k and the re-annotation of SWE-bench Verified.
Evaluating AI agents: what "works better" actually means
"It works much better" carries no information: missing are what against, measured how, and with how much confidence. This piece builds the two missing objects. The rubric: how binary, verifiable criteria are derived from the golden set, with blocking criteria, trajectory evaluation and the null agent test. The baseline: the five mandatory reference points, including the ablation almost nobody runs. And the minimum statistics needed to know whether a three-point difference is real. With τ-bench and pass^k, "AI Agents That Matter", HealthBench, Agent-as-a-Judge, Anthropic's work on error bars and METR's trials.
LLMOps: the real path to production for an LLM system, and how to keep it alive
LLMOps is not MLOps with a different model. In MLOps you version weights; in LLMOps you version a context bundle — prompts, tools, retrieval, guardrails, thresholds — on top of a third-party model that changes without you touching anything. This is the operational version of the cycle: the nine stages, the two quality gates (offline evaluation in CI and online evaluation over a canary), the metrics that decide (cost per useful task, groundedness, containment) and the four failures that only appear with real traffic. Includes a four-level maturity model and the first two weeks' plan.
How to start with AI: the seven layers of maturity
The maturity model Gartner published in 2025 organises AI into seven dimensions — strategy, value, organisation, people, governance, engineering and data — but it describes what arriving looks like, not what to do on Monday morning with 300 people and a small team. This piece translates that map into a route for the mid-sized company, and argues a thesis that runs against the model's own order: AI strategy gets written last, after the first measured case. It covers the seven layers in business language, why most companies enter through engineering and die two layers down, the sequence we propose, and three questions to tell which layer you are stuck in.
AI value: choosing and measuring your first use case
The first layer of the maturity series. Without a baseline, the after means nothing: this layer does not measure the technical quality of the solution but whether your organisation can tell a case that works from one that does not. It covers the three criteria for picking the first case (repetitive, verifiable, and with an owner who wants it), why the most valuable case is not the first one, the three diagnostic stages, the 90-day plan with the decision written before the result is seen, and what not to do yet. With the thesis behind the piece: many of the pilots the market counts as failures did not fail, they simply could not be demonstrated.
Data for AI: the minimum viable preparation
The second layer of the maturity series. This layer does not ask whether you have a data lake: it asks whether the judgement your company applies is written down anywhere or lives only in people's heads. Gartner attributes 60% of abandonments through 2026 to inadequate data foundations, and puts at 63% the organisations without — or unsure of — fit practices. What follows is not fixing the company's data but the data one concrete case needs. Includes the question that orders the layer, the three stages, the ten-case sufficiency test, and which data project not to run yet.
AI governance: the minimum required in 2026
The third layer of the maturity series. The test here is blunt: when an AI output reaches a client, who answers for it? The required minimum fits on two pages and a spreadsheet: a short forbidden list, a name per decision, and a register. Includes the real EU AI Act timeline after the Digital Omnibus — transparency obligations apply from 2 August 2026 and were not postponed; Annex III high-risk moved to December 2027 — the three diagnostic stages, and why not to set up a committee for one pilot.
AI engineering: from pilot to production
The fourth layer of the maturity series. A demo proves something can happen; production requires it to happen every time, on unforeseen inputs, at predictable cost. This layer resolves into four decisions — build or buy judged on exit cost, an evaluation set, cost per task, and observability — and none of them is model choice, which consumes most committee time. Includes why it comes fourth rather than first, the three stages, the 90-day plan, and why not to start with autonomous agents.
People and culture: the AI nobody uses
The fifth layer of the maturity series. This layer measures neither training delivered nor licences deployed: it measures sustained use at eight weeks without reminders. The definitive test is simpler: if the tool were withdrawn tomorrow, would that team have an operational problem? Includes why training everyone equally produces no use cases, the correct unit of work (the task, not the job, following Susskind), the three stages, the 90-day plan for a group of five to fifteen people, and why not to put AI use into performance reviews.
AI organisation: who decides and who operates
The sixth layer of the maturity series. This layer measures the ability to choose between cases, not how many people you have dedicated to AI. The practical threshold for structure is low and concrete: three live cases competing for the same budget or the same people. Includes why importing the enterprise centre-of-excellence model fails in mid-market, the role that does work — a named function with allocated time and authority to stop — when outsourcing it makes sense, and its limit: operations can be outsourced, prioritisation cannot.
AI strategy: why you write it last
The seventh and final layer of the maturity series, and the one that carries its thesis. In Gartner's model strategy heads the seven dimensions; our experience in mid-sized companies places it at the end of the starting sequence, with a stated exception: in banking, life and health insurance or healthcare it moves to the front. Includes why it gets written so early and what to take to the board instead, the five elements of a useful strategy, the three stages, and why not to commission it from the vendor who will execute it — even if it is free, especially if it is free.
AI does not replace professions: it redistributes tasks
Drawn from Daniel Susskind's talk at Exponential Day in July 2026. His thesis: current automation reaches non-routine tasks in skilled professions, and does so without reasoning the way humans reason — Deep Blue and Watson won by routes that were not ours. The practical consequence is that the unit of analysis stops being the profession and becomes the task. We add a field caveat almost no summary carries: the thesis is right in aggregate and harsh at the individual level, and using it as a sedative is expensive. With three actions for Monday.
AI for public tenders: buy the tool or build the capability?
Spain's public sector tendered 204 contracts with AI at their core in the last full year on record, and more than half of all tenders score quality above price: the quality of your proposals is a revenue lever, not paperwork. Almost every offering on the market solves it the same way — Gober, LICAI, TendersTool, AIverso and the rest sell access to a capability that lives outside your company. This piece is not a tool comparison. It separates the three layers of the real work (finding, retrieving and adapting your own material, deciding what to commit to), gives the three questions that settle the decision, a table to take to your committee, the exit cost that never appears in the budget, the cases where buying is the right answer, and the four assets that stay in-house when you choose to build.