Skip to main content
onext technology
Transformación 2 August 2026 - 9 min read

AI value: choosing and measuring your first use case

Most companies a year into AI cannot say whether their pilot worked. They have impressions. They do not have the number from before.

Jordi García
Tech Lead at onext
Operations manager comparing a printed process report against figures on her laptop, annotating by hand

For your board (60 seconds)

Without the before, the after means nothing. This layer does not measure the technical quality of the solution: it measures whether your organisation can tell a case that works from one that does not. It comes first in the sequence because it is the only layer that unlocks budget for the rest, and because a large share of the projects the market counts as failures did not fail: they simply could not be demonstrated.

What this layer actually measures

One thing only: whether your organisation can tell a case that works from one that does not.

That sounds trivial until you test it. Most companies a year into AI cannot say whether their pilot worked. They have impressions — "the team likes it", "we save time" — and no number from before. Without the before, the after means nothing.

What this layer does not measure: the technical quality of the solution. An excellent system on an irrelevant case scores zero here, and rightly so.

Why mid-sized companies get it wrong

One structural reason and one cultural.

The structural one: in a mid-sized company nobody has "measure processes" in their job description. The process you want to improve with AI — drafting proposals, triaging tickets, preparing reports — is not timed anywhere. Nobody knows how many hours a month it consumes because nobody ever needed to. And when the vendor asks for a baseline, the honest answer is "we do not know", so one gets invented.

The cultural one: measuring before starting feels like lost time. There is board pressure to show something, and two weeks of instrumentation read as two weeks of delay. They are the two weeks that decide whether the project survives month nine.

The market data shows the result. Gartner puts generative AI projects abandoned after proof of concept above 50%. RAND documented in 2024 that more than 80% fail to deliver the intended business value. And MIT's 2025 report — whose design deserves a careful look: 153 executives, 52 interviews, 300+ initiatives reviewed between January and June, with well-founded criticism of its sampling — puts pilots with no measurable P&L impact at 95%.

Our read from the field: a large share of those "failures" are not. They are projects that cannot be demonstrated. That is not the same thing, and the difference is decided in this layer.

Diagnosis: where you are

Stage 1 · Anecdote. There is enthusiasm and there are testimonials. Nobody has a prior number. If the sponsor changes role, the project dies with them. Observable signal: when you ask about the saving, you get shown a demo.

Stage 2 · Measured case. There is a case with a baseline, a declared target and a decision date. Everyone knows what would justify continuing and what would justify stopping. Observable signal: someone can tell you "today it is 340 hours a month" without going to check.

Stage 3 · Governed portfolio. Several cases, compared on the same yardstick, with budget pulled from the ones that do not perform. Observable signal: in the last twelve months an AI case was killed on purpose. If none ever has been, there is no portfolio: there is accumulation.

Most mid-sized companies that call us are at stage 1 and believe they are at stage 2.

The first 90 days

Weeks 1-2 · Pick the case. Three criteria, in this order: the process must be repetitive (if it happens four times a year there is no signal to measure), its output must be verifiable by someone (if nobody can say whether the result is good, there is no case), and its owner must want to do it (an imposed case sabotages itself).

One criterion deliberately absent from that list: that it be the most valuable. The first case is not the one that saves most, it is the one that teaches most.

Weeks 3-4 · Baseline. Measure the process as it stands today: monthly volume, time per unit, rework rate, cost of the people involved. It is dull spreadsheet work and it is what separates this layer from marketing.

Weeks 5-10 · Bounded pilot. Against the baseline, with the same measurement. No widening the scope midway: if the scope moves, the comparison is lost.

Weeks 11-12 · Declared decision. Continue, adjust or stop, with the criterion written down before seeing the result. Writing it afterwards means picking the criterion that confirms what you already wanted to do.

Deliverable for the quarter: one page with the before, the after and the decision. If it does not fit on one page, it is not clear.

What NOT to do yet

  • Do not build an AI dashboard. With one case there is nothing to chart. AI dashboards in companies with a single pilot are decoration.
  • Do not try to calculate the ROI of AI as a whole. It does not exist. The return on a specific case does.
  • Do not buy by volume. Licences for the whole workforce before you have a measured case is the fastest way to spend the next four quarters of budget.
  • Do not promise the board a percentage. Promise a decision date. It is the only thing you can deliver with certainty.

What skipping it costs

It costs the whole project, slowly. A pilot without a baseline cannot be defended when the budget review arrives, and it cannot be cleanly cancelled either: it sits in an expensive limbo, consuming licences and attention, until somebody gets tired.

And it costs something worse: the second chance. An organisation that has lived through an AI project nobody could demonstrate takes a long time to authorise the next one. The real cost is not the failed pilot; it is the year and a half of paralysis behind it.

This is the first layer of the series on the seven layers of AI maturity. Next comes data: what material the case you just picked actually needs.

Frequently asked questions

How do you measure the ROI of an AI project?

Against a baseline taken before you start: volume, time per unit, rework and the cost of the people involved. Without that prior measurement there is no calculable ROI, only estimates. Instrumentation takes about two weeks and is what lets you defend or cancel the project on evidence.

What should a company's first AI use case be?

One that is repetitive, has a verifiable output, and has an owner who wants to do it. Not the most valuable one: the one that teaches most. The first case buys you knowledge about your own organisation, and that is what makes the second one land.

How long before AI shows a return?

A bounded case gives signal in a quarter: six weeks of piloting against a baseline plus the decision. Consolidated financial return depends on the case and the volume, and any figure quoted before someone has seen your data is a sector benchmark, not a forecast.

Why are so many AI pilots abandoned after proof of concept?

Gartner puts that figure above 50%. In our experience a significant share do not fail technically: they fail at demonstration. Without a baseline the pilot cannot prove it worked, and what cannot be proven does not survive a budget review.

See how we work
Jordi García
Written by
Jordi García
Tech Lead at onext

Jordi García is Tech Lead at onext. He works on bringing AI into governed production across development and product teams —with Spec-Driven Development, context engineering and human verification at every step— and authors onext's technical insights on the method, quality and cost of applied AI.

LinkedIn →