What this layer actually measures
It measures context availability, not data volume.
That distinction is hard to accept because the industry has spent fifteen years measuring the opposite. A mid-sized company can hold twenty years of ERP history and have nothing usable for the case it wants to solve, because what that case needs is not transactions: it is the reasons behind them. The "why" is rarely in the database.
What it does not measure: the sophistication of your architecture. We have seen cases work on a well-organised folder of two hundred documents, and cases die on seven-figure data platforms.
Why mid-sized companies get it wrong
Overcorrection. The company hears "AI needs data", concludes it needs a data project, and launches a cross-company initiative that eats the budget and the patience before any use case reaches production.
It is the mirror image of the engineering mistake. There, a tool gets bought with no problem defined; here, ground gets prepared with no destination.
The market numbers point at this layer more insistently than at any other. Gartner estimates 60% of AI projects will be abandoned through 2026 due to inadequate data foundations, and that 63% of organisations do not have — or do not know whether they have — the data management practices AI requires. The second figure is the more revealing one: do not know. The problem is not only the gap; it is that the gap is undiagnosed.
But what follows is not that you must fix the company's data. It is that you must fix the data one concrete case needs.
The question that orders this layer
If tomorrow you wanted a system to answer with your company's judgement — not the generic judgement of the internet — where would that judgement come from?
Answers we accept:
- From a history of documents where that judgement was applied and written down.
- From a rule set someone maintains that reflects how decisions are actually made.
- From past examples verified by someone who knows whether they were right.
Answers that are a warning:
- "From our people." Then the judgement is not data: it is tacit knowledge, and it has to be externalised before automating anything.
- "From the ERP." The ERP holds the what, almost never the why.
- "The vendor knows." Then the output will carry the sector's judgement, not yours, and there will be no differentiation at all.
Diagnosis: where you are
Stage 1 · Tacit knowledge. What makes your company good sits in people. It works while those people are there. Observable signal: when someone leaves, a way of working is lost, not just capacity.
Stage 2 · Bounded corpus. A body of material exists — proposals, reports, case files, resolved tickets — organised and sufficient for one concrete case. Not everything; the material for that case. Observable signal: someone can point at a folder and say "this is how we do it, and it is current".
Stage 3 · Maintained context. The corpus has an owner, is updated as part of normal work, and it is known what has gone stale. Observable signal: there is a last-reviewed date and someone answers for it.
Almost no mid-sized company is at stage 3, and that is fine. The jump that changes things is 1 to 2, and it is far cheaper than the industry suggests.
The first 90 days
Weeks 1-3 · Targeted inventory. Not an inventory of corporate data: an inventory of what the case chosen in the value layer needs. Which documents, where they live, who maintains them, which is the good version.
Weeks 4-6 · Sufficiency test. Take ten real past cases and check whether the inventoried material would have resolved them. It is the cheapest and most informative test in the whole layer, and almost nobody runs it. If ten historical examples do not work, a thousand will not either.
Weeks 7-10 · Minimum clean-up. Only what the test flagged: confusing duplicates, stale versions that contradict, gaps that force invention. Nothing else.
Weeks 11-12 · Owner and cadence. Who maintains this and how often. Without that decision the corpus degrades within six months and the system starts answering with judgement that is no longer the house's.
What NOT to do yet
- Do not start a data lake. Nor a data mesh, nor a data governance platform. These are good decisions when there is a portfolio; they are ways of deferring when there is one case.
- Do not hand-label thousands of examples. First check with ten whether the approach holds.
- Do not confuse data quality with judgement quality. Impeccable data on the wrong judgement produces impeccable errors.
- Do not outsource the corpus. The material encoding how your company works is the asset; if it lives at the vendor, the differentiation is the vendor's.
What skipping it costs
It costs a particularly demoralising kind of failure: the system works technically and answers badly. It states with confidence things your company would not do that way. The team loses trust within two weeks, and winning it back costs more than doing it right would have.
That, in our experience, is the real meaning of Gartner's 60%. These are not projects that crash into a technical error: they are projects that produce plausible, foreign-sounding output that nobody wants to use.
This is the second layer of the series on the seven layers of AI maturity. The previous one is the value layer, which decides which case you are feeding.
Frequently asked questions
What data does a company need to start with AI?
The data for the specific case it wants to solve, not the whole company's. Usually a bounded corpus of documents where its judgement was applied in writing: proposals, reports, case files or resolved tickets. Ten historical examples are enough to know whether the approach holds.
Do you need a data lake before using AI?
Not for the first case. A data lake is a good decision when there is a portfolio of cases and volume to justify the platform. Before that it is an expensive way to defer: Gartner attributes 60% of abandonments to inadequate data foundations, but the remedy is targeted preparation, not total preparation.
What if our knowledge lives only in people's heads?
Then there is a step before automation: externalising it. While judgement lives only in heads, any system will answer with someone else's. That step is not technical and is usually the most profitable part of the whole project, because it also reduces dependence on specific individuals.
How much data quality is actually needed?
Enough for ten historical cases to resolve correctly. That is a threshold verifiable in weeks, against the unreachable goal of "clean data". Perfect quality is not an AI requirement: it is a common excuse for not starting.

Jordi García is Tech Lead at onext. He works on bringing AI into governed production across development and product teams —with Spec-Driven Development, context engineering and human verification at every step— and authors onext's technical insights on the method, quality and cost of applied AI.
LinkedIn →