When a CIO, a CTO or a Head of Data opens an artificial intelligence RFP in 2026, the template questions haven't changed in years: number of certifications with the hyperscaler, partner tier (Select, Advanced, Premier), revenue over the last three years, published case studies, number of professionals dedicated to AI. They're legitimate questions — but they've stopped discriminating.
Today almost any reasonably serious provider answers all of them well, and yet the failure rate of enterprise AI projects remains above 40% according to Gartner. There's a mismatch between what gets asked and what predicts success. This article proposes the five criteria that are actually correlated with an AI project reaching production and generating impact — and that almost never get asked in an RFP.
The AI market context in 2026
Public data on consolidation and failure of enterprise AI projects
Why the classic criteria have stopped discriminating
There are two structural reasons. The first is technical: the capability frontier is no longer about "having a team that understands AI." In 2026, practically any mid-sized-and-up technology consultancy has an AI practice, experience with the leading commercial models, and teams that can stand up a pilot on a hyperscaler without much friction. The technology commodity has shifted. What distinguishes a project that works from one that stalls in PoC is no longer the stack — it's how the business context is understood, how the system is operated in production, and how cost is contained when it scales.
The second reason is about the market. Over the last 18 months the ecosystem of AI-native providers has gone through accelerated consolidation. Leading European and American boutiques — Faculty in the UK, NeuraFlash, RANGR and Halfspace across different geographies, Keepler in Spain — have been absorbed by global integrators. In the Spanish market the pattern adds to earlier deals such as Bluetab inside IBM or Synergic Partners inside Telefónica Tech.
The result is that when a buyer asks "what partner tier do you have with AWS?" or "what's your revenue?", the answer no longer tells them whether the team that will show up on Monday to their project is stable, whether the delivery architecture is still the same as before the integration, or whether the rate structure will shift in six months. Questions that once carried informational value now return data that no longer predicts anything relevant.
That's where the five criteria that follow come from. They're not magic and they're not the only ones — they're the ones that, in our experience and that of several of our clients, most correlate with an AI project reaching production on time and on budget.
The five criteria that do predict success
Not of the total workforce. Of the team named by name in the proposal, with declared FTE and an approval clause for substitutions.
The best partners turn use cases down. If it has never happened, that's a bad sign.
Verifiable deliveries on at least two hyperscalers and two model providers over the last 18 months.
Read access to the repo, a reproducible evals pipeline, a ticket board and a post-mortem archive. Open from day one.
Historical churn, a maximum-turnover SLA, a minimum handover window and % of in-house vs. subcontracted staff.
Criterion 1 · Senior/junior ratio of the team actually assigned
Not of the provider's total workforce. Not of the "AI team" in the abstract. Of the specific team that is going to work on your project, named by name in the proposal. In an enterprise AI project, seniority decides almost everything: the ability to say no to a wrong approach, the speed of iteration, the judgment to choose a small model over a large one when cost matters, the discipline not to over-engineer. A team with a 1:5 senior/junior ratio moves more slowly and with more risk than one with a 1:2 ratio — even if on paper the price per day looks better.
Concrete question for the RFP: "Provide the named list of the assigned team, with years of experience on AI projects in production (not pilot), the senior/mid/junior ratio, and the percentage of dedication that will be on the project (FTE)." Ask for it to appear in the contract. Ask for explicit client approval for substitutions.
Criterion 2 · Demonstrated ability to say no to a poorly defined use case
This one is less intuitive but perhaps the one that correlates most with success. The best AI partners turn use cases down. They do it because they know that starting a project without a clear value hypothesis, without access to the right data, or without a sponsor with decision-making power guarantees an ending in PoC and an invoice paid with no impact. A partner that accepts everything you ask of it isn't more collaborative: it's a worse partner. It's charging you to learn something it already knew.
Concrete question for the RFP: "Describe a recent case (last 12 months) in which you recommended that a client not start a project they wanted to launch. What was the alternative recommendation? What happened next?" If the answer is "it's never happened to us," that's a bad sign.
It's the same reasoning we explain in why most LLM projects fail: the problem usually isn't the model, but starting without having validated that the use case had a value hypothesis, data and a sponsor. A partner with judgment closes those three gaps before writing the first line of code.
Criterion 3 · Real vendor neutrality, not declared
In enterprise AI, the ability to operate without getting locked into a specific platform is a technical asset — not a slogan. There are two tests.
The first: has the team delivered over the last 18 months on at least two different hyperscalers (AWS, Azure, GCP) with projects of comparable scale? The second: has it delivered cases with models from at least two providers (Anthropic, OpenAI, Mistral, Google, open models on its own infrastructure)? If the two answers aren't clearly affirmative, the partner has a technical bias (not necessarily bad, but a bias) worth knowing before signing.
The practical corollary is that in 2026 neutrality isn't "we have no preferences" — it's "we have reasoned preferences and the ability to change them when the case calls for it".
Concrete question for the RFP: "What platform or model architectural decision have you reviewed in the last six months on a live project, and what data made you change it?"
It's the same logic we describe in proprietary vs. open-source LLMs: the right model decision depends on the specific use case, not on a structural bet by the partner. Whoever has delivered with several models has data to argue with; whoever has delivered with only one has a hypothesis.
Criterion 4 · Delivery traceability
This is the question that makes providers most uncomfortable and informs the buyer the most:
"Can I have read access to the project repository, the evals pipeline, the ticket board and the post-mortem archive?"
In a serious AI project in 2026, all of that exists (if it doesn't, that's a problem). A solid partner opens it without friction: it understands that traceability is part of the value delivered, not an added cost.
The RFP question, broken down:
- Is the project code delivered versioned in a repository accessible by the client from day one?
- Are model evaluations run in a reproducible pipeline (with seeds, frozen datasets, logged metrics)?
- Are incident post-mortems documented and shared with the client?
- Are operational runbooks delivered as part of the product, not as a "next phase"?
If the provider resists opening this up, it's saying something about how it works. Partners that operate with traditional software discipline tend to feel comfortable here; those coming from other backgrounds, less so.
It's a point close to what we develop in AI agents in production: the quality gap: observability without traceability is theater. If the partner can't show how it evaluates and how it learns iteration after iteration, the client is funding a cycle that repeats.
Criterion 5 · Team stability over the next 12-24 months
The last criterion is the hardest to measure and the one that weighs most. In professional services, turnover of the assigned team tends to be invisible to the client until it's felt — and by the time it's felt, you're already mid-project and switching partners costs more than tolerating the degradation. The smart buyer asks up front:
- What is the annual churn of the AI-dedicated workforce over the last 24 months?
- Is there a contractual SLA for maximum turnover on the team assigned to the project?
- Is there a substitution-approval clause for key roles (architect, lead engineer, lead data)?
- Is there a minimum handover window if someone leaves the team (proposed: 20 business days)?
- What percentage of the assigned team is the provider's own staff versus subcontracted?
This criterion becomes especially relevant when the partner has recently been acquired. The literature on mergers and acquisitions in professional services agrees that the 12-24 months following an integration concentrate the highest turnover rate, with ranges usually reported between 15% and 30% of senior talent depending on sector and context (for example, Deloitte Human Capital and Bain reports on post-acquisition integration). It isn't a criticism of any provider in particular: it's a sector constant worth anticipating.
What to ask for in writing when your partner has just been acquired
This past year the European map of data and AI boutiques has consolidated fast. Accenture closed the acquisition of Keepler Data Tech in Spain on April 8, 2026, adding it to a chain that already included Faculty, NeuraFlash, RANGR and Halfspace. In the Spanish market the deal isn't isolated: Bluetab is inside IBM, Synergic inside Telefónica Tech, and other local providers are moving with the same logic. From the client's side, that means the shortlist of partners considered 18 months ago probably no longer reflects the current picture.
Given this scenario, it's worth asking in writing for four things from a partner going through an integration — or from any partner in a renegotiation:
With a client-approval clause for substitutions in key roles (architect, lead engineer, lead data).
For the next 12-24 months, with contractual consequences if breached.
And which move to the integrator's standards, with concrete transition timelines.
Applicable over the next 6 months. It's a reasonable request when the provider has substantially changed its ownership and operating conditions.
These four things are asked for with respect and answered in writing. A partner that sees them as an aggression is saying more about itself than about the client.
Where each type of partner fits
A final note, because this article isn't trying to push any particular profile. The enterprise AI market in 2026 rests on three distinct partner archetypes, and each makes sense in its context.
Independent technical boutique
Specialist firms with a stable senior team, not dependent on a single hyperscaler.
Fits: projects where the decisive criterion is the specific composition of the assigned team, multi-cloud/multi-model neutrality, speed of iteration and the ability to accompany the shift from pilot to operation with predictable cost. In Spain in 2026 several remain with a relevant profile; the informed buyer identifies them and keeps them on the shortlist.
Platform provider with a network of delivery partners
Databricks, Snowflake, Anthropic via certified partners.
Fits: decisions strongly anchored to a specific platform and a relatively predictable implementation, where the partner adds delivery capacity on top of a stack the client has already chosen.
Every project has its profile. The most common mistake is trying to make the same type of partner work for all three contexts. A regulated multi-country program with global scope isn't executed the same way as a 12-week sprint to take a recommendation model to production with predictable cost. Confusing the archetypes usually ends in over-cost (if you hire a global integrator for something small) or in operational deadlock (if you hire a small boutique for something massive).
The operational insight: the first decision in an AI RFP in 2026 isn't who to call, but which partner archetype fits your project. Only after that do the five selection criteria make sense. Reversing the order is what makes poorly built shortlists turn into impossible evaluations between providers that aren't competing for the same thing.
Closing
Choosing an AI partner in 2026 isn't easier than it was three years ago — it's different. The criteria that used to discriminate (certifications, partner tier, revenue) now return noisier information than useful information. The five criteria proposed here — the senior/junior ratio of the assigned team, the ability to reject poorly defined cases, verifiable vendor neutrality, delivery traceability and team stability — are neither secret nor proprietary. Any buyer with judgment can add them to the template of their next RFP and turn a commercial conversation into a genuinely technical one.
The best AI partner isn't the one with the most certifications. It's the one that will be delivering on Monday the same thing it promised on Friday — with the same team, the same traceability and the same rejection criteria it applied in the proposal.
Sources and references: Gartner Research — AI Project Success Rates 2026; MIT NANDA — State of AI in Business 2026; public press releases on M&A deals (Accenture-Keepler April 2026, IBM-Bluetab, Telefónica Tech-Synergic Partners, Accenture-Faculty, Accenture-NeuraFlash, Accenture-RANGR, Accenture-Halfspace); Deloitte Human Capital and Bain — reports on turnover in post-acquisition integrations in professional services.
Further reading: Why most LLM projects fail | Proprietary vs. open-source LLMs | AI agents in production: the quality gap | Compliance-First AI Design (EU AI Act)
onext methodology: onext AI-Accelerated Development is the methodology with which we accompany medium and large companies through the shift from pilot to production. A stable senior team named by name, multi-cloud and multi-model neutrality, delivery traceability from day one and predictable cost. Without stopping deliveries.

Jordi García is Tech Lead at onext. He works on bringing AI into governed production across development and product teams —with Spec-Driven Development, context engineering and human verification at every step— and authors onext's technical insights on the method, quality and cost of applied AI.
LinkedIn →