Skip to main content
onext technology
Leadership May 2, 2026 - 9 min read

Context engineering: 7 questions for your next AI RFP when headcount is no longer the criterion

The large integrators are starting to sell «context engineering» at scale — with announcements of thousands of hireable engineers and editorial campaigns that will land in your committee before the summer. When that proposal reaches your RFP under the same label, headcount stops being the criterion. It's seven concrete questions that separate the real engineer from the marketing label.

Jordi García
Tech Lead at onext
CIO and CTO in a bright corporate boardroom reviewing AI RFP proposals with a translucent panel behind them displaying seven context engineering evaluation criteria

Quantity is trivial to match. Any large integrator can hire a thousand people and give them the «context engineer» title in six months. What isn't trivial to match — and what decides whether your AI investment produces competitive advantage or a budget hole — is the following: certifiable methodology, verifiable flagship cases, and context ownership in production.

These seven questions are the ones we recommend taking into your next AI RFP. Each one includes why it matters, the red flag to watch for in the vendor's answer, and the preferable alternative. At the end of the article there's a downloadable one-page PDF with the seven condensed, ready to print and take to the committee.

The 2026 editorial context

Why the category will saturate in a matter of months

1,000+ context engineers announced by a single integrator in one campaign
3 consolidated editorial layers: context engineering · context advantage · ethos
Q2-Q3 window in which the category lands in Spanish with several large integrators
7 questions that separate method from headcount in any proposal

Why headcount stops discriminating

When a category is new, the first defensive move by the large integrators is to relabel the team they already have with the category's name. It requires no methodology, no internal certification and no flagship cases: it requires an email from the COO to HR so that «context engineer» appears on LinkedIn next to the «AI engineer», «cloud engineer» or «data engineer» they already had.

The consequence is operational, not semantic. If the RFP asks «how many context engineers do you have», the integrator that sent its press release yesterday answers with the same figure as the integrator that has been doing this for three years. And both charge the same. The real asymmetry — who delivers and who doesn't — shifts to the first quarter of the project, once it's already signed.

The operational insight: headcount is a noisy data point because it's the one most quickly matched across vendors. The following seven questions are designed to restore differentiating criteria through questions that aren't easy to match in the short term: documented method, verifiable cases, auditable certification, context ownership, governance, portability and reproducible eval.

The 7 questions, at a glance

1 Is there a documented, replicable methodology, or «we adapt it to each client»?

Without a document, the methodology is the individual judgment of the assigned engineer. If that person leaves, the context leaves with them.

2 Verifiable flagship cases with a named client and a measurable metric, or anonymous references?

The difference between a pilot and a flagship case: named client, business metric and an auditable period in production.

3 Auditable internal certification, or the engineer's individual judgment?

Cloud certifications (GCP/AWS/Azure) prove the platform, not the discipline.

4 Is the context layer owned by the client or by the vendor in year 3?

The most important question in the RFP. If the context lives on the vendor's platform, it isn't yours when the conditions change.

5 Who decides in production what gets retained and what gets discarded from the context?

Operational decisions have cost and business impact. If they're made outside your organization, you've outsourced your governance.

6 What happens to the context when the contract ends — exportable, portable or trapped?

Without an annual portability test, contractual ownership is theoretical. Any proprietary format is non-portable by definition.

7 Is there an eval/test framework over the context, or «we trust the engineer's eye»?

Context quality is measurable (CLEAR, pass@k, retrieval P/R, drift). Without instrumentation, you're selling trust.

Question 1 · Is there a documented, replicable methodology?

Why it matters: a methodology that isn't documented is, in practice, the individual judgment of the assigned engineer. If that person moves to another project, another client or another company, the methodology goes with them. The context — which is what you're buying — is not replicable backward or forward.

Red flag to watch for: answers like «we customize it for each client» or «it depends on the assigned engineer». That flexibility sounds professional but it means there's no method: there are craftspeople. And craftspeople don't scale.

Preferable alternative: ask to see the methodology document before signing. If it doesn't exist in written, reproducible form, you're taking on the risk that the project's quality depends on whoever happens to be available the week of the kick-off.

The logic is the same one we defend in Spec-Driven Development: controlled, predictable AI. If a system's behavior depends on ad hoc prompts or individual decisions, the system is non-replicable and non-traceable. If it depends on structured specifications, the system is both.

Question 2 · Flagship cases with a named client and a measurable metric?

Why it matters: context engineering is a young category in which almost every vendor has «success stories» that are really non-auditable pilots. The difference between a pilot and a flagship case is that the latter has a named client, a measurable business metric and a period in production long enough for the figures to be real.

Red flag to watch for: «for confidentiality reasons we can't share names». Sometimes it's legitimate — but if all the cases are anonymous, most likely there are no real flagship cases and only projects in pilot state or in recent production.

Preferable alternative: ask for at least two cases with a named client, a concrete metric (not «efficiency improved» but «reduced the support agent's response time by 34% over six months across 12,000 tickets/month») and a direct reference with the IT person who signed off on the project.

Question 3 · Auditable internal certification, or individual judgment?

Why it matters: this extends question 1 to the human piece. A documented methodology doesn't on its own guarantee that the assigned engineer knows how to apply it. Internal certification — an auditable internal process that validates that each person on the team knows and applies the methodology — does guarantee it.

Red flag to watch for: generic vendor certifications (Google Cloud Professional, AWS Certified Solutions Architect, Azure AI Engineer) presented as proof of context engineering competence. They're proof of competence in the platform, not in the discipline.

Preferable alternative: ask whether there's a context-engineering-specific internal certification, what it evaluates, how often it's renewed and who audits it. If the answer is «all our engineers are senior», you're taking on the risk that «senior» means different things depending on the manager.

Question 4 · Is the context layer owned by the client or by the vendor in year 3?

Why it matters: this is the most important question in the entire RFP. The context layer — the curated data, the documented relationships, the codified domain knowledge — is what differentiates your company from the competition. If that layer lives on the vendor's platform, in its trained model, in its internal repository, it isn't yours. And when the situation changes (a price increase, a shift in the vendor's strategy, M&A on their side), you have no leverage.

Red flag to watch for: contract clauses that reserve rights over the «trained model», «the derived weights», «the embeddings built on client data», «the extracted knowledge». Any of those formulations translates to «the context is ours».

Preferable alternative: the contract must explicitly state that the entire context layer — data, embeddings, indexes, curated prompts, structured knowledge — is owned by the client, exportable in open formats and portable to any other platform at no additional cost. If the vendor resists, you've identified the lock-in before signing it.

The discussion parallels the one we had in Proprietary vs. open source LLMs: the model decision can be reversed; the decision about who owns the context layer is much harder to reverse three years later.

Question 5 · Who decides in production what gets retained from the context?

Why it matters: the day-to-day operation of a context engineering system involves constant decisions about what information to retain, what to discard, when to refresh the context, when to invalidate the cache. Those decisions have operational impact (latency, cost) and also business impact (which automated decisions are possible). If they're made outside your organization, you've outsourced your operational governance capability.

Red flag to watch for: an SLA where retention decisions are made off-shore, in a vendor Center of Excellence, or in an «AI Operations» team that reports to the vendor and not to your IT. It sounds efficient; it means you've lost control.

Preferable alternative: the client must have a named role (it can be a profile shared with the vendor) that makes the operational decisions about the context, with full visibility of the logs and authority to change policies. If the vendor offers you «turnkey operations» without that figure, what it's offering you is operational dependency.

Question 6 · What happens to the context when the contract ends?

Why it matters: it extends question 4 to the moment of exit. Even if the contract says the context is your property, if the format it lives in is proprietary and non-exportable, in practice it's trapped. Portability is the difference between having a real option to switch vendors and having a theoretical option that costs six months of migration.

Red flag to watch for: answers like «the context lives on our platform» with no further detail, or «we export to a proprietary format that any vendor can import» (any proprietary format is non-portable by definition).

Preferable alternative: ask for an annual portability test included in the contract. Once a year, export the context, import it into a test platform (yours or a third party's) and verify that the results are equivalent. If it can't be done, you aren't portable. And if you aren't portable, the lock-in is real even if the contract says it isn't.

It's the same logic we apply when evaluating platforms in Claude Managed Agents: the make/buy dilemma isn't about cost. The monthly TCO matters, but the exit cost is the data point almost no one puts on the spreadsheet. In context engineering, the equivalent of exit cost is the real portability of the context.

Question 7 · Is there a reproducible eval framework, or «we trust the engineer's eye»?

Why it matters: context quality is measurable. There are frameworks (CLEAR, pass@k, retrieval precision/recall over golden datasets, drift detection over embeddings, A/B testing of responses with humans in the loop) that produce reproducible quality metrics. If the vendor uses none of them, what it's selling is its team's opinion. The opinion may be excellent — and it can also degrade without anyone detecting it until an end client complains.

Red flag to watch for: a total absence of reproducible context-quality metrics in production. «We review it manually every quarter» is not an eval framework.

Preferable alternative: ask that the contract include an explicit eval framework, run at a minimum monthly frequency, with dashboards accessible to the client and automatic alert thresholds when quality drops below an agreed threshold. If the vendor doesn't have that codified, they're selling you trust without instrumentation.

It's exactly the gap we describe in AI agents in production: the quality gap: observability without traceability is theater. If there's no way to measure quality iteration after iteration, what you're financing is an opaque cycle.

Summary table: real engineer vs marketing label

Dimension Real engineer Marketing label
Methodology Documented, replicable «We adapt it to each client»
Flagship cases Named client + measurable metric Anonymous references
Certification Specific, auditable, internal Generic cloud vendor
Context ownership Client, open format Vendor, «trained model»
Operation Client with visibility and authority Off-shore «turnkey»
Portability Annual test included in contract «We export to a proprietary format»
Eval Framework with dashboards and alerts «We review it manually»

The four-row test: if your next «context engineering» proposal falls into four or more of the right-hand columns, what you're buying isn't context engineering. It's a label. The probability that the project ends in a pretty pilot with no sustained impact is high.

What to do in your next RFP

Three concrete actions for this very week:

1
Add the seven questions to your AI RFP template

None adds friction for the serious vendor; all filter out the one that only carries the label. The one-page PDF at the end of this article is the condensed version.

2
Review the vendors you already have under contract

Not to break relationships, but to understand which columns you're in. Renewals are the best moment to introduce the ownership and portability clauses.

3
Brief your evaluation committee

Fifteen minutes of briefing before the next evaluation change the outcome. The difference between a committee that asks these seven questions and one that values «experience + number of engineers» is the difference between buying context and buying a label.

The first bell has already rung. Other large integrators will follow between Q2 and Q3 of 2026, with the category fully landed in Spanish before the summer. The window to introduce criteria into the evaluation closes every week. These seven filters are our way of saying «no», «not yet» or «yes, with conditions» — before signing.

PDF · 1 page

The 7 questions on one page, ready to print

Condensed version with the seven questions, their red flags and the preferable alternative. Designed to take to the evaluation committee.

Download PDF

Closing

Quantity gets matched. The label gets bought. What isn't matched in six months nor bought with a press release is the replicable method, the verifiable flagship case and the real ownership of the context in production. When the next «context engineering» proposal reaches your desk, the seven questions above are the difference between signing criteria or signing a label.

The best context engineering vendor isn't the one with the most context engineers. It's the one that delivers on Monday what it promised on Friday — with documented methodology, verifiable cases, context that stays yours and an eval framework that doesn't depend on anyone's opinion.

Sources and references: analysis of the «context engineering / context advantage / ethos» editorial narrative consolidated by large integrators between Q1 and Q2 of 2026; public announcements of hiring campaigns with four-digit figures of context engineers; Spanish-language narratives forecast for Q2-Q3 of 2026; onext's Spec-Driven Development methodology.

Further reading: Context Engineering: the discipline | How to choose an AI partner in 2026 | Proprietary vs. open source LLMs | AI agents in production: the quality gap

onext methodology: onext AI-Accelerated Development is the methodology we use to support medium and large companies in the move from pilot to production. Spec-Driven Development as the method layer, multi-cloud and multi-model neutrality, and context ownership and portability from day one.

Jordi García
Written by
Jordi García
Tech Lead at onext

Jordi García is Tech Lead at onext. He works on bringing AI into governed production across development and product teams —with Spec-Driven Development, context engineering and human verification at every step— and authors onext's technical insights on the method, quality and cost of applied AI.

LinkedIn →

Do you have an AI RFP with a «context engineering» label on the table?

In 30 minutes we test the seven questions against your specific RFP: which clauses to add to the contract, how to read the vendor's answers and what to ask for in writing before signing. No commitment.

See how we work

12 teams transformed. 0 sprints lost.