Skip to main content
onext technology
AI July 19, 2026 - 11 min read

AI for insurers: why claims and underwriting pilots don't reach production

Mid-market insurers fill up with AI pilots that impress in the demo and stall before production. It's not a model problem: it's context, cost control and compliance. What automating claims, underwriting and document response really requires.

Jordi García
Tech Lead at onext
Insurance operations professional reviewing claims and underwriting documents on screen at dusk, illustrating why insurance-sector AI pilots don't reach production

For your leadership team (60 seconds)

  • What's happening: your insurer has AI pilots in claims, underwriting or RFP response that work in the demo and never quite reach production. It's the sector's pattern, not a one-off failure.
  • What it means for your company: document-heavy processes stay manual, the cost of AI grows without the savings scaling, and the risk and compliance team blocks every release because there's no traceability or control. The business case dilutes.
  • What you can do: don't change the model, change the method. One concrete process, with the context of your insurance business, a verifiable quality criterion, cost measured per case and governance from day one. That's how a pilot becomes an operation, not another shelved POC.

If you're looking for AI for your insurer —or already have pilots running—, you probably recognise the scene: a vendor shows you a demo that reads a claim report and summarises it in seconds, or classifies an underwriting risk instantly. It's impressive. Six months later, that pilot is still a pilot. The real automation of claims and underwriting hasn't reached production, and nobody can explain why. The reason is the same at almost every mid-market insurer, and it has nothing to do with the AI model you chose.

Why insurers accumulate pilots that don't scale

Insurance is, at its core, a business of documents and regulated decisions: reports, policies, terms and conditions, expert appraisals, histories, tender responses. It's exactly the kind of work where generative AI seems to fit perfectly —and that's why pilots multiply—. But a demo is done with the clean file and the happy path; real operations have the casuistry of thousands of policyholders, with odd formats, exceptions and regulatory consequences. Between the demo and that reality there are three walls that any insurer CIO or head of operations recognises:

  • The AI doesn't know your insurance business. A generic assistant doesn't know what a deductible is in your line, how you tell fraud from a legitimate claim, or which clause of your terms applies. It answers generically and gets it wrong right at the edge where the error is expensive.
  • There's no reproducible quality criterion. Without an automatic way to check whether the AI classified a risk correctly or extracted a datum from the report correctly, every release is a bet. And in a process that affects the premium or the payout of a claim, betting isn't an option.
  • Cost and traceability aren't under control. Without measuring the cost per case processed and without leaving a trail of why the AI decided what it decided, the token bill grows, the compliance team blocks the deployment, and the business case that justified the project disappears.

The three processes where AI promises the most (and stalls the most)

Not all processes fail for the same reason, but the pattern repeats. These are the three where mid-market insurers concentrate their pilots —and where the leap to production is decided:

Claims

Automating claims with AI —reading the report, extracting the data, cross-checking against the policy, proposing a resolution— is the star case. And the one that breaks most when scaling: each line has its documentation, each claim its exception, and an error in the extraction or in applying a coverage has a direct cost. Production requires that the AI understand your claims typology and that every decision be verifiable and auditable, not that it "usually gets it right".

If this is your case, the piece that drills into the detail —the six phases of a claim, what the AI Act requires after the July 2026 reform and how to start with one narrow claim type— is automating insurance claims with AI.

Underwriting

Automating insurance underwriting with AI —assessing the risk, classifying, requesting missing information— promises to cut issuance times. But underwriting is deciding on the premium and on who you insure: a model that biases or hallucinates isn't a bug, it's a regulatory and commercial problem. Without the context of your underwriting rules and without control over why it decides, it won't pass the risk filter.

Document response and RFP

Brokers and insurers spend hundreds of hours responding to tenders and producing technical documentation over a huge corpus of terms and precedents. A RAG over those documents is one of the highest-return uses —and one of the easiest to underestimate—: if the system cites a clause wrong or mixes up terms, the error travels to a client or a binding offer. The value is real; so is the demand for accuracy.

The detail of why context —and not the prompt— is what makes AI answer like your business: context engineering vs. prompt engineering.

What production in a regulated environment requires (that the POC doesn't)

Taking AI from an insurance pilot to production isn't a problem of choosing a better model: it's a problem of consistency, governance and cost, made worse because the sector is regulated. Four things the demo skips and production doesn't forgive:

  • Context engineering of your business: capturing how your lines really work, your underwriting rules, your claims typology and your terms and conditions, and turning it into the context the AI uses. It's what makes it answer like your insurer and not like a generic chatbot.
  • Verification at every step: reproducible evaluations that check whether the output is correct before it affects a case. Without evals there's no possible production in a process that touches premiums or payouts.
  • Traceability and compliance (AI Act): compliance-first design —decision logging, explainability, human control— isn't a later add-on; it's what lets risk and compliance approve the deployment. In insurance, a system that can't justify its decisions doesn't reach production.
  • Cost per case, measured: the real cost of AI isn't the model's price, it's what it costs to process a useful case end to end. Measuring it from the first prototype is what avoids the bill that multiplies when scaling and keeps the business case alive.

How to start without a "big AI project"

The expensive mistake is launching a cross-cutting AI programme that promises to transform the whole company and that, six months later, hasn't put a single process into production. The way that works is the opposite:

  1. One process, not a platform. Pick the process with the clearest and most measurable pain —usually claims in a specific line, or document response—. One, done well, teaches more and convinces more than ten done halfway.
  2. Quality criterion before automation. Define how you'll know the AI is right (the eval) before building anything. If you can't measure it, you can't put it in front of a policyholder.
  3. Governance and cost from day one. Instrument traceability and cost per case from the first prototype, not when the audit or the bill arrives. In a regulated sector, this isn't optional.
  4. From the proven pilot, to expanding. When a process works in production with metrics, you have the internal case to bring the same method to the next one. Your insurer's corporate AI is built process by process, not in a big bang.

An honest note

At onext we work the method —making AI understand your business and reach governed production— in a sector-agnostic way; we don't claim to be actuarial specialists or to have a public insurance case we haven't published yet. What we do bring is the rigour of taking document-heavy processes to production with verification, compliance and cost under control. If your insurance process is the first, we approach it with that honesty up front.

Frequently asked questions

What is AI for in a mid-market insurer?

Above all, to automate document- and decision-heavy processes: claims management (reading reports, extracting data, cross-checking against the policy), underwriting (assessing and classifying risk), and document or RFP response over your corpus of terms and conditions. The value isn't in a generic chatbot, but in applying AI to a concrete process with the context of your business and with cost and compliance control.

Why isn't my claims or underwriting AI pilot reaching production?

Because the demo uses the clean file and production uses the real casuistry of thousands of policyholders. Without context engineering (so the AI knows your lines, rules and terms), without reproducible evals (that check whether it gets it right) and without traceability for compliance, the system fails at the edge, the risk team blocks the deployment and the cost spirals. It's not the model: it's the method that separates a POC from a governed operation.

How does an insurance AI system comply with the AI Act?

With compliance-first design: logging of every decision, explainability of why the AI proposed what it proposed, and human control at the points that affect premium, coverage or payout. Traceability isn't added at the end; it's designed from the start, and it's precisely what lets the system pass the risk and compliance filter and reach production.

Conclusion

Insurers' AI pilots don't stall for lack of talent or for choosing the wrong model: they stall by skipping the method that turns a demo into a regulated, profitable operation. Context of your insurance business, verification at every step, traceability for compliance and cost per case measured — that's what makes claims, underwriting and document response reach production and hold. And it's done through concrete processes, not with a grand project that promises everything.

If you have an insurance pilot stuck between the demo and production —or you want to start off right— begin with a diagnostic: in a few weeks you'll know what context and what evals are needed, what your process requires to comply, and what it really costs.

All of this is part of a broader idea: building your company's AI intelligence, process by process, with the AI that understands how you work.

Jordi García
Written by
Jordi García
Tech Lead at onext

Jordi García is Tech Lead at onext. He works on bringing AI into governed production across development and product teams —with Spec-Driven Development, context engineering and human verification at every step— and authors onext's technical insights on the method, quality and cost of applied AI.

LinkedIn →

Do you have an AI pilot stuck between the demo and production?

An onext diagnostic tells you, in a few weeks, what context and what evaluation your process —claims, underwriting or document— needs to reach production with compliance and cost under control.

See how we work

A universal method, applied to document-heavy processes: context of your business · verification at every step · cost per useful task measured.