---
title: "Agentic AI for insurance claims: eight of nine agents talk, one investigates"
description: "Of nine named agentic claims deployments catalogued from 2025 and 2026, eight sit in intake, summarization, or estimating. One sits at the investigation layer, and it calls itself a support tool."
date: "2026-08-24"
lastModified: "2026-08-24"
author: "Pankaj Dhariwal"
tags: ["Guides"]
canonical: "https://gethesperai.com/blog/agentic-ai-insurance-claims-beyond-fnol/"
---

# Agentic AI for insurance claims: eight of nine agents talk, one investigates

> **TL;DR** Of the nine named agentic claims deployments and product launches catalogued in this post from 2025 and 2026, eight sit in the conversation, estimating, or data-access layers. One, Thomson Reuters CLEAR Investigate, sits at the investigation layer, and it describes itself as a support tool. The industry is spending its agentic capacity on the parts of a claim that get talked about, not the parts that get contested.
>
> - 8 of 9 catalogued agentic claims launches sit outside investigation
> - Allianz runs seven agents on one claim type, capped at AUD $500
> - The UCSPA governs investigation, the least-automated layer of all

- **1 of 9** - catalogued agentic claims launches sit at the investigation layer (Our own count of named 2025-2026 deployments, not a market share)
- **7%** - of insurers say they have achieved scalable AI success (Sedgwick report, via Insurance Journal, March 2026)
- **AUD $500** - ceiling on the claim type Allianz runs with seven agents (Allianz media center, November 2025)
- **14+ days** - manual SIU investigation, per case (Hesper internal benchmarks)

In July 2026, Insurance Australia Group, the largest general insurer in Australia, named OpenAI as its partner for agentic claims work. Read the published task list back: lodge, guide, inform, update, absorb surge. Every verb is a conversation verb. Not one adjudicates and not one investigates. One of the largest general insurers in a G20 market went to one of the most heavily capitalised AI labs in the world, and the first thing the two of them built was a better phone line.

IAG never claimed otherwise, which is what makes it evidence rather than a gotcha. But run the same read across every agentic claims deployment announced in the last two years and the pattern holds. Of the nine catalogued in this post, eight sit in the conversation, estimating, routing, or data-access layers. One sits at the investigation layer: Thomson Reuters CLEAR Investigate, launched in early 2026, which describes itself as a support tool. That count is our own catalogue of dated 2025 and 2026 announcements, not a market survey and not a census. It is still the clearest available picture of where the industry is pointing its first agents.

What follows is the vocabulary to tell those layers apart during a vendor demo: a conversation-layer versus decision-layer split, a table of all nine deployments with what each agent actually does, a six-rung autonomy model grounded in Shift Technology's published ARISE framework, the four properties that separate an agent from a workflow with a language model inside it, the five things that break when you aim an agent at a contested claim, and what the regulators have actually written down. This is a cluster post under our guide to [autonomous AI claims investigation](/blog/autonomous-ai-claims-investigation-pillar/).

## IAG, OpenAI, and the shape of the first big agentic claims deal

OpenAI launched [Presence](https://openai.com/index/introducing-openai-presence/) on July 22, 2026. [VentureBeat](https://venturebeat.com/orchestration/openai-unveils-presence-a-new-platform-that-lets-enterprises-launch-and-manage-realtime-voice-agents-and-chatbots) described it as a platform that "lets enterprises launch and manage realtime voice agents and chatbots," and named IAG as one of three launch customers alongside BBVA in Mexico and SoftBank. IAG [confirmed the partnership](https://www.iag.com.au/newsroom/innovation/iag-partners-with-openai-on-agent-solutions) five days later, on July 27, 2026.

The delivery target is the first half of IAG's FY27, which [iTnews](https://www.itnews.com.au/news/iag-turns-to-openai-presence-for-faster-disaster-claims-response-627690) specifies as "in the first half of its 2027 financial year, which runs from July to December 2026." The initial scope is natural perils surge capacity and simple claims intake. [insuranceNEWS.com.au](https://www.insurancenews.com.au/corporate/open-ai-agentic-platform-to-answer-iags-claim-calls) lists the agent tasks as answering calls from customers wanting to lodge claims, guiding motor claim submissions, providing policy information, answering customer questions, keeping customers updated, and managing surge events. iTnews reports the agents will handle "simple, non-event related claims" over the phone, and adds the detail that matters most for calibration: IAG's engagement "has so far centred on solution design, rather than live deployment."

Julie Batch, IAG's CEO of Retail Insurance Australia, framed it in support terms in remarks carried by [fintech.global](https://fintech.global/2026/07/27/iag-teams-up-with-openai-to-boost-insurance-claims/): "We are excited to partner with OpenAI as it will provide our people with extra support to help them deliver higher value service to customers." She said the initial focus "will be on where the need is greatest," which she tied to high-volume natural perils. An IAG spokesperson told insuranceNEWS.com.au that "Customers will be clearly informed when they are interacting with an AI-enabled service and will have the option to speak with a human representative if they prefer." OpenAI's head of go-to-market for Australia and New Zealand described the goal as exploring how customer agents can "help resolve issues faster, expand service coverage and deliver more consistent experiences, while reducing the operational burden on frontline teams."

The architecture is scoped the same way the language is. iTnews describes Presence's design as giving each agent "only the knowledge and system access needed for that job," with a handoff "to a person when a case falls outside its remit." That is a correctly built conversation agent. It also tells you what the agent is not: a system that expands its own tool access when the file gets complicated, which is precisely what an investigation requires.

One more number, at the weight it deserves. OpenAI reports that its own English-language phone support line now resolves 75% of inbound issues without human assistance. VentureBeat notes those figures are "company-reported and have not been independently verified." Read it as a strong conversation-layer result and nothing more. Issues resolved on a call are not claims resolved on a file.

> **What IAG claimed, and what it did not**
>
> IAG picked the right first use case for a voice agent: high volume, low judgment, surge-sensitive, and conduct-visible, so failure shows up fast and shows up loudly. Nothing in the announcement claims investigation, adjudication, or coverage determination. This post is not an argument that IAG chose wrong. It is an argument about where the industry's agentic capacity is being allocated, and the largest agentic claims deal announced in 2026 landing squarely in intake is the cleanest evidence available.

## The conversation layer and the decision layer

Every claims agent sits on one of two layers. At the conversation layer, the agent's counterparty is a person and the output is an interaction. At the decision layer, the counterparty is a file and the output is a determination that somebody else acts on. Intake agents live on the first. Detection and investigation live on the second.

That axis cuts across the one we usually draw. The claims fraud stack has three layers: prevention blocks bad claims before they are filed, detection flags suspicious claims after first notice of loss, and investigation takes a flagged claim and resolves it. Detection is upstream; investigation is downstream. That model is the spine of [our three-layer breakdown of prevention, detection, and investigation](/blog/insurance-fraud-prevention-vs-detection-vs-investigation/), and it still holds. The conversation-versus-decision split is the second axis, and it is the one carriers are actually buying along in 2026.

Take the conversation layer first. Success means the person was understood, the record was captured, and the handoff happened when the case left the agent's remit. Failure is a bad customer experience, which is visible within minutes and correctable within one call. Intake, status updates, policy questions, and surge absorption all live here.

At the decision layer, the agent's counterparty is a file and the output is a determination that somebody else acts on. Success means the determination survives a second reader: the adjuster who pays or denies on it, the SIU lead who signs the referral, the examiner who pulls the file two years later. Failure is not visible within minutes. It is visible in a market conduct exam, an examination under oath, or a bad-faith complaint. Detection sits at the decision layer and produces a score. Investigation sits at the decision layer and produces a finding.

The two axes are independent, and conflating them is how a demo wins a meeting it should not. An agent can be highly autonomous within the conversation layer and still never touch a decision. A rung-3 conversational task agent that runs an entire FNOL dialogue to completion, picks the right follow-up questions, and escalates cleanly is a genuinely agentic system. It is also not doing decision-layer work, and it will not become an investigator by being given more autonomy on the axis it already runs on.

There are two questions that separate the layers in any demo, and an SIU director can ask both in thirty seconds. What does this agent produce? And who reads it next? If the answer to the second question is a person on the phone, you are looking at a conversation agent. If the answer is an adjuster, an SIU lead, a defense attorney, or a state Department of Insurance examiner, you are looking at a decision-layer product and every claim it makes needs to be checked against that standard.

*Figure: The largest agentic claims deal of 2026 built a better phone line. The flagged claim still waits 14+ days for a human. Six rungs of claims agent autonomy, and the nine catalogued deployments plotted across them.*

## Nine deployments and the layer each one occupies

Here is the catalogue. Nine dated, sourced agentic claims deployments and product launches from 2025 and 2026, with what each agent actually does rather than what the category label says it does.

| Deployment | Date announced | What the agent actually does | Layer |
| --- | --- | --- | --- |
| IAG with OpenAI Presence | July 27, 2026 | Answers claim lodgement calls, guides motor submissions, provides policy information, keeps customers updated, absorbs surge volume | Conversation / intake |
| Guidewire Agentic FNOL | August 3, 2026 | Guides claimants through first notice of loss using conversational AI voice, capturing key claim details | Conversation / intake |
| Guidewire Claim Summarization | August 3, 2026 | Gives adjusters claim summaries so they can skip manual note review | Summarization |
| Shift Claims | September 16, 2025 | Assesses, triages, and advises; automates tasks or entire claims while the insurer retains control of the process | Triage / handler assist |
| Sedgwick Omni | May 4, 2026 | Document and call summarization, digital triage, severity modeling, automated reserving, fraud detection, quality oversight | Triage, reserving, detection |
| CCC Intelligent Solutions | August 26, 2025 | Proactively notifies customers of repair delays, digitizes and reconciles invoices, coaches human staff during live interactions | Estimating / back office |
| Verisk MCP connectors in Claude | May 5, 2026 | Gives natural-language access to underwriting indications and restoration estimating data for someone else's agent | Data access |
| Allianz Project Nemo | November 3, 2025 | Seven agents run food spoilage claims under AUD $500 from filing to human-review-ready; the payout decision is excluded by design | Bounded straight-through settlement |
| Thomson Reuters CLEAR Investigate | April 28, 2026 | Interprets an investigator's input, builds a plan, and runs research across premium records and curated open-web sources | Investigation (research agent) |

> **How to read the one-of-nine count**
>
> This is our own catalogue of named, dated agentic claims deployments and product launches from 2025 and 2026, assembled from vendor announcements and trade coverage. It is not a market share, not a census, and not a claim about deployment volume. Plenty of unannounced work exists inside carriers. The count is useful for the same reason a press-release survey is useful: it shows what the industry is willing to put its name on, and where it is choosing to point its first agents.

### Allianz Project Nemo is the most autonomous thing shipping, and it is capped at AUD $500

Project Nemo launched in July 2025 and was [described publicly by Allianz](https://www.allianz.com/en/mediacenter/news/articles/251103-when-the-storm-clears-so-should-the-claim-queue.html) on November 3, 2025. Seven specialised AI agents coordinate the full workflow for food spoilage claims under AUD $500 caused by severe weather on home contents policies. One of the seven is a fraud agent that checks for signs of fraud before human review. Allianz reports an 80% reduction in claim processing and settlement time, with the full seven-agent workflow completing in less than five minutes from filing to human-review-ready. Thomas Baach, Managing Director of Core Insurance Platforms at Allianz Technology, said processing time for those claims fell from several days to one day or even hours.

> By design, payout decisions are never automated.
>
> - Maria Janssen, Chief Transformation Officer, Allianz Services, November 2025

Now look at every constraint on Nemo at once: one claim type, a hard dollar ceiling, a weather-event trigger that can be checked against an external record, a single policy form, and an explicit carve-out of the payout decision. Those constraints are not timidity. They are the conditions under which the claim has a checkable answer. Was there a storm. Was there an outage. Is the policy in force. Is the amount under the ceiling. Each of those resolves to a fact the agent can retrieve, and the seven agents work because every question they ask has a retrievable answer. That is the most useful thing in this entire catalogue, and it is a lesson about what autonomy requires rather than a preview of where it goes next.

### Thomson Reuters CLEAR Investigate is the one genuine investigation-layer entrant

CLEAR Investigate launched in early 2026 and was described by Michael Purcell, Senior Solution Consultant at [Thomson Reuters](https://legal.thomsonreuters.com/blog/clear-investigate-ushering-in-the-era-of-agentic-ai-for-corporate-risk-fraud-professionals/), on April 28, 2026. The product "Interprets your input, builds a plan," and guides the process through to an answer. Thomson Reuters says it "thinks in workflows, not isolated steps" and is "a support tool that plans, executes, and adapts." It "Queries premium Thomson Reuters content alongside curated open-web sources, synthesizing results into concise, actionable insights." The named audience includes KYC, CIP, AML and BSA, TPRM, and Special Investigation Units.

That is real multi-tool planning aimed at investigative work, and it belongs on this list at full weight. A flat claim that the investigation layer is empty would be wrong. The accurate claim is narrower and still holds. CLEAR Investigate is an investigator-facing research agent over a records corpus plus the open web. It is horizontal across financial crime, third-party risk, and insurance rather than P&C-claim-native, and it names itself a support tool. That is a real rung on the ladder. It is a different thing from an agent that takes a flagged P&C claim as input and returns a resolved, provenance-complete finding that a carrier can file under its antifraud plan. Both halves of that sentence are true and the post needs both.

### Guidewire shipped the framework without shipping the investigation agent

Guidewire's [Qusar release](https://www.guidewire.com/about/press-center/press-releases/20260803/guidewire-introduces-qusar-release-to-help-insurers-build-and-control-ai-agents) landed on August 3, 2026 with an Agentic Framework and three pre-built agents. Claim Summarization "Provides adjusters claim summaries, allowing them to focus on complex resolutions instead of manual note review." Policy Change is "an embedded assistant" for underwriters and service representatives. Agentic First Notice of Loss "Guides claimants through the first notice of loss using conversational AI voice, capturing key claim details to improve the customer experience." Diego Devalle, Chief Product Development Officer, framed the framework as "giving developers and AI builders the tools to engineer and safely deploy insurance-aware AI agents into their daily operations."

Two of the three shipped agents are conversation or summarization; one is policy servicing. Zero are investigation. That is the correct division of labour for a claims system of record, and it is what makes Guidewire an integration target rather than a competitor at this layer. The platform vendor supplies the runtime, the model routing, and the secure access to policy, claims, and billing data. The investigative methodology is a different asset and it does not come in the box.

### Shift, Sedgwick, CCC, and Verisk repeat the pattern

Shift Technology launched Shift Claims on September 16, 2025, with agents that "assess, triage, and advise, as well as automate tasks or entire claims while insurers retain full control of the process." Eric Sibony, Chief Scientist and Chief Product Officer, described it as increasing automation when appropriate and providing "the right kind of advice, guidance, and recommendations when a human claim handler needs to be in the loop." Shift reports early-adopter figures of 3% lower claims losses, 30% faster claims handling, a 60% overall automation rate, and greater than 99% accuracy in claims assessment; those are vendor-reported and not independently verified. This is a well-built product and it is deliberately handler-assist. The product-level comparison lives in [Shift Claims versus Hesper](/blog/shift-claims-agentic-ai-vs-hesper/) rather than being re-argued here.

Sedgwick announced Omni at RISKWORLD on May 4, 2026, with document and call summarization, digital triage, severity modeling, automated reserving, fraud detection, and quality oversight. CEO Mike Arbour called it "expert-led, AI-assisted, and relentlessly outcome-focused," and Chief Transformation Officer Vishy Padmanabhan said it was "designed to enhance examiner effectiveness." Note where that capability list stops. Fraud detection is on it. Fraud investigation is not. That is the layer boundary drawn by the largest TPA in the market, in its own launch materials, with no reason to draw it.

CCC's agentic work, described by Vice President of Product Management Mark Fincher in August 2025, covers proactive customer notification when cars will be late, invoice digitization and reconciliation without accountant review, and monitoring live customer interactions to coach human employees. Fincher's framing of oversight is worth keeping: "It's not going to guess and the human knows." Verisk, on May 5, 2026, shipped two Model Context Protocol connectors into Anthropic's Claude, giving natural-language access to Verisk Underwriting Intelligence and Verisk XactRestore. Neither connector is claims investigation. One is underwriting, one is restoration estimating. Even the industry's data utility is aiming its agent-readiness at underwriting and estimating first.

## Six rungs of claims agent autonomy

Claims agent autonomy runs six rungs, from a scripted bot that chooses nothing to an autonomous investigator that resolves a contested claim. The four middle rungs map cleanly onto Shift Technology's published ARISE framework. The sixth does not map onto anything, and that gap is the argument of this post. The six-rung model below is Hesper's framing, not an industry standard.

The regulator built the autonomy spectrum into its definitions before the vendors built products along it. The [NAIC Model Bulletin on the Use of Artificial Intelligence Systems by Insurers](https://content.naic.org/sites/default/files/inline-files/2023-12-4%20Model%20Bulletin_Adopted_0.pdf), adopted December 4, 2023, defines an AI System as "a machine-based system that can, for a given set of objectives, generate outputs such as predictions, recommendations, content (such as text, images, videos, or sounds), or other output influencing decisions made in real or virtual environments," and closes the definition with a sentence that does a lot of work: "AI Systems are designed to operate with varying levels of autonomy."

The one published industry framework along that spectrum is Shift Technology's [ARISE framework](https://www.shift-technology.com/resources/reports-and-insights/arise-the-5-levels-of-insurance-ai-autonomy), published in June 2026. Five levels: Answers, Recommends, Initiates, Solves, Exceeds. At Answers, "The agent operates within a conversational interface, answering specific inquiries from claim handlers." At Recommends, the agent "proactively analyzes data to highlight insights and recommend the best next steps for complex claims without being prompted." At Initiates, "While keeping a human in the loop, the agent prepares all necessary actions," leaving the handler a single click to validate. At Solves, "the agent processes and solves claims in full autonomy," which Shift ties to over 99% accuracy. At Exceeds, the agent "mimics top-tier claims handlers by knowing when to safely deviate from standard operating procedures to optimize the final claim outcome and maximize customer satisfaction."

ARISE is a genuine contribution and carriers should use it. Two things about the page as published matter here, and both are fair: it makes no explicit comparison to any automotive autonomy standard, and fraud investigation is not named as a use case in the framework itself. We wrote the long-form response when it published, in [ARISE levels of autonomy versus investigation-specific AI](/blog/arise-framework-investigation-autonomy-rebuttal/), so the argument is not repeated here. What follows extends it into a rung model for agentic claims products generally.

| Rung | Name | What it does | Nearest ARISE level | Shipped example |
| --- | --- | --- | --- | --- |
| 0 | Scripted bot | Fixed decision tree, IVR, form logic. No model chooses anything. | Below Answers | Legacy claims IVR |
| 1 | Retrieval assistant | Answers questions over a corpus. Informs, does not act. | Answers | ARISE Level 1 as described |
| 2 | Single-tool agent | One tool, one fixed path: extract, summarize, transcribe, draft. | Answers / Recommends | Guidewire Claim Summarization |
| 3 | Conversational task agent | Runs a bounded dialogue to completion, captures structured output, escalates when the case leaves its remit. | Recommends / Initiates | IAG with Presence, Guidewire Agentic FNOL |
| 4 | Multi-tool planner on bounded claims | Chooses the tool sequence and executes to a settlement-ready state on a claim type with a checkable answer and a dollar ceiling. The human keeps the payout decision. | Initiates / Solves | Allianz Project Nemo |
| 5 | Autonomous investigator | Takes a flagged, contested claim. Plans and runs many evidence phases in parallel, reconciles conflicting evidence, handles adversarial inputs, and emits a provenance-complete finding an SIU lead can defend to a regulator or in a deposition. | No ARISE level describes this | Hesper AI |

The load-bearing point about this table is not that rung 5 is higher. It is that rung 5 is not on the same axis. ARISE measures autonomy over a cooperative counterparty on a claim whose correct answer is discoverable, and it measures it well. Rung 5 starts from a claim where the counterparty may be constructing the record and the correct answer is contested. That is why the top of ARISE does not describe an investigator: an agent that deviates from standard operating procedure to maximise customer satisfaction is doing the opposite of what a contested claim requires. Turning up the autonomy dial on a handler agent produces a more autonomous handler agent.

> A conversation agent is graded by the person it talked to. An investigation agent is graded by the second reader: the SIU lead, the defense attorney, the DOI examiner. Those are different products with different failure modes, and no amount of autonomy on the first one produces the second.
>
> - Hesper AI product research

## Agent or workflow: four properties to hold a vendor to

An agent selects its own tools and sequences them toward a goal. A workflow runs a language model along a path a developer fixed in advance. Both get badged agentic in 2026, and most claims demos carrying the label are the second thing. Four properties tell them apart, and an SIU director can check all four without a technical background.

The engineering definition is cleaner than the marketing one. Anthropic's [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents), published December 19, 2024, draws the line: "Workflows are systems where LLMs and tools are orchestrated through predefined code paths," while "Agents are systems where LLMs dynamically direct their own processes and tool usage, maintaining control over how they accomplish tasks." Agents "begin their work with either a command from, or interactive discussion with, the human user," then "plan and operate independently, potentially returning to the human for further information or judgement," and must "gain 'ground truth' from the environment at each step (such as tool call results or code execution)."

Here are the four, in the order they matter in a demo.

1. Tool use the model itself selects. Ask which tools were available on this claim and which ones the system chose not to call. If the answer is that it always calls the same five, it is a script.
2. Planning that decomposes a goal into steps the model chose. Ask to see the plan for one specific claim and the plan for a different claim, and check whether they differ.
3. Memory across steps within a task, and across tasks where it matters. Ask what the system knew at step nine that it did not know at step two.
4. A loop with environmental feedback, not a single forward pass. Ask what happens when a tool call returns something that contradicts the working hypothesis.

A product that runs one prompt over a claim note and returns a paragraph is a workflow with a language model in it, by Anthropic's own definition. That is a useful thing to own and there is nothing wrong with buying one. It is just not an agent, and most demos badged agentic in 2026 sit at rung 2 or rung 3. Anthropic also stresses "transparency by explicitly showing the agent's planning steps," which is good engineering practice for a coding agent and a hard requirement at the investigation layer. The build-level version of all of this is in [AI agent architecture for claims investigation](/blog/ai-agent-architecture-claims-investigation/).

## What breaks when you aim an agent at a contested claim

Five things break, and they break for structural reasons rather than because the models are not good enough yet. This is why the investigation layer is thin, and it is also the specification for anything that intends to occupy it.

### 1. There is no ground truth at decision time

Anthropic's requirement that an agent gain ground truth from the environment at each step is satisfiable for a code test or a database read. It is not satisfiable for a flagged claim. A tool call on a contested claim returns evidence, not truth. Allianz's food spoilage claim under AUD $500 after a named storm has a checkable answer, which is why seven agents can close it in under five minutes. A suspected staged accident does not have one. The reason Nemo works is the same reason it cannot be extended by turning a dial: the ceiling and the weather trigger are what make the ground truth retrievable.

### 2. The input is adversarial

In an intake agent, the counterparty wants to be understood. In an investigation, some fraction of counterparties are constructing the record the agent is reading. Kayla McCallum, Associate Attorney at Swift Currie, argued in [Claims Journal](https://www.claimsjournal.com/news/national/2026/05/15/337321.htm) on May 15, 2026 that investigators can no longer start from the premise that most of a claim file is genuine, and gave the practical tell: "An AI-fabricated claim can sustain a general narrative, but it struggles with specificity." Every architectural assumption that holds when the input is cooperative inverts when it is not. An agent that trusts its inputs is not doing investigation; it is doing transcription.

### 3. Conflicting evidence has to be reconciled, not summarized

Summarization collapses a set of documents into their agreement. Investigation is interested in exactly the places where they disagree: the repair invoice date against the tow record, the recorded statement against the geolocation, the medical narrative against the treatment code. An agent optimised to produce a coherent summary is optimised against the signal. This is why Hesper runs 15+ investigation phases in parallel on every flagged claim: document forensics, OSINT, statement cross-reference, timeline reconstruction, and financial pattern analysis are not stages in a pipeline, they are independent readings of the same claim that only mean something when set against each other. The parallelism is an architectural requirement for reconciliation, not a throughput statistic.

### 4. Provenance has to survive a second reader

A conversation agent's output is consumed and discarded within the call. An investigation agent's output is read later by an SIU lead, a defense attorney, a state DOI examiner, and possibly a jury. Every assertion needs its source, its timestamp, and the reasoning that connected them, reconstructable months after the agent that produced it was retired or retrained. That is a different engineering problem from producing a good answer, and it cannot be retrofitted with logging after the fact. The full standard is set out in [the defensibility standard for fraud investigation AI](/blog/fraud-investigation-ai-defensibility-standard/).

### 5. The regulatory stake is asymmetric

The intake layer touches conduct and disclosure, which is real and is why IAG tells customers when they are speaking to an AI service. The investigation and disposition layer touches the Unfair Claims Settlement Practices Act. The layer with the highest statutory exposure has the least agentic capacity deployed against it. That inversion is the shape of the market in 2026, and it is worth reading the regulators' actual text to see how wide the gap is.

## What the regulators have actually written

The NAIC Model Bulletin on the Use of Artificial Intelligence Systems by Insurers was adopted by the Innovation, Cybersecurity, and Technology (H) Committee on December 1, 2023 and by Executive (EX) Committee and Plenary on December 4, 2023. Its most useful line for anyone deploying agents sits in Section 1, after the descriptions of the Unfair Trade Practices Act and the Unfair Claims Settlement Practices Act.

> Actions taken by Insurers in the state must not violate the UTPA or the UCSPA, regardless of the methods the Insurer used to determine or support its actions.
>
> - NAIC Model Bulletin: Use of Artificial Intelligence Systems by Insurers, Section 1, adopted December 4, 2023

That is method-neutral. An agent is held to the standard an adjuster is held to. Autonomy is not a defence and it is not a prohibition either. The bulletin then describes the model act itself: "Unfair Claims Settlement Practices Model Act (#900): The Unfair Claims Settlement Practices Act, [insert citation to state statute or regulation corresponding to Model #900] (UCSPA), sets forth standards for the investigation and disposition of claims arising under policies or certificates of insurance issued to residents of [insert state]." The model act's own subject is the investigation and disposition of claims. Intake, status updates, and summarization are not what it names.

Section 3 sets a proportionality test that scores autonomy directly, which is why a rung-3 conversation agent and a rung-5 investigation agent do not carry the same control burden.

> The controls and processes that an Insurer adopts and implements as part of its AIS Program should be reflective of, and commensurate with, the Insurer's own assessment of the degree and nature of risk posed to consumers by the AI Systems that it uses, considering: (i) the nature of the decisions being made, informed, or supported using the AI System; (ii) the type and Degree of Potential Harm to Consumers resulting from the use of AI Systems; (iii) the extent to which humans are involved in the final decision-making process; (iv) the transparency and explainability of outcomes to the impacted consumer; and (v) the extent and scope of the insurer's use or reliance on data, Predictive Models, and AI Systems from third parties.
>
> - NAIC Model Bulletin, Section 3, adopted December 4, 2023

Clause (iii) is the autonomy dial and clause (iv) is the audit trail, both written into the governance test in 2023. The bulletin also scopes the program across the lifecycle, saying the AIS Program "should address the use of AI Systems across the insurance life cycle, including areas such as product development and design, marketing, use, underwriting, rating and pricing, case management, claim administration and payment, and fraud detection." On third-party systems, Section 4.1 asks insurers to ensure that decisions "made or supported from such AI Systems that could lead to Adverse Consumer Outcomes will meet the legal standards imposed on the Insurer itself," and Section 4.2 contemplates contract terms that provide audit rights and require the vendor to cooperate with regulatory inquiries and investigations. Item 3.6 covers data and record retention. Section 4 also warns that in the context of an investigation or market conduct action an insurer "can expect to be asked about its development, deployment, and use of AI Systems, or any specific Predictive Model, AI System or application and its outcomes (including Adverse Consumer Outcomes) from the use of those AI Systems, as well as any other information or documentation deemed relevant by the Department."

On adoption: the NAIC Big Data and Artificial Intelligence (H) Working Group [implementation map](https://content.naic.org/sites/default/files/cmte-h-big-data-artificial-intelligence-wg-map-ai-model-bulletin.pdf), status as of April 1, 2026, shows 25 jurisdictions have adopted the model bulletin, 24 states plus the District of Columbia. Four more jurisdictions have insurance-specific AI regulation or guidance instead: California, Colorado, New York, and Texas.

Two of those four get cited in claims-AI conversations where they do not apply, so state the scope plainly. [New York DFS Insurance Circular Letter No. 7 (2024)](https://www.dfs.ny.gov/industry-guidance/circular-letters/cl2024-07), issued July 11, 2024, is scoped to underwriting and pricing, not claims investigation. Its own AIS definition is limited to systems used to supplement or proxy traditional underwriting or pricing. It is not claims-binding authority. It is worth reading anyway as the template regulators are converging on for documentation: "Insurers should maintain comprehensive documentation for their use of all AIS, including all ECDIS relied upon for such AIS, whether developed internally or supplied by third parties consistent with 11 NYCRR 243, and be prepared to make such documentation available to the Department upon request." On vendors, the circular states that insurers "retain responsibility for understanding any tools, EDCIS [sic], or AIS used in underwriting and pricing for insurance that were developed or deployed by third-party vendors and ensuring such tools, EDCIS [sic], or AIS comply with all applicable laws, rules, and regulations."

Colorado is the other one. 3 CCR 702-10 took effect November 13, 2023, with amendments effective October 15, 2025 that extend the external consumer data governance and risk-management framework beyond life insurers to private passenger automobile and health benefit plan insurers. Auto insurers using such data filed a progress narrative by December 1, 2025 and a compliance report on July 1, 2026, annually thereafter. The regime targets unfair discrimination in underwriting and pricing through quantitative testing of external consumer data, the quantitative testing standards for the newly covered lines remain under development at the Division, and it is not a claims-investigation regime. Cite it for what it is.

The 2026 activity that actually touches agentic decisioning is the [NAIC AI Systems Evaluation Tool multistate pilot](https://www.fenwick.com/insights/publications/naic-expands-ai-systems-evaluation-tool-pilot-program-to-12-states-key-updates-for-insurers-and-ai-vendors-supporting-insurers), running March 2, 2026 through September 2026 across California, Colorado, Connecticut, Florida, Iowa, Louisiana, Maryland, Pennsylvania, Rhode Island, Vermont, Virginia, and Wisconsin. Four exhibits cover AI usage quantification, governance risk assessment frameworks, details on high-risk AI systems, and AI data particulars, used inside market conduct exams and financial examinations across P&C, life, and health. The working group's stated position is that the tool "does not create new requirements for AI governance risk assessments," and that regulators will "prioritize examining high-risk AI systems that could cause serious consumer or financial issues, while paying less attention to low-risk back-office systems." Fenwick's note to the market is direct: "AI vendors working with insurance companies should also take note of ongoing developments here." Tool updates are expected September to October 2026, with adoption expected at the NAIC fall national meeting in November 2026.

The line to draw from all of it is proportionality. The compliance burden scales with the decision the agent makes. A voice agent that captures an FNOL is a low-risk system under this framing and should be treated as one. An agent that produces the finding a denial rests on is not, and a carrier deploying at rung 5 should expect its AIS Program documentation read closely and should require in the contract the audit rights the model bulletin's Section 4.2 contemplates. On top of that sit California's 10 CCR 2698.36 documented-decision requirement and the antifraud plan filings under NAIC Model Act #680, both of which assume a reconstructable record of how a determination was reached.

## Five objections, answered

### The conversation layer is where the volume is, so that is where the ROI is

Correct on cost-to-serve, and the evidence supports it. Sedgwick research reported by [Insurance Journal](https://www.insurancejournal.com/magazines/mag-features/2026/03/23/862425.htm) in March 2026 found intake automation cut average claim processing times from 10 days to 36 hours, and that 82% of carriers use AI for routine tasks such as data extraction and automated customer interactions. A [PwC survey of 136 insurance executives](https://www.insurancejournal.com/news/national/2025/11/14/847326.htm) reported in November 2025 found 57% listing generative and agentic AI among their top technology investment priorities for 2026, and 54% calling them the most transformative technology over the next three years, 22 points clear of the next option.

But volume-weighted cost is not loss-weighted cost. Roughly 10% of P&C claims involve fraud, the Coalition Against Insurance Fraud puts the annual US cost at $308.6 billion per [Triple-I](https://www.iii.org/fact-statistic/facts-and-statistics-insurance-fraud), and the same page notes that participants in the 2022 Insurer SIU Benchmarking Study saw SIU staff grow 1.4% from 2021 to 2022, down from previous growth of 2.5%. An investigator carries 200+ cases and completes roughly 10 investigations a month. The conversation layer optimises the cost of handling a claim. The investigation layer optimises what the claim pays out. Those are different lines on the P&L and both are real. For the stage-by-stage version of that map, see [what is automated in insurance claims in 2026](/blog/insurance-claims-automation-2026-whats-automated/).

| Insurer AI adoption versus AI maturity (Sedgwick research via Insurance Journal, March 2026) | Value | Share |
| --- | --- | --- |
| Use AI for routine tasks such as data extraction and automated customer interactions | 82% | 82% |
| Claims professionals who believe AI needs human oversight | 75% | 75% |
| Say they have fully mature AI capabilities | 12% | 12% |
| Say they have achieved scalable AI success | 7% | 7% |

### Allianz already runs seven agents end to end, so investigation is just next

Project Nemo is the most substantial shipped example in this survey and it includes a fraud agent, so the objection is a serious one. Look again at the scope Allianz chose: food spoilage, under AUD $500, triggered by a severe weather event, on home contents, with payout decisions never automated by design. Every constraint exists because the claim has a checkable answer. Contested claims do not become tractable by adding agents. They become tractable by adding evidence and a method for reconciling it when the evidence disagrees. Nemo is a correctly-scoped rung 4. Rung 5 is a different problem, not a bigger one.

### Regulators will not permit autonomous investigation

Nothing in the NAIC framework prohibits it. The model bulletin is method-neutral by its own text. What the framework requires is documentation, record retention, third-party audit rights, and the ability to answer a market conduct examiner's questions about a specific AI system and its outcomes. That is a bar, not a wall, and it favours systems built with provenance from the first line of code over systems retrofitted with logging. State the converse just as plainly: a carrier that cannot reconstruct an agent's reasoning for a DOI examiner has a problem its vendor cannot fix after the fact.

### We will build it on Guidewire's Agentic Framework

Qusar genuinely supplies the runtime, per-task model choice, and secure real-time access to policy, claims, and billing data. That is the hard infrastructure and it is now solved, which is a real change from two years ago. What it does not supply is the investigative methodology: which phases to run on a suspected staged accident, how to weight a contradicted recorded statement, what an evidence chain has to contain to survive an examination under oath. Build-versus-buy economics are a separate analysis and this post is not it.

### An agent that talks to claimants is riskier than one that reads a file

Partly right, and it explains the sequencing well. Voice agents are conduct-visible, and [Insurance Business Australia](https://www.insurancebusinessmag.com/au/news/technology/iag-bets-on-agentic-ai-where-conduct-risk-is-highest-583756.aspx) reported ASIC flagging agentic AI as a potential source of consumer harm and APRA noting that governance has not matured at the same pace as deployment. Carriers reasonably started where the failure mode is a bad customer experience rather than a bad coverage decision. But the regulatory stake runs the other way. The Unfair Claims Settlement Practices Act sets standards for the investigation and disposition of claims. The layer with the highest statutory exposure has the least agentic capacity deployed against it, and that is a description of the market, not a criticism of any carrier in it.

Which is the whole argument, compressed. Detection is upstream; investigation is downstream. Detection vendors flag, claims systems route, conversation agents intake, and every one of those is a solved or solving problem with named vendors and shipped products. The flagged claim still waits 14+ days for a human, and roughly 25% of flagged claims get a full investigation at all. Hesper takes the flagged claim and runs the full playbook, 15+ phases in parallel, in hours rather than weeks, with the evidence chain assembled as it goes rather than reconstructed afterwards. Built-in detection means it runs standalone; FRISS, Shift Technology, and Verisk remain complementary rather than replaced. From fraud detection to fraud resolution. The investigator's role shifts from execution to decision-making.

## Key takeaways

- Of the nine named agentic claims deployments catalogued here from 2025 and 2026, eight sit in conversation, summarization, estimating, or data-access layers, and that count is our own catalogue rather than a market share.
- Thomson Reuters CLEAR Investigate, launched in early 2026, is the one genuine investigation-layer entrant, and it is a horizontal investigator-facing research agent that describes itself as a support tool.
- Allianz Project Nemo shows what autonomy requires rather than where it goes next: seven agents, food spoilage claims under AUD $500, a weather trigger that can be externally checked, and payout decisions never automated by design.
- Four properties separate an agent from a workflow with a language model in it, per Anthropic's definition: model-selected tool use, planning, memory across steps, and a feedback loop, and most 2026 claims demos sit at rung 2 or rung 3.
- The NAIC model bulletin is method-neutral and scales control burden with the decision the agent makes, which means the layer with the highest statutory exposure under the UCSPA is also the layer with the least agentic capacity deployed against it.

## Frequently asked questions

### What is agentic AI in insurance claims?

Agentic AI in claims means software that selects its own tools and sequences them toward a goal, rather than running a fixed script. Anthropic's engineering definition is the cleanest available: workflows orchestrate models and tools through predefined code paths, while agents dynamically direct their own processes and tool usage. Four properties follow from that: the model itself picks which tools to call, it plans by decomposing a goal into steps it chose, it carries memory across those steps, and it runs in a loop against environmental feedback rather than one forward pass. In claims specifically, the important follow-up question is which layer the agent operates in, because an intake agent and an investigation agent share the word and share almost nothing else.

### What is the difference between agentic AI and a chatbot in claims?

A chatbot answers within a script; an agent selects tools and sequences them toward a goal. The boundary case is worth knowing because it is genuinely agentic. Guidewire's Agentic First Notice of Loss, shipped in the Qusar release on August 3, 2026, is described as guiding claimants through the first notice of loss using conversational AI voice and capturing key claim details. That is a voice interface, and it is still an agent in the sense that it manages a bounded dialogue to completion, decides what to ask next, and escalates when the case leaves its remit. What separates it from a chatbot is the decision-making inside the dialogue. What separates it from an investigation agent is that its output is a captured record, not a determination.

### Which insurers are using agentic AI for claims in 2026?

Several, at different layers and different stages. IAG confirmed its OpenAI Presence partnership on July 27, 2026, targeting delivery in the first half of its 2027 financial year, which runs July to December 2026; iTnews reported the work has so far centred on solution design rather than live deployment. Allianz has run Project Nemo since July 2025, with seven agents on food spoilage claims under AUD $500. AXA Switzerland is named as an early adopter of Shift Claims, launched September 2025. On the platform and services side, Sedgwick announced Omni in May 2026, Guidewire shipped its Agentic Framework in August 2026, and Verisk shipped Model Context Protocol connectors into Anthropic's Claude in May 2026.

### Can AI agents settle insurance claims without a human?

On narrow claim types, yes, and Allianz does it. Project Nemo runs seven agents across triage, coverage verification, fraud screening, and settlement preparation for food spoilage claims under AUD $500 caused by severe weather, completing the workflow in less than five minutes from filing to human-review-ready. Allianz's Chief Transformation Officer at Allianz Services, Maria Janssen, states the limit directly: "By design, payout decisions are never automated." What makes a claim type eligible is not the technology. It is that the claim has a checkable answer, a verifiable trigger such as a named weather event, and a dollar ceiling that caps the downside of being wrong. Remove any one of those and the design stops working.

### Is agentic AI allowed under US insurance regulations?

Yes, with governance. The NAIC Model Bulletin on the Use of Artificial Intelligence Systems by Insurers, adopted December 4, 2023, is method-neutral: actions taken by insurers must not violate the Unfair Trade Practices Act or the Unfair Claims Settlement Practices Act regardless of the methods used to determine or support them. As of April 1, 2026, 25 jurisdictions had adopted the bulletin and four more had their own insurance-specific AI regulation or guidance. The NAIC AI Systems Evaluation Tool pilot runs March 2 through September 2026 across 12 states, with regulators prioritising high-risk AI systems over low-risk back-office ones. Control burden scales with the decision the agent makes.

### What are the levels of AI autonomy in insurance claims?

The one published industry framework is Shift Technology's ARISE, released in June 2026, with five levels: Answers, Recommends, Initiates, Solves, and Exceeds. It runs from an agent answering handler questions in a conversational interface up to one that deviates from standard operating procedures to optimise the claim outcome. Hesper extends it with a six-rung model, labelled as our framing rather than an industry standard: scripted bot, retrieval assistant, single-tool agent, conversational task agent, multi-tool planner on bounded claims, and autonomous investigator. The distinction that matters is that the sixth rung is not a higher setting on the ARISE axis, it is a second axis. Our full response is in the ARISE rebuttal post.

### Can agentic AI investigate insurance fraud, or only detect it?

Both exist, and they are different products. Detection scores a claim and hands off; investigation resolves it. Thomson Reuters CLEAR Investigate, described publicly in April 2026, is a genuine investigation-layer agent that interprets an investigator's input, builds a plan, and queries premium records alongside curated open-web sources, and it names Special Investigation Units in its audience. It is horizontal across KYC, AML, SIU, and third-party risk, and Thomson Reuters describes it as a support tool. Hesper is the P&C-claim-native version: it takes a flagged claim and returns a resolved, audit-ready finding. Detection is upstream; investigation is downstream, and FRISS, Shift Technology, and Verisk remain complementary rather than replaced.

### Why is agentic AI easier at FNOL than at claims investigation?

Four reasons, and none of them is model quality. Ground truth is available at intake and not at investigation, because a tool call on a contested claim returns evidence rather than truth. The counterparty is cooperative at intake and sometimes adversarial at investigation, where part of the record may have been constructed to pass. Intake output is consumed within the call and discarded, while an investigation finding is read later by an SIU lead, a defense attorney, and possibly a state DOI examiner, so every assertion needs its source and timestamp. And the Unfair Claims Settlement Practices Model Act sets standards for the investigation and disposition of claims, not for intake.
