---
title: "NAIC AI risk evaluation v5.0: your fraud model now needs a page number"
description: "NAIC did not write a fraud rule. Exhibit A already had a Fraud/Waste & Abuse row. What v5.0 changes is that regulators now ask for the model by name, with a document name and page number attached."
date: "2026-09-14"
lastModified: "2026-09-14"
author: "Pankaj Dhariwal"
tags: ["Guides"]
canonical: "https://gethesperai.com/blog/naic-ai-risk-evaluation-supplement-claims/"
---

# NAIC AI risk evaluation v5.0: your fraud model now needs a page number

> **TL;DR** NAIC exposed version 5.0 of its AI Risk Evaluation Supplement on August 31, 2026. Comments close September 29, and adoption is considered in November. The instrument already carried a Fraud/Waste and Abuse row, so fraud models were never out of scope. What changed is that regulators now ask for a model inventory by name, with a document name and page number behind every answer.
>
> - v5.0 comment window closes close of business September 29, 2026
> - 12 states piloted the instrument from March to September 2026
> - Exhibit B now wants the document name and the page number

- **Sept 29** - Comment deadline on AI Risk Evaluation Supplement v5.0 (Close of business, per NAIC's exposure notice)
- **12** - States that piloted the instrument, March to September 2026 (NAIC Pilot Project Summary)
- **29** - US jurisdictions with an insurer-AI instrument in force (25 Model Bulletin adopters plus CA, CO, NY, TX (NAIC map, April 1, 2026))
- **$308.6B** - Lost to US insurance fraud every year (Coalition Against Insurance Fraud)

Three lines in an NAIC staff memo changed what a carrier has to be able to prove about its fraud model. On August 31, 2026, the NAIC Big Data and Artificial Intelligence (H) Working Group exposed version 5.0 of the AI Risk Evaluation Supplement, an instrument state regulators use to ask an insurer what AI it runs and what documents back the answers. Regulators now ask for a model inventory explicitly, and the Exhibit B checklist asks for the document name and the page number behind every governance answer. That reads like a compliance burden. It is the same artifact that makes a fraud investigation defensible, which is why a carrier that cannot produce it for an examiner cannot produce it for a bad-faith plaintiff either.

Read it this month, not in November. The comment window closes at the close of business on Tuesday, September 29, 2026, and NAIC's own pilot timeline puts adoption consideration at the Fall National Meeting, which runs November 14 to 17, 2026 in Dallas. That is roughly ten weeks from exposure to a vote, on an instrument twelve states have already used on live companies.

NAIC did not write a fraud rule, and this post will not pretend otherwise. The Supplement is general insurer-AI governance. It also already carried a line item called Fraud/Waste and Abuse in Exhibit A and named fraud detection as a claims AI use case in its definitions, so the fraud model was never outside the scope. What version 5.0 changes is that regulators now ask for it by name, in an inventory, with a citation attached.

What follows is the news spine, an exhibit-by-exhibit map of where a claims fraud model lands, the four-step inference chain from the new agentic AI definition to your Exhibit A row, and the one omission worth stating plainly. The argument underneath is the same one in our post on [AI fraud investigation and court admissibility](/blog/ai-fraud-investigation-court-admissibility/): the record that satisfies an examiner is the record that survives a deposition, and most carriers have neither.

## NAIC exposed version 5.0 on August 31, and comments close September 29

The NAIC AI Risk Evaluation Supplement is the instrument state insurance regulators use to ask an insurer what AI it runs, how it governs those systems, and what documents back the answers. NAIC exposed version 5.0 on August 31, 2026, for a 30-day comment period that closes at the close of business on September 29, 2026.

The notice sits on the Working Group page and is unambiguous about what is wanted and by when. [NAIC's exposure notice](https://content.naic.org/committees/h/big-data-artificial-intelligence-wg) reads: "Following up from the August 31, 2026, Big Data and Artificial Intelligence (H) Working Group meeting, we are exposing the AI Risk Evaluation Supplement version 5.0 for written comments and feedback, for a 30-day comment period ending Tuesday, September 29, 2026." Comments go to Scott Sobel or Miguel Romero at NAIC. The exposure package carries the version 5.0 document and a two-page staff Summary of Changes against version 4.0.

The name is new too. The instrument spent its earlier life as the AI Systems Evaluation Tool, and version 5.0 is the first published as the AI Risk Evaluation Supplement. NAIC gives its own reason in the [Summary of Changes](https://content.naic.org/sites/default/files/inline-files/summary-of-changes.pdf): feedback in prior comment periods concerned how broadly the document would be used, so "additional text was added to emphasize the document as a supplement." The practical reading is that it supplements the Market Regulation Handbook, the Financial Condition Examiners Handbook and the Financial Analysis Handbook rather than replacing any of them, and that the question of who receives an inquiry under it gets answered by those handbooks rather than by the Supplement.

One artifact of how recent the rename is: trade coverage published on September 11, 2026 still called it the AI Systems Evaluation Tool, and NAIC's own pilot documentation still uses the old name because those documents predate the change. That is not an error worth dwelling on. It is a useful signal for anyone searching an internal drive for prior analysis, because both names refer to the same instrument.

> **Two names, one instrument, and where each fact in this post comes from**
>
> On first mention, write it as the AI Risk Evaluation Supplement (formerly the AI Systems Evaluation Tool). Structural details in this post that describe rows, fields and question numbers are drawn from the published version 4.0 document, because version 5.0 was posted only as a .docx and its text is not extractable. Every version 5.0 change claim traces to NAIC's own Summary of Changes, which is a primary NAIC staff document. No definition wording from version 5.0 is quoted anywhere in this post, because that wording has not been verified against a readable source.

## The pilot behind version 5.0 ran in twelve states for six months

The pilot behind version 5.0 ran in twelve states from March 2026 to September 2026. NAIC's pilot project summary says participating regulators used the Tool in market conduct exams and reviews, financial analysis, and financial exams, across property and casualty, life, and health insurers, focusing on their domestic companies.

The twelve are California, Colorado, Connecticut, Florida, Iowa, Louisiana, Maryland, Pennsylvania, Rhode Island, Vermont, Virginia and Wisconsin. [NAIC's pilot project summary](https://content.naic.org/sites/default/files/inline-files/Pilot%20Project%20Summary_1.pdf) also describes how they were told to work: focus on domestic insurers, apply proportionality by spending more review time on high-risk AI systems that could cause serious consumer or financial harm and less on low-risk back-office systems, and modify the Tool where a jurisdiction needs to. The pilot was therefore not uniform across all twelve, which matters when a multi-state carrier compares the questions it got in one state against another.

Ten of the twelve pilot states already had a state-level insurer-AI instrument before the pilot began. That figure is a derivation rather than an NAIC statistic, so treat it as one: cross-reference the pilot list against [NAIC's Model Bulletin implementation map as of April 1, 2026](https://content.naic.org/sites/default/files/cmte-h-big-data-artificial-intelligence-wg-map-ai-model-bulletin.pdf) and eight of the pilot states appear as Model Bulletin adopters, while California and Colorado run their own insurance-specific AI regulation or guidance. Florida and Louisiana appear on neither list. Nationally the same map counts 25 adopted jurisdictions plus four insurance-specific ones, which is 29 with something in force.

The sequencing is the point. Version 5.0 is not the first draft of an idea. It is a redraft published after twelve regulators spent six months using the previous version on real companies. Read the Summary of Changes with that in mind and it reads like field notes: the requests examiners kept having to make verbally are now printed in the document.

## Where the Supplement already names fraud, and where it says nothing at all

The Supplement is general insurer-AI governance rather than a fraud rule, and it already names fraud twice. The published version 4.0 instrument carries an Exhibit A operational-area row labeled Fraud/Waste & Abuse, separate from its Claims/Adjudication row, and its Operations definitions name fraud detection as an example of a claims AI use case.

Now the omission, stated plainly, because it is the part most internal summaries will get wrong. Read all eighteen pages of the published instrument and the term SIU does not appear. Neither does special investigation, special investigations unit, or any equivalent. The Supplement governs models and the program that governs models. It says nothing about how an investigative workflow should run once a flag is raised. The relevance to an investigations unit is real, but it is inferential, and any memo that goes to a domestic regulator should label it that way.

The structural details are worth having in front of you before the first exhibit lands. Exhibit A lists 14 named operational areas plus an Other row, with Claims/Adjudication and Fraud/Waste & Abuse as separate lines and a footnote attaching salvage and subrogation to the claims row. Exhibit C collects 13 model-level fields on high-risk models, one of which asks how the model is reviewed for compliance with unfair claims settlement laws. Exhibit D covers 25 data element types. The instructions themselves cite unfair claims settlement practices among the reasons the inquiry exists at all.

| Exhibit | What it asks for | What changed in v5.0 | What a carrier must produce for a fraud model |
| --- | --- | --- | --- |
| A - Quantify use of AI Systems | Counts of AI models by operational area, including rows for Claims/Adjudication and Fraud/Waste & Abuse | Columns revised for model type; counts requested separately for Direct Consumer Impact and Material Financial Impact models; Model Inventory now asked for explicitly | A named, counted inventory entry for every fraud and claims model, in-house and vendor |
| B - AIS Program governance | How the governance framework addresses unfair trade practices, adverse consumer outcomes, ERM, ORSA, vendor procurement and consumer complaints | Renamed to focus on the AIS Program; new 3o on materiality, new 3p on third-party model oversight, new 3d on explainability and transparency in the narrative version; checklist now asks for document name and page # | The governing document, by name, and the page where the fraud model's controls are actually written down |
| C - High-risk model detail | 13 model-level fields: name and version, model type, implementation date, internal vs third party with vendor name, risk classification, risks and limitations, AI type (automate, augment, support), drift and output testing, last test date, use case and purpose, financial-statement effect, legal review, regulatory actions | Field 2 moved up to use case and purpose; field 3 gained model-type examples; field 4 expanded on how the model was developed; risks and limitations split into two fields with a high, moderate and low scale | A per-model dossier, including whether the fraud model automates a decision or augments a human |
| D - Data details | 25 data element types with type of AI system, how the data is used, internal source, and third-party vendor name | New field linking data sets to the specific AI Models that rely on them; regulators may request a data dictionary; regulators may customize by line of business | Named data sources behind the fraud score, including contributory and third-party data the carrier does not own |

Two rows, two answers, and in most carriers two different vendors. A company that buys a detection score and then works the flagged queue with a human team has one model in the Fraud/Waste and Abuse row and a manual process sitting behind it. The manual process is not a model and does not go in the inventory, which is precisely why the coverage question this post ends on never shows up in Exhibit A at all.

## What putting agentic AI in the definitions actually does

Version 5.0 adds a definition for agentic AI. NAIC's Summary of Changes lists it alongside added or revised definitions for AI Model, AIS Program, Direct Consumer Impact, Material Financial Impact, generalized linear models and third parties, and it removes the prior Degree of Potential Harm to Consumers definition.

Adding a definition to a document that is ten weeks from an adoption vote is a decision about scope durability. The Summary of Changes also notes where the borrowed language comes from: the AI Model definition draws on NIST and White House Executive Order wording, the AIS Program definition comes from the NAIC Model Bulletin on the Use of Artificial Intelligence Systems by Insurers, and the third-party language comes from the Third Party Registration Framework. Borrowing definitions is how a regulator avoids re-litigating scope every time the technology moves.

NAIC has not published extractable text for its own agentic AI definition, so this post does not quote one. For the general meaning, [NIST describes agentic AI](https://www.nist.gov/agentic-ai) as artificial intelligence systems that function as autonomous agents capable of independently making decisions, learning from interactions, and adapting to changing environments. That is NIST's wording, not NAIC's. It is useful here only because NAIC says its AI Model definition leans on NIST and Executive Order language, which makes NIST the nearest available reference point.

### The inference chain from a definition to your Exhibit A row

Four steps, stated as a chain so a compliance officer can mark exactly where the reasoning stops being NAIC's and starts being ours.

1. An autonomous claims-investigation agent takes multi-step action without a human at each step, which puts it inside the new agentic AI definition and inside the AI Model definition drawn from NIST and Executive Order language. If it is an AI Model, it belongs in the Model Inventory that Exhibit A now asks for explicitly.
2. A fraud finding that supports a denial is a decision that adversely affects a consumer, which makes it a Direct Consumer Impact model, the category Exhibit A now counts separately from Material Financial Impact. A fraud model that also feeds reserving sits in both columns.
3. The Exhibit B checklist asks for the document name and page number supporting each governance answer. That is an evidentiary standard rather than a narrative one. For a vendor-supplied score with no per-claim record on the carrier's side, there is frequently no document and no page to name.
4. Nothing in the Supplement tells an investigations unit how to investigate. The obligations attach to the model and to the program that governs it. What reaches the investigation is the documentation burden, not a workflow rule.

Steps one and two are readings of published NAIC structure. Step three is verbatim from the Summary of Changes. Step four is the boundary. Keeping those three categories separate is the difference between a memo a regulator finds credible and one that overstates what NAIC said.

This is also where the Exhibit C field asking whether a model automates, augments or supports a decision stops being a formality. An architecture that runs 15+ investigation phases in parallel on a flagged claim, including document forensics, OSINT, statement cross-reference, timeline reconstruction and financial pattern analysis, has a different answer to that field than a score that ranks a queue, and a different answer again from a handler-assist tool that drafts next steps for an adjuster. Exhibit C also asks for model risks and limitations. A phase-decomposed system can answer that at phase granularity rather than as one opaque paragraph about a number.

## The model inventory is the part carriers will feel

A model inventory is a register of every model an organization runs, with enough metadata for a reviewer to locate, classify and evaluate each one. Version 5.0 promotes the request from implied to explicit, which converts a narrative exam answer into an asset register a company has to maintain between exams.

The sentence itself is short. NAIC's Summary of Changes says: "There were several implied requests in the prior version of the Supplement. These are now explicitly listed. Most notably, regulators now ask for a Model Inventory." The published version 4.0 had already gestured at it, noting in Exhibit C that regulators may request a company's risk assessment and a model inventory if that information has not otherwise been provided. Version 5.0 moves the request forward into Exhibit A, which is the exhibit a company fills in first and the one that sets how far the rest of the inquiry goes.

*Figure: Exhibit A promotes the model inventory from implied to explicit. The detection score, the claims platform ML and the investigation agent each become their own row, with an owner, a vendor and a citation. A score is a number, and a number has no page.*

Banking has had this artifact for fifteen years, which is the comparison worth using with anyone inside a carrier who has to justify the work. On April 4, 2011 the Federal Reserve and the OCC issued [SR 11-7, Guidance on Model Risk Management](https://www.federalreserve.gov/boarddocs/srletters/2011/sr1107.pdf). It told banking organizations to "maintain an inventory of models implemented for use, under development for implementation, or recently retired." It said documentation should be "sufficiently detailed to allow parties unfamiliar with a model to understand how the model operates, as well as its limitations and key assumptions." And it said validation "applies equally to models developed in-house and to those purchased from or developed by vendors or consultants." Insurance is arriving at the same three requirements through a supplement to an exam handbook rather than a supervisory letter.

|  | Federal Reserve and OCC, SR 11-7 (April 4, 2011) | NAIC AI Risk Evaluation Supplement v5.0 (exposed August 31, 2026) |
| --- | --- | --- |
| Instrument type | Joint supervisory guidance | Supplement to existing exam handbooks, not a standalone framework |
| Inventory requirement | Organizations should maintain an inventory of models implemented for use, under development for implementation, or recently retired | Regulators now ask for a Model Inventory, made explicit in Exhibit A in v5.0 |
| Vendor models | Validation applies equally to models developed in-house and to those purchased from or developed by vendors or consultants | Exhibit C field 4 names the vendor; new Exhibit B 3p asks how the company oversees third-party models |
| Documentation standard | Documentation sufficiently detailed to allow parties unfamiliar with a model to understand how it operates, its limitations and key assumptions | Exhibit B checklist asks for the document name and page number supporting each answer |
| Materiality | Materiality is an important consideration in model risk management | New materiality definition aligned to the Financial Condition Examiners Handbook; the Exhibit A threshold is set and then disclosed |
| Who it binds | Banking organizations supervised by the Federal Reserve | Insurers, through whichever states adopt it; twelve states have already piloted it |

Exhibit A's split between Direct Consumer Impact and Material Financial Impact is the classification most carriers will get wrong on the first pass, because the categories are not exclusive. A model whose output augments or automates a decision affecting a consumer, such as a denial or a coverage determination, has direct consumer impact. A model that moves numbers material to the company's financial position, such as reserving or capital allocation, has material financial impact. A fraud model that drives denials and also feeds reserve setting belongs in both columns. NAIC's stated reason for the split is to let a regulator see whether a company's AI use is concentrated in one type, which can then limit how far the inquiry runs.

The second classification problem is that detection and investigation are different layers and get answered by different documents. A detection score that flags a claim and a system that works the flag are two inventory rows with two owners, two vendors and two sets of validation evidence. Our [guide to insurance fraud detection methods, tools and gaps](/blog/insurance-fraud-detection-pillar/) lays out that stack layer by layer. Hesper AI sits downstream of detection and also carries its own detection layer, which has one practical consequence for Exhibit A: the row can name a single model family with a single evidence trail rather than an undocumented chain of vendor score, then adjuster judgment, then a case note nobody can cite by page.

## Document name and page number is an evidentiary standard, not a documentation request

The Exhibit B checklist now asks companies to provide a document name and page number when they respond to inquiries. NAIC's stated reason is that it lets regulators connect company responses to the documents provided. Version 4.0 already carried a page-number column, so the delta in version 5.0 is the document name plus two new questions.

NAIC states the reason without decoration. The Summary of Changes says the checklist "added request to provide document name and page # when companies are responding to inquiries. This will allow regulators to connect the company responses to documents provided." Precision matters on what is new: version 4.0 already had a Page # column against governance questions 3a through 3n. Version 5.0 adds the document name alongside it, adds 3o on materiality and 3p on how a company oversees third-party models, and adds a new 3d on explainability and transparency to the narrative version. Fourteen governance sub-questions become sixteen, which is a derived count from the published v4.0 checklist plus the two additions NAIC lists.

Now take that standard to a fraud model. An examiner asks how the carrier validates the system that decides which claims get worked. The honest answer in most carriers is that the validation lives with the vendor, the governance language lives in a policy written for a different purpose, and the per-claim record lives partly in the claims system and partly nowhere. None of that is a scandal. It is the structural consequence of buying a score rather than an investigation. A score is a number, and a number has no page.

> The document name and the page number are not a compliance deliverable. They are the byproduct of a system that writes down what it did while it was doing it.
>
> - Hesper AI product research

### What the detection layer does and does not hand the carrier

Take the vendors one at a time, without snark, because each is doing the job it was built for. [FRISS](https://www.friss.com/) scores claims at FNOL and through the lifecycle. Under Exhibit A that score is a third-party AI model with an arguable direct consumer impact when it routes a claim toward denial. Under Exhibit C field 4 the carrier names the vendor. Under the new Exhibit B 3p the carrier describes how it oversees that third-party model. Note who owns every one of those answers: the carrier, not the vendor.

[Shift Technology](https://www.shift-technology.com/) ships detection plus handler-assist agentic AI for adjusters, which makes it the most visible product in the market inside the category NAIC just defined. That is worth saying plainly and it is not a criticism. Shift routes to a human handler, so the carrier's Exhibit C answer to the automate, augment or support field is augment. An autonomous investigation agent is a different answer to the same field. What the exhibit does is force carriers to write that distinction down, and most have never had to.

[Verisk ClaimDirector](https://www.verisk.com/products/claim-scoring/) scores on ISO ClaimSearch contributory data, which creates an unusual Exhibit D problem: the exhibit asks for the third-party data source or vendor name behind each of 25 data element types, and cross-carrier contributory data is not the carrier's to document. One more sentence for the claims platforms, because it gets missed. Embedded machine learning inside [Guidewire ClaimCenter](https://www.guidewire.com/products/claimcenter) or [Duck Creek Claims](https://www.duckcreek.com/products/claims/) is still an AI model in the inventory even when nobody inside the carrier thinks of it as an AI project.

Complementary to FRISS, Shift Technology, and Verisk - not a replacement. Those are detection-layer answers to Exhibit A. The investigation layer is a separate row, and the incumbent there is a manual team whose evidence trail is a case note. Hesper logs every one of its 15+ phases with sources, reasoning and timestamps, which is what from fraud detection to fraud resolution means in an exam context: the document and the page exist because the system wrote them while it ran, not because a compliance project reconstructed them afterwards. Hesper holds SOC 2 Type I, which is the kind of evidence Exhibit B's question on standards for procuring and engaging AI vendors is reaching for. On the practical question of who inside a carrier fills out Exhibit B and what they need from a vendor before they can, our [compliance officer's guide to AI claims investigation deployment](/blog/compliance-officer-ai-investigation-deployment/) is the working document.

## A regulator asked about approved claims, not denied ones

Regulators can look at claims a carrier approved, not only claims it denied. That reframes AI claims oversight. Most carrier attention points at wrongful denial, because that is where consumer harm and market conduct exposure are obvious. Examining approved claims looks for the opposite failure: claims that were flagged and paid anyway.

The sharpest line in the trade coverage is not about technology at all. J.P. Wieske, co-founder of the American InsurTech Council and a former deputy insurance commissioner of Wisconsin, described what came up in the pilot discussions [to Digital Insurance on September 11, 2026](https://www.dig-in.com/news/naic-to-make-insurers-show-ai-use-in-claims-and-models).

> One of the regulators brought up that they like to also take a look at approved claims, not just denied claims, and how those are looking, because that can certainly impact their solvency.
>
> - J.P. Wieske, co-founder, American InsurTech Council, to Digital Insurance, September 11, 2026

Read what that implies for an exam. An examiner pulling approved claims is not looking for wrongful denial. They are looking for claims that were flagged and paid anyway, and for whether anything was documented in between. This is the leakage argument the claims side has been making for years, arriving from the regulator's chair instead of the loss-cost line, and with solvency rather than consumer harm as the stated motive.

The arithmetic behind that gap is not mysterious. Rules-based detection produces false positive rates in the 60-85% range, so the flagged queue is large by construction. A manual investigator carries 200+ cases and completes roughly 10 investigations per month, because manual investigation runs 14+ days per case. The result across US property and casualty carriers is that roughly 25% of flagged claims get a full documented investigation. The other three quarters are paid, denied without full work, or queued until they age out. For scale, the [Coalition Against Insurance Fraud](https://insurancefraud.org/fraud-stats/) puts US insurance fraud at $308.6 billion a year and finds fraud in about 10% of property and casualty losses.

| Partial line-level breakdown of US insurance fraud (CAIF 2022 figures via the Insurance Information Institute). These four lines do not sum to the $308.6B total. | Value | Share |
| --- | --- | --- |
| Life insurance | $74.7B | 100% |
| Property and casualty | $45B | 60% |
| Workers compensation | $34B | 46% |
| Auto theft | $7.4B | 10% |

Those line-level figures come from the [Insurance Information Institute](https://www.iii.org/article/background-on-insurance-fraud), citing the same Coalition Against Insurance Fraud study. They are a partial breakdown, not a full decomposition, which is worth saying before anyone puts them in a board deck.

An investigation layer that works every flagged claim in hours, not weeks, changes the answer available at exam time from roughly 25% of flagged claims to 100% of them, and it changes what a documented decision looks like on the ones that get paid. That is the part worth carrying into a comment letter before September 29. If an examiner pulls an approved claim and asks what the carrier did with the flag, we did not have capacity is an answer with a document name and a page number attached. It is just the wrong page.

## Key takeaways

## Key takeaways

- NAIC exposed version 5.0 of the AI Risk Evaluation Supplement on August 31, 2026, the comment window closes at the close of business on September 29, 2026, and NAIC's pilot timeline puts adoption consideration at the Fall National Meeting in Dallas from November 14 to 17, 2026.
- The Supplement is general insurer-AI governance rather than a fraud rule, but the published instrument already carries an Exhibit A operational-area row labeled Fraud/Waste & Abuse and names fraud detection as a claims AI use case, so a fraud model was never outside its scope.
- The instrument contains no SIU or special-investigation language anywhere in its eighteen published pages, which means the relevance to an investigations unit is inferential and should be labeled that way in anything that goes to a domestic regulator.
- Version 5.0 makes the model inventory request explicit and adds the document name to the existing page-number request in the Exhibit B checklist, which turns a narrative exam answer into an evidentiary one and puts insurance on the same artifact banking has kept since SR 11-7 in April 2011.
- The phase-by-phase evidence trail that answers Exhibit B and Exhibit C is the same record that makes an investigation defensible later, which is why the requirement is better treated as an asset than as a cost.

## Frequently asked questions

### What is the NAIC AI Risk Evaluation Supplement?

It is the instrument state insurance regulators use to ask an insurer how it uses artificial intelligence, how it governs those systems, and what documentation backs the answers. NAIC exposed version 5.0 on August 31, 2026, with a 30-day comment window closing at the close of business on September 29, 2026. It is organized into four exhibits: Exhibit A quantifies how many AI models a company runs and in which operational areas, Exhibit B covers the AI Systems governance program, Exhibit C collects model-level detail on high-risk systems, and Exhibit D covers the data behind those models. NAIC describes it as a supplement to existing market conduct, financial analysis and financial examination review procedures rather than a new standalone framework.

### Why did NAIC rename the AI Systems Evaluation Tool?

Because stakeholders kept asking how broadly the document would be used. NAIC's own Summary of Changes says feedback during prior comment periods concerned the scope of use, so additional text was added to emphasize the document as a supplement. The rename from AI Systems Evaluation Tool to AI Risk Evaluation Supplement carries that message in the title. Practically, it means decisions about who receives inquiries under the document are expected to run through existing handbook guidance, meaning the Market Regulation Handbook, the Financial Condition Examiners Handbook and the Financial Analysis Handbook, rather than through the Supplement itself. It supplements exam procedures; it does not replace them. Trade coverage published as recently as September 11, 2026 still used the old name.

### When does the NAIC AI Risk Evaluation Supplement take effect?

It has not been adopted yet. The comment window on version 5.0 closes at the close of business on Tuesday, September 29, 2026. NAIC's pilot project timeline says November is when regulators consider the updated Tool for adoption at the Fall National Meeting, and the 2026 Fall National Meeting runs November 14 to 17 in Dallas, Texas. Adoption at the NAIC level does not by itself bind any insurer; individual states decide whether and how to use it in their exams. Twelve states already used it during a pilot that ran from March 2026 through September 2026, so in practical terms some carriers have been answering these exhibits for six months before any formal adoption vote.

### Which states are in the NAIC AI evaluation tool pilot?

Twelve: California, Colorado, Connecticut, Florida, Iowa, Louisiana, Maryland, Pennsylvania, Rhode Island, Vermont, Virginia and Wisconsin. NAIC's pilot project summary says those states used the Tool from March 2026 to September 2026 across market conduct exams and reviews, financial analysis, and financial exams, with participating insurers drawn from property and casualty, life, and health. States focused on their domestic insurers and applied proportionality, spending more time on high-risk AI systems that could cause serious consumer or financial harm and less on low-risk back-office systems. Each jurisdiction retained authority to modify the Tool for its own needs, so the pilot was not uniform across all twelve.

### Does the NAIC AI Risk Evaluation Supplement apply to fraud detection models?

Yes, though the Supplement is general AI governance rather than a fraud rule. The published instrument's Exhibit A includes an operational-area row labeled Fraud/Waste & Abuse alongside a separate row for Claims/Adjudication, and the definitions section names fraud detection as an example of a claims AI use case. What the document does not contain anywhere is SIU-specific language; the term does not appear in its eighteen published pages. So a carrier's fraud scoring model is squarely inside the inventory, but nothing in the Supplement dictates how a special investigations unit should operate. The obligations attach to the model and its governance, not to the investigative workflow that follows a flag.

### What is a model inventory and what does an insurer have to include?

A model inventory is a register of every model an organization runs, with enough metadata for a reviewer to locate, classify and evaluate each one. Version 5.0 makes the request explicit for the first time: NAIC's Summary of Changes states that several requests were implied in the prior version and are now listed, and that most notably, regulators now ask for a Model Inventory. Based on the exhibits, a workable entry covers model name and version, the operational area it serves, whether it was developed internally or supplied by a third party and by whom, its implementation date, its risk classification, its risks and limitations, whether it automates or augments a decision, and the data sources behind it. Banks have maintained equivalents since the Federal Reserve and OCC issued SR 11-7 in April 2011.

### What is the difference between Direct Consumer Impact and Material Financial Impact models?

Version 5.0 adds or revises definitions for both and splits Exhibit A so a company reports counts separately for each. The distinction is about who absorbs the consequence of the model being wrong. A model whose output augments or automates a decision affecting a consumer, such as a denial, a coverage determination or a rate, has direct consumer impact. A model whose output moves numbers that matter to the company's financial position, such as reserving or capital allocation, has material financial impact. The categories are not exclusive. A fraud model that drives denials and also feeds reserving belongs in both columns. NAIC's stated reason for the split is to let regulators see whether a company's AI use is concentrated in one type, which may then limit how far the inquiry goes.

### Why does the NAIC supplement ask for document names and page numbers?

So regulators can connect an answer to the evidence behind it. NAIC's Summary of Changes says the Exhibit B checklist added a request to provide the document name and page number when companies respond to inquiries, and gives the rationale directly: it will allow regulators to connect the company responses to documents provided. The prior version already had a page-number column against its governance questions; version 5.0 adds the document name and two new questions, one on materiality and one on how a company oversees third-party models. The practical effect is that a narrative answer written for the exam is no longer sufficient on its own. The control has to already exist, in a named document, on a findable page.

### How does the NAIC AI Risk Evaluation Supplement relate to the NAIC AI Model Bulletin?

The Model Bulletin sets expectations; the Supplement is how a regulator checks them. As of NAIC's April 1, 2026 implementation map, 25 jurisdictions had adopted the Model Bulletin on the Use of Artificial Intelligence Systems by Insurers, with four more, California, Colorado, New York and Texas, operating their own insurance-specific AI regulation or guidance. That is 29 jurisdictions with something in force. Version 5.0 of the Supplement draws its AIS Program definition from the Model Bulletin's language, which keeps the vocabulary consistent between the expectation a state published and the exhibit an examiner hands a carrier. Note that New York's circular letter addresses underwriting and pricing rather than claims.

### Is an AI claims investigation agent covered by the new agentic AI definition?

Very likely, though this is an inference rather than something NAIC states. Version 5.0 adds a definition for agentic AI and a definition for AI Model drawn from NIST and White House Executive Order language. NIST characterizes agentic AI as systems that function as autonomous agents capable of independently making decisions, learning from interactions, and adapting to changing environments. A system that takes a flagged claim, runs multi-step analysis across documents and data sources, and produces a recommendation fits that shape. If it does, it belongs in the Exhibit A model inventory, and Exhibit C's field asking whether the model automates, augments or supports a decision becomes a question a carrier has to answer precisely rather than approximately.
