The seat that sets the deployment date for AI claims automation at a carrier is not the Claims VP who wants it or the CIO who integrates it. It is the third-party risk reviewer, who cannot approve the purchase and can hold it indefinitely. But that review is not a maze. It is a document request, and the request is finite: roughly sixteen artifacts, eight AI contract clauses, and one data-flow diagram. Treat it as a known list and it closes in weeks. Discover the list one item at a time and it takes two quarters.
Most of the delay is wait state, not scrutiny. A reviewer asks for a document, the vendor takes two weeks to produce it, the document raises a follow-up, and the round trip queues behind an entire vendor portfolio carried by one or two people. The 2025 Venminder and Ncontracts State of Third-Party Risk Management survey, a ninth annual study fielded from November 2024 to January 2025 across financial services, insurance, fintech, healthcare, retail, and IT respondents, found that getting timely, accurate documentation from vendors is the number one daily challenge in third-party risk, placed in the top three by 45% of respondents. Analyzing the documents that do arrive, including SOC reports, financials, and contracts, was named by another 20%. The bottleneck is paper, not judgment.
This is written for the composite seat that runs that review - the third-party risk owner sitting under risk management, information security, or compliance - and for the CIO and compliance officer who sign next to them. It covers who owns the gate, why the AI review got heavier in 2026, what four regulatory documents actually require of a carrier buying third-party AI, the artifact list, the eight contract clauses that decide the deal, and how to compress the calendar without reducing scrutiny. It is the review-side companion to the complete guide to claims automation, which covers what you are actually buying.
The seat that cannot say yes but can say no
Third-party risk management is the carrier function that reviews and approves outside vendors before they touch carrier data or carrier decisions. At most organizations it does not sit in procurement. It sits under risk management, information security, or compliance, and it holds a veto rather than a budget. That is the seat that decides your calendar.
Third-party risk reports to risk management or a risk committee at 37% of organizations, to information security at 11%, to compliance at 10%, to IT at 8%, and to procurement at 6%, per the 2025 Venminder and Ncontracts survey. Executive management, finance, legal, and operations split most of the remainder. The consequence for anyone selling into a carrier is that sending the security packet to the sourcing contact usually means it sits until someone forwards it. The consequence for the carrier is that the person who will decide how long this takes is often not in the commercial thread yet.
That person is also thin on capacity. 48% of third-party risk programs run on one or two dedicated full-time employees, up from 43% the prior year, and 22% run on three to five. Programs staffed with six to ten FTEs dropped by 60%, from 10% of respondents to 4%. The portfolios moved the other way: 28% of programs manage 100 to 300 vendors, 18% manage more than a thousand, and only 7% manage fewer than fifty. Every incomplete answer a vendor sends re-enters that queue at the back.
The pressure on the seat is external and documented. 70% of respondents reported feeling pressure to improve their third-party risk program, with 34% of that pressure coming from auditors, regulators, and examiners and 30% from management or the board. At their most recent exam or audit, 29% were told improvements were required, and only 37% came away with no findings at all. A reviewer who slows down on an AI vendor is not being obstructive. They are responding to a finding they have already received or expect to receive next cycle.
Vendor AI now ranks second on the third-party risk concern list, behind cyberattacks at vendors and ahead of pending regulatory changes. That ordering explains most of what happens to an AI purchase inside a carrier in 2026. For the wider map of who else sits in the carrier buying committee and what each seat blocks, see the carrier buying center map for AI claims investigation. This post is the detailed treatment of the one seat that map identifies as a veto without a budget.
Why the AI vendor review got heavier in 2026
Three forces converged on the AI vendor review, and none of them is about any single product. Third-party breach exposure roughly tripled in two years. AI-specific insurance supervisory expectations reached 29 of the 56 NAIC jurisdictions. And vendor AI moved into second place on the third-party risk concern list, behind only cyberattacks at vendors.
The breach number changed the tone of every questionnaire. The Verizon 2026 Data Breach Investigations Report, published May 19, 2026, found that breaches involving a third party now account for 48% of all breaches, with third-party involvement up 60% year over year. The 2025 edition, published April 23, 2025, reported that the share had doubled to 30% across more than 22,000 security incidents and 12,195 confirmed breaches. Two editions ago it was 15%. A reviewer working from that trend line will not accept a vendor security posture on assertion.
The regulatory line moved at the same time. The NAIC adopted its Model Bulletin on the Use of Artificial Intelligence Systems by Insurers on December 4, 2023. As of the NAIC Big Data and Artificial Intelligence Working Group implementation map dated April 1, 2026, 25 jurisdictions had adopted the bulletin and four more run insurance-specific AI regulation or guidance of their own: California Bulletin 2022-5, Colorado 3 CCR 702-10, New York Circular Letter No. 7, and Texas Bulletin B-0036-20. That is 29 of 56 jurisdictions with a published AI supervisory expectation, from a standing start two and a half years earlier.
The third force is experience. 49% of organizations experienced a third-party vendor cyber incident in 2024. Among moderate-impact incidents, 66% carried a monetary cost, 50% caused reputation damage, and 33% drew regulatory scrutiny, with 64% recovering in under 60 days. A reviewer who has worked one of those does not need convincing that the questionnaire matters.
None of this is bureaucracy. A carrier that hands a vendor claim files, adjuster notes, medical records, and recorded statements has extended its own regulatory perimeter to that vendor. The review is the mechanism that makes the extension defensible. The productive vendor response is not to argue the review down. It is to arrive with the evidence already assembled.
What regulators put in writing about your vendor
Four documents define the US expectation set for third-party AI in insurance: the NAIC Model Bulletin, New York DFS Circular Letter No. 7 (2024), Colorado Amended Regulation 10-1-1, and the NIST AI Risk Management Framework the bulletin points at. All four converge on one rule. The carrier stays responsible for a system it did not build.
The NAIC Model Bulletin puts three obligations on the carrier
Section 4.0 of the NAIC Model Bulletin on the Use of Artificial Intelligence Systems by Insurers is three paragraphs long and it sets the entire vendor agenda. Section 4.1 covers "due diligence and the methods employed by the Insurer to assess the third party and its data or AI Systems acquired from the third party to ensure that decisions made or supported from such AI Systems that could lead to Adverse Consumer Outcomes will meet the legal standards imposed on the Insurer itself." Section 4.2 asks for contract terms that "provide audit rights and/or entitle the Insurer to receive audit reports by qualified auditing entities" and that "require the third party to cooperate with the Insurer with regard to regulatory inquiries and investigations." Section 4.3 covers the performance of those rights, which means holding the clause is not enough. The carrier has to use it.
The scope answers the question a vendor gets asked next. Section 1.6 says the AI systems program should address use "across the insurance life cycle, including areas such as product development and design, marketing, use, underwriting, rating and pricing, case management, claim administration and payment, and fraud detection." Section 1.8 says it applies to systems used in regulated insurance practices "whether developed by the Insurer or a third-party vendor." Claims investigation AI is squarely inside both. And the bulletin's examination provisions tell a carrier what a department may ask for, including information and documentation relating to the insurer's pre-acquisition and pre-use diligence, monitoring, oversight, and auditing of AI systems developed by a third party. Pre-use diligence is a documentation obligation that attaches before the contract is signed, which is why this review happens when it happens.
Section 1.5 says the program "may adopt, incorporate, or rely upon, in whole or in part, a framework or standards developed by an official third-party standard organization, such as the National Institute of Standards and Technology (NIST) Artificial Intelligence Risk Management Framework, Version 1.0." That is the through-line from a federal framework to a state insurance expectation. NIST released AI RMF 1.0 on January 26, 2023 with four core functions - Govern, Map, Measure, and Manage - and two subcategories written for exactly this situation. GOVERN 6.1 calls for policies and procedures addressing AI risks associated with third-party entities. GOVERN 6.2 calls for contingency processes to handle failures or incidents in third-party data or AI systems deemed high-risk. If a carrier says its AI program follows NIST, those two subcategories are what it owes on your file.
New York and Colorado, read precisely
New York DFS Insurance Circular Letter No. 7 (2024), issued July 11, 2024, is the clearest published articulation of what a US insurance regulator expects a carrier to hold in a vendor contract. It states that insurers retain responsibility for understanding any tools, external consumer data sources, or AI systems developed or deployed by third-party vendors. It requires written standards, policies, procedures, and protocols for acquiring or relying on vendor-developed AI. And it says contract terms should, where feasible, provide audit rights or entitle the insurer to receive audit reports by qualified auditing entities, and require the vendor to cooperate with the insurer on regulatory inquiries. Read the scope line carefully: CL 7 is scoped to underwriting and pricing, not claims. It does not govern claims AI. It matters anyway, because carrier third-party risk teams run one internal AI template across every purchase, and CL 7 is where much of that template came from.
Colorado Amended Regulation 10-1-1, 3 CCR 702-10, is the most granular published list of what a US insurance regulator wants documented about a third-party model. Section 5.A.9 requires a documented, up-to-date inventory with version control of every external consumer data source, algorithm, and predictive model, including a detailed description of each, its clearly stated purpose, and the outputs generated. 5.A.10 requires a documented explanation of any material change to that inventory and the rationale. 5.A.12 requires a documented description of ongoing monitoring of model performance including accounting for model drift. 5.A.13 requires a documented description of the process used for selecting external resources including third-party vendors. 5.A.14 requires documented comprehensive annual reviews of the governance structure. The Colorado Division of Insurance notice of adoption covers the amended version, effective October 15, 2025. The original took effect November 14, 2023.
Two precision points. First, 10-1-1 is an unfair-discrimination regulation. Its trigger is the use of external consumer data and information sources, and its covered lines are individually issued life insurance, private passenger automobile insurance, and health benefit plans. It does not directly regulate commercial property and casualty claims investigation, and any vendor who says otherwise has not read it. Second, and worth its own sentence: Section 5.B keeps responsibility with the insurer for everything in 5.A when third-party vendors are involved, and then adds that insurers may satisfy requests for documentation and information by having third-party vendors provide the requested documents or information directly to the Division on the insurer's behalf. That is an explicit regulatory endorsement of a vendor answering the state directly. For private passenger auto and health benefit plan insurers, the compliance report is due July 1, 2026 and annually after that, so this is live rather than prospective.
Scope lines are the credibility test
New York Circular Letter No. 7 (2024) is scoped to underwriting and pricing, not claims. Colorado 10-1-1 is triggered by external consumer data use and covers life, private passenger auto, and health benefit plans, not commercial property and casualty claims. Neither directly governs a P&C claims investigation system. Both still shape the file, because carrier third-party risk teams run one internal AI template across every purchase and these two documents are where much of that template's content originated. A vendor that states the scope limits before being asked buys more credibility than one that claims broad applicability and gets corrected by a reviewer who has read the text.
The structure of the review is borrowed from banking
The Interagency Guidance on Third-Party Relationships: Risk Management, issued jointly by the Federal Reserve, the FDIC, and the OCC on June 6, 2023, defines five life-cycle stages: planning, due diligence and third-party selection, contract negotiation, ongoing monitoring, and termination. Carriers are state-regulated rather than bank-regulated, so the guidance does not bind them. Most carrier third-party risk programs copy the five stages anyway, because the examiners and consultants who staff them came from the same talent pool. If you know the five stages, you can predict the order of the questions and the document each one is waiting on.
Read that table as a vendor and one thing stands out. Most of the rows are satisfied by documentation rather than by a control. Hesper's position on the documentation rows is structural rather than procedural: every decision the investigation agent makes is logged with sources, reasoning, and timestamps, because an SIU lead reviewing the file and a state examiner sampling it both need to read the same record. The audit trail is not a compliance feature bolted onto the product. It is the product's output format.
The artifact list a carrier will request
The request is finite and knowable: sixteen artifacts across four blocks, which are security, privacy and data, AI governance, and contract terms. The security block stalls young vendors on the age of their evidence. The AI-governance block stalls almost everyone, because those documents have rarely been written for an insurance reviewer before.
The sharpest number in the third-party risk data measures exactly that gap. Between 2024 and 2025, the share of organizations assessing vendor AI by adding usage language to the contract rose from 11% to 40%. Documenting AI risks internally went from 6% to 39%. Verbal communication with the vendor went from 2% to 38%. Sending questionnaires went from 15% to 32%. But collecting vendor documentation moved only from 8% to 14%. Carriers are writing AI clauses roughly three times more often than they are collecting the documentation those clauses point at. That is not indifference. It is that the documents mostly do not exist on the vendor side yet. Organizations not monitoring vendor AI usage at all fell from 37% to 23% over the same period, so the population asking is growing fast.
Two cadence facts shape what lands in your inbox. 55% of organizations updated their vendor risk questionnaire and due diligence document requirements within the past year, so the list your last carrier used is probably already stale. And 92% complete inherent risk assessments before diligence begins, with 47% tying reassessment frequency to the vendor's risk tier and 33% reassessing annually. The tier assigned in week one determines how heavy the rest of the review is, which is why the data-flow answer belongs at first contact rather than at week six.
The security block is the half of the packet with an established grammar. SOC 2 reports against the AICPA trust services criteria - security, availability, processing integrity, confidentiality, and privacy - and a Type II attests to whether controls operated effectively over an observation period, typically three to twelve months, rather than to control design at a single point in time. That distinction is where young AI vendors stall: a Type 1, or a Type 2 with a two-month window, reads as thin to a reviewer accustomed to twelve-month reports from incumbents. The depth on that half of the packet, including the seven data-handling controls that sit on top of SOC 2, is in SOC 2 and data handling for AI fraud investigation.
The AI-governance block has no established grammar yet, which is why it stalls. Most AI vendors have never authored a model card for an insurance reviewer, have never written down where the human decision boundary sits, and have never committed in writing to a model-change notice period. ISO/IEC 42001:2023, the first AI management system standard, is becoming the shorthand answer to this block, though certification is still uncommon. Note also that this is not the CIO's review. The third-party risk owner is assessing whether the carrier can defend the relationship; the CIO is assessing whether the system integrates and where the data physically sits. The two run in parallel and ask different questions, and the integration and MSA side is covered in the CIO checklist for an AI claims investigation rollout.
Detection vendors and investigation vendors get reviewed differently
The artifact a product emits determines the review it gets. A score is reviewed as a model, which creates a standing model-governance obligation the carrier carries for as long as the model is in production. A documented investigation record is reviewed as evidence, and the explainability the reviewer is asking for is the deliverable itself rather than a separate document.
Start from the layer map, because it explains the review difference without any marketing in it. Detection is upstream; investigation is downstream. FRISS, Shift Technology, and Verisk sit at the detection layer and output a score or a flag on a claim. Hesper sits at the investigation layer and outputs a documented record of how a conclusion was reached, produced by 15+ investigation phases run in parallel on every flagged claim. Complementary to FRISS, Shift Technology, and Verisk - not a replacement. The modal carrier deployment runs both layers, and the move from fraud detection to fraud resolution is what changes the shape of the review.
A carrier reviewing a scoring model asks about training data, drift, bias testing, and the explainability of a number. That is precisely the list Colorado 10-1-1 formalizes across 5.A.9 through 5.A.12: inventory with version control, stated purpose and outputs, material-change explanation, and drift monitoring. It is a real and ongoing regulatory obligation, not vendor overhead, and it recurs on every model version for the life of the deployment. A detection vendor is not doing anything wrong by creating it. That obligation is the price of occupying that layer.
Two concessions belong here. First, incumbency is a genuine procurement advantage and pretending otherwise is not useful. Verisk's ISO ClaimSearch is contributory industry infrastructure most carriers papered years ago, so the third-party risk file already exists and a review is a refresh rather than an origination. A new AI entrant has no such precedent and should expect the longer path. Second, the re-review trigger row above is identical on both sides on purpose. Any material model change re-opens the file for an investigation vendor exactly as it does for a scoring vendor. A comparison table where one side wins every row is marketing, and an experienced reviewer discounts the whole document when they see one.
One further distinction saves argument. Security certifications answer the security half of a third-party risk packet and say nothing about the AI-governance half. A vendor can hold every certification on the list and still have no model documentation, no stated human-review boundary, and no model-change notice commitment. That is the half carriers are newly writing into contracts, and it is the half where the packets are empty.
The escalation path above the third-party risk owner runs to the chief risk officer, who will ask whether the system belongs in the model inventory at all and what the concentration exposure looks like if it fails. That framing, including where an investigation agent lands inside a model risk taxonomy built for pricing models, is worked through in the chief risk officer's view of AI claims investigation.
The AI addendum: eight clauses that decide the deal
The security packet rarely kills an AI deal in insurance. The contract does. Eight clauses carry the weight: training-data use, model change notification, human-in-the-loop attestation, subprocessor disclosure, retention and deletion, exit and data return, audit rights, and regulatory cooperation. Carriers now pre-draft most of them, so expect redlines rather than blanks.
Expect them because carriers started writing them. The share of organizations adding AI usage language to vendor contracts went from 11% to 40% in a single year in the Venminder and Ncontracts data. A vendor that arrives with the eight already drafted is negotiating from its own paper. A vendor that waits for the carrier's addendum is negotiating from someone else's and has added a full redline cycle to the calendar.
- Training-data use. State plainly whether carrier data is used to train or fine-tune any model, and if it is not, say so as a contractual commitment rather than a policy statement. The carrier CIO's stated shutdown triggers are pre-SOC-2 vendors, training-on-customer-data clauses, and vague data-flow diagrams. This clause covers two of the three.
- Model change notification. Define what counts as a material change and how much notice the carrier gets. This is the clause that maps to Colorado 5.A.10's material-change explanation requirement, and the carrier cannot satisfy that requirement if it learns about your change from a release note.
- Human-in-the-loop attestation. Write down what the system decides and what a person decides. For claims work the boundary is not ambiguous: the agent investigates and produces the record, and the SIU lead or adjuster makes the determination. Put that sentence in the contract, not only in the deck.
- Subprocessor list and change notification. Name the model providers, the hosting infrastructure, and any data enrichment sources, then commit to a notice period before the list changes. 58% of organizations review their third parties' own third-party risk, so this list gets read carefully.
- Data retention and deletion. Not a policy stating that data is retained no longer than necessary. A schedule with actual periods on it, per data class, with a deletion attestation available on request.
- Exit, termination assistance, and data return. Format, timeline, and who bears the cost. This is the clause most often discovered at signature because nobody drafted it, and it is the fifth stage of the interagency life cycle.
- Audit rights. NAIC Model Bulletin 4.2(a) asks for audit rights and/or the right to receive audit reports by qualified auditing entities. Offering reports alone is a common vendor position and often an acceptable one. State which you are offering rather than leaving the carrier to discover it in redlines.
- Regulatory cooperation. NAIC 4.2(b) calls for the vendor to cooperate with regulatory inquiries and investigations. Colorado 5.B goes further and permits the vendor to provide documents directly to the Division on the insurer's behalf. Offer a named regulatory-response contact and a response commitment before it is asked for. It costs nothing and a compliance officer will remember it.
None of the eight is exotic and each is answerable in a page. The reason they stall deals is that they get answered late, one at a time, by different people at the vendor, each answer arriving after the reviewer has already moved the file to the back of the queue.
Compressing the review from two quarters to a few weeks
The review is sequential by default and parallel by design. Compression comes from removing wait states, not from reducing scrutiny. Map the five life-cycle stages onto a carrier AI purchase, identify what each stage is waiting on, and pre-stage the document that clears it before the stage begins.
Five life-cycle stages, the wait state inside each one, and the document that removes it. Four of the five are waiting on paper that could have been sent at first contact.
Read the wait-state column and the pattern is plain. Four of the five stages are waiting on a document that could have been produced earlier at no cost. The prescription follows directly.
- Send the complete document packet unrequested at first contact, indexed, with a cover page mapping each artifact to the question it answers.
- Pre-redline the AI addendum, conceding up front on the clauses that cost you nothing.
- Name one vendor-side owner for the questionnaire so the answers share a vocabulary. Three authors produce three vocabularies and one follow-up round per author.
- Give the reviewer a one-page data-flow diagram rather than a narrative. Diagrams set the criticality tier; narratives generate questions.
- Offer the regulatory-cooperation commitment and a named response contact before anyone asks for it.
- Never ask a reviewer to accept an assertion that a document could settle. Every unpapered assertion becomes a round trip, and every round trip queues behind an entire portfolio carried by one or two people.
What the carrier waits on while the review runs is the real cost of every avoidable round trip. Manual investigation takes 14+ days per case and one investigator carries 200+ cases. Across US P&C carriers, manual teams fully investigate roughly 25% of flagged claims, and the rest are paid, denied without full work, or queued indefinitely. Cost sits near $2,500 per manually investigated case against roughly $150 with automated investigation, and coverage moves from 25% to 100%. Automated investigation returns a reviewable, audit-ready record in hours, not weeks. Against a fraud loss the Coalition Against Insurance Fraud puts at $308 billion a year in the United States, with fraud present in roughly 10% of property-casualty losses, a review that runs one quarter longer than it needed to is a quarter of that gap left open.
The compression argument is not that the carrier should scrutinize less. It is that scrutiny and calendar are separable. A reviewer who receives a complete, indexed packet on day one spends their time analyzing rather than chasing, which is the 20% problem in the survey data rather than the 45% one. That is a better use of the one or two people carrying the program, and it leaves a better-documented file at the end.
What the EU AI Act does and does not do to a US carrier's file
Less than most vendor decks claim. Annex III point 5(c) of the EU AI Act classifies AI used for risk assessment and pricing in relation to natural persons as high-risk in the case of life and health insurance. It does not name property and casualty claims handling. Point 5(b) expressly carves fraud detection out of the creditworthiness category.
The exact language matters here, because this is a section where a vendor either earns credibility or loses it. Annex III point 5(b) classifies AI systems "intended to be used to evaluate the creditworthiness of natural persons or establish their credit score" as high-risk, "with the exception of AI systems used for the purpose of detecting financial fraud." Point 5(c) covers "AI systems intended to be used for risk assessment and pricing in relation to natural persons in the case of life and health insurance." A US property and casualty carrier running claims investigation AI with no EU exposure sits outside the Act entirely.
The dates a reviewer will have in front of them: the Act entered into force on 1 August 2024; prohibitions and AI literacy obligations applied from 2 February 2025; general-purpose AI model rules, governance provisions, notified bodies, and penalties applied from 2 August 2025, with GPAI models already on the market before that date given until 2 August 2027; Annex III high-risk obligations become applicable on 2 August 2026; and Article 6(1) product-embedded high-risk obligations follow on 2 August 2027.
The honest framing for a vendor is that a carrier with EU operations should still expect the AI Act question in a third-party review this year, and a vendor that has already mapped its system against Annex III answers it in one sentence instead of opening a research project. Overstating the Act is how a vendor loses a reviewer who has read it. ISO/IEC 42001:2023 usually sits on the same page of the questionnaire, and it specifies requirements for establishing, implementing, maintaining, and continually improving an AI management system for organizations that provide or use AI-based products and services. Between the two, the AI-governance block of the packet is answerable today.
Key takeaways
- The seat that sets the deployment date for AI claims automation is the third-party risk reviewer, and that seat reports to procurement at only 6% of organizations while reporting to risk management or a risk committee at 37%.
- Third-party involvement in breaches reached 48% in the Verizon 2026 DBIR, up from 30% a year earlier and 15% two editions ago, which is why every vendor review got heavier regardless of what the vendor sells.
- The NAIC Model Bulletin, adopted in 25 jurisdictions with four more running their own AI rules as of April 2026, requires the carrier to hold pre-use diligence evidence, audit rights, and a vendor commitment to cooperate with regulatory inquiries.
- Carriers are writing AI clauses into contracts roughly three times faster than they are collecting the AI documentation those clauses point at, so the vendor that arrives with model documentation already written clears the AI-governance block while others open a drafting project.
- A scoring vendor's explainability evidence has to be authored separately from the product and maintained on every model change, while an investigation vendor's audit trail is the deliverable itself, which is why the two clear a review on different timelines under identical scrutiny.