The duty is symmetric, and almost nobody in claims is treating it that way. A carrier that denies a claim on a black-box AI output and a carrier that denies a claim on a twenty-minute human file review are failing the same duty: the unfair-claims-practices obligation to conduct a reasonable investigation before acting on it. The regulation that creates that duty does not ask which method produced the thin file.
That framing changes what to do with the defense bar's warnings about AI. The checklist counsel are circulating - methodology, data sources, reproducibility, audit log, a named human decision-maker - is usually read as a list of reasons to slow down. Read it instead as a specification. Build to it and the AI-assisted investigation comes out more defensible than the manual file it replaced, for one unglamorous reason: the manual file has no audit log at all.
This post walks the test as it stands on September 10, 2026. The Advisory Committee on Evidence Rules' May 2026 report on proposed Rule 707 and its five-factor reliability analysis. The discovery orders in Estate of Lokken v. UnitedHealth Group and Huskey v. State Farm, which is the closest thing in the case law to a property carrier's fraud-screening algorithm being put under a microscope. The regulator scope map, including two corrections worth getting right. And the symmetry argument that follows from all of it.
Two things this post deliberately does not re-cover. Privilege and work product are handled in our general counsel's guide to AI fraud investigation, which is the companion piece and the better starting point if the anticipation-of-litigation line is your immediate problem. What an audit-ready record has to contain, field by field, is in the defensibility standard for fraud investigation AI. That post is the what. This one is the test it has to survive.
The event to design for is a deposition, not a Daubert hearing
An AI-assisted fraud investigation almost never faces a Daubert or Frye hearing. It faces discovery and a corporate designee under oath. The carrier is not offering the model as expert evidence; it is offering the facts the investigation surfaced. What gets tested is whether anyone can explain how those facts were produced.
Picture the sequence, because the sequence is the whole risk model. A claim is flagged. Something automated contributes to the decision to deny, or to the decision to refer, or to the decision to close without a full look. Eighteen months later the claimant's counsel serves a document request covering every system that touched the file, then notices a deposition of the person most knowledgeable about how it works. That witness is a claims operations lead or an SIU manager, not a data scientist and not the vendor. They will be asked what the system considered, what it discarded, whether the same claim run again produces the same finding, who reviewed the output, and what that person changed.
None of those questions requires the model to be opened. All of them require a record to exist. And the industry's own numbers say the record usually does not.
In the Grant Thornton 2026 AI Impact Survey, which included 100 insurance respondents, only 24% of insurance executives said they were very confident they could pass an independent AI governance review within 90 days. In the same survey, 61% said their boards had established AI governance policies, and 44% said governance or compliance challenges had contributed to an AI project failing or underperforming. Insurance Journal's report on the survey contains the line that matters most: even where controls exist, the evidence of them is fragmented across teams and tools.
Fragmented across teams and tools is a description of a deposition going badly. It means the policy is in a compliance drive, the model documentation is with the vendor, the per-claim record is partly in the claims system and partly nowhere, and the reviewer's actual reasoning was never written down. Each fragment is defensible on its own. The set is not reconstructable under oath by one witness in one sitting, which is the only form in which it will ever be asked for.
These are three different surveys with three different samples and they should not be read as one dataset. The 58% to 82% range, the 75% human-oversight figure, the 12% maturity number and the 7% scalable-success number come from Sedgwick's "Future-ready property claims" report, reported by Insurance Journal in March 2026. The 41% figure is from an AM Best survey of more than 150 rated insurers and MGAs released in April 2026. The governance figures are Grant Thornton's. Read across them and one shape holds in all three: usage is broad, maturity is narrow, and the thing that lags hardest is the evidence that the usage was governed.
The 75% number deserves separate attention because it is the one claims professionals themselves produced. Three quarters of them say AI needs human oversight. That is not a hedge. It is the operating model the rest of this post argues is also the legally correct one, and it is worth noticing that practitioners reached it before the rulemakers did.
What insurance defense counsel are actually saying
The insurance defense bar's guidance on AI is narrow, specific and correct: understand how the system operates, and never let it be the sole decision-maker. Jessica Gross and Brett Carey of RumbergerKirk wrote it that way in Claims Journal in June 2026. Their frame is lawyers using AI in litigation. The harder question sits one step upstream.
Gross and Carey are worth quoting at length because their piece is unusually free of both hype and panic. Writing in Claims Journal on June 17, 2026, they set the rule as: "it is important for lawyers to understand how these systems operate and how they are structured but not to rely on AI as the sole decision-maker." They add that "Attorneys must also be considerate of data sources, potential bias, waiving privileges by using AI hosted through public platforms and the general credibility of the information." Both are practicing insurance defense attorneys - Gross in complex insurance defense and casualty litigation, Carey in insurance coverage and bad faith, representing carriers in first-party and third-party matters.
That sentence is the whole problem stated once. A file in which human judgment and machine output are indistinguishable cannot be defended, because you cannot defend a conclusion whose author is unknown. Their remedy is also right: they write that "AI will not replace attorney judgment but rather will drastically improve strategy decisions at the onset." Judgment stays with the person; the machine does the assembly.
The extension this post makes is a timing point, not a disagreement. Gross and Carey are writing about the moment a lawyer picks up a tool. By then the carrier has usually already used AI to build the claim file: to summarize the recorded statement, to flag the invoice, to score the claim, to decide whether the file was worth a full look. That work happened months earlier, was performed by people who were not thinking about litigation, and is the part that will be produced in discovery. The lawyer's use of AI is disclosable and controllable. The carrier's use of AI is historical and fixed by the time anyone asks.
The same conclusion arrives from the discovery side. Reviewing the Lokken discovery order in July 2026, Cory Chipman and Katie Graham of Freeman Mathis & Gary drew three practical lessons for insurers, the first being that using AI in claims handling may open the door to broad discovery about how the tool is implemented, supervised and used in practice. Their remedy is documentation: insurers should be prepared to show that AI is used as a support tool rather than a substitute for required human review. Two firms, two vantage points, one rule.
It is also the rule a legislature has now written down. California SB 1120, the Physicians Make Decisions Act, took effect January 1, 2025 and requires that where a health or disability insurer uses AI or algorithms in utilization review, the ultimate medical necessity determination be made by a licensed physician or other competent licensed professional. The statute still allows an AI tool in utilization review, subject to conditions including that it not supplant health care provider decision-making, but it provides that the tool shall not deny, delay or modify services based in whole or in part on medical necessity. The bill was signed on September 28, 2024, per the bill author's announcement. Note the scope carefully: this is health and disability utilization review. There is no property and casualty analogue in California as of today. Its value here is as the shape of the emerging rule, which is that a licensed human makes the call and the AI does the assembly.
Rule 707 did not pass, and it still tells you what to build
Proposed Federal Rule of Evidence 707 would extend Rule 702's reliability requirements to AI output offered without a sponsoring expert. In May 2026 the Advisory Committee on Evidence Rules declined to recommend action and sent a revised draft to its Fall 2026 meeting. The rule is not law. Its draft committee note is still the clearest published specification available.
The timeline is short and worth having exactly right. The Advisory Committee voted to publish proposed Rule 707 for comment on May 2, 2025. The Standing Committee approved publication on June 10, 2025. Public comment ran from August 15, 2025 to February 16, 2026. The Advisory Committee met on May 7, 2026 and reported to the Standing Committee for its June 3 and 4, 2026 meeting. In the Report of the Advisory Committee on Evidence Rules dated May 17, 2026, Chair Judge Jesse M. Furman wrote: "The Committee does not recommend action on the proposed Rule 707 at this time."
The report explains why. Public comment was described as mixed but helpful. The Committee revised the draft, agreed that the revised version would require re-publication if it went forward, and decided to have it vetted at the Committee's Fall meeting by technology experts and others working in AI and law. It took the same approach to deepfakes, which are handled below. Nothing was adopted. Nothing takes effect.
What the revised draft says
The redraft is more interesting than the original. The title changed from machine-generated evidence to "Evidence Produced by Artificial Intelligence and Presented at Trial Without an Expert," because commenters said machine-generated was overbroad and would sweep in ordinary instrumentation. Subsection (a) then tracks Rule 702 almost word for word: if evidence is a product of artificial intelligence, is offered without an expert witness, but would be subject to Rule 702 if a witness testified to it, the proponent must establish that the evidence will help the trier of fact, is based on sufficient facts or data, is the product of reliable principles and methods, and reflects a reliable application of those principles and methods to the facts of the case.
Subsection (b) is the operational one. Admissibility under the rule "ordinarily requires the proponent to provide an expert to explain how the system of artificial intelligence reliably produced the evidence," with an escape valve for exceptional circumstances where the court may rely on other proof. Subsection (c) adds a notice requirement: the evidence is admissible only if the proponent gives an adverse party reasonable notice of the intent to offer it, so the party has a fair opportunity to meet it.
Subsection (e) defines artificial intelligence, borrowing from the National Artificial Intelligence Initiative Act of 2020, as "a machine-based system that can, for a given set of human-defined objectives, make predictions, recommendations or decisions influencing real or virtual environments." Read that definition slowly against your own stack. A fraud score is a prediction. A referral recommendation is a recommendation. An agent that decides which investigative checks a claim needs is making decisions influencing a real environment. Every one of them is squarely inside the definition. There is no architecture that gets a carrier outside it, and any vendor claiming otherwise is reading a different document.
The five-factor analysis, which is the useful part
Buried in the proposed committee note is the closest thing any US authority has published to a checklist for interrogating an AI system. The Committee wrote that a Rule 707 analysis will usually involve the following, among other things.
- Considering whether the inputs into the process are sufficient for purposes of ensuring the validity of the resulting output. The Committee's own example is whether training data is sufficiently representative to render an accurate output for the population involved in the case at hand.
- Considering whether the process has been validated in circumstances sufficiently similar to the case at hand.
- Reviewing information about what records the system keeps, what is deleted, and whether the output can be reproduced or audited.
- Considering whether the process or system has been subject to evaluation by entities independent of the developer, and whether the opponent and independent evaluators have access to the system, for example through meaningfully available research licenses.
- Considering whether the opponent has received sufficient information about how and why the process works, including information analogous to that which would be disclosed if the conclusion were that of a human expert.
Factor 3 is an audit-log specification written in judicial language. What records does the system keep. What does it delete. Can the output be reproduced or audited. That is a retention policy and a logging architecture, not a model architecture. Factor 5 is a disclosure specification: the opponent must get information analogous to what a human expert would have disclosed, which for a human investigator means what they looked at, in what order, and why they concluded what they did. Neither factor requires opening the model. Both require the system to have been built with the record in mind before the first claim ran through it.
This is the point at which the legal document and the product roadmap converge. Logging that gets bolted on after a demand letter arrives is not the same artifact as logging that runs as a design property of the system. In Hesper's case the 15+ investigation phases each write their own sourced, timestamped record as they run - inputs, tool calls, what each returned, and the reasoning that connected evidence to conclusion - because a record assembled afterward is a reconstruction, and a reconstruction is exactly what factor 3 is designed to catch.
The sentence a black-box output cannot survive
The full passage is more careful than the pull quote. The Committee notes that an AI process can sometimes develop in such a way that nobody is able to explain how the system reached a result, because the machine has developed the ability to program itself. Then it draws the consequence above. Be evenhanded about what this does and does not mean. It does not say proprietary systems are inadmissible; the case law says the opposite, as the Wakefield discussion below shows. It says that inexplicability is usually fatal, and the Committee's own escape route makes the point sharper: a proponent may overcome inexplicability by showing through an expert how the machine got trained and establishing, for example through validation studies, that the process leads to a low rate of error. Explicability is a property of the deployment, not of the model family.
The Committee also flagged a remedy that survives even a win on admissibility: where machine-generated evidence comes in, a court should consider a limiting instruction telling the jury that such evidence is subject to error and should not be assumed reliable simply because a machine produced it. A carrier that wins the admissibility fight and then hears that instruction read has still lost something. It is a practical argument for human-authored conclusions supported by machine-assembled evidence, rather than machine-authored conclusions a human signed off on.
One more piece of the May 2026 report is relevant to fraud specifically, and it points the other way. The Committee declined to amend Rule 901 to address deepfakes, concluding that an amendment "is not warranted" for now and citing a Federal Judicial Center survey of District, Magistrate and Bankruptcy judges in which only fifteen respondents reported having dealt with deepfake issues. A working draft of a new Rule 901(c) is being held in abeyance. That is the courts saying the fabricated-evidence problem has not reached them yet. It has already reached claims intake, which we covered in the playbook on AI-generated evidence in fraud investigation. The asymmetry is worth sitting with: carriers are being asked to prove their own AI is reliable at the same time as they are receiving claimant evidence that AI made up.
The rule that already applies while 707 waits
Rule 702 already carries the burden. As amended effective December 1, 2023, it requires the proponent to demonstrate to the court that it is more likely than not that expert testimony rests on sufficient facts, reliable principles and methods, and a reliable application of those methods to the facts of the case. Reliability is an admissibility question.
The operative text is worth reading against your own workflow, not in the abstract. A qualified expert may testify in the form of an opinion or otherwise if the proponent demonstrates to the court that it is more likely than not that (a) the specialized knowledge will help the trier of fact, (b) the testimony is based on sufficient facts or data, (c) the testimony is the product of reliable principles and methods, and (d) the expert's opinion reflects a reliable application of the principles and methods to the facts of the case. The 2023 amendment moved the burden into the rule text itself. Reliability stopped being something a jury weighs and became something a judge decides before the jury hears it.
For a carrier the practical channel runs through its own experts. An SIU manager testifying about why a claim was referred, a retained accounting expert relying on an AI-assembled timeline, a document examiner whose opinion incorporates an automated forensic finding: all of them are within (c) and (d). If the underlying process cannot be described, the expert who relied on it inherits the problem.
Authentication is a different and much easier question, and conflating the two is the most common error in this area. Rule 901(a) requires only that the proponent produce evidence sufficient to support a finding that the item is what the proponent claims it is. Rule 901(b)(9) gives the software-specific route: evidence describing a process or system and showing that it produces an accurate result. Carriers reach for 901(b)(9) constantly, because it sounds like it settles the question.
Authenticating the log is not the same as proving the method reliable
The Advisory Committee addressed this directly in its May 2026 report. Rule 901(b)(9)'s requirement that a process or system produces an accurate result is subsumed by the reliability requirements that must be established by a preponderance of the evidence under Rule 702 or 707. Because the threshold for authenticity is significantly lower than the threshold for reliability, evidence qualified under 702 or 707 automatically satisfies 901(b)(9). The reverse does not hold. Satisfying Rule 901(b)(9) does not suffice for admissibility. The sentence "we can authenticate the audit log" answers the easy question and leaves the hard one untouched.
There is a structural observation here that is easy to state and hard to argue with. A numeric fraud score and a documented investigation are different evidentiary objects. A score of 812 on a 0 to 999 scale has nothing to explain under 702(c) and (d) unless the methodology behind it is disclosed, and the methodology is usually the vendor's core intellectual property, which is precisely why it is not disclosed. A documented investigation has a methodology by construction: these sources were consulted, in this order, they returned this, and a named person concluded that.
This is not an attack on scoring. Scoring is a good technology doing the job it was built for, which is deciding where to look. The point is narrower and it is an evidentiary one: the score is a reason to look, not a reason to deny. What supports the action is the investigation. Carriers get into trouble when the score is asked to do the second job, and detection vendors get unfairly blamed for a use their own product documentation never claimed.
Two ways AI evidence fails, and one way it survives
In every reported case where AI output has been kept out, the proponent could not explain the process. In the case where comparable proprietary software came in, the proponent could. Matter of Weber and State v. Puloka are the failures. People v. Wakefield is the counterexample, and its source code was never produced.
Start with Weber, because it is the most instructive failure in the literature and because the facts are almost comically clean. In Matter of Weber, 2024 NY Slip Op 24258, decided October 10, 2024 by Surrogate Schopf in Saratoga County, an objectant's damages expert had used Microsoft Copilot to support a supplemental damages calculation. On cross-examination he could not recall what input or prompt he had used. He could not state what sources Copilot relied on. He could not explain any details about how Copilot works or how it arrives at a given output.
The court then did something no vendor deck prepares you for. It ran the query itself. It entered the prompt asking for the value of $250,000 invested in the Vanguard Balanced Index Fund from December 31, 2004 through January 31, 2021 into Copilot on a Unified Court System computer and got $949,070.97. It ran the same query on two more court computers and got $948,209.63 and a little more than $951,000.00. The court's conclusion is the reason this case matters to anyone building a claims system: while the variations were not large, the fact that there were variations at all called into question the reliability and accuracy of the tool to generate evidence to be relied upon in a court proceeding.
Non-determinism, on its own, sank it. Not bias, not hallucination, not a wrong answer. Three runs, three numbers. That maps directly onto Rule 707 factor 3, which asks whether the output can be reproduced or audited, and it is the single most under-engineered property in commercial claims AI. Reproducibility is not a model property you can buy; it is a system property you design for, by pinning versions, capturing inputs, storing intermediate outputs, and making the record rather than the model the thing you reproduce from.
Weber failed on general acceptance too. The expert was adamant that using Copilot or other AI tools for drafting expert reports is generally accepted in the field of fiduciary services, but he could not name any publication or source confirming it. That is a Frye failure, and it is worth flagging early because it is the factor an audit trail does not fix.
Weber also created a disclosure duty, and it is the part most likely to spread. The court held that due to the rapid evolution of AI and its inherent reliability issues, counsel has an affirmative duty to disclose the use of artificial intelligence before AI-generated evidence is introduced, and that the evidence should properly be subject to a Frye hearing whose scope the court determines. Weight matters here: this is one trial court, in Surrogate's Court practice, describing the question as one of first impression. It binds nobody statewide. It is also the kind of holding other trial judges reach for when they have nothing else, which is why it has been cited well beyond its jurisdiction.
One definitional detail from Weber is directly useful for claims. The court defined AI as any technology using machine learning, natural language processing or other computational mechanism to simulate human intelligence, including document generation, evidence creation or analysis, and legal research, and split it into generative and assistive. Assistive materials are any document or evidence prepared with the assistance of AI technologies but not solely generated by them. An investigation record where a named human investigator reviews the evidence and signs the conclusion is assistive under that definition. A finding produced and issued with no human author is generative. The categories carry different burdens, and which one you are in is a workflow decision, not a technology decision.
The second failure is one line and it is a criminal case, so treat it as illustration only. In State v. Puloka, cause no. 21-1-04851-2 in King County Superior Court, Washington, the court held a Frye hearing in 2024 and excluded expert testimony based on video that had been upscaled with Topaz Labs AI, holding that the relevant community for Frye purposes was the forensic video analysis community and that this community had not accepted the tool. Greenberg Traurig covered the ruling in May 2024. The lesson is about who counts as the relevant field, which is a question nobody has answered for AI-assisted claims investigation.
Now the counterexample, which is the most important case in this post. In People v. Wakefield, 38 N.Y.3d 367 (2022), New York's Court of Appeals affirmed the admission of DNA evidence produced by TrueAllele, a proprietary probabilistic-genotyping program, after a full Frye hearing. It also affirmed the denial of the defendant's request for the source code. The Weber court described what made the difference: Wakefield involved a full Frye hearing that included expert testimony explaining the mathematical formulas, the processes involved, and the peer-reviewed published articles in scientific journals.
Hold those two cases side by side. Same category of technology: proprietary software producing a probabilistic output that a human then relies on. Opposite results. In one the proponent produced someone who could explain the method and point to independent literature. In the other the proponent could not name a prompt. The variable is the explanation, not the algorithm, and the defendant in Wakefield never got the source code. That is the single most reassuring fact available to a carrier deploying proprietary AI, and it comes with an obligation attached: somebody has to be able to explain the method.
A note on weight before anyone builds a compliance memo on this table. Weber is one trial court. Puloka is one criminal case in one county in a Frye state. Wakefield is a state high court. New York applies Frye, Washington applies Frye, and federal courts apply Rule 702. Anyone who tells you a specific number of states follow one test or the other should be asked for the citation, because no authoritative count is publicly maintained.
Inside the Lokken and Huskey discovery orders
Trade-secret protection holds for the model and fails for the record of how you used it. In Estate of Lokken v. UnitedHealth Group a federal court refused to order production of source code and underlying data, then ordered policies, procedures, training materials and governance records produced. Huskey v. State Farm went further into the fraud-screening question itself.
Lokken is health insurance, not property and casualty, and the distinction should be kept visible in any internal memo. It is nonetheless the most detailed public map of what a court will make an insurer produce about an AI claims tool. A federal magistrate judge in the District of Minnesota entered a discovery order on March 9, 2026, cited as Estate of Gene B. Lokken v. UnitedHealth Group, Inc., 2026 WL 658883 (D. Minn. 2026).
Per the ArentFox Schiff alert on the order, the categories ordered produced included the following, with documents reaching back to January 2017.
- Policies and procedures for post-acute care claims, and the employee training materials that accompany them.
- Documents analyzing or discussing the nH Predict tool.
- Business records relating to the naviHealth acquisition, including projected cost savings.
- Materials on government investigations into the use of AI in claims assessment.
- Performance reviews and compensation data for care coordinators and medical directors.
- Internal AI review board materials, including the identities of the board's members.
- Contact information for the medical directors and care coordinators involved in denials, provided to 300 putative class members.
Read that list as a specification for what has to exist and be findable. Not one item on it is a model artifact. They are organizational artifacts: policies, training, governance minutes, acquisition rationale, compensation. The court denied the requests for the AI system's source code, underlying data and embedded medical guidelines, recognizing the need to protect proprietary technical information, while permitting discovery into policies, procedures, governance, regulatory oversight and how the tool was used in practice. Freeman Mathis and Gary's third lesson from the order states the split cleanly: proprietary AI materials may receive protection, but operational and governance-level information likely will not.
The split every carrier should plan around
The model is shielded. The record of how you used it is not. Trade-secret protection, protective orders and vendor confidentiality provisions all operate on the technical artifact. None of them reach the carrier's own policies, procedures, training materials, governance minutes, vendor contracts, or the per-claim record of what the system did on a specific file. Assume every one of those is producible, and build them as though a stranger will read them in three years, because that is the realistic scenario.
Huskey is the property and casualty one, and it is the closest analogue in the case law to a carrier's fraud-screening system being put under discovery. Huskey v. State Farm Fire & Casualty Co., No. 1:22-cv-07014 (N.D. Ill.), was filed December 14, 2022. On September 11, 2023 Judge Kendall granted in part and denied in part a motion to dismiss, allowing a Fair Housing Act disparate-impact claim to proceed. In February 2024 Magistrate Judge Harjani entered an order that, per the Civil Rights Litigation Clearinghouse docket, provided that Phase I discovery "shall be limited to algorithmic decision-making tools employed by State Farm to screen out potentially fraudulent or complex claims from straightforward homeowners insurance claims."
Sit with the phrasing. Not underwriting models. Not pricing models. Algorithmic decision-making tools employed to screen out potentially fraudulent or complex claims. That is a fraud-triage system, in homeowners, made the express subject of a phase of discovery. The docket has continued to move: on January 15, 2025 State Farm's motion for a protective order was granted without prejudice and the plaintiffs' motion to compel was denied, and discovery remained ongoing as of February 19, 2026.
Be precise about what Huskey is and is not. The theory is Fair Housing Act disparate impact, not insurance bad faith. These are allegations; nothing has been proven, and the disposition of the case says nothing about whether any particular carrier's tooling is sound. What it establishes is narrower and more durable: a plaintiff can get a court to open a discovery phase aimed squarely at a P&C carrier's fraud-screening algorithms, and the carrier will be answering questions about them for years.
Two other cases round out the picture and neither needs more than a clause. Kisting-Leung v. Cigna, No. 2:23-cv-01477 (E.D. Cal.), survived a motion to dismiss in part in March 2025 and is still in discovery, with a scheduling order entered on August 26, 2026. In Barrows v. Humana, No. 3:23-cv-00654-RGJ (W.D. Ky.), an August 14, 2025 order let breach of contract, breach of the implied covenant of good faith and fair dealing, unjust enrichment and common law fraud proceed while dismissing the statutory unfair claims settlement and bad-faith counts with prejudice. The pattern across both is worth noting: contract and good-faith theories are proving more durable than statutory bad-faith theories in AI claims litigation.
The operational conclusion for a claims legal lead is a single sentence. Assume that every policy, every procedure, every governance artifact, every training deck, every vendor contract and every per-claim tool-call log is producible, and build the program so that producing them is a query rather than a project. That is what audit-trail-native means as a procurement requirement, not a marketing phrase, and it is laid out field by field in our defensibility standard.
Which regulator actually reaches your claims file
Most US insurance AI regulation reaches underwriting and pricing, not claims. The instruments that do reach claims are the NAIC Model Bulletin, which names fraud detection explicitly, the unfair claims settlement statutes, and California's SIU regulations. New York's circular letter does not reach claims. California SB 1120 covers health and disability utilization review only.
Start with the count, from the primary source, not a secondary summary. The NAIC's adoption tracking map, status as of August 6, 2026, records 25 jurisdictions that have adopted the Model Bulletin on the Use of Artificial Intelligence Systems by Insurers: Alaska, Arkansas, Connecticut, Delaware, the District of Columbia, Hawaii, Illinois, Iowa, Kentucky, Maryland, Massachusetts, Michigan, Nebraska, Nevada, New Hampshire, New Jersey, North Carolina, Oklahoma, Pennsylvania, Rhode Island, Vermont, Virginia, Washington, West Virginia and Wisconsin. Four more - California, Colorado, New York and Texas - have their own insurance-specific AI regulation or guidance. That is 29 jurisdictions with something on the books.
The bulletin itself, adopted December 4, 2023, closes the most common objection to its relevance. AIS Program guideline 1.6 states that the program should address the use of AI Systems across the insurance life cycle, "including areas such as product development and design, marketing, use, underwriting, rating and pricing, case management, claim administration and payment, and fraud detection." Fraud detection is named. Claim administration is named. Any argument that the bulletin is an underwriting document is answered by its own text.
What the regulator will ask for, and why it looks familiar
Section 4 of the bulletin sets out what a regulator may request in an investigation or market conduct examination. The written AIS Program. Information about data used in the development and oversight of the specific model or AI System, including data source, provenance, data lineage, quality, integrity, bias analysis and minimization, suitability, and data currency. Documentation related to validation, testing and auditing, including evaluation of model drift. And for third-party systems: due diligence records, the contracts themselves including terms on representations, warranties, data security and privacy, data sourcing, intellectual property, confidentiality and cooperation with regulators, and any audits or confirmation processes.
Put that list next to the Advisory Committee's five factors and the resemblance is not a coincidence, it is convergence. Both ask about inputs and data provenance. Both ask about validation. Both ask what documentation exists about how the system behaves. One is a court testing evidence and the other is a state examiner testing conduct, and they arrived at nearly the same document set from opposite directions. A carrier that assembles this once for the regulator has substantially assembled it for the deposition, and vice versa. That is the strongest practical argument for treating this as one program rather than two.
Guideline 2.2 makes the timing explicit: the insurer's documentation requirements should be developed with Section 4 in mind. Document it now for the examination later. Guideline 4.1 puts the standard on the carrier, requiring due diligence to ensure that decisions made or supported by third-party AI Systems that could lead to adverse consumer outcomes will meet the legal standards imposed on the insurer itself. Guideline 4.2 then gives procurement two clauses to insist on: terms that provide audit rights or entitle the insurer to receive audit reports by qualified auditing entities, and terms that require the third party to cooperate with the insurer on regulatory inquiries and investigations related to the insurer's use of the vendor's product.
Those two clauses are the single most actionable thing in the bulletin and almost no fraud-technology contract currently contains them. Every authority in this section - the NAIC bulletin, Colorado's governance regulation, New York's circular letter - puts the production obligation on the insurer, not on the vendor. If the contract does not give the carrier audit rights and a regulator-cooperation commitment, the carrier has accepted an obligation it has no contractual mechanism to discharge. Ask for both in the next renewal, from every vendor in the stack, including this one.
Two scope corrections worth getting right
The first correction is New York, and it is misreported often enough that repeating the error in a compliance memo is a real risk. Insurance Circular Letter No. 7 (2024), issued July 11, 2024, states expressly that it "is not intended to address phases of the insurance product lifecycle other than underwriting and pricing." It does not reach claims handling. It does not reach SIU. It remains instructive as direction of travel, because it holds that insurers retain responsibility for understanding any tools developed or deployed by third-party vendors and ensuring they comply with applicable law, and that an insurer must be able to explain at all times how its AI system operates. Useful principle, wrong lifecycle phase.
The second correction is California SB 1120, covered above. Health and disability utilization review only. There is no P&C analogue in California as of today, and describing SB 1120 as a general "AI cannot deny claims" law is wrong in a way a claims legal lead will catch immediately.
Colorado is the state where the claims side is already in statute. SB21-169 defines the covered insurance practices to include marketing, underwriting, pricing, utilization management, reimbursement methodologies and claims management in the transaction of insurance, per the Division of Insurance stakeholder materials. Precision matters on what has been adopted underneath it. The governance regulation at 3 CCR 702-10, effective November 13, 2023 with amendments effective October 15, 2025, requires a documented governance and risk management framework including a documented process for selecting external resources and third-party vendors, and provides that the insurer remains responsible for meeting the requirements including production of any documents or information the Division deems necessary. Its current reach is ECDIS use in life, private passenger automobile and health benefit plans. The statute covers claims management; the Division is building out regulations line by line. Do not say Colorado regulates AI in claims. Say the statutory hook exists and the regulations are arriving.
The four words that decide this whole question
The NAIC Model Bulletin's legislative authority section names the Unfair Claims Settlement Practices Act (Model #900), which sets standards for the investigation and disposition of claims, and then adds the sentence carriers should print and hang somewhere: actions taken by insurers in the state must not violate the UTPA or the UCSPA, regardless of the methods the insurer used to determine or support its actions. Regardless of the methods. The claims-handling duty is method-neutral. It was method-neutral before AI existed and it will be method-neutral after. That cuts against a bad AI investigation and against a thin manual one with exactly equal force.
The row that gets skipped in most vendor compliance decks is the second one, and it is the row that actually decides cases. Unfair claims settlement practices statutes predate every AI rule on this list, apply in every state, and are the theory a plaintiff's lawyer will plead. The AI-specific instruments are how a regulator examines you. The claims-handling statute is how a claimant sues you.
For a carrier building against this map, the useful design property is that the Section 4 document list is satisfied by construction rather than by reconstruction. A system that writes its own record as it works can produce data lineage, per-claim inputs, tool calls, retention behavior and the human review record on request, because those artifacts are outputs of normal operation, not a special project commissioned when a letter arrives.
A thin human file fails the same duty
The duty does not distinguish between methods. California's 10 CCR 2695.7(d) requires every insurer to conduct and diligently pursue a thorough, fair and objective investigation. An unexplainable AI denial fails objective. A denial with almost no investigation behind it fails thorough. Carriers have spent two years worrying about the first while living with the second.
Carriers are being asked to prove their AI is reliable. Nobody has ever asked them to prove the manual file was. The Advisory Committee's five factors applied to both methods, including the two rows where the documented record does not win.
The regulation is worth reading in full because the three adjectives are doing separate work. 10 CCR 2695.7(d): "Every insurer shall conduct and diligently pursue a thorough, fair and objective investigation and shall not persist in seeking information not reasonably required for or material to the resolution of a claim dispute." Thorough is about coverage and depth. Fair is about even-handedness. Objective is about the basis being something other than the adjuster's or the algorithm's unexplained impression. A file can fail any one of them independently.
California then specifies what an SIU investigation has to contain, which almost nobody outside California reads and everybody should. 10 CCR 2698.36(a) provides that an investigation shall include a thorough analysis of the claim file, application or transaction with consideration of the factors indicating fraud; identification and interviews of potential witnesses; use of one or more industry-recognized databases identified by the SIU as appropriate for the line of insurance in question; preservation of documents and other evidence obtained during the investigation; and a concise and complete summary of the entire investigation answering five specified questions and stating whether or not the investigation is complete.
That is a report schema written by a regulator in 2011-era language, and it maps almost one to one onto what a modern investigation agent produces: analysis of the file against fraud indicators, witness identification, database checks, evidence preservation, and a structured summary. The regulation does not care whether a person or a system performed each step. It cares that each step happened and that the file shows it.
Then comes subsection (c), which is the coverage-gap regulation and the one that should change how carriers think about detection tooling. The SIU shall investigate each credible referral of suspected insurance fraud it receives from integral anti-fraud personnel, including automated or system-generated referrals, and a credible referral is one that includes a red flag or red flags. The SIU need not open an investigation only where a preliminary review makes it reasonably clear that the red flags are not the result of suspected insurance fraud, and the reasons for that conclusion are documented in the claim file or the SIU investigation file.
Read that against a modern detection deployment. A scoring engine produces automated referrals at volume. Each one that carries a red flag is a credible referral, and a credible referral creates a duty. The only lawful way to close one without a full investigation is a documented preliminary review. So the more sensitive the detection layer, the larger the documentation obligation it generates downstream. Better detection does not reduce exposure. It manufactures exposure, and the exposure is discharged by investigation and documentation, not by a lower score threshold.
This is where the operating numbers stop being an efficiency story and become a compliance story. On Hesper internal benchmarks, manual SIU investigation runs 14+ days per case and a single investigator carries 200+ cases. The arithmetic of those two numbers is the reason carriers fully investigate roughly 25% of flagged claims, with the remainder paid, denied without full work, or sitting in a queue. Every one of those closures is a file that, under 2698.36(c), needs a documented basis. Most of them have a disposition code. A disposition code is not a documented basis. We walked through what the full record looks like in our piece on automated claims investigation.
The pool this is drawn from is not marginal. The Coalition Against Insurance Fraud puts US insurance fraud at $308 billion a year and reports that fraud occurs in roughly 10% of property and casualty losses. The 75% of flagged claims that never get a full investigation are not a rounding error in that figure. They are where most of it lives.
That is the symmetry, and it is the argument this post exists to make. Apply factor 3 to a human investigation: what records does the process keep, what is deleted, can the output be reproduced or audited. The honest answer for most manual SIU files is a completion date, a disposition code, and some notes. Apply factor 5: did the opponent receive information analogous to what a human expert would disclose about how and why the conclusion was reached. Usually not, because it was never captured. The human file passes today only because nobody thinks to apply the test to it.
The remedy that follows is not less AI. It is AI that keeps a record, plus a human whose judgment is documented, not assumed. That is a specific architecture: built-in detection so the referral itself is explainable instead of inherited as an opaque number, 15+ investigation phases run in parallel with each writing its own sourced and timestamped record, and a named SIU lead who reviews the assembled evidence and signs the conclusion. The coverage effect is the part that matters to a claims VP: every flagged claim gets a full documented investigation, not one in four, in hours rather than weeks. From fraud detection to fraud resolution. The investigator's role shifts from execution to decision-making, which is also, precisely, what the defense bar and the Advisory Committee are both asking for.
One severity note, clearly labeled as not an AI case. In December 2023 the South Carolina Court of Appeals affirmed an award of $27,339,535 against Penn National on five policies providing $500,000 each, more than ten times the carrier's maximum exposure had it defended the insured, with the court describing the conduct as exactly the type South Carolina bad faith law seeks to deter. That was a duty-to-defend dispute with no algorithm anywhere near it. It is included for one reason: extracontractual exposure is not proportional to the claim, and the thing that produces it is a file that cannot be defended, whatever built it.
What an audit trail answers, and what it does not
An audit-ready evidence chain fully answers two of the Advisory Committee's five factors, partly answers two more, and does not answer the fifth at all. It settles what records exist and whether the output reproduces. It does not establish independent evaluation, general acceptance in a field, or validation on your own book of business.
Any vendor who tells you their audit trail closes this question is selling. Here is the honest accounting, laid out factor by factor across the three kinds of AI output a carrier is likely to have in the building. The third column is not all green, and that is the point of the table.
Take the gaps one at a time, because each has a different owner.
General acceptance in the relevant field
Frye asks whether a methodology is generally accepted in the relevant scientific community. Weber's expert lost on exactly this: adamant that AI-assisted report drafting was accepted in fiduciary services, unable to name a single publication saying so. There is no equivalent literature for AI-assisted SIU investigation, no peer-reviewed body of work, and no settled answer to what the relevant field even is. Puloka shows how much that definition matters, since the court there chose forensic video analysis rather than computer vision. An audit trail does not touch this. Time, published evaluation work and professional-body engagement do.
Independent evaluation and opponent access
Factor 4 asks specifically about evaluation by entities independent of the developer, and about whether the opponent and independent evaluators can access the system, giving meaningfully available research licenses as the example. That is a vendor obligation, not a carrier one, and almost no fraud-technology vendor in this market currently meets it. It is a fair thing to ask for in a procurement process and a fair thing to hold vendors to over time, including us.
Validation on your book, and a documented error rate
Factor 2 asks whether the process was validated in circumstances sufficiently similar to the case at hand. A vendor's benchmark on somebody else's data is not that. Validation on your lines, your geography and your claim mix is carrier-side work: a sampled human review program, a measured override rate, and a written record of when the validation ran and what it found. Nobody can hand you this. It is also the item most likely to be missing when a 30(b)(6) witness is asked how you know the system works on your claims.
Bias and disparate impact
An audit log records what a system did. It does not tell you whether what it did fell unevenly across protected classes. Huskey is pleaded as Fair Housing Act disparate impact precisely because that theory does not require proving intent or opening the model, only showing an effect. Bias analysis is a separate discipline with separate artifacts, and the NAIC bulletin's Section 4 asks for it by name under data provenance and bias analysis and minimization. Do not let a strong audit trail create the impression this box is ticked.
Whether the human actually exercised judgment
This is the uncomfortable one. A trail showing that a reviewer opened a 40-page finding and clicked approve four seconds later is worse than no trail, because it documents the absence of the judgment the whole structure depends on. Human-in-the-loop is a claim about behavior, and an audit trail is very good at testing claims about behavior. Design the review step so that it takes real time and produces real artifacts: what the reviewer changed, what they questioned, what they added. Measure override rates and be able to produce examples. If the override rate is zero, the review is decorative and a deposition will establish that in about ten minutes.
The summary position is that an audit trail is necessary and insufficient. It converts the questions that were unanswerable into questions that are answerable, which is a large gain, and it leaves a defined remainder that is carrier-side and vendor-side work. We wrote separately about what makes a fraud report audit-ready in the first place; treat that as the floor rather than the finish line.
There is also a limit worth naming that has nothing to do with the model. The Advisory Committee's suggested limiting instruction - that machine-generated evidence is subject to error and should not be assumed reliable simply because a machine produced it - is available to a court even where everything else goes your way. No amount of logging prevents it. What reduces its sting is a file whose conclusions were authored by a named human being who can take the stand and explain them, supported by machine-assembled evidence, rather than the reverse.
What to answer before you deny on AI-assisted findings
Nine questions, derived from the Advisory Committee's own factor list, the NAIC bulletin's Section 4 examination requests, and California's SIU regulations. No firm publishes this list; it is assembled here from primary sources. If a claims legal lead cannot answer all nine about a specific file today, the exposure is the record, not the AI.
- Name the human who made the decision, and state what they reviewed before making it. Not the team. The person, and the artifacts in front of them.
- Reproduce the finding from the stored record. Inputs, sources, timestamps, and the same conclusion. If two runs give two answers and you cannot explain why, you have Weber's problem.
- State what the system retains and what it deletes, and for how long. This is Rule 707 factor 3 verbatim and it is a retention policy, not a technical question.
- Produce the tool-call log for this specific claim. Which checks ran, in what order, what each returned, and what was concluded from the combination.
- Produce the written AI governance program and point to the paragraph that covers claims and SIU use. If the program describes underwriting only, the claims deployment is ungoverned on its face.
- Produce the vendor contract clauses on audit rights and regulator cooperation. NAIC guideline 4.2 asks for both. Most contracts have neither.
- State what validation was performed on your own book of business, when it ran, and what it found. A vendor benchmark on other data does not answer factor 2.
- State the override rate and produce examples of overrides. A zero override rate is evidence that the human review is not functioning, not evidence that the system is perfect.
- State whether the finding was disclosed as AI-assisted, to whom, and when. Weber imposed an affirmative disclosure duty on counsel; the direction of travel on notice is one way, and revised Rule 707(c) would require pre-trial notice.
Hand items 1 through 4 and 7 through 9 to the SIU and claims operations. Hand items 5 and 6 to compliance and procurement. They can be worked in parallel and most of them are answerable within a quarter, which is a useful contrast with the 90-day governance review that only 24% of insurance executives are very confident of passing.
The closing argument is the one this post opened with. The defense bar's checklist is not a case against AI in claims investigation. It is a specification, and it happens to be a specification that a manual file has never met and cannot meet, because a person working alone under a 200+ case load does not produce a reproducible record of their reasoning. Build the AI investigation to the checklist and you end up with something a manual process never gave you: a file where every source is named, every step is timestamped, every conclusion traces to evidence, and a human being's judgment is recorded rather than inferred.
The economics that made partial coverage rational are also gone, which is the part that turns this from a legal argument into an operating decision. When a full documented investigation was expensive and slow, investigating one flagged claim in four was a defensible allocation of scarce capacity. It is harder to describe that way once the same investigation can be run on every flagged claim, in hours rather than weeks, with the record written as it goes. The question a regulator or a plaintiff's lawyer will eventually ask is not why you used AI. It is why the other three files have nothing in them.
Key takeaways
- The unfair-claims-practices duty is method-neutral: the NAIC's AI Model Bulletin states that claims-handling standards apply regardless of the methods the insurer used, so an unexplainable AI denial and a denial with no real investigation behind it fail the same duty under rules like California's 10 CCR 2695.7(d).
- Proposed Federal Rule of Evidence 707 was not adopted; the Advisory Committee on Evidence Rules declined to recommend action in its May 17, 2026 report and sent a revised draft to its Fall 2026 meeting, but its five-factor reliability analysis is the clearest published specification of how a court will test AI output, and factor 3 is an audit-log requirement in judicial language.
- Case law splits on explanation rather than technology: proprietary probabilistic-genotyping software was admitted in People v. Wakefield after experts explained the method and the source code was never produced, while in Matter of Weber an expert who could not recall his prompt was excluded, partly because the same query returned three different numbers on three court computers.
- Discovery is the real exposure and the split is now visible: Estate of Lokken v. UnitedHealth Group protected the source code and underlying data while ordering production of policies, procedures, training, governance minutes and internal analyses, and Huskey v. State Farm opened a discovery phase aimed specifically at a homeowners carrier's fraud-screening algorithms.
- An audit trail is necessary and insufficient: it answers what records exist, whether the output reproduces, and what the opponent gets, but it does not establish general acceptance in a field, independent third-party evaluation, validation on your own book, bias testing, or that the human reviewer actually exercised judgment.