Hesper AI
ITC Vegas 2026We're on the floor at Mandalay Bay, Sep 29 - Oct 1Meet us there
BlogGuides
GuidesSeptember 21, 2026·27 min read·Nitish Badu, COO

Human-in-the-loop AI fraud investigation: which decisions stay human, and what they get handed

The objection to autonomous claims AI is an objection to opacity. Separate the axes and the question becomes decision rights: which decisions stay human, and what evidence they are handed.

NB
Nitish Badu · COO and Co-founder
September 21, 2026·27 min read
GUIDESHesper AI7%Correct expert calls overturnedto follow wrong AI adviceROSBACH ET AL. 2024, 28 PATHOLOGY EXPERTS, ARXIV:2411.00998
The numbers behind this
(iii) of 5Where human involvement sits in the NAIC risk-calibration testNAIC Model Bulletin on AI Systems, Section 3, adopted December 4, 2023
3 of 3PA DOI claims violations that were missing status lettersPennsylvania Insurance Department market conduct exam, docket MC25-05-001
~25% -> 100%Flagged-claim investigation coverage, manual vs HesperHesper internal benchmark

Twenty-eight trained pathology experts were given an AI assistant that was sometimes wrong. Seven percent of the time they overturned an evaluation they had already gotten right and went with the machine. Nobody had skipped the human in that experiment. The human was in the seat, looked at the case, and turned a correct answer into a wrong one.

Keep a human in the loop is the remedy the insurance trade press prescribes for AI risk in claims, and that finding is the hole in it. A human in the loop names a seat. Whether the seat is a control depends on which decisions are actually reserved to it and what evidence it gets handed. Neither of those is a question about how much of the work the machine did.

The published case against autonomous claims AI is a case against opaque scoring, and the two got welded together by accident. Autonomy describes how much of the work a machine does end to end. Opacity describes whether anyone can reconstruct how it got there. Those are separate dials. The industry has spent two years arguing about the first one while the defect everyone is actually describing sits on the second.

Take the clearest recent statement of the objection. Tom Rasmussen, VP of Product for Claims at Carpe Data, writing in Digital Insurance on August 4, 2026, says black-box scores "don't provide an evidence trail explaining what triggers concern," and that "A fraud score alone is not defensible when AI is being used to fuel claims fraud." Both sentences are correct and we sign both of them. He goes on to argue that "AI should support claims teams, not replace them," because "Humans still need to evaluate the full context of the claim, make judgment calls, and ensure decisions are fair and consistent." Read the piece closely and what is absent from it is any claim that autonomy is impossible. What is present is a claim about judgment, and judgment is a much narrower category than work.

The disagreement that remains is narrow. The remedy for a missing evidence trail is to produce one. Capping autonomy does not produce an evidence trail. It produces less work done, and it relocates the reasoning back into a person's head, which is where the three-line file note problem started. An agent that logs every tool call, query, source, timestamp and inference hands a reviewer a denser record than a human-run claim file usually contains.

So the design question is not how much human. It is which decisions are reserved to a human and on what legal basis, and what that human is handed, dense enough that the review is a control instead of a signature block. The rest of this post works through the NAIC's five-factor sliding scale, the scope of every United States rule that names a required human decider, a decision-rights taxonomy of eleven steps across one flagged claim, and what the automation-bias literature says a checkpoint has to contain before it does any work. The other half of the two-axis argument, why a vendor autonomy ladder cannot describe evidence depth, is made separately in ARISE levels of autonomy vs investigation-specific AI. This post leaves the ladder alone and works on decision rights.

Autonomy and opacity are two different axes

Autonomy and opacity are independent properties of a claims AI system. Autonomy measures how much of an investigation runs end to end without a human step. Opacity measures how little of that work the system records and can explain afterwards. A low-autonomy black-box score is opaque. A high-autonomy logged agent is not.

The NAIC already separated them. Its Model Bulletin on the Use of Artificial Intelligence Systems by Insurers, adopted by the Executive Committee and Plenary on December 4, 2023, defines AI Systems as "designed to operate with varying levels of autonomy." That is a spectrum written into the definition section. Section 3 then lists the extent of human involvement in final decision-making at (iii) and the transparency and explainability of outcomes at (iv) as two separate factors out of five. Two factors, listed independently, in the governing document 25 jurisdictions have now adopted.

Cross the two dials and four configurations fall out. A well-kept adjuster file is manual and legible, which is what the claims regulations assume exists. A phone call to a body shop that produces a view and never produces a note is manual and opaque, and a market conduct examiner cannot read a memory. A bare score with a reason code and nothing underneath either is autonomous and opaque, and that quadrant is exactly what the objection describes. A logged investigation that writes down every source queried, every source discarded and why, every inference and every timestamp is autonomous and legible. It is also the only quadrant where the quality of the record improves as volume goes up, because the record is produced by the work, not recalled after it.

The objection to autonomous claims AI is an objection to opacity. The industry has spent two years arguing about the wrong axis. Autonomy runs left to right, legibility of the record runs bottom to top, and they move independently. Axis separation follows the NAIC Model Bulletin on the Use of Artificial Intelligence Systems by Insurers, adopted December 4, 2023, Section 3, which lists the extent of human involvement in final decision-making at (iii) and the transparency and explainability of outcomes at (iv) as two separate factors of five. Caseload and coverage figures are Hesper internal benchmarks.

The conflation persists for a structural reason, not a rhetorical one. The dominant artifact of the detection layer is a score, so for most carriers the only machine output that has ever touched a claim file is a number with a reason code attached. Generalizing from that sample gives you the view that machine output is inherently thin. It is a fair generalization from the evidence most claims organizations have. It is also the wrong axis: what makes a score indefensible is that nothing sits behind it, not that it arrived without a human pressing a button.

What the NAIC bulletin says about human involvement

The NAIC Model Bulletin does not mandate human review of any individual claims decision. Section 3 makes the extent of human involvement in final decision-making factor (iii) of five that calibrate how much control an insurer needs, sitting directly beside factor (iv), the transparency and explainability of outcomes. It is a sliding scale, not a floor.

The controls and processes that an Insurer adopts and implements as part of its AIS Program should be reflective of, and commensurate with, the Insurer's own assessment of the degree and nature of risk posed to consumers by the AI Systems that it uses, considering: (i) the nature of the decisions being made, informed, or supported using the AI System; (ii) the type and Degree of Potential Harm to Consumers resulting from the use of AI Systems; (iii) the extent to which humans are involved in the final decision-making process; (iv) the transparency and explainability of outcomes to the impacted consumer; and (v) the extent and scope of the insurer's use or reliance on data, Predictive Models, and AI Systems from third parties.

NAIC Model Bulletin: Use of Artificial Intelligence Systems by Insurers, Section 3, adopted December 4, 2023

Read that as a formula, not a list. Five inputs set the required level of control. Less human involvement at (iii) can be offset by more transparency and explainability at (iv). Nothing in the sentence makes (iii) a minimum. The word doing the work is commensurate, and what it is commensurate with is the insurer's own assessment of consumer risk across all five factors at once. A carrier that reduces the human touch on evidence-gathering steps and simultaneously raises the transparency of the record is doing exactly what the paragraph contemplates.

Claims and fraud are in scope, which matters because the bulletin is sometimes read as an underwriting document. Section 1.6 says the AIS Program should address AI use across the insurance life cycle, including case management, claim administration and payment, and fraud detection. The bulletin's legislative authority rests partly on the Unfair Claims Settlement Practices Model Act, number 900, which sets forth standards for the investigation and disposition of claims. An SIU running AI on a flagged claim is inside the bulletin, not outside it.

Now look at what the bulletin actually requires, because every named control is a record control. Section 3.1 calls for the identification of constraints and controls on automation and design to align and balance function with risk. Section 3.3(c) calls for assessments such as interpretability, repeatability, robustness, regular tuning, reproducibility, traceability, model drift, and the auditability of these measurements where appropriate. Section 3.6 covers data and record retention. Traceability, reproducibility, auditability, retention. Those are properties a logged autonomous investigation has by construction and a human-run file has only if somebody typed them. Our walkthrough of the NAIC AI risk evaluation supplement covers what the filing side of this looks like in practice.

Adoption is broad enough that this is the operative national standard. The NAIC's own implementation map, current as of April 1, 2026, records 25 jurisdictions that have adopted the model bulletin and four more - California, Colorado, New York and Texas - that have insurance-specific AI regulation or guidance instead. Twenty-nine jurisdictions with something on the books. Not one of them requires a human to review each claims decision.

The rules that name a required human are health rules

Every United States rule that reserves an insurance claims decision to a licensed human is a health rule. Colorado 3 CCR 702-10 Section 5.A.5 covers health benefit plan authorization. California SB 1120 covers medical necessity. New York DFS Circular Letter No. 7 covers underwriting and pricing only. None of the three reaches property and casualty claims investigation.

Start with Colorado, because it is the one most often cited loosely. 3 CCR 702-10, Regulation 10-1-1 took effect November 13, 2023, with amendments effective October 15, 2025. Section 3 scopes it to individually issued life insurance, private passenger automobile insurance, and health benefit plans, and it governs external consumer data and information sources and the models built on them. It is a governance framework requirement: risk management, documentation, testing, reporting. It is not a per-decision human review requirement, and it is not about claims investigation.

The one provision that names a required human decider is Section 5.A.5, and it reads as utilization review. Health benefit plan insurers must ensure that a provider acting on behalf of the insurer is ultimately responsible for the decisions made when external consumer data, or algorithms or predictive models that use it, are used to inform decisions to modify or deny requests for authorization prior to or concurrent with the provision of health care services. A provider. Prior authorization. Health benefit plans. Nothing in that sentence touches a property claim or an SIU referral.

Colorado does carry one obligation worth borrowing, and it points the same way this post does. Section 5.A.7 requires documented processes that provide the applicant, policyholder, beneficiary or covered person with information necessary to take meaningful action in the event of an adverse decision. That is an evidence-handover obligation. The regulator is specifying what the affected person receives, not how autonomous the system that produced it was.

California SB 1120, chaptered September 28, 2024, is the cleanest example of a statutory human-decider mandate in United States insurance law: "A determination of medical necessity shall be made only by a licensed physician or licensed health care professional competent to evaluate the specific clinical issues involved in the health care services requested by the provider." It amends Health and Safety Code section 1367.01 and Insurance Code section 10123.135, which govern health care service plans and disability insurers. It is worth quoting precisely because of how narrowly it was drafted. The legislature knew how to write "only a licensed human may decide," wrote it once, and scoped it to utilization review.

New York DFS Insurance Circular Letter No. 7 (2024), issued July 11, 2024, is the other citation that travels further than its text allows. Its governance section requires board and senior management oversight of AI use at the insurer, plus comprehensive documentation. It contains no per-decision human review requirement. And its scope paragraph forecloses the claims reading directly: "This Circular Letter also is not intended to address phases of the insurance product lifecycle other than underwriting and pricing." If a trade piece cites CL 7 as a claims-AI human-review mandate, that citation is wrong.

The pattern holds outside the United States too. The EU AI Act's human oversight obligations attach only to high-risk systems, and Annex III point 5(c) limits the insurance category to "AI systems intended to be used for risk assessment and pricing in relation to natural persons in the case of life and health insurance." Property and casualty claims handling is not on the high-risk list. Four separate drafting bodies - Colorado, the California legislature, New York DFS and the European Parliament - wrote a version of the reserved-human sentence. Every one of them scoped it to health, or to underwriting and pricing. None wrote it for property and casualty claims investigation.

What this post does not claim

It does not claim any of these rules mandates human-in-the-loop review for property and casualty SIU investigation, because none of them does. Colorado 3 CCR 702-10 Section 5.A.5 and California SB 1120 are utilization review. NY DFS Circular Letter No. 7 disclaims anything beyond underwriting and pricing in its own scope paragraph. EU AI Act Article 14 does not attach to P&C claims handling because Annex III point 5(c) does not list it, and the Act does not bind United States carriers in any case. The argument here runs the other way: since no statute sets the human floor for investigation work, a carrier has to set it deliberately, and the place to set it is at the decision acts that other rules do reserve.

P&C claims regulation regulates the record, not the reasoner

United States property and casualty claims regulation regulates the record and the clock, not the reasoner. California 10 CCR 2695.7 sets a 40-day accept-or-deny deadline and requires a written denial listing every factual and legal basis. 31 Pa. Code 146.6 sets a 30-day investigation deadline. Neither names who performed the reasoning.

California 10 CCR 2695.7, the Fair Claims Settlement Practices regulation, is the load-bearing example. Subsection (b) requires the insurer to accept or deny a claim within 40 calendar days of receiving proof of claim. Subsection (b)(1) requires that a denial be in writing and contain "a statement listing all bases for such rejection or denial and the factual and legal bases for each reason given." Read that requirement as a specification and it is a demand for sourcing. Every reason has to carry its factual basis. An agent that produces a source-anchored factual basis for each reason given is serving 2695.7(b)(1) more directly than a file note reading "denied per SIU."

Subsection (d) is the one that decides the argument. Every insurer shall conduct and diligently pursue a thorough, fair and objective investigation, and shall not persist in seeking information not reasonably required for or material to the resolution of a claim dispute. Three adjectives, all of them about the investigation itself. The duty runs to the insurer as an entity. The regulation is silent on the substrate: it does not say a licensed adjuster must have performed the searches, does not say a human must have read each document, and does not say which individual signs. What it regulates is completeness.

10 CCR 2698.36 is the closest thing in United States property and casualty to a decision reserved to a named function. The SIU must investigate each credible referral containing red flags, and the required investigation elements are specific: a thorough analysis of a claim file, application or insurance transaction, identification and interviews of potential witnesses, use of industry-recognized databases appropriate to the line, preservation of documents and other evidence, and a concise and complete written summary of the entire investigation. If the SIU declines to open an investigation, it must document in the claim file or SIU investigation file the reasons supporting its conclusion. Who decides is the SIU. What the SIU owes is a documented reason. Our 10 CCR 2698 compliance walkthrough goes element by element.

California Insurance Code 1872.4(a) adds a clock with an unusual starting gun. A company that has determined, after the completion of the insurer's special investigative unit investigation, that it reasonably suspects or knows an act of insurance fraud may have occurred, has 60 days after that determination to send the file to the Fraud Division. Two things follow. The determination is an insurer act, not a machine act. And because the 60-day clock starts only after the SIU investigation completes, throughput at the investigation step sets how much of the statutory window is left for everything after it. That second point is a derivation from the statute, not an industry finding.

Pennsylvania supplies the tightest clock of the four. 31 Pa. Code 146.6 requires every insurer to complete investigation of a claim within 30 days after notification, unless the investigation cannot reasonably be completed within 30 days, and every 45 days thereafter the insurer must provide the claimant a reasonable written explanation for the delay and state when a decision may be expected. Set that against our internal benchmark of 14+ days per manual SIU case and the arithmetic is uncomfortable: a single manual investigation consumes roughly half the regulatory investigation window before any adjudication step begins. That is a derivation from one regulation and one Hesper internal benchmark, not an industry statistic.

The density test: an agent's log against a file note

The density test asks which artifact a market conduct examiner can actually read. A logged autonomous investigation records every source queried, every source discarded, every timestamp and every inference as the work happens. A manual claim file records what an adjuster remembered to type, and the reasoning behind it usually never gets written down.

There is a public answer to what an examiner inspects, and it comes from a state DOI examining an AI-forward carrier. The Pennsylvania Insurance Department's market conduct examination of Lemonade Insurance Company, docket MC25-05-001, covered the experience period January 1 through December 31, 2023. The report issued April 1, 2025 and the consent order followed on May 22, 2025. The claims section found three violations. All three were of 31 Pa. Code 146.6, and all three were the same failure: the company did not provide a 30/45-day status letter. One homeowner file out of 40 sampled from a universe of 80. Two tenant homeowner files out of 75 sampled from a universe of 921. Zero violations across all 23 condominium claims reviewed. A 3% error ratio on each sampled population.

Note what is absent from the findings. Nothing about whether a human or a machine reasoned about the claim. Nothing about model logic, thresholds or scores. Three findings about a letter that did not go out on time. In the same experience period, 24 of 31 consumer complaints, or 77%, were claims related, and the complaint handling itself drew no violation. The compliance surface a market conduct examiner touches is documentation and timeliness. That is the empirical answer to the question of what a regulator actually inspects, and it is an answer about the record.

What an examiner asks forManual claim fileLogged autonomous investigation
Every source consultedWhatever the adjuster noted. No enumerated list existsEnumerated, with the query that produced each one
Sources checked that returned nothingAlmost never recordedLogged as negative results, which is what 2695.7(d) thoroughness looks like
TimestampsDiary entries and letter datesPer step, written as the work happens, not recalled afterwards
Reasoning behind the conclusionA file note, often three linesThe inference chain, each step anchored to a source
Contradictions across statements and documentsFound if noticed, recorded if notedMapped to document, page and line
Reproducibility of the findingNot testable after the factThe run replays from the logged inputs
What the reviewer changedRarely capturedEdits and overrides logged as first-class evidence
Attribution of each stepAt the file levelAt the step level

Our own configuration runs 15+ investigation phases in parallel on a flagged claim, each phase logging its sources, its queries and its timestamps as it goes. That is a Hesper internal benchmark, not an industry figure, and the number itself is less interesting than the property it creates: the file is a byproduct of the work instead of a task competing with the work. An adjuster documenting thoroughly is spending minutes that could have gone into the investigation. An agent documenting thoroughly is spending nothing, because writing the log is how it holds state.

This is the same argument that decides admissibility, and it cuts symmetrically. A carrier denying on a black-box output nobody can explain and a carrier denying on a thin twenty-minute file review are failing the same duty. We worked that through in AI fraud investigation and court admissibility, where the finding is that the manual claim file fails the reliability test the defense bar is aiming at AI. The density test is the market conduct version of the same point.

A human in the loop is not automatically a control

A human review step is an input specification, not a control. It works only when the decision is genuinely reserved, the reviewer holds independent evidence to disagree with, and accountability attaches to the review itself. The EU AI Act names the failure mode, automation bias, inside the same article that mandates oversight.

Article 14(1) requires that high-risk AI systems be designed so that they can be effectively overseen by natural persons during the period in which they are in use, including through appropriate human-machine interface tools. Then paragraph 4(b) tells the deployer what the overseer has to be protected against.

to remain aware of the possible tendency of automatically relying or over-relying on the output produced by a high-risk AI system (automation bias), in particular for high-risk AI systems used to provide information or recommendations for decisions to be taken by natural persons

EU AI Act, Article 14(4)(b)

The lawmakers who mandated the human also wrote down why the human might not work, in the same article, two paragraphs later. That is the most candid statement of the problem available in any statute. The scope caveat is mandatory and stated plainly: the Act does not bind United States carriers, and Annex III point 5(c) confines the insurance high-risk category to risk assessment and pricing for life and health insurance, so Article 14 does not attach to property and casualty claims handling by that route. Cite it as a drafting reference. Do not cite it as an obligation.

The effect has been measured in expert hands. Rosbach and colleagues (2024), working at Technische Hochschule Ingolstadt and Flensburg University of Applied Sciences, ran 28 trained pathology experts through an AI-assisted decision task. AI assistance improved overall accuracy, and the paper reports the cost alongside the benefit: "it also resulted in a 7% automation bias rate, where initially correct evaluations were overturned by erroneous AI advice." These were domain experts who had already reached the right answer and reversed themselves because the machine disagreed. The discipline is computational pathology and the paper is a preprint, so read it as a mechanism finding rather than an insurance statistic. The mechanism is what transfers.

Time pressure is the aggravating condition, and the same paper measured it. Under time pressure the authors found that pressure did not exacerbate the occurrence of automation bias but appeared to increase its severity, evidenced by heightened reliance on the system's negative consultations and a subsequent performance decline. Now map that onto an SIU desk. Our internal benchmark puts a single investigator at 200+ open cases. That is the textbook time-pressured reviewer, and a review checkpoint designed without accounting for it is decorative. The reviewer under that load is not adjudicating the machine's reasoning. They are deciding how much of the day one file gets.

This is not one finding. Goddard, Roudsari and Wyatt's systematic review in the Journal of the American Medical Informatics Association (2012;19(1):121-127) screened 13,821 papers and included 74 that met inclusion criteria. They grouped the mediators into user factors, attitudinal factors and environmental factors, with workload and task complexity sitting in the environmental group. They grouped the mitigators into implementation factors, naming training and emphasis on user accountability, and decision-support design factors such as where advice sits on the screen, whether confidence levels are updated, and whether the system supplies information or a recommendation. Those are the design brief for the checkpoint section below.

The industry is already running the shallow version of the checkpoint, and it has the data on itself. The Verisk State of Insurance Fraud Study, published March 2026, ran two surveys - 1,000 United States consumers aged 18 and over, and 300 United States insurance claims professionals at manager level or above at P&C, life and reinsurance companies plus TPAs - with data collected December 2025 through January 2026.

The human checkpoint as the industry currently runs it (Verisk, March 2026)

Insurers relying on manual review in their digital-media-fraud defense44%
Of those, requiring manual review for all submissions50%
Of those, applying it only to suspicious or high-risk submissions42%
Insurers who believe digital media fraud goes undetected often or very often66%

Read the four bars together. Manual review is part of the defense at 44% of insurers. Inside that group, 42% apply it only to submissions already flagged as suspicious or high-risk, which means the checkpoint is conditioned on the very triage it is meant to backstop. And 66% of insurers think digital media fraud goes undetected often or very often. A human checkpoint applied to a subset of a subset, by reviewers carrying a full caseload, against a threat the same respondents say is getting past them, is not the control the objection assumes it is. Keeping a human in the loop is a statement about inputs. Whether the control works is a separate empirical question, and the industry's own survey answers it unfavorably.

Decision rights across one flagged claim

Decision rights split a flagged claim into two blocks. Evidence-gathering steps, which no United States property and casualty statute reserves to a human, and decision acts, which specific rules reserve to the insurer or its special investigation unit with a documentation duty attached. Autonomy belongs in the first block. Judgment belongs in the second.

That split is the whole operating model, and it is worth stating as a sentence a claims VP can repeat in a governance meeting: the investigator's role shifts from execution to decision-making. Nobody is removed from any decision the law reserves to a person. What gets removed is the querying, cross-referencing, document review and file assembly that currently sit between a referral and a decision, and which are the reason most referrals never reach a decision at all. Our internal benchmark for that stretch of work is 14+ days per manual SIU case.

Decision or stepWho must decideLegal basisWhat the agent hands the human
Order public-records, OSINT and database checksNo reserved human. Evidence-gatheringNone. 10 CCR 2698.36 requires use of industry-recognized databases and does not reserve who runs the queryEvery query run, source queried, retrieval timestamp and result, including the nulls
Reconstruct the claim timelineNo reserved humanNoneEvent-by-event timeline with a source citation per event
Detect inconsistencies across statements and documentsNo reserved humanNoneContradiction map anchored to document, page and line
Analyze documents for alteration or fabricationNo reserved humanNonePer-document findings with the specific artifact relied on
Open or decline an SIU investigation on a flagged referralThe SIU10 CCR 2698.36. The SIU must investigate each credible referral carrying red flags, and must document the reasons in the claim file or SIU file if it declinesRed-flag findings plus a drafted documented rationale the SIU edits and signs
Complete the investigation and issue the status letterThe insurer, via the claim representative31 Pa. Code 146.6. Thirty days to complete, written explanation every 45 days thereafter. The PA DOI cited exactly this three times in the Lemonade examCompleted investigation plus a drafted written explanation. The human issues it
Accept or deny the claim in whole or in partThe insurer, in writing10 CCR 2695.7(b), 40 days. 2695.7(b)(1) requires all bases and the factual and legal basis for each reason given. No individual is namedEach factual basis with its source. The human applies policy language and decides
Demand an examination under oathThe insurer, under the policy conditionContractual. A policy EUO provision, not a general statuteThe specific unresolved contradictions that would justify the demand
Determine reasonable suspicion of fraud and refer to the state fraud bureauThe insurerCal. Ins. Code 1872.4(a). The determination follows completion of the SIU investigation, then 60 days to fileThe completed investigation package mapped to the prescribed form fields
Set or change reserves, exercise settlement authorityThe insurer, per its authority matrixInternal governance, not statuteExposure analysis and the evidence behind it
Conduct a thorough, fair and objective investigationThe insurer as an entity10 CCR 2695.7(d). Silent on the substrate. The duty is completeness, not who performed itThe full trail, including what was checked and came back clean

The shape of the table is the argument. The first four rows are evidence-gathering with no reserved human anywhere in the legal basis column. The next six are decision acts, each with a named decider and a documentation duty. The last row is the duty that covers the whole run and says nothing about who performed it. Autonomy belongs entirely in the top block. Judgment belongs entirely in the bottom block. And the bottom block gets better the denser the top block's output is, which is the part the one-axis debate has no way of expressing.

There is a public, regulator-accepted instance of exactly this split, filed by an AI-forward carrier under examination. Recommendation 3 of the Pennsylvania exam directed Lemonade to review and revise internal control procedures to ensure compliance with the claims handling requirements of 31 Pa. Code Chapter 146.6. The company's written response, signed by its head of compliance and attached to the report as Exhibit A, describes what it built.

On December 20, 2023, the Claims team implemented an automated workflow to ensure compliance with the claims handling requirements of 31 Pa. Code, Chapter 146.6, Unfair Claims Settlement Practices moving forward. This workflow generates a task for the claim representative six days before the required date of the status letter (i.e. 24 days, 54 days, 84 days, etc). The task is completed once the claim representative issues the status letter to the appropriate parties.

Lemonade Insurance Company, Exhibit A response, Pennsylvania Insurance Department market conduct examination, docket MC25-05-001

Read that as a decision-rights design, not a compliance note. A carrier caught on a claims-handling failure automated the tracking and the evidence production, and reserved the issuing act to a named human claim representative. Six days of buffer at 24, 54 and 84 days. The machine owns the clock and the artifact; the person owns the signature. That is the split this post argues for, executed in public, in a state DOI filing, and accepted by the regulator. It is also a useful reminder that autonomy is not a single switch but a ladder, which is the argument we run against a vendor taxonomy in the ARISE efficiency-curve analysis.

What the human should be handed at the checkpoint

The checkpoint is real only if the reviewer is handed enough independent evidence to disagree with the finding. That means the finding itself, every fact with its source and retrieval timestamp, the queries that were run, the negative results, the contradictions anchored to document and page, and a stated confidence.

A recommendation is not a packet. Handing a time-pressured reviewer a conclusion and a confidence score reproduces the exact configuration Article 14(4)(b) warns about: information or recommendations for decisions to be taken by natural persons, with nothing independent for the person to check them against. The packet below is specified so that a reviewer can reach the opposite conclusion from the same materials.

  • The finding, stated as a claim about the world, not a score. Not a 0.86; a sentence that can be true or false, with the evidence that makes it one or the other.
  • Every fact with its source and retrieval timestamp. A date of retrieval is not decoration. It is what lets a reviewer eighteen months later know whether the record they are reading is the record that existed when the decision was made.
  • The tool calls and queries that ran, verbatim. The reviewer needs to see the question that was asked, not just the answer that came back, because a wrong question produces a clean-looking answer.
  • The negative results. What was checked and came back clean, listed with the same prominence as the hits. This is the part almost no system surfaces and it is the part that carries the most regulatory weight.
  • Contradictions anchored to document, page and line, with both sides shown. A contradiction summarized is a conclusion. A contradiction anchored is evidence.
  • The stated confidence and what would change it. A named condition - one more document, one interview, one database that was unavailable - turns a number into an instruction.
  • A drafted rationale the reviewer edits rather than authors from scratch, with the edits captured. The reviewer's changes are evidence about the review and should be stored as such, not discarded on save.

The negative results deserve their own paragraph because they are the item most often cut and the item that does the most work. 10 CCR 2695.7(d) requires a thorough, fair and objective investigation. Thoroughness cannot be demonstrated from hits alone: a file containing four adverse findings and no record of what else was checked is indistinguishable from a file where only four things were checked. 10 CCR 2698.36 makes the same demand from the other direction by requiring preservation of documents and other evidence plus a concise and complete summary of the entire investigation. Field by field, the defensibility standard for fraud investigation AI sets out what the record has to contain to survive review.

One test tells you whether a packet is real. Hand it to a reviewer and ask them to argue the opposite finding using only what is in it. If they can build the counter-case, the checkpoint is a control. If the only material available supports the conclusion the system already reached, the reviewer's agreement carries no information and the signature is ceremonial.

Designing the checkpoint so it is not a signature block

A review checkpoint is a control that has to be tested, not an input that can be assumed to work. The automation bias literature points to three mitigators: training, accountability attached to the review itself, and system design. Each translates into a specific move in a special investigation unit workflow.

  1. Attach accountability to the review, not only to the outcome. Goddard, Roudsari and Wyatt name emphasis on user accountability as a mitigator across the 74 studies they included. In practice that means the reviewer's edits, overrides and confirmations are logged as first-class evidence with a timestamp and an author, and are retained, never overwritten. A reviewer who knows their review is itself part of the record reviews differently.
  2. Surface disconfirming evidence next to the finding. The default output of any system tuned to find something is the supporting case. Putting the contrary evidence in the same view, at the same weight, is a system design mitigator and it is also what 10 CCR 2695.7(d)'s objectivity standard describes.
  3. Sample reviewed decisions for audit instead of trusting the review rate. Pull a fixed percentage of approved findings each month and have a second reviewer work them blind from the packet. This is how you find out whether the first review was reading or clicking.
  4. Track override rate as a health metric and read a near-zero rate as a warning. A checkpoint where the human never changes anything is either a system that is never wrong or a review that is not happening, and only one of those is plausible. Trend the rate, and trend the distribution of what kinds of findings get overridden.
  5. Calibrate review depth to consequence instead of applying it uniformly. This is NAIC Section 3 factors (i) and (ii) operationalized: the nature of the decision and the degree of potential harm to consumers. A declination to open an SIU file on a low-severity claim and a fraud-bureau referral do not warrant the same review, and spreading a fixed review budget evenly across both guarantees the second one gets less attention than it needs.

The override rate deserves emphasis because it is the metric most likely to be reported the wrong way round. A vendor showing a very low override rate as evidence of accuracy is showing you a number that is equally consistent with a review step nobody is performing. The diagnostic value sits in the movement: a rate that falls steadily over the first two quarters of a deployment, with no corresponding change in model behavior, is a rate measuring reviewer fatigue, not system quality.

None of this is testable by assertion, which is why the checkpoint belongs inside the evaluation harness rather than beside it. The same discipline that measures whether an agent found the right contradiction should measure whether the human caught the one it got wrong. We set out the measurement approach in evaluations for fraud investigation agents.

What this does to the SIU operating model

Coverage is the variable that moves, not headcount. The number of investigators stays the same and their job changes: from running queries and assembling files to deciding what the evidence means and signing what leaves the building. Our internal benchmark puts manual coverage of flagged claims at roughly 25%.

The pressure on that 25% is not easing. The Coalition Against Insurance Fraud puts the annual cost at at least $308.6B stolen every year from American consumers, and finds that fraud occurs in about 10% of property-casualty insurance losses. Detection keeps improving, which enlarges the flagged queue rather than shrinking it. Our State of Insurance Fraud Detection 2026 report sets out the SIU capacity benchmarks behind that queue. The Verisk study measures the shape of that pressure directly.

Detection pressure and caseload expectations (Verisk, March 2026)

Insurers agreeing AI editing tools are fueling a rise in digital media fraud98%
Insurers very confident they can detect edits to real photos or video58%
Insurers that issued new adjuster guidance on suspicious media in the past year51%
Insurers very confident they can identify deepfakes32%
Insurers expecting increased SIU caseloads over the next three to five years29%

The gap between 58% and 32% is the interesting one. Insurers are roughly twice as confident about catching an edit to a real photograph as they are about catching a wholly synthetic one, and 98% agree the tools producing both are getting better. Meanwhile 51% have responded by writing new guidance for adjusters, which is a training intervention aimed at a throughput problem, and 29% expect their SIU caseloads to rise over the next three to five years. More guidance handed to the same desks, which our internal benchmark puts at 200+ open cases each, does not change how many claims get investigated.

The consumer side of the same study reorders the risk conversation. 36% of consumers named a legitimate claim being denied because documents or photos were mistakenly flagged as suspicious among their concerns, against 22% who named insurers relying too heavily on AI instead of human review when deciding whether media is authentic. Consumers are more worried about false positives than about automation, by a margin of 14 points. A false positive is a function of investigation depth. The way to reduce one is to finish the investigation before acting on the flag, not to add a review step to a flag nobody investigated.

The detection layer is not the opponent here, and its own messaging says so. FRISS states on its homepage that "Every score carries its rationale, so a reviewer can always see why a file was flagged, and overrule it," and that it "automates the repetitive screening work - straight-through processing at underwriting, fast-tracking at claims - so your teams spend their time on the cases that actually need a person." That is not an argument that humans must touch everything. It is an argument that machine output must carry its rationale, and it concedes the automation of routine work in the same breath. We agree with the principle and extend it: what holds for a score's rationale holds for an investigation's full evidence chain.

Shift Technology makes a different design choice, and it is a choice, not an error. Its homepage frames Coverage and Liability agents as providing "guidance for final determination" and Fraud and Risk agents as agents that "automatically assign cases, guide investigations, and autonomously adapt fraud strategies." The autonomy in that sentence sits at the strategy layer. On the individual case, the machine is in the advisory seat and the human is in the doing seat. It is the configuration most exposed to the Rosbach finding, because a recommendation handed to a time-pressured reviewer with no independent evidence attached is precisely the setup that produced the 7% reversal rate. Verisk, for its part, is the source of the study the whole trade-press objection rests on, which is worth saying out loud: the same research that supplies the 98% and 32% figures also found 44% of insurers relying on manual review while 66% believe fraud is getting past them.

What changes when the evidence-gathering block runs autonomously is coverage, and coverage is the number that decides everything downstream. Our internal benchmarks: 14+ days per manual SIU case, 200+ open cases per investigator, 15+ investigation phases running in parallel, and flagged-claim coverage moving from roughly 25% to 100%. Investigations complete in hours rather than weeks, which is what turns the 30-day clock in 31 Pa. Code 146.6 from a constraint into slack. The reason three quarters of flagged claims never get a full look is labour, not intent, which is the argument in why 75% of flagged insurance claims are never fully investigated. Hesper carries built-in detection and runs the full investigation behind it - from fraud detection to fraud resolution - and it is complementary to FRISS, Shift Technology and Verisk, not a replacement for them.

Where the vendor comparison lives

This post takes a position on an argument, not on a product. Carpe Data appears here once, as the clearest published statement of the opacity objection, and the concession is genuine: a bare fraud score is not an evidence trail and is not defensible on its own. The vendor-level comparison - what it means to order an investigation versus to run one - is handled separately in Carpe Data alternatives, and nothing in this post should be read as a claim about any vendor's product capabilities.

Key takeaways

  • The objection to autonomous claims AI confuses autonomy with opacity: an agent that logs every query, source and inference is more inspectable than a human-run claim file where the reasoning surfaces as a three-line note.
  • The NAIC Model Bulletin, adopted in 25 jurisdictions as of April 1, 2026, treats human involvement in final decision-making as one of five risk-calibration factors that can be offset by transparency and explainability, not as a floor.
  • Every United States rule that genuinely reserves a decision to a licensed human, including Colorado 3 CCR 702-10 Section 5.A.5 and California SB 1120, is a health utilization review rule, and no equivalent mandate exists for property and casualty claims investigation.
  • Research on automation bias, including a 7% rate at which trained experts overturned their own correct judgments to follow erroneous AI advice, shows that a nominal human checkpoint is an input specification, not a working control.
  • The design question worth answering is which decisions are reserved to a human and what evidence that human is handed, which is why California 10 CCR 2695.7(d) regulates the thoroughness of the investigation and says nothing about who performed it.

Not for property and casualty claims investigation in the United States. The NAIC Model Bulletin on the Use of AI Systems by Insurers, adopted in 25 jurisdictions as of April 1, 2026, treats the extent to which humans are involved in the final decision-making process as one of five factors that calibrate how much control an insurer needs, alongside the transparency and explainability of outcomes. It is a sliding scale, not a floor. Specific decision acts are reserved by other rules: California 10 CCR 2698.36 reserves the decision to open or decline an SIU investigation to the SIU, and California Insurance Code 1872.4 makes fraud-bureau referral an insurer determination with a 60-day clock. Evidence-gathering steps carry no such reservation.

It does not mandate human review of individual decisions. Section 3 says controls should be commensurate with the insurer's own risk assessment, considering five factors: the nature of the decision, the degree of potential harm to consumers, the extent to which humans are involved in the final decision-making process, the transparency and explainability of outcomes, and reliance on third-party models. Human involvement is factor three of five and can be offset by factor four. The bulletin also defines AI systems as designed to operate with varying levels of autonomy, and Section 1.6 places claim administration and payment and fraud detection squarely in scope. What it does require is a written AIS Program covering governance, traceability, auditability and record retention.

In health, medical necessity determinations. California SB 1120, chaptered September 28, 2024, requires that a determination of medical necessity be made only by a licensed physician or licensed health care professional competent to evaluate the specific clinical issues involved in the health care services requested by the provider, and Colorado 3 CCR 702-10 Section 5.A.5 requires that a provider acting for a health benefit plan insurer be ultimately responsible for prior-authorization denials informed by external consumer data. In property and casualty there is no equivalent licensed-human mandate. Decisions are reserved to the insurer as an entity or to its SIU: opening or declining an SIU investigation under 10 CCR 2698.36, denying a claim in writing with all factual and legal bases under 10 CCR 2695.7(b)(1), and referring suspected fraud under Insurance Code 1872.4.

Not by itself. The EU AI Act names the problem in the same article that mandates oversight: Article 14(4)(b) requires overseers be kept aware of the possible tendency of automatically relying or over-relying on the output produced by a high-risk AI system, which it calls automation bias. A 2024 study of 28 trained pathology experts measured a 7% automation bias rate, where initially correct evaluations were overturned by erroneous AI advice, and found that time pressure increased the severity of the effect. A reviewer with a heavy caseload, handed a conclusion and no independent evidence, is not a control. Defensibility comes from the record: what was checked, what the sources were, and what came back clean.

No. It does not bind United States carriers, and even within the EU the insurance high-risk category in Annex III point 5(c) is limited to AI systems intended to be used for risk assessment and pricing in relation to natural persons in the case of life and health insurance. Property and casualty claims handling is not on the high-risk list, so the Article 14 human oversight obligations do not attach to it by that route. Article 14 is still worth reading as a drafting reference, because it is the most developed statutory treatment of human oversight anywhere and because it is unusually candid about automation bias as the failure mode of the oversight it requires.

Automation bias is the tendency to treat automated output as correct and to stop seeking independent information. It produces commission errors, where a reviewer follows a machine recommendation that contradicts better evidence, and omission errors, where a reviewer misses a problem the system did not flag. A 2012 systematic review in the Journal of the American Medical Informatics Association screened 13,821 papers and included 74, identifying workload and task complexity among the environmental mediators. It matters for SIU because caseload pressure is the aggravating condition. Hesper's internal benchmark is 200+ open cases per investigator. A review step designed without accounting for that pressure degrades toward a signature block.

Enough independent evidence that the reviewer can disagree. That means the finding stated as a checkable claim, every fact with its source and retrieval timestamp, the tool calls and queries that were run, the negative results showing what was checked and came back clean, contradictions anchored to document, page and line, a stated confidence with the conditions that would change it, and a drafted rationale the reviewer edits rather than authors from scratch. The negative results matter more than they look: California 10 CCR 2695.7(d) requires a thorough, fair and objective investigation, and thoroughness cannot be demonstrated from hits alone. The reviewer's own edits and overrides should be logged as evidence too, not discarded.

The referral is an insurer act, not a machine act. California Insurance Code 1872.4(a) applies to a company that has determined, after the completion of the insurer's special investigative unit investigation, that it reasonably suspects or knows an act of insurance fraud may have occurred, and requires filing with the Fraud Division within 60 days of that determination. Two design consequences follow. The determination stays with a human in the SIU. And because the 60-day clock starts only after the SIU investigation completes, investigation throughput sets how much of the statutory window remains. An agent's job is to complete the investigation and hand over a package mapped to the prescribed form fields.

Treat the checkpoint as a control that has to be tested, not an input that is assumed to work. The automation bias literature points to three mitigators: training, accountability attached to the review itself, and system design. In practice that means logging the reviewer's edits and overrides as first-class evidence, surfacing disconfirming evidence next to the finding instead of only the supporting case, sampling reviewed decisions for blind second review, and calibrating review depth to the consequence of the decision instead of applying one uniform review. Watch the override rate as a health metric. An override rate near zero is a warning sign, not a success.

← More articles on the Hesper AI blog

See Hesper AI on your documents

Request a demo and we'll run an analysis on your real document samples.