Automated claims investigation is the step that absorbs the risk every other claims automation creates. In 2024 the New York Insurance Frauds Bureau received 41,686 reports of suspected healthcare-related fraud, 93% of them no-fault auto. Of those, 89% were designated by the insurer as intelligence-only or as still under investigation, leaving 4,585 recommended for further review by the Department. In 2020 the same count was 21,015. The regulator publishes the series itself, and what it documents is a referral pile that nearly doubled in four years while the share reaching a conclusion did not move.
Referrals are not the scarce resource. Resolution is. A carrier can add detection capacity on a procurement cycle and, until recently, could not buy investigative capacity at any price, because no vendor sold it. Meanwhile every increment of straight-through processing moves more claims past the last person who would have looked at them, so the flagged residue gets denser and more adverse rather than smaller. Automating everything upstream is what made the investigation gap expensive.
This post defines automated claims investigation as a capability test rather than a product category, shows why the flagged lane grows as the clean lane accelerates, and walks the settlement clocks that put a price on the difference. The lifecycle view of what automates at each stage is in our guide to claims automation. Two companion posts cover the adjacent questions and are worth reading as a set: which layer of the claims stack has which vendor, and what full investigation autonomy requires architecturally. Neither makes the causal claim this one does.
Claims automation worked, and that is what made investigation expensive
Claims automation has produced real results at every stage that reduces to a repeatable operation: intake, triage, coverage verification, damage estimating, payment. Investigation does not reduce that way. So the pipeline got faster on both sides of investigation and stayed the same in the middle, which is a load imbalance rather than an oversight.
The clearest evidence that the front of the pipeline automated comes from the US government rather than from a vendor. The Bureau of Labor Statistics puts employment of claims adjusters, appraisers, examiners and investigators at 389,700, and projects a decline of 6 percent from 2025 to 2035, with about 21,600 openings a year over the decade. The reason it gives is explicit: "Technology is expected to automate some of the tasks that these workers currently perform. For example, computer software can evaluate photographs of damaged property and calculate an estimated claim amount."
Photo estimating is the named mechanism, and it is a fair one. It also describes the whole class of work that automates cleanly: a bounded task, one input type, and an answer checkable against a repair cost. Investigation has none of those three properties.
Inside auto, the channel mix moves the same direction. CCC Intelligent Solutions reports in Crash Course 2026, which is vendor-published industry data rather than a regulatory dataset, that photo estimating now accounts for 26.4% of repairable appraisals, up 0.8 points, while staff appraisers account for 16.5%, down 0.8 points. Direct repair programs are 46.7%. Repairable claim volume fell 9.7% in 2025.
Then the half of that dataset that matters here. Bodily injury paid claim frequency is up 11% over two years. The average paid BI claim is up 10.3% in one year and 32% over four. BI now accounts for 52.4% of total liability dollars paid. The automatable part of the book is shrinking and the judgment-heavy part is growing on both frequency and severity, so the mix is moving toward exactly the claims a photograph cannot settle.
Two fraud numbers, two different scopes
The Coalition Against Insurance Fraud puts total US insurance fraud at $308.6 billion a year across all lines, and reports that fraud occurs in about 10% of property-casualty insurance losses. The property-casualty figure from the same 2022 study is $45 billion, derived as 10% of the $450.8 billion in P&C losses and loss adjustment expenses incurred in 2020. Use $45 billion when the sentence is about P&C. The $308.6 billion figure is all lines, and pairing it with the words property-casualty overstates the P&C exposure by roughly seven times.
The honest version of the claims automation story, then, is not that investigation was forgotten. It is that everything around investigation got cheaper and faster, which raised both the volume arriving at the investigation step and the cost of not doing it. Detection is upstream; investigation is downstream. Speed up the upstream and the downstream inherits a bigger queue.
What automated claims investigation actually means
Automated claims investigation is software that takes a claim already flagged as suspicious and runs the investigation itself: deciding what that specific claim requires, gathering evidence across independent sources, reconciling the contradictions between them, and producing a documented finding with an explicit basis that a human reviews, challenges and signs.
The term is currently loose enough to mean almost anything, so the definition below is written as a capability test rather than a product description. Each item is anchored to something a regulator already requires of a human investigation, which means the bar is not one a vendor gets to set.
The eight capabilities that make it count
- It takes a flagged claim as input and decides what that specific claim requires. A fixed script run on every referral is a checklist. California 10 CCR § 2698.36(a)(3) contemplates procedure selection rather than a uniform script: it requires using one or more industry-recognized databases identified by the SIU as appropriate for the particular line of insurance in question.
- It acquires evidence across independent sources without a human commissioning each one. Claim file and loss history, policy and underwriting record, cross-carrier data, public and court records, provider and billing history, imagery and document metadata, statements. Commissioning one source at a time is what makes the manual process serial, and serialisation is the constraint the whole category exists to remove.
- It reconciles conflicts between sources rather than concatenating them. A report that lists five sources and leaves the contradictions to the reader has moved the work, not done it.
- It produces a finding with an explicit basis. California 10 CCR § 2698.36(a)(5)(A)-(F) is the practical template: what facts caused the belief that fraud occurred, what the suspected misrepresentations are and who made them, how they are material, who the pertinent witnesses are, what documentation exists, and whether the investigation is complete.
- It emits a reconstructable record. NAIC Model Regulation #902 § 4.B requires detailed documentation in each claim file sufficient to permit reconstruction of the insurer's activities relative to each claim. Applied to an automated system, the insurer's activities include every source queried, every tool call, every timestamp and every inference the system made.
- It runs on the whole flagged population rather than the top of the queue. Coverage is definitional. A system that clears a queue faster is a productivity tool. An automation layer is one where the number of flagged claims stops governing how many get worked.
- It finishes inside the settlement clock. NAIC Model #902 sets 21-day and 45-day checkpoints on first-party claims. A finding delivered after the pay decision has already been made is an audit artifact, not an investigation.
- It hands adjudication to a human. Nothing in this definition licenses auto-denial. NAIC Model #900 § 4.F names refusing to pay claims without conducting a reasonable investigation as an unfair claims practice, and the pay, deny or refer call stays with the adjuster and the SIU lead who signs it.
The last clause of capability four is the tell. A system that cannot say that the investigation is not complete, and name what is missing, is not investigating. It is summarising.
What automated claims investigation is not
- Not a fraud score. A score is a prior about a claim. An investigation is a posterior conditioned on evidence gathered for that specific claim.
- Not a rules-engine alert or a red flag firing. Under California 10 CCR § 2698.36(c), a red flag is the input that obligates an investigation. It is definitionally not the output.
- Not an OCR or intelligent document processing pass. Extracting the fields on a document is not testing whether they are true.
- Not a case-management workspace. Queues, routing, notes and SIU case files are the container, not the contents. In the Coalition Against Insurance Fraud's 2024 State of Insurance Fraud Technology Study, only 25.7% of the 35 responding carriers said case management was not incorporated. The container is well served.
- Not a summarisation layer over a file a human already assembled. That compresses reporting time, not investigation time.
- Not a tool-ordering API. Being able to order surveillance, a records pull or an OSINT sweep is commissioning. Investigation is deciding what to commission, reading what comes back, and resolving it against everything else in the file.
- Not adjudication, and not auto-denial. The investigator's role shifts from execution to decision-making, which is a change in what the human does rather than a removal of the human.
For an SIU director the test collapses to a single question: can an investigator see what the system did, override it, and hand the trail to a state examiner. Capabilities four and five are the whole answer, and they are the reason this definition is written as capabilities rather than as outcomes.
Faster settlement moves more claims past the only person who would have looked
Straight-through processing does not reduce investigation load. It concentrates it. Paying the clean lane without a human touch is the point of STP and it works, but the residue left behind is denser in exactly the fraud type that is hardest to catch, because a fast lane selects for claims that look ordinary.
Detection vendors describe this architecture plainly, which is worth crediting rather than criticising. FRISS, one of the established detection platforms, describes its own product this way:
That is the correct design, and it is also a complete statement of the problem. The clean lane runs untouched at machine speed and the flagged lane terminates in a person. FRISS reports, on vendor-published figures, that 90% of honest claims are fast-tracked. The other lane inherits everything else, and its throughput is set by headcount.
The composition of what survives the fast lane is the part carriers underweight. Deloitte data reported by Insurance Journal in May 2025 splits fraud in two. Soft fraud - inflation, buildup, exaggeration - is about 60% of all incidents and is detected at rates of 20% to 40%. Hard fraud is about 40% of claims fraud and is detected at 40% to 80%. The majority category is the least-detected one, and it is the category a fast lane pays, because nothing is categorically wrong with the claim, only with its size.
The flagged lane, meanwhile, fills with a high proportion of claims that are not fraud at all. Rules-based and hybrid scoring runs a 60-85% false positive rate against Hesper's internal benchmarks, so most of what reaches the queue is clearable and every item still costs something to clear. The lane model underneath this is in our claims triage framework, which sets out five lanes and why only four of them get measured.
Put the two together and the arithmetic is unfriendly. Upstream automation raises throughput into the flagged lane. Detection improvements raise the flag rate. Neither touches the number of claims one investigator can work in a month.
The flag volume is compounding, and the referral data says so
Fraud referral volume is growing considerably faster than the capacity to resolve it, and one state regulator publishes the series that proves it. Suspected healthcare fraud reports filed with the New York Insurance Frauds Bureau went from 21,015 in 2020 to 41,686 in 2024, a 98% increase over four years.
Scoping matters here. These are healthcare-related reports, New York only, and 93% of them are no-fault auto, which is property-casualty business. The no-fault subset alone went from 19,153 to 38,846 over the same period. Healthcare fraud reports accounted for approximately 80% of all fraud reports the Department received in 2024, so the whole-book referral flow is larger still. The report is worth reading in the original.
What happens to those reports is the finding, and the Department states it in its own words:
That sentence is easy to over-claim from, so hold it precisely. It is not a carrier SIU coverage rate. The Department's own criminal caseload for the year, 71 healthcare fraud investigations opened and 42 arrests, is a separate number that should not be divided into 41,686. What the sentence does establish is that the overwhelming majority of fraud referrals arrive at a regulator in an unresolved state, and that the insurers filing them said so themselves.
Carriers describe the same dynamic in their own survey responses. The Coalition Against Insurance Fraud's 2024 State of Insurance Fraud Technology Study, published December 10, 2024, surveyed 35 carriers with 69% of respondents from SIU, so treat every percentage from it as a small-sample indication. On one slide, respondents list the benefits of anti-fraud technology as "More referrals," "Increased speed of detection," "Higher quality referrals," "More consistent claims investigations" and "Straight through processing." On the same slide, among the challenges, they list "SIU cannot handle the volume of potentially fraudulent claims."
Same carriers, same page, naming more referrals as a benefit and the inability to work them as a challenge. Nothing in this post states the thesis more plainly, and it comes from the industry rather than from a vendor.
The arithmetic underneath fits on a whiteboard. Against Hesper's internal benchmarks, a manual SIU investigation runs 14+ days per case, an investigator carries 200+ open cases, and throughput lands near 10 investigations per investigator per month. At that rate an investigator needs roughly 20 months to work a standing caseload once through, before a single new referral arrives. That is why about 25% of flagged claims get a full manual investigation and the rest are paid, denied without complete work, or queued until they age out. The mechanics are in why most flagged claims are never investigated.
The regulation prices the gap
Claims regulation puts a price on the investigation gap in a way an operations dashboard does not. Settlement clocks run on the promptness duty, the fraud carve-out from those clocks is conditional on producing investigative output, and at least three states collect or examine numbers that make the gap visible from outside the carrier.
NAIC Model #900, the Unfair Claims Settlement Practices Act, was adopted as a free-standing act in June 1990. Section 4 defines the practices, and four of them frame this problem: failing to adopt and implement reasonable standards for the prompt investigation and settlement of claims; not attempting in good faith to effectuate prompt, fair and equitable settlement of claims in which liability has become reasonably clear; refusing to pay claims without conducting a reasonable investigation; and failing to affirm or deny coverage within a reasonable time after having completed its investigation.
Read 4.D and 4.F together and they form a two-sided obligation. Straight-through processing is a direct answer to 4.D. Nothing in an STP pipeline answers 4.F. A carrier that automates only the pay side improves its posture on one clause, leaves the other exactly where it was, and increases the number of claims transiting both.
NAIC Model #902 is where the clocks live. Section 3.G defines investigation as all activities of an insurer directly or indirectly related to the determination of liabilities under coverages afforded by an insurance policy. Section 7.A requires that a first party claimant be advised of acceptance or denial within twenty-one days of properly executed proofs of loss. Section 7.B requires a letter at forty-five days, and every forty-five days after that, setting out the reasons more time is needed for investigation. The stage-by-stage version of that timeline is in our walkthrough of the claims process from FNOL to settlement.
Then the carve-out, which is the sentence this section exists for:
The relief is real and it is conditional. The phrase "specific information available for review by the insurance regulatory authority" describes investigative output, and there is nothing else it could describe. So the carve-out is available in law to every carrier and reachable in practice only by a carrier whose investigation capacity matches its flagged volume. A carrier investigating roughly a quarter of what it flags holds the carve-out on paper for all of those claims and in fact for a quarter of them. The rest either settle on the clock or generate a 45-day letter that says, in substance, we have not looked yet.
One fairness note before the state law. Model #902 § 2 states that nothing in it creates or implies a private cause of action for violation of the regulation, and the NAIC models are templates that states adapt rather than binding law in themselves. The exposure here is market conduct examination and the state-adopted analogues, not a claimant suing on a model regulation.
California extends the duty to investigate to system-generated referrals
California 10 CCR § 2698.36(c) requires the SIU to investigate each credible referral of suspected insurance fraud it receives from integral anti-fraud personnel, including automated or system-generated referrals. A credible referral is one that includes a red flag. The only lawful way not to open an investigation is a preliminary review concluding it is reasonably clear the red flag is not the result of suspected fraud, and that conclusion has to be documented in the claim file with the reasons supporting it. Buying a detection engine in California therefore does not just produce more flags. It produces more legally obligated investigations, and every declined one is its own written work product. The state walkthrough is in our guide to 10 CCR 2698 SIU compliance.
Florida turns the gap into a filed statistic. Fla. Stat. § 626.9891, titled Insurer anti-fraud investigative units; reporting requirements; penalties for noncompliance, requires each insurer to report annually by March 1, for each line of business, the number of claims received, the number of claims referred to the anti-fraud investigative unit, the number of claims investigated or accepted by that unit, and the number of cases referred to the Division of Investigative and Forensic Services. Received, referred and investigated are three separate counts. The distance between referred and investigated is a number the state collects every year, by line, which means investigation coverage is not an internal metric a carrier can choose whether to track.
New York adds two constraints that shape what an automated layer can look like. 11 NYCRR § 86.6(b)(1), Insurance Regulation 95, requires a full-time Special Investigations Unit separate from the underwriting or claims functions of the insurer, which forecloses the tidy fix of folding investigation back into an automated claims workflow. Investigation has to be its own layer, by regulation, in New York. And § 86.6(b)(3) requires the fraud prevention plan to state the rationale for SIU staffing, which may include "an assessment of optimal caseload which can be handled by an investigator on an annual basis," alongside "volume of suspected fraudulent New York claims currently being detected."
That second clause changes the register of the caseload number. Measured against a figure a carrier has already filed with the state, 200+ cases per investigator is not an operations statistic. It is a regulatory position that has to be defensible, and a detection upgrade mechanically changes one of the inputs to the staffing rationale on file.
None of this forbids automation. The NAIC adopted its Model Bulletin on the Use of Artificial Intelligence Systems by Insurers on December 4, 2023, and per the NAIC's own adoption map, status as of August 6, 2026, 25 jurisdictions have adopted it, with four more - California, Colorado, New York and Texas - running insurance-specific AI regulation or guidance instead. The operative expectations are that decisions supported by AI comply with existing insurance law and that regulators can request information about the system in an examination. Combine that with Model #902 § 4.B and California's completeness requirement and the practical bar is one sentence long: an automated investigation has to be able to explain, per claim, what it did, what it found, what it could not resolve, and why.
Speed up the clean lane and the flagged lane does not shrink. It gets denser, and its throughput is still set by headcount.
Manual, scored, automated: the three ways a flagged claim gets handled
A flagged claim gets handled one of three ways today. A human investigator works it end to end. A model scores it and a human works whatever the score elevates. Or an automated investigation layer runs the evidence work on every flagged claim and a human adjudicates the output. The three differ most on coverage, not on speed.
The row that decides procurement is coverage. Cycle time is the number that gets quoted and it is the less important one, because a faster manual process is still bounded by the number of investigators. Coverage is bounded by nothing once per-case attention stops being the constraint, which is why the ~25% to 100% shift rather than the day count is the loss-cost lever.
The row worth arguing about is the last one. There are named vendors in the scoring column and, until recently, none in the third. That is a statement about the market rather than about the technology.
Which layer each vendor actually occupies
Every layer of the claims stack has an incumbent except one. Prevention, detection, estimating, payments, case management and claims administration all have named vendors with real products behind them. Investigation execution has manual SIU teams. That is the gap, and the industry's own investment plans describe it without meaning to.
Take the vendors in order, fairly. Verisk ISO ClaimSearch is the contributory claims data infrastructure most of the US P&C market runs on, and Xactimate is the property estimate standard; both are inputs to an investigation rather than the investigation itself. FRISS scores and routes, with the architecture quoted above. Shift Technology's published AXA Switzerland deployment screened over a million claims and, on the carrier's account, stopped more than EUR 12 million in fraud, which is evidence that detection works at scale. CCC Estimate-STP is the clearest public example of genuine straight-through processing in P&C, described as automatically initiating and populating detailed estimates in seconds, and announced as in market with four national insurers including USAA on October 28, 2021. Tractable's Tokio Marine deployment was announced on the carrier's own statement that AI can cut remote claim review from days to minutes.
Each of those does its own job well, and the observation that ties them together is structural rather than critical. An estimate answers how much. A score answers how likely. A document extraction answers what the page says. None of them answers whether this claim is true, because none of them was built to.
Clearspeed positions as a trust layer that screens at the point of decision, returning minimal or elevated risk in seconds. Its published customer results are vendor-reported, including a 200% uplift in fraud findings at Allianz. Read that metric in the context of this post: an uplift in fraud findings is more flags, which is precisely the load the investigation layer has to absorb.
The agent wave has the same shape. As we found in a survey of shipped claims agents, deployments cluster in intake, communication and summarisation, which are the places where a language model has an obvious bounded job. Investigation is where the work is unbounded, which is why it comes last.
The 2024 CAIF study makes the point twice. First on capability, where the most basic detection function is close to universal while the analytical ones are half-deployed. These are the printed "not incorporated" percentages from the 35 responding carriers:
Read the inverse and 88.6% of respondents have automated red flags and business rules in place, while the analytical capabilities trail well behind. Second, on intent. Asked what they plan to invest in over the next 12 to 24 months, respondents named AI and machine learning, predictive modeling, automated red flags, generative AI, link analysis, reporting and data visualisation, case management, text mining, exception reporting and geographic data mapping. Every option on that list is detection, analytics or case management. Investigation execution does not appear, because at the time of the survey it was not something a carrier could buy.
That is the layer Hesper AI occupies. It takes a flagged claim, runs 15+ investigation phases in parallel - document forensics, OSINT, statement cross-reference, timeline reconstruction, financial pattern analysis - and returns an audit-ready finding with its sources, its reasoning and its timestamps, in hours rather than weeks. It has built-in detection, so it runs standalone, and it is complementary to FRISS, Shift Technology and Verisk rather than a replacement for them. From fraud detection to fraud resolution is the shortest description of what changes.
What to do with the next automation dollar
If FNOL through payment is already automated and a detection platform is already in place, the marginal loss-cost lever is not more front-end speed. It is coverage of the flagged population, because that is the one number in the stack a decade of fraud-tech investment has not moved.
Deloitte's April 2025 analysis projects that P&C insurers could reduce fraudulent claims and save between US$80 billion and US$160 billion by 2032. It is a range, and it is frequently miscited as the top of the range. More telling is where executives say they are aiming. In a Deloitte survey of 200 US insurance executives conducted in June 2024 and reported by Insurance Journal, 35% selected fraud detection as a top-five priority for developing or implementing generative AI over the next year. Detection: the layer that is already the best served in the stack.
The unit economics of the alternative are straightforward. Against Hesper's internal benchmarks a manual investigation costs roughly $2,500 per case and an automated one roughly $150, and throughput moves from about 10 investigations per investigator per month to 800+. Cost per case is the wrong headline, though. The headline is that the coverage constraint disappears: flagged volume stops governing how many claims get worked, which is what moves ~25% to 100%.
This is not a headcount argument and it should not be sold as one. SIU staffing does not shrink in this model, it gets re-aimed. The investigator's role shifts from execution to decision-making: reviewing findings, challenging them, overriding the thin ones, and owning the outcome on the record. What changes is that the flagged claims nobody had time for get worked at all.
Make every flagged claim investigable. That is the whole proposition, and the test of whether a carrier has done it is not a demo. It is whether the number of claims flagged in a month still predicts the number investigated that month. If it does, the investigation layer is still manual, whatever else has been automated around it.
Key takeaways
- Automated claims investigation is the step that absorbs the risk every other layer of claims automation creates, because faster straight-through processing moves more claims past the last person who would have looked at them.
- New York's Insurance Frauds Bureau received 41,686 suspected healthcare fraud reports in 2024, up 98% from 21,015 in 2020, and 89% of them were logged by the filing insurer as intelligence-only or as still under investigation.
- A system only counts as automated claims investigation if it selects procedures per claim, gathers and reconciles evidence across independent sources, states whether the investigation is complete, and emits a record reconstructable under NAIC Model #902 § 4.B, which a score, a red flag, an OCR pass and a case-management queue are not.
- NAIC Model #902 § 7.A relieves an insurer of the 21-day settlement clock only where there is a reasonable basis supported by specific information available for review by the regulator, so a carrier that automated settlement but not investigation runs the promptness duty at machine speed and the fraud carve-out at human speed.
- Every anti-fraud technology the 35 carriers in the 2024 CAIF study said they plan to buy in the next 12 to 24 months is detection, analytics or case management, which is the clearest available evidence that investigation execution is the one layer of the claims stack with no software incumbent.