ITC Vegas 2026We're on the floor at Mandalay Bay, Sep 29 - Oct 1Meet us there
Blog›Guides
GuidesSeptember 24, 2026·20 min read·Pankaj Dhariwal, CEO

Claims leakage audit: a stage-by-stage checklist, and the files it never opens

A stage-by-stage claims leakage audit checklist anchored to the NAIC claims standards, with the sample-size math and the reason a sampled leakage rate describes only the files you opened.

PD
Pankaj Dhariwal · CEO and Co-founder
September 24, 2026·20 min read
GUIDESHesper AI7-14%Leakage as a share ofcarriers' total claims spendEY PROPERTY AND CASUALTY CLAIMS TRANSFORMATION,LATEST CLAIMS QUALITY ASSESSMENTS
The numbers behind this
$637.8BUS P&C losses and loss adjustment expenses incurred, 2025Sum of two NAIC line items: $551.8B net losses plus $86.0B loss expenses
1 in 4P&C insurers that make no salvage or subrogation recovery effortBisco and Fier, NAIC Journal of Insurance Regulation, firm-level 1996 to 2021
7%Claim-practice error rate that presumes a general business practice violationNAIC market conduct sampling guidance, 2014 exposure draft

A claims leakage audit measures the gap between what was paid and what the claim file supports. Not the gap between what was paid and what a second reviewer would have paid, which is an opinion. Not the gap between your loss ratio and a peer's, which is a ratio with a hundred inputs. Per file, with the evidence cited, coded to a stage, or the finding does not survive the first conversation with the adjuster who handled it.

The size of the gap is not in dispute. EY's Property and Casualty Claims Transformation practice puts leakage at 7% to 14% of a carrier's total claims spend, based on the team's latest assessments. In a case study of one top US P&C insurer, built from an examination of hundreds of claims, leakage ran to 10% of total paid, nearly two-thirds of the files reviewed had some degree of leakage assessed, and more than 85% of the assessed leakage sat in three areas: coverage determination, litigation prevention, and evaluation and resolution. EY projected at least $50 million a year in payment-accuracy-only benefits across seven initiatives at that one insurer.

The checklist below is the deliverable. The argument underneath it is the part that rarely gets said out loud: sampling is a capacity constraint that has been reclassified as a methodology. A carrier opens the files it has reviewer hours for, extrapolates, and calls the output its leakage rate. But NAIC's own market conduct sampling guidance says results cannot be generalized beyond the field of files the sample was drawn from, and that tolerable error levels are not a safe harbor. Therefore a sampled leakage rate is a statement about the files you opened, and nothing else. Everything outside the sample is not clean. It is unmeasured.

Which leads somewhere specific. If the evidence standard you apply in the audit is the right standard, it is the right standard at the moment the decision is made, not eleven months later. What follows is the stage-by-stage checklist anchored to the NAIC Market Regulation Handbook claims standards, NAIC Model Regulation #902, California's claim file regulations and the Texas and Pennsylvania claim-handling clocks, plus the sample-size math and a sizing method that states what the number is not. The definitional companion to this piece is how uninvestigated claims drain profitability, and the stage map underneath the whole thing is the complete guide to claims automation.

What a claims leakage audit measures

Answer

What does a claims leakage audit measure?

It measures the variance between what was paid on a closed claim and what the file evidence supports, priced per file and coded to the stage where the gap opened. The output is a financial and quality measure rather than a compliance test, which is what separates it from a state market conduct examination and its penalties.

The definition has a hard edge, and three common audits fail it. The hindsight audit prices what a reviewer with the full file and no clock would have paid, which is unfalsifiable and gets argued down to nothing by the handling unit. The benchmark audit prices the variance against a peer table, which measures your mix as much as your handling. The ratio audit starts from the loss ratio and works backwards, which cannot tell you which file to fix. Only the file-evidence version produces a finding with a named cause, a named stage and a remediation that someone can actually own.

EY names four root causes in the litigated-claims work, and it is worth reading them as file tests rather than as categories. Evaluation: general damages were not accurately assessed. Prevention: a pre-suit settlement opportunity was missed. Investigation: injury causation and liability were not properly established. Litigation: strategy was not matched to the claim facts. Every one of those is a question about what is in the file, answerable from the file, and none of the four is fraud.

EY root causeWhat it looks like in the fileWhere the checklist catches it
EvaluationA settlement amount with no damages model, comparables or authority trail behind itSettlement stage
PreventionA pre-suit window that closed with no recorded offer, evaluation or decision not to offerSettlement stage
InvestigationA conclusion on causation or liability with no dated record of what was checkedInvestigation stage
LitigationA defense strategy that does not reference the facts the file actually establishesSettlement stage, escalated

Leakage is not fraud, and conflating the two produces an audit that points at the SIU and stops. The Coalition Against Insurance Fraud puts fraud at about 10% of property and casualty losses, with at least $308.6 billion stolen every year from American consumers. Separate problem, separate workflow. In EY's case study the leakage concentrated in coverage determination and in evaluation and resolution, which are handling defects on claims nobody alleged were fraudulent. The two share a cause - nobody had the hours to build the file properly - but they do not share a remediation plan.

Scoping the audit: population, sample size and two thresholds

Answer

How do you scope a claims leakage audit?

Scope is set by three decisions: the population of files the finding will apply to, the confidence level and tolerance the sample has to hit, and the error rate that counts as a failure. NAIC market conduct sampling guidance sets 95% as the minimum first-stage confidence level and says a second-stage sample should never fall below 90%.

Define the population first, and define it narrowly enough to be honest. Line of business, state, handling unit, closed-date band, paid-amount band. The reason to be precise is not methodological neatness. It is that the population you define is the only thing your number will ever describe.

Generalization or extrapolation of results beyond the field of files from which the sample is selected is not acceptable... A sample can only be representative of the population from which it was drawn - and no other.

NAIC Market Conduct Examination Standards (D) Working Group, sampling guidance (exposure draft, September 2014)

That language comes from the NAIC working group's exposure draft of the market conduct sampling chapter, circulated in September 2014 as revisions to Section D of Chapter 14, so cite it as sampling guidance rather than as the current handbook. The constraint it describes is not controversial and it is not new. It is the standard reason a regulator will not let an examiner turn 220 files into a statement about 11,300.

Sample size then falls out of the expected population proportion, not out of a rule of thumb. The same guidance shows 1,067 files for a plus or minus three percentage point interval at 95% confidence when the population proportion is 50%, and 203 files for the same interval when the proportion is 5% or 95%. The arithmetic is simply that variance peaks at 50%: the less you know about the error rate going in, the more files you have to open to say anything precise. Any vendor guidance that hands you a fixed file count per line of business is skipping the step that determines the answer.

Two thresholds then decide what the result means. The sampling guidance records that a benchmark error rate of 7 percent has historically been established for auditing claim practices and 10 percent for other trade practices, and that error rates exceeding those benchmarks are presumed to indicate a general business practice contrary to the law. The second threshold is the one that gets misread: the guidance states plainly that tolerable error levels do not establish a safe harbor such that violations below them are considered tolerable or legally permissible. A 4% error rate is not a pass. It is a number below a presumption trigger.

Design choiceWhat the NAIC sampling guidance saysWhat it means for the audit finding
First-stage confidence level95% minimumBelow it, the output is an indication, not a measurement
Second-stage confidence levelDiscretionary, but never less than 90 percentA follow-up sample can be smaller without being loose
Sample size, 50% population proportion1,067 files for a +/-3 point interval at 95%The no-prior-estimate case, and the reviewer-hour bill that comes with it
Sample size, 5% or 95% proportion203 files for the same intervalDefensible only when a prior pass establishes the proportion
Benchmark error rate, claim practices7 percentAbove it, a general business practice contrary to the law is presumed
Benchmark error rate, other trade practices10 percentSame presumption, applied to a different practice set
ExtrapolationNot acceptable beyond the field of files sampledYour rate describes the stratum you opened, not the book
Tolerable errorDoes not establish a safe harborFindings below the benchmark are still findings

The evidence test: what must be in the file before a finding counts

Answer

What evidence does a claims leakage audit finding need?

A finding holds up when the file itself shows the decision, the basis for it, and the date. NAIC Model Regulation #902 requires documentation detailed enough to permit reconstruction of the insurer's activities on each claim, and California requires all documents, notes and work papers in enough detail that pertinent events and their dates can be reconstructed.

The documentation standard an internal audit needs is already written, by regulators, and there is no reason to invent a second one. NAIC Model Regulation #902, the Unfair Property/Casualty Claims Settlement Practices model, section 4.B: "Detailed documentation shall be contained in each claim file in order to permit reconstruction of the insurer's activities relative to each claim." Section 4.C adds the timestamps: "Each relevant document within the claim file shall be noted as to date received, date processed or date mailed." Section 4.A sets retention at the current year plus the two preceding years.

California goes further on both counts. 10 CCR section 2695.3(a) requires that claim files "shall contain all documents, notes and work papers (including copies of all correspondence) which reasonably pertain to each claim in such detail that pertinent events and the dates of the events can be reconstructed," and subsections (b)(1) and (b)(3) push retention to the current year plus four preceding years. Pennsylvania carries its own file and record documentation standard at 31 Pa. Code section 146.3. And the NAIC Market Regulation Handbook general examination standards include, as Claims Standard 5, five words that are the whole audit: claim files are adequately documented.

A leakage finding and a market conduct violation are often the same fact

An undocumented coverage position is a leakage finding to the claims executive and a documentation violation to the examiner. A denial with no recorded factual and legal basis is a reopened file to one reader and an unfair claims practice to the other. The audit team and the compliance team are usually looking at the same defect through two different price lists, which is an argument for running one checklist rather than two.

The price list on the regulatory side is public. In May 2026 the California Department of Insurance announced a market conduct examination of one carrier's handling of 2025 Los Angeles wildfire claims. Examiners reviewed a sample of 220 claims out of roughly 11,300 residential claims, and alleged 398 violations across 114 of the sampled files, with 34 more drawn from consumer complaints. California Insurance Code section 790.035 makes penalties of up to $5,000 per violation available, and up to $10,000 where a violation is willful. These are allegations in an examination rather than adjudicated findings, and the carrier has its own process to answer them.

Read those numbers as sampling arithmetic, not as a story about one carrier. A 220-file sample against a roughly 11,300-file population is under 2% of the population, and more than half of the sampled files carried at least one alleged violation. Two conclusions follow and they pull in opposite directions: a small sample was enough to establish a pattern worth acting on, and the same small sample says nothing, by the regulator's own rule, about the other 98%. The practical answer is to make the trail a byproduct of doing the work rather than an artifact assembled later, which is the argument in how to generate an audit-ready investigation report in hours.

The stage-by-stage checklist

Answer

What are the stages of a claims leakage audit checklist?

The checklist runs six stages plus one cross-cutting test: coverage, investigation, estimation and reserving, settlement, payment and closure, recovery, and statistical coding. Each stage gets a pull rule for which files to sample, an evidence test for what has to be in the file, and a regulatory anchor that makes the finding defensible outside the audit team.

The anchors below come from two published lists: the eleven general claims standards in the NAIC Market Regulation Handbook general examination standards, and the three property and casualty claims standards alongside them. Cite them from the 2017 Examination Standards Summary, since numbering may have moved since. Anchoring an internal audit to a regulator's own list means a finding written this way answers to two audiences at once.

StageWhat to pullWhat must be in the fileRegulatory anchor
CoverageClosed paid files with any endorsement, exclusion, sublimit or reservation of rights; every denial and closed-without-payment fileThe policy version in force on the date of loss, the specific provision relied on, a written coverage position, the reservation of rights or excess-of-loss letter where required, and a dated denial letterNAIC Ch.16 Claims Standards 6 and 9; Ch.17 Claims Standard 1; Model Reg #902 section 5.C; 10 CCR 2695.3(a)
InvestigationFiles with a fraud or severity flag; files where liability was contested; files closed inside the median cycle time for their bandA dated record of what was investigated, by whom, what was found and what was ruled outNAIC Ch.16 Claims Standards 2 and 5; 31 Pa. Code 146.6; Tex. Ins. Code 542.055
Estimation and reservingFiles with supplements or re-inspections; reserve changes above a set threshold; total lossesThe estimate of record, every supplement with its trigger, the reserve history with a reason per change, and the basis for the final valuationNAIC Ch.16 Claims Standard 8
SettlementFiles settled outside the benchmark band for their injury or damage profile; every file with attorney representation; every file that went into suitThe damages model behind the offer, the comparables or benchmarks used, the negotiation history and the authority trailNAIC Ch.16 Claims Standards 3, 4 and 11; 31 Pa. Code 146.7
Payment and closureDuplicate payee names, payments dated after the close date, payments to non-matching addresses or accounts, voided and reissued draftsThe payee verification record, the calculation the payment reconciles to, the release or settlement document, and the dates issued and clearedNAIC Ch.16 Claims Standard 10; Tex. Ins. Code 542.058 and 542.060
RecoveryEvery closed file with a third party, a defective product, a premises owner or salvageable property; every file where a deductible was collectedThe subrogation screen with a dated result, the referral or the documented decision not to pursue, and the deductible reimbursement recordNAIC Ch.17 Claims Standard 2
Statistical coding (cross-cutting)Coding on every sampled fileCause-of-loss, line and coverage codes that match the file narrativeNAIC Ch.16 Claims Standard 7; Ch.17 Claims Standard 3

Coverage

Coverage determination is one of the three areas that held more than 85% of the assessed leakage in EY's case study, and it is the stage where a bad outcome is cheapest to prevent and most expensive to unwind. The most common finding is not a wrong coverage call. It is a right call with nothing behind it: a paid file where the endorsement that limited the loss was never read, or a denial where the provision relied on appears only as a phrase in an adjuster note. Model Regulation #902 section 5.C is specific on the denial side, barring denial for failure to exhibit property unless there is documentation of a breach of the policy provisions in the claim file. The mechanics of producing that record as the claim moves are on the coverage verification page.

Investigation

EY's third root cause is that injury causation and liability were not properly established. In file terms that is almost never a wrong conclusion. It is an absent one: a file that records what was decided without recording what was checked. NAIC Claims Standard 2 requires timely investigations and Standard 5 requires adequate documentation, and the statutory clocks put dates on both. Pennsylvania requires the investigation to be completed within 30 days of notification, with a written status every 45 days after that. Texas requires the insurer to acknowledge the claim, commence the investigation and request the items it needs within 15 days of notice.

The pull rule that finds the most is the counterintuitive one. Do not only pull the slow files. Pull the files that closed inside the median cycle time for their severity band, because a complex liability claim that closed fast either had an unusually clean fact pattern or skipped a phase. Manual SIU investigation runs 14+ days per case against a caseload of 200+ cases per investigator (Hesper internal benchmarks), and that arithmetic decides which files get the work: roughly 25% of flagged claims get a full manual investigation, the subject of why most flagged claims are never fully investigated. An audit sample does not restrict itself to that 25%. Neither should the handling standard. The workflow is on the fraud investigation page.

Estimation and reserving

NAIC Claims Standard 8 asks whether claim files are reserved in accordance with the entity's own established procedures, which makes this the one stage where the audit tests you against your own written rule rather than against a statute. Pull files with supplements, re-inspections, total losses, and any reserve change above a threshold you set before you look. The finding that recurs is a reserve history that is a number series: four movements, four dates, no reasons. A reserve that moved $40,000 with no recorded trigger is not a reserving error on its own, but it removes any ability to tell whether the final paid amount was a valuation or a negotiation outcome, and it makes the settlement-stage test in the next section unanswerable.

Settlement

Settlement is where EY puts three of its four root causes, and where the NAIC list contains the sharpest standard on the book. Claims Standard 11 asks whether claim handling practices compel claimants to institute litigation, in cases of clear liability and coverage, by offering substantially less than is due under the policy. That standard cuts both ways for a leakage audit: underpayment and overpayment are both variances from what the file supports, and an audit that only counts overpayment is a cost-reduction exercise wearing an audit's clothes. Severity is moving underneath all of it. EY, citing CCC research, reports average indemnity of about $27,000 per injured party in third-party bodily injury claims, up 8.3% since 2023 and 38% since 2020. Demand handling mechanics are on the settlement and demand review page.

Payment and closure

This stage is tested with queries rather than with file reads, which makes it the cheapest section of the audit and the one most often skipped. Run duplicate payee names, payments dated after the close date, payments to addresses or accounts that do not match the claimant record, and voided drafts that were reissued. NAIC Claims Standard 10 covers canceled benefit checks and drafts as evidence of claim handling practice, and Texas Insurance Code section 542.060 prices non-compliance at interest on the claim amount at the rate of 18 percent a year as damages, together with reasonable and necessary attorney's fees. On direction of error, the GAO estimate of federal improper payments for fiscal 2025 is about $186 billion across 64 programs, roughly 82% of it ($153 billion) overpayments. That is federal government spending, not insurance: a labelled analogue only.

Recovery

Recovery is the only stage on the checklist where the money is still collectable at the moment you find it. Everywhere else the audit produces a lesson. Here it produces a demand. Work by Bisco and Fier in the NAIC Journal of Insurance Regulation computes $51.6 billion in salvage and subrogation recovered by US insurers in 2021 across auto physical damage, commercial auto liability and personal auto liability, from the authors' calculations on 2021 NAIC annual statements. The firm-level finding is the one to audit against: over three-quarters of sample firms recover some amount of salvage and subrogation, which means roughly one out of four property-liability insurers make no recovery efforts at all.

Salvage and subrogation as a share of net claims paid (Bisco and Fier, NAIC Journal of Insurance Regulation)

All sample firms, 1996 to 20214.5%
Firms with any positive recovery6.2%
Highest firm in the sample~90%

The spread is the finding. Across the full sample, recoveries run 4.5% of net claims paid; among firms that recover anything at all, 6.2%; at the top of the sample, nearly 90%. A carrier near the sample average is not necessarily doing recovery badly, but it sits in a distribution wide enough that the question deserves a file test rather than a benchmark comparison. NAIC Chapter 17 Claims Standard 2 adds the obligation most often missed on the way back: deductible reimbursement to insureds upon subrogation recovery, made in a timely and accurate manner. Most subrogation is not lost at the referral desk, as covered in why missed subrogation is lost at intake, not at closure.

Statistical coding

This one reads like an afterthought and runs first in practice. NAIC Chapter 17 Claims Standard 3 requires loss statistical coding to be complete and accurate, and Chapter 16 Claims Standard 7 covers whether claim forms are appropriate for the product. Miscoded cause of loss does not just distort reporting. It corrupts the audit's own stratification, because the population defined in the scoping step was defined using those codes. Wrong coding means the field of files you sampled is not the field you think you sampled.

Everything outside the sample is not clean. It is unmeasured. Three review regimes, one axis: the share of a file population that gets a full evidence-standard review, and what sets each number.

Sizing the finding in dollars

Answer

How do you calculate a claims leakage percentage?

Price each sampled file as paid amount minus what the file evidence supports, sum those variances across the stratum you sampled, and divide by that stratum's paid losses plus allocated loss adjustment expense. Report an interval rather than a point estimate, and apply the rate only to the population the sample was drawn from.

Two choices in that sentence do most of the work. The denominator is paid losses plus allocated loss adjustment expense for the stratum, not premium and not the whole book, because premium mixes in pricing decisions the claims department does not make. And the output is an interval, because the sample size you chose in the scoping step already fixed the width of that interval, and hiding it behind a point estimate is the single fastest way to lose the room when the number gets challenged.

StepThe inputThe discipline that keeps it defensible
1. Define the populationLine, state, handling unit, closed-date band, paid bandWrite it down before drawing files, not after seeing results
2. Draw the sampleConfidence and tolerance set from the NAIC guidance1,067 files with no prior proportion; 203 when a prior pass sets it near 5%
3. Price each filePaid amount minus what the file evidence supports, with the evidence citedCode every variance to one of the seven stages, or it cannot be remediated
4. Compute the rateSum of variances over paid losses plus allocated LAE for that stratumReport the interval, not the midpoint
5. Apply itOnly to the stratum sampledGeneralization beyond the field of files sampled is not acceptable
6. Report the residualEverything outside the sampleIt carries an unmeasured rate, not a zero rate. Say so in the deck

Industry aggregates are useful for scale, not for benchmarking your own book. The NAIC property and casualty industry analysis report for full year 2025 shows net losses incurred of $551,755 million and loss expenses incurred of $86,003 million, a 66.5% net loss ratio, a 25.8% expense ratio and a 92.9% combined ratio, against $68,742 million of net underwriting gain. Direct premiums written reached $1.105 trillion, up 4.2%, with direct losses incurred of $617.6 billion and a 57.0% pure direct loss ratio, down 4.7 points.

Put the two sources together and the illustrative scale is easy to state and easy to over-claim. Net losses incurred plus loss expenses incurred sum to $637.8 billion for 2025, and applying EY's 7% to 14% range to that sum gives $44.6 billion to $89.3 billion. That range is arithmetic performed here on two published inputs. It is not an EY estimate, not an NAIC estimate, and EY's range describes total claims spend at the carriers its team assessed rather than an industry constant. Use it to size a conversation, never a business case. The business case comes from your own stratum, with your own interval attached. To put your own numbers against the same arithmetic, the claims leakage calculator runs it on your paid losses rather than on the industry total.

Sampling is a capacity constraint, not a methodology

Answer

Why do carriers only audit a sample of claim files?

Carriers sample because reviewer hours are finite, not because a sample is the right way to measure leakage. NAIC guidance needs 1,067 files for a plus or minus three point interval at 95% confidence when no prior proportion exists, and bars extrapolating past them. US claims adjuster employment is projected to decline 6 percent through 2035.

Start with the reviewer-hour bill, the thing nobody puts on the slide. A defensible first-stage sample with no prior estimate of the error rate is 1,067 files, and a real leakage review of one litigated bodily injury file is not a checklist tick: it is the policy, the medical records, the estimate history, the reserve movements, the negotiation log and the defense file, then a priced and evidenced variance. Multiply by a thousand and the audit is a program with a budget and a calendar, which is why most carriers run it once a year, on one line, in one state.

Hiring out of the problem is getting harder, not easier. The US Bureau of Labor Statistics counts 389,700 claims adjusters, appraisers, examiners and investigators in 2025 and projects employment to decline 6 percent from 2025 to 2035, with about 21,600 openings a year attributed largely to replacement rather than growth, at a May 2025 median wage of $78,020 for the occupation as a whole. The answer to a sampling constraint cannot be to hire more reviewers when the occupation itself is shrinking.

So the method adapts to the constraint and then gets described as if the constraint were a choice. Athenium Analytics sells claims review and quality audit software and is a real incumbent on this workflow. It describes a leakage study on its own site as sampling "a large, random sample of files" with results "extrapolated to provide an indication of the financial impact." That is an accurate description of how the industry runs these studies, and a capable product doing the standard thing well. It also sits against the NAIC guidance quoted earlier. EY is in the same structural position: the 7% to 14% range is the most defensible published sizing available, and it comes from periodic, sampled engagements whose own third root cause is that investigations did not properly establish causation and liability.

A sampled leakage rate is a measurement of the files you had the hours to open. Everything outside the sample is not clean, it is unmeasured. Closing that gap is not a better sampling design. It is applying the audit's evidence standard at the moment the decision is made.

Hesper AI product research

Which is what the whole checklist has been building toward. The evidence test above is a good test, good enough that a regulator wrote it, a consultant prices findings against it, and an internal audit team uses it to overturn decisions. But it is applied to a small share of files, months after the money left, by people whose hours are the binding constraint. Therefore the question worth asking is not how to sample better. It is what changes if the same test runs on every file, at the time the decision is made, by something without an hourly capacity ceiling.

That is the shape of the Hesper argument, and it is a capacity argument rather than a cleverness argument. Agents run 15+ investigation phases in parallel on a single file and complete the work in hours, not weeks, against a manual baseline of 14+ days per case (Hesper internal benchmarks). Throughput moves from roughly 10 investigations per investigator per month to 800+ cases per investigator per month, which is what takes full-workup coverage of flagged claims from about 25% to 100% (Hesper internal benchmarks). Speed and leakage are the same problem, and the sampling constraint is where that stops being a slogan and starts being arithmetic.

Running the checklist on every file

Answer

Can a claims leakage audit run on every file instead of a sample?

Applying the checklist at the time of the decision rather than a year later turns the audit from a retrospective into a control. The same evidence a finding needs - the provision relied on, the dated record of what was checked, the calculation a payment reconciles to - is what the file should have carried before it closed.

Stage by stage, this is what the checklist looks like when it runs as a control instead of as a review. The coverage position is written with the policy version in force on the date of loss and the specific provision cited, before the payment authority is granted. The investigation records what was checked and what was ruled out, on the file, dated, whether or not anything was found. The valuation carries its comparables. The payment reconciles to a calculation that is attached to it. Every file is screened for subrogation and salvage, rather than the subset an adjuster had time to flag. Evidence behind every decision, not only behind the suspicious ones.

That last distinction is the whole design. Detection vendors score suspicious claims and hand off. Lifecycle vendors move claims through stages and leave the evidence to the adjuster. Hesper runs both paths on the same engine: clean claims resolve straight through, suspicious claims get an investigation-grade workup, and the evidence depth built for SIU work sits behind the coverage position, the settlement valuation and the payment check as well. That is also why the recovery stage is not bolted on at the end, and why the platform is described as first notice to final recovery rather than as a set of point tools.

The audit trail is the primary output rather than a log of it. Every agent action is recorded with its sources, its reasoning and its timestamps, attached to the claim record, which is the same reconstruction standard Model Regulation #902 and 10 CCR 2695.3(a) already require and the same artifact a capacity partner or a client auditor opens. Adjusters keep decision authority throughout: coverage, reserve, settlement and denial stay with people, and the split is set out in which decisions stay human. Adjusters review evidence-backed files instead of building them, and the per-file review workflow is on the claim file review page.

For technology and compliance reviewers the answers are narrow. Hesper integrates with the claims system of record by API and can pick a claim up at any stage rather than replacing the administration system. It holds SOC 2 Type I. It does not train on customer data. Fraud detection is built in, so it runs standalone or alongside FRISS, Shift or Verisk. For a TPA the same trail is what a client auditor opens across every program, on the TPA page; for an MGA it is what a capacity partner opens at the delegated-authority audit, on the MGA page.

What changes in the reporting is not that the leakage number gets better. It is that the number acquires a second half. A claims executive still reports an interval on a defined population, because that is what a sample can support. What is new is the control statement alongside it: here is the evidence standard now applied to every file in the book, here is the date it started, here is the reconstructable record behind each of those decisions. The audit stops being the only place the standard exists.

Key takeaways

  • A claims leakage audit prices the gap between what was paid and what the claim file supports, per file and coded to a stage, which is a different exercise from benchmarking a loss ratio or second-guessing a settlement in hindsight.
  • EY's claims quality assessments put leakage at 7% to 14% of carriers' total claims spend, and in its US insurer case study leakage ran to 10% of total paid with nearly two-thirds of reviewed files carrying some and more than 85% of it concentrated in coverage determination, litigation prevention, and evaluation and resolution.
  • NAIC market conduct sampling guidance sets a 95% minimum first-stage confidence level, needs 1,067 files for a plus or minus three point interval when the population proportion is 50%, bars extrapolation beyond the field of files sampled, and states that tolerable error is not a safe harbor.
  • The documentation standard for a defensible finding is already written into NAIC Model Regulation #902 and California's 10 CCR 2695.3(a), both of which require the file to permit reconstruction of the insurer's activities and the dates of the events.
  • Recovery is the only stage where the money found is still collectable, and roughly one out of four property-liability insurers make no salvage or subrogation recovery efforts at all, against a sample average of 4.5% of net claims paid.

Frequently asked questions

A claims leakage audit is a structured review of closed claim files that prices the difference between what the insurer paid and what the file evidence supports. The variance is calculated per file, coded to a stage such as coverage, investigation, estimation, settlement, payment or recovery, and rolled up against paid losses plus allocated loss adjustment expense for the population reviewed. EY's Property and Casualty Claims Transformation practice puts leakage at 7% to 14% of carriers' total claims spend based on its latest assessments. It is an internal financial and quality exercise, and it differs from a state market conduct examination in purpose and consequence: the audit is voluntary and priced in dollars, the examination is compulsory and priced in penalties. The underlying file defect is frequently identical.

There is no published industry total, and any figure presented as one deserves a check of its source. The most defensible sizing is EY's, whose Property and Casualty Claims Transformation practice puts leakage at 7% to 14% of a carrier's total claims spend based on its latest claims quality assessments. For scale, NAIC reported $551.8 billion of net losses incurred and $86.0 billion of loss expenses for US property and casualty insurers in full year 2025. Applying EY's range to that $637.8 billion sum gives $44.6 billion to $89.3 billion. That is arithmetic performed on two published inputs rather than a published estimate, and EY's range describes the carriers its team assessed, so treat it as a scale check and size your own book from your own stratum.

It depends on the expected error rate in the population, not on a fixed file count. NAIC market conduct sampling guidance shows 1,067 files for a plus or minus three percentage point interval at 95% confidence when the population proportion is 50%, and 203 files for the same interval when the proportion is 5% or 95%. The guidance sets 95% as the minimum confidence level for a first-stage acceptance sample and says a second-stage sample should never fall below 90%. Treat vendor guidance that prescribes a flat number of files per line of business with caution, because it skips the step that actually determines the answer. Whatever size you land on, the result describes only the field of files you drew from.

Enough to reconstruct the claim without talking to the adjuster. NAIC Model Regulation #902 section 4.B requires detailed documentation in each claim file in order to permit reconstruction of the insurer's activities relative to each claim, and section 4.C requires each relevant document to be noted as to date received, date processed or date mailed. California's 10 CCR 2695.3(a) requires all documents, notes and work papers, including copies of all correspondence, in such detail that pertinent events and the dates of the events can be reconstructed. Pennsylvania carries its own file and record documentation standard at 31 Pa. Code 146.3. The NAIC Market Regulation Handbook states it most simply as Claims Standard 5: claim files are adequately documented.

In EY's US insurer case study, more than 85% of the assessed leakage sat in three areas: coverage determination, litigation prevention, and evaluation and resolution. That points at coverage and settlement rather than at fraud, which is a separate problem with a separate workflow. Recovery deserves its own place on the list for a different reason: it is the only stage where the money you find is still collectable. Research published in the NAIC Journal of Insurance Regulation computes $51.6 billion of salvage and subrogation recovered in 2021 across three auto lines, and finds that roughly one out of four property-liability insurers make no recovery efforts at all, against a sample average of 4.5% of net claims paid.

One is voluntary and priced in dollars, the other is compulsory and priced in penalties. A leakage audit is commissioned internally and measures the gap between paid amounts and what the file supports. A market conduct examination is run by a state insurance department and tests statutory compliance. In May 2026 the California Department of Insurance announced an examination of one carrier's 2025 Los Angeles wildfire claims in which examiners reviewed 220 claims out of roughly 11,300 residential claims and alleged 398 violations across 114 sampled files. California Insurance Code section 790.035 makes penalties of up to $5,000 per violation available, and $10,000 where a violation is willful. The underlying file defect is frequently the same fact in both reviews.

Annual closed-file review is the common cadence, and it has a structural flaw worth naming: a sample-based annual audit measures leakage roughly a year after it left, on the small share of files reviewer hours could cover, and by the NAIC's own rule it cannot speak to the files outside the sample. Raising the cadence runs into the same constraint that set it. The US Bureau of Labor Statistics counts 389,700 claims adjusters, appraisers, examiners and investigators in 2025 and projects employment to decline 6 percent through 2035, so the answer cannot be more reviewers. The workable version is to keep the periodic audit for measurement and move the evidence standard itself into handling, so every file is built to pass.

On the Hesper platform

Hesper AI runs every stage of a claim, from first notice of loss to subrogation recovery, with investigation-grade evidence behind every decision.

The platformFNOL intakeClaims triageFraud investigationSubrogation & recovery
← More articles on the Hesper AI blog

See Hesper AI on your documents

Request a demo and we'll run an analysis on your real document samples.