Admiral's fraud team published a list of what it caught in 2025. One entry: the same damaged car, re-edited with a different number plate, submitted again as a separate claim. No deepfake, no synthetic image. One real photograph with one field changed - cheaper to produce than a generated image and harder for any detector to score, because everything else in the frame is genuine.
That is the shape of the problem, and Kayla M. McCallum diagnosed it correctly in Claims Journal in May 2026, then prescribed the right checks. The trouble is not the checklist. Every item on it is a manual step, and at 200+ cases per investigator and 14+ days per manual case, a US carrier can run that checklist on roughly a quarter of the claims it flags. The other three quarters are paid, denied without full work, or queued.
The second problem is deeper and it is not hers to solve. The pixel-level test does not close. SWGDE, the forensic standards body whose guidance is cited in US courts, states that manipulations can be performed in a single image which a trained forensic practitioner may not adequately detect. Peer-reviewed detector performance falls from 0.98 AUC inside its training distribution to 0.65 on a dataset it has not seen. Any control that rests on "is this JPEG fake" rests on a signal that degrades every time a new generator ships.
The durable defense is corroboration, not detection. Fabricated evidence is cheap to make and expensive to corroborate, and that asymmetry runs in the carrier's favor - but only if the carrier actually runs the corroboration. What follows maps the published playbook step by step onto agent tool calls, marks the two steps that stay human and should, and sets out what the resulting claim file has to survive. For the scale of the problem and the ways generative tools are being used against carriers, start with the scale piece: deepfake insurance claims and AI-generated fraud in 2026. This post is about the investigation.
The playbook Claims Journal published, and the step it stops at
An AI-generated evidence playbook is the ordered set of checks that establishes whether submitted proof came from the loss it claims to document. Kayla M. McCallum set that list out in Claims Journal on May 15, 2026: native files, metadata, visual artifact review, direct verification with the issuing source, reverse image search, site inspection, examination under oath, and a good-faith standard behind any denial.
McCallum is an associate attorney at Swift Currie in Atlanta, practising automobile litigation, first-party property, insurance coverage and premises liability, per her firm biography. That matters for how the list should be read. It is written by a litigator, for the people who will have to defend the file, and it is close to what a court will expect to see behind a denial that turns on the authenticity of a photograph. Her piece, Adapting Claim Investigations for AI-Driven Fraud, opens on the Coalition Against Insurance Fraud figure of $308.6 billion in annual US insurance fraud losses, and works down from there to what an adjuster should do differently on Monday morning.
Read as a specification rather than as advice, the playbook has eight steps:
- Request native or original files rather than screenshots and forwarded attachments.
- Examine metadata: creation date, capture device, GPS coordinates, edit history.
- Look for generation artifacts in the image itself: misaligned shadows, impossible reflections, corrupted or nonsensical text.
- Verify documents directly with the issuing source - the vendor, contractor or repair shop named on the invoice.
- Run reverse image search to trace photographs to prior incidents, salvage listings or stock libraries.
- Conduct in-person site inspections and witness interviews where the file warrants them.
- Use the examination under oath, structured broad to narrow, when documentary inconsistencies exist.
- Hold the good-faith investigation standard: support any denial with concrete evidence and a documented basis.
Nothing in that list is new and nothing in it is wrong. Steps 1 through 5 were in SIU practice long before diffusion models, and step 8 is the law. What generative editing changed is not the method, it is the hit rate: checks that used to be worth running on the handful of claims that smelled wrong are now worth running on everything, because the cost of producing a convincing fake fell to roughly zero while the cost of checking it did not move at all.
That is the unpriced part of the prescription. A manual SIU investigation runs 14+ days per case and a single investigator carries 200+ cases, so full forensic work lands on about 25% of flagged claims. The standard SIU process is not skipping these steps out of ignorance. It is skipping them out of arithmetic. A checklist that only runs when someone has time is not a control; it is a hope.
What is actually in the claims pipeline in 2026
The encounter rate for AI-altered claim material is now effectively universal and the incidence rate is unknown. Verisk's 2026 State of Insurance Fraud study found that 99% of surveyed US claims professionals had already seen manipulated or AI-altered documentation. No published study says what share of claims actually contains it, and the two numbers are not the same thing.
The Verisk study is the largest and most recent US-scoped dataset on this question, and it is worth citing precisely: 1,000 US consumers plus 300 insurance claims professionals at manager level and above, fielded December 2025 to January 2026, reported by Digital Insurance in April 2026 and published by Verisk. Alongside the 99% encounter rate, 98% of insurers agree AI-powered editing tools are driving a rise in digital fraud and 76% say altered submissions became more sophisticated over the prior year. On the consumer side, 36% said they would consider digitally altering a claim image or document, and that willingness is sharply generational.
Encounter rate is not incidence rate
99% of claims professionals having seen AI-altered material does not mean 99% of claims contain it. It means almost every professional has encountered at least one instance across a career of thousands of files. Nobody has published the incidence number, and no honest model of your exposure can be built from the encounter number. Treat any vendor claim of the form "X% of claims now contain AI-generated evidence" as unsourced until they name the study, the sample and the field dates. The defensible planning assumption is that the material is present, growing, and unquantified.
What is quantified sits on the carrier side of the ledger. Admiral's fraud team reported £86.8m of detected fraudulent motor, home and travel claims in 2025, up 71% from £50.9m in 2024, and named the cases: a designer watch manipulated to appear damaged, rear vehicle damage added in Photoshop, internet photos of luggage submitted as the claimant's own, and the duplicate-claim number plate this post opened on. Note what is absent from that list. No cinematic deepfakes, no synthetic humans. Four ordinary edits to four real photographs, which is the version of this problem that actually reaches a claims desk.
Two other lanes are moving. RGA's Colin M. DeForge and Jennifer Johnson put synthetic identity fraud at $8 billion in 2020 rising past $30 billion by June 2025, roughly 400% in five years, and estimate synthetic identities at 80-85% of all identity fraud. On the voice channel, Pindrop's 2025 Voice Intelligence and Security Report, built on analysis of more than 1.2 billion customer calls in 2024, recorded a 475% increase in synthetic voice attacks against insurance, higher than banking at 149% or retail at 107%. And in the European financial sector specifically, Signicat's February 2025 survey of 1,206 fraud decision-makers across seven countries found the deepfake share of detected fraud attempts moving from 0.1% to about 6.5% over three years, a 2,137% rise. That last figure gets quoted as a US insurance number constantly. It is not one.
None of this is a 2026 surprise. Zurich UK's head of fraud Scott Clayton was describing registration numbers digitally implanted onto salvage-site photographs, software-generated impact damage on buildings and wholly fabricated repair invoices in January 2024. Two and a half years later the tooling is better and the volumes are larger. For the full picture on scale, our State of Insurance Fraud 2026 report collects the industry numbers in one place. The rest of this post assumes the material is in your pipeline and asks what to do about it.
The pixel-level arms race does not close
Pixel-level detection is a probabilistic signal whose accuracy falls as generators change and as images pass through ordinary compression. A detector trained and tested on the same dataset can score 0.98 AUC and drop to 0.65 on a dataset it has not seen. That single gap is the reason a claim denial cannot rest on a detector score.
The number comes from peer-reviewed work rather than vendor marketing. Yang and colleagues, writing in Frontiers in Big Data in 2025, document that the detector proposed by Qian et al. in 2020 achieves 0.98 AUC when trained and tested within the FaceForensics++ dataset and falls to 0.65 when evaluated on Celeb-DF under cross-dataset protocols. In plain terms: strong in-distribution, close to a coin flip once the generator changes. Every new image model shipped is a distribution shift.
Robustness under ordinary handling is a separate and equally unsolved problem. The NTIRE 2026 Robust Deepfake Detection Challenge report, submitted in April 2026 with 337 participants and 57 final leaderboard submissions, exists because, in the organizers' words, detection performance is nearly worthless in the real world if it suffers under exposure to even slight image degradation. Now consider the path a claim photo actually takes: phone camera, messaging app or claims portal upload, server-side resize, re-encode, thumbnail generation, sometimes a screenshot in between. Every one of those steps is image degradation. The research community is still competing on this in 2026; the claims pipeline has been shipping it into detectors for years.
The forensic standards position is the one to put in front of a general counsel. SWGDE 18-I-001 v2.0, published March 3, 2025, states that "the state of the art in digital imagery is such that in a single image, manipulations can be performed which a trained forensic practitioner may not adequately detect." The same document recommends authenticating against a series of images or video rather than a single still. That is a standards body, not a vendor, saying the single-image test does not close - and it is a considerably stronger statement than any detection vendor would make about its own product.
The generator-detector asymmetry has been measured formally in an adjacent modality. NIST's 2024 GenAI pilot study, published June 25, 2025, found that some generators could deceive most discriminators while some discriminators could detect content from almost all generators. That was the text-to-text track, not images, and it should be cited as such. The structural point carries: in a generator-versus-detector race, the population of both is heterogeneous and the ranking is unstable. You cannot procure your way to a permanent answer.
The market has already tested this empirically. Detection tools are widely deployed and confidence has not followed them.
Read those four bars together. Roughly two-thirds of insurers have bought detection and half have built their own, and still only 32% are very confident they could identify a deepfake. Confidence also collapses as the artifact gets more synthetic: 58% on edits to a real photograph, 32% on a fully generated one. That is not a vendor failure and it should not be sold as one. It is the pixel-level ceiling showing up in a survey.
Riedman is describing the gap between a signal and a resolved claim file, and he is describing it accurately from inside the vendor that produced the dataset this post leans on. That gap has a name in the layered model we use: detection scores the claim, investigation resolves it. Verisk, FRISS and Shift Technology all operate at the detection layer and do it well; the investigation layer beneath them has no software incumbent, only manual SIU teams. A carrier running detection plus autonomous investigation is the modal deployment, not a replacement decision. For where each tool sits and what each layer can and cannot answer, see our complete guide to insurance fraud detection methods, tools and gaps.
The operating rule that follows is short. A detector score is an input to an investigation, never a substitute for one. Write it into the file as one cited finding among many, with the model, the version and the confidence recorded, and then go find out whether anything outside the image agrees with it.
Metadata and provenance are necessary, not sufficient
Metadata and Content Credentials are worth reading on every claim image, and neither can carry a denial. Both are one-directional signals. Present, valid provenance is strong evidence of an honest capture path. Absent provenance is evidence of nothing, because messaging apps, screenshots and ordinary edits strip it from legitimate photographs too.
What EXIF can and cannot establish
The standards body is explicit. SWGDE 17-I-001 v1.1, section 6.2, states that metadata "may be edited, intentionally or unintentionally, or lost," and that this loss "can impact the ability to establish the provenance of imagery after the fact." The companion image-authentication document adds that metadata used to establish source or processing history "can be limited, absent, or altered." Metadata is writable. Anyone who can fabricate the photograph can fabricate the EXIF block attached to it, and the tools to do so are older and simpler than the tools to generate the image.
That does not make it worthless, because it is asymmetrically cheap. Extraction costs nothing and occasionally it is decisive: a GPS coordinate 400 miles from the reported loss location closes a file in one step. The right posture is to treat metadata as a lead generator, not as proof. It produces hypotheses - this image was created eleven days before the reported loss, this device does not match the one on the last three submissions from this claimant, this file has been through two encoders - and each hypothesis then gets tested against a record the claimant does not control.
The same logic applies on the document side, where the artifacts are different but the reasoning is identical. A PDF carries an incremental save history, producer strings, object structure and font subsetting that are all readable and all forgeable. We walk that in detail in how to tell if a PDF has been edited or tampered with. The conclusion there is the same as here: the container tells you where to look, and the outside world tells you what is true.
Content Credentials are a positive signal only
C2PA is the piece of provenance infrastructure most likely to matter to claims intake, and its own specification is precise about its limits. The C2PA technical specification 2.4 lists as a guiding principle that the specs "SHOULD NOT provide value judgments about whether a given set of provenance data is 'good' or 'bad,' merely whether the assertions included within can be validated as associated with the underlying asset, correctly formed, and free from tampering." It also acknowledges that "an asset can become separated from its C2PA Manifest due to removal or corruption of asset metadata."
Two operational consequences follow, and carriers get the second one wrong. First, a present and valid manifest chaining back to a known capture device is a genuinely strong positive signal, and a claims intake path that preserves manifests instead of stripping them is worth building. Second, an absent manifest means nothing at all. Most consumer capture paths, most editing tools and most messaging apps do not preserve credentials. Any control that treats a missing manifest as evidence of fabrication will deny honest claims at volume, and a bad-faith exposure created by a provenance rule is a worse outcome than the fraud it was written to stop.
The asymmetry: cheap to fabricate, expensive to corroborate
The asymmetry that favors the carrier is not detection accuracy. It is cost. A generator produces a photograph of damage for free in seconds. It cannot produce a repair shop that answers the phone, a valid contractor licence, a clean prior-claim history across carriers, a GPS track, a weather record and a policy-inception date that all agree with each other and with the story.
That is the whole argument, and it inverts the usual framing. The pixel question - is this image synthetic - is a question the fraudster gets to keep answering, because they control the generator and they can iterate until the artifact score comes back clean. The corroboration question - does anything outside this document agree with it - is a question the fraudster does not control, because it is answered by records held by third parties who have no idea a claim was filed. Fabricating one photograph is free. Fabricating a consistent world around that photograph is expensive, and it gets more expensive with every independent source checked.
The corroboration surfaces are enumerable, and none of them require looking at a single pixel:
- Source-of-record verification on every named entity: repair shop, contractor, medical provider, adjuster, tow operator. Business registration, licence status, address, phone reachability, and whether the entity existed before the loss date.
- Invoice-format consistency against that vendor's known documents, and against other invoices from the same vendor across the carrier's own book.
- Prior-claim and cross-carrier matching on the claimant, the vehicle, the property and the vendor, including the same loss submitted twice with one field changed.
- Perceptual-hash and reverse image search across prior claim submissions, salvage listings, marketplace photographs and stock libraries.
- Loss-date reconciliation: image timestamp against reported loss date, GPS against reported location, and both against the weather record for that place and hour.
- Timeline consistency across the whole file - FNOL statement, recorded statement, medical records, repair estimate and photographs all describing the same sequence of events.
- Policy-inception-to-loss interval, coverage-change history, and prior-claim cadence on the policy.
Every item on that list is a database lookup, a document comparison, a phone verification or a cross-reference. All are individually unremarkable. What makes them a control rather than a wish list is running them together, on every flagged claim, without a human deciding which ones there is time for. That is the mechanism: 15+ investigation phases executing in parallel rather than a sequential checklist gated by one investigator's attention. We describe the architecture in parallel processing in SIU: running 15 investigation phases simultaneously.
There is also a structural reason the corroboration layer keeps getting skipped, and it is not laziness. Straight-through processing pays a clean claim in minutes, and it does that by reading a loss photograph for severity rather than for provenance. That is a design property of touchless estimating, not a criticism of any particular estimating vendor: a pipeline optimized to turn an image into a dollar figure in seconds has no step at which anyone asks where the image came from. Generative editing made the input to that pipeline free. The gate has to sit somewhere, and the only place it can sit without giving back the cycle-time win is inside an automated investigation layer that runs faster than the payment does.
The checklist as parallel tool calls
The published playbook automates almost completely, and what automates is the execution rather than the method. Six of the eight steps map onto tool calls an investigation agent runs in parallel on every flagged claim. Two do not automate and should not. The table below is the mapping, step by step, with the manual cost of each.
A generator can fake a photograph for free. It cannot fake a repair shop that answers the phone. The same claim through a touchless path and through an investigation layer - six standard forensic checks, run on every claim instead of on the ones somebody has time for.
None of steps 1 through 5 and 8 are new. McCallum did not invent them and neither did we. What changed is their marginal cost. At 200+ cases per investigator and 14+ days per manual case, a carrier can afford this checklist on roughly 25% of flagged claims. As parallel tool calls, the same checklist runs on 100% of them, in hours rather than weeks. That coverage shift is the entire argument, and it is the thing no amount of additional detection accuracy delivers - a better score on a claim nobody investigates changes nothing. We took the coverage math apart in why 75% of flagged insurance claims are never fully investigated.
For the person holding the budget, the unit economics are the short version. A manually investigated case costs on the order of $2,500 in loaded investigator time; an automated one runs closer to $150, and coverage moves from about 25% of flagged claims to all of them. Those are Hesper internal benchmarks and the ratio matters more than the decimal. Cost per case is the wrong line to bring to a loss-ratio conversation anyway. Bring the other one: the claims you were never going to investigate are now investigated, and that is where the leakage sits.
This is what the phrase from fraud detection to fraud resolution means in practice on this specific problem. Detection answers whether the photograph looks manipulated. Resolution answers whether the repair shop exists, whether the same vehicle appeared on a salvage listing three months ago, whether the image predates the loss, and whether the resulting conclusion is written down with sources. The investigator's role shifts from execution to decision-making: the agent runs the phases, the investigator reads the assembled contradictions and makes the call.
What stays human
Two steps in the playbook stay human: the in-person site inspection and the examination under oath. So does the coverage decision itself. What changes is that the human arrives with the timeline, the conflicts and the exhibit index already assembled, so the visit or the EUO tests a specific hypothesis rather than starting cold.
This is not a hedge and it is not a limitation to be engineered away later. An EUO is a legal proceeding with a transcript, conducted by counsel, in which a person answers under oath. A site inspection is a licensed human standing in a room forming a judgment about what happened there. Neither is a data problem. The value automation adds to both is upstream: an EUO built from a documented list of inconsistencies, with each exhibit already tied to the finding it contradicts, is a substantially different proceeding from one built the night before from a claim file nobody has fully read.
The industry's own read lands in the same place. Kedar Kamalapurkar, Managing Director in Deloitte Consulting's insurance sector claims practice, told Insurance Journal in June 2025 that "AI plus human is going to be better than human alone or AI alone." The same piece reports Deloitte projecting $80 billion to $160 billion in P&C savings by 2032 from AI across the claims lifecycle, and notes that 35% of insurance executives named fraud detection a top-five AI application area in a June 2024 survey.
The same article gives the baseline that makes the case concrete. Soft fraud accounts for roughly 60% of incidents and is detected at 20-40%; hard fraud is about 40% of incidents and is detected at 40-80%. The weakest detection rate sits on the largest category, and soft fraud is precisely where AI-altered evidence lives: a real loss with one inflated or fabricated element. That is the case a pixel detector is least useful on, because most of the image is genuine, and the case corroboration resolves cleanly, because the fabricated element has no supporting record behind it.
For an SIU director, ignore the accuracy claim and ask three questions instead. Can my investigator see every step the agent took. Can they override any finding. Can I hand the resulting file to a state examiner or opposing counsel without translation. If the answer to any of those is no, the tool is a liability regardless of how well it scores images. An investigation product that cannot be cross-examined is not an investigation product.
What the file has to survive
An AI-generated-evidence denial gets litigated, so the record has to be reconstructable by someone who was not there. The operative US requirements already exist independent of any new AI rule: California 10 CCR 2698.36 requires a documented decision trail on SIU referrals, and NAIC Model Act #680, adopted in 48 states, governs antifraud-plan filing.
The federal evidentiary rules are still in motion and should be described carefully. The Judicial Conference's Committee on Rules of Practice and Procedure approved a proposed new Federal Rule of Evidence 707, addressing machine-generated evidence, for publication on June 10, 2025, and the public comment period ran from August 15, 2025 to February 16, 2026, per the US Courts published amendments page. It is not adopted. Anyone telling you the rules for AI evidence are settled has not read the docket.
Rule 707 is proposed, not adopted
Proposed FRE 707 covers machine-generated evidence offered without a human expert. It was published for comment on June 10, 2025 and the comment window closed February 16, 2026. It has not been adopted, and no carrier should build a control on the assumption that it has. The practical implication is the opposite of waiting: until the standard settles, an AI-assisted denial is litigated under existing authentication and good-faith standards, which means the carrier's protection is the completeness of its own record rather than a rule that blesses the method.
That cuts both ways, and the second direction is the one carriers underweight. A claimant can submit AI-generated evidence, and a carrier can also generate evidence about the claimant with AI. Both are machine-generated material entering a file. If the reasoning behind a denial exists only as a model output with no cited sources and no timestamps, the carrier is in the same evidentiary position as the fraudster whose photograph it rejected: holding an artifact nobody can trace back to a source.
Which is why the audit trail cannot be a report generated after the fact. Every check has to write a cited, timestamped finding as it runs, with the source consulted, the query issued, the result returned and the reasoning that connected it to a conclusion. Built that way, the file that satisfies the SIU lead, the file that satisfies a state DOI examiner sampling antifraud-plan compliance, and the file that goes into discovery are the same document. Built as a summary written afterwards, they are three different documents and two of them will not hold.
One more structural point for compliance. Fraud is a specific crime in 48 states per the Coalition Against Insurance Fraud, which also puts annual US insurance fraud at $308.6 billion and finds fraud in roughly 10% of property-casualty losses. The NAIC repeats the $308.6 billion figure and carries the FBI estimate that fraud adds $4,000 to $7,000 to a family's premiums over ten years. A referral that goes to a state fraud bureau becomes someone else's evidentiary problem, and the quality of what you hand over determines whether it is prosecutable. The litigation-side treatment of all of this is in the general counsel's guide to AI fraud investigation.
The through-line from the top of this post: you cannot win the pixel argument, and you do not have to. You have to be able to show, on any claim, what was checked, what was found, what disagreed with what, and who decided. That record is cheap to produce when it is a by-product of the investigation and nearly impossible to reconstruct when it is not.
Key takeaways
- The manual playbook published in Claims Journal in May 2026 prescribes the right checks - native files, metadata, artifact review, source verification, reverse image search, inspection, EUO and a documented good-faith basis - and every one of them is a manual step that a carrier running 200+ cases per investigator can afford on roughly a quarter of flagged claims.
- Pixel-level detection cannot carry a denial: peer-reviewed detector performance falls from 0.98 AUC in-dataset to 0.65 cross-dataset, SWGDE states that a trained forensic practitioner may not adequately detect single-image manipulation, and only 32% of insurers are very confident they could identify a deepfake despite 65% already using third-party detection tools.
- Metadata and Content Credentials are one-directional signals, so present provenance is a strong positive and absent provenance proves nothing - any control that treats a missing C2PA manifest as evidence of fabrication will deny honest claims at volume.
- The durable defense is corroboration: a generator produces a photograph for free but cannot produce a repair shop that answers the phone, a licence record, a clean cross-carrier claim history, a matching GPS track and a weather record that all agree, and that cost asymmetry runs in the carrier's favor only if the corroboration is actually run.
- Six of the eight playbook steps map onto tool calls that run in parallel on 100% of flagged claims in hours rather than weeks; the site inspection and the examination under oath stay human, and the human arrives with the timeline, the conflicts and the exhibit index already assembled.