---
title: "Shift's ARISE ladder has two columns. Efficiency climbs 8x, the money column moves 3 points"
description: "Shift's ARISE ladder pairs an efficiency gain with an indemnity impact at every level. Across the five rungs efficiency multiplies 8x and the outcome column moves 3 points. Buy on the second column."
date: "2026-09-18"
lastModified: "2026-09-18"
author: "Pankaj Dhariwal"
tags: ["Vendor Comparison"]
canonical: "https://gethesperai.com/blog/arise-autonomy-efficiency-curve-investigation-roi/"
---

# Shift's ARISE ladder has two columns. Efficiency climbs 8x, the money column moves 3 points

> **TL;DR** Shift Technology's ARISE framework attaches an efficiency gain to each of five autonomy levels, and the same published table attaches an indemnity impact. Across the ladder the efficiency column multiplies eightfold while the outcome column moves three percentage points. A carrier buying a fraud-investigation capability is purchasing the second column, and almost no RFP scores it.
>
> - ARISE efficiency runs 10, 20, 30, 50, 80 percent
> - Indemnity impact moves 3 points across the same five rungs
> - 43% of US insurers never capture Impact Ratio at all

- **50% / 2%** - Efficiency gain and indemnity impact at ARISE Level 4 (Shift Technology, ARISE report, June 3 2026)
- **43%** - US insurers that do not capture Impact Ratio (Coalition Against Insurance Fraud and PwC, 2023)
- **13%** - Carriers very confident they can develop an AI ROI measurement (AM Best survey via Insurance Journal, May 2026)
- **~25% -> 100%** - Flagged-claim investigation coverage (Manual SIU vs an investigation layer (Hesper internal benchmark))

Shift Technology's ARISE framework attaches an efficiency gain to each of its five autonomy levels: 10% at L1 Answers, 20% at L2 Recommends, 30% at L3 Initiates, 50% at L4 Solves and 80% at L5 Exceeds. Those five figures are now visible on Shift's [short ARISE explainer](https://www.shift-technology.com/resources/reports-and-insights/the-arise-framework-explained-in-under-a-minute), dated August 5, 2026, and they are set out in full in the [June 3, 2026 ARISE report](https://www.shift-technology.com/resources/reports-and-insights/arise-a-standard-framework-for-ai-agent-autonomy-in-insurance). The report states the figures are drawn from production deployments rather than aspiration. Take that at face value. Nothing in this post depends on the numbers being soft.

The same table carries a second column, and it is the one that does not travel. Indemnity impact runs from no stated figure at L1 to 1% at L2, 1% at L3, 2% at L4 and 3% at L5. Read the two columns down the ladder and the shape is hard to miss: across five rungs the efficiency figure multiplies eightfold while the outcome figure moves three percentage points. Shift built that table, published both columns, and stands behind both as production measurements. The question this post is about is which column a carrier buying a fraud-investigation capability is actually buying, and the answer is the second one.

This is a procurement argument, not a product takedown. ARISE is scoped by Shift to claims processing and scoped correctly: the worked-example column in the report is headed "Example: claims processing," and every example is a handling decision, a coverage question, a repair initiation, a pre-filled payment. The failure mode is what happens when a single efficiency figure travels out of that scope and into a fraud-investigation RFP, where it becomes the number everyone scores because it is the only number on the page. That is a buyer error, and it is the buyer this post is written for.

What follows is the published table in full with both columns, the definitional reason an efficiency ratio cannot describe coverage, the outcome metric the US industry already named and mostly does not capture, the arithmetic a CFO needs to turn an autonomy claim into loss-cost, three vendor postures on outcome reporting, and an eight-metric set to put in the RFP. The taxonomy argument, that autonomy level is one axis and investigation needs a second, is made separately in [ARISE levels of autonomy vs investigation-specific AI](/blog/arise-framework-investigation-autonomy-rebuttal/). This post leaves the taxonomy alone and reads the table.

> **The posture of this post**
>
> Shift published a two-column table and said the numbers come from production deployments. Both of those are to its credit, and this post concedes both in the first hundred words. It does not argue that the efficiency figures are inflated, illustrative or marketing, and it never says Shift promises 80% savings, because 80% belongs to L5, the top of the taxonomy, and Shift places its own production capability at L4. The argument lands on a buyer who carries a claims-handling metric into a fraud-investigation procurement, never on Shift's construction of the framework.

## What the ARISE ladder publishes, in full

ARISE is Shift Technology's five-level taxonomy for AI agent autonomy in insurance, running from L1 Answers through L5 Exceeds. Its published table pairs every level with two numbers rather than one: an efficiency gain of 10, 20, 30, 50 or 80 percent, and an indemnity impact of no stated figure, 1, 1, 2 or 3 percent.

Two documents matter and they carry different dates. The full report, "ARISE: A standard framework for AI agent autonomy in insurance," is dated June 3, 2026 and contains the complete table, the level definitions, the efficiency column and the indemnity column. A short explainer dated August 5, 2026 carries the five efficiency percentages on their own. The report is the citable artifact for everything below. The explainer is where most readers will have first met the percentages, stripped of the second column and of the definitions that scope them.

| Level | Shift's gloss | Definition as published | Efficiency gain | Indemnity impact |
| --- | --- | --- | --- | --- |
| L1 Answers | Intelligent Information Retrieval | "The AI agent responds to direct questions, retrieving and synthesizing relevant information from policy documents, claims records, and regulatory sources." | 10% | No figure stated |
| L2 Recommends | Situational Analysis and Recommendation | "The agent analyzes the full situation ... and recommends best next steps with clear rationale." | 20% | 1% |
| L3 Initiates | Human-Validated Autonomous Execution | "The agent initiates all required checks, pre-fills decision parameters, and presents a fully validated action package ready for one-click human approval." | 30% | 1% |
| L4 Solves | Full Straight-Through Processing at 99%+ Accuracy | "The agent acts end-to-end without human intervention, achieving 99%+ accuracy by applying contractual, regulatory, and insurer-specific logic consistently at scale." | 50% | 2% |
| L5 Exceeds | Superhuman Performance and Process Innovation | "The agent not only operates autonomously but surpasses the outcomes of the top 1% of human performers, proactively identifying process inefficiencies and deviating intelligently to optimize results." | 80% | 3% |

Read the table before reacting to it. The progression is coherent and the efficiency figures rise with it in a way that makes operational sense: the more of a handling step the agent takes, the less human time the step consumes. An information-retrieval agent saves a lookup. A straight-through agent saves the whole task. Ten to eighty percent is a reasonable shape for that curve, and the level definitions are specific enough to argue with, which is more than most autonomy marketing offers.

One feature of the table deserves attention because it is the guardrail that gets dropped in transit. The report's worked-example column is headed "Example: claims processing," and each level's example is a handling decision: retrieving a policy provision, recommending a next step with rationale, pre-filling a decision package for one-click approval, executing a payment under contractual and regulatory logic. Shift scoped the curve to claims handling and labelled the scope on the face of the table. By the time a single percentage reaches a slide in a procurement deck, the column header is gone.

The obvious line of attack on any vendor-published percentage is that it is a projection dressed as a measurement. That line is not available here, because the report closes it explicitly.

> These metrics are not aspirational. They are drawn from Shift's current production deployments.
>
> - Shift Technology, "ARISE: A standard framework for AI agent autonomy in insurance," June 3, 2026

Concede it and move on. If the efficiency figures are production measurements, they are better evidence than most of what circulates in claims-technology marketing, and the interesting question shifts. It is no longer whether the number is real. It is what the number is a measurement of.

### Where Shift places itself on its own ladder

The report also says where Shift sits, which matters because the 80% figure gets quoted loosely. Shift states that it has achieved L4 Solves capability in production for auto glass and APD repairs, property electronic devices, property building and content losses, workers' compensation coverage, medical bill review, and travel trip interruption claims, and that its L1 agents are deployed in production across auto, property, workers' compensation and travel lines today. L4 pairs 50% efficiency with 2% indemnity impact. The 80% figure sits at L5, the top of the taxonomy, which is a definition of a capability ceiling rather than a statement about a shipped product.

That distinction is worth protecting in both directions. A buyer who quotes 80% at a vendor is quoting the ceiling of a framework, not a promise anyone made. A buyer who quotes 50% is quoting Shift's stated production position, and should quote the 2% that sits beside it in the same row. Half a row is not a citation.

*Figure: Shift published two columns. Only one of them travels. Across the same five rungs the efficiency column multiplies 8x and the indemnity column moves 3 points, both from Shift's own ARISE report of June 3, 2026, which states the figures are drawn from production deployments. Shift's stated production position is L4, pairing 50% with 2%; the 80% belongs to L5, the top of the taxonomy. The ~25% flagged-claim coverage figure is a Hesper internal benchmark.*

Two derived quantities follow directly from the two columns, and both are arithmetic on Shift's published figures rather than claims of Shift's. From L1 to L5 the efficiency column multiplies by eight, from 10% to 80%. Over the same five rungs the indemnity column moves three percentage points, from no stated figure to 3%. Autonomy scales what a claims operation can do to its own handling cost far faster than it scales what the operation pays out. That is not a flaw in the framework. It is a fair description of what climbing an autonomy ladder buys, published by the vendor that built the ladder.

### The fraud figure Shift does publish

ARISE connects to fraud once, and the connection deserves conceding in full: the report cites $2B in fraud uncovered in the US in 2025 as evidence of the financial materiality of L4 Solves-level fraud detection at scale. That is a detection figure and Shift labels it as one. It is also a large number and it belongs in any fair account of what Shift does. Detection is upstream; investigation is downstream. A carrier sitting behind a detection capability operating at that scale has more flagged claims to resolve, not fewer, which is the setup for everything that follows.

## An efficiency gain is a ratio over work already in scope

An efficiency gain is a ratio. It compares cost or time per unit of work handled against the same work handled before, which means its denominator is the work already in scope. Nothing in the metric can reach work that was never in scope. That is a definitional property of the measure, not a criticism of anyone who publishes one.

Spell the denominator out, because the denominator is where the whole argument lives. If a claims operation handles 100,000 claims a year and an agent removes half the handling cost from a subset of them, the efficiency gain is computed over that subset. Claims that were never opened, never worked past a disposition code, or never assigned to anyone are in neither the numerator nor the denominator. They sit outside the measurement entirely, and improving the ratio does not move them one inch. A percentage cannot describe something it does not range over.

Shift's own L4 definition tells you what class of work the ratio is computed over, and the wording is precise enough to be load-bearing.

> The agent acts end-to-end without human intervention, achieving 99%+ accuracy by applying contractual, regulatory, and insurer-specific logic consistently at scale.
>
> - Shift Technology, ARISE report, Level 4 "Solves" definition, June 3, 2026

### Deterministic work and contested work

Applying contractual, regulatory and insurer-specific logic describes a class of problem where the correct answer is knowable in advance. The policy says what it says, the jurisdiction says what it says, the schedule of benefits says what it says, and the agent's job is to apply that consistently at speed. Cost per claim handled is the right metric for that class of work, because throughput is what varies while correctness is largely determined by the inputs. A 99%+ accuracy threshold is only a coherent target when a right answer exists to measure against.

Fraud investigation is not that class of problem. The answer is not in the policy document. The counterparty is actively assembling a record designed to pass review. The evidence has to be built from documents, public records, prior claims, statements and timelines that do not agree with each other by default, and the finding then has to survive an examination under oath, a state DOI review, or litigation. There is no ground truth to hit 99% against at the moment the decision is made. There is only a record that either holds up later or does not.

This is where the layered model earns its keep as a measurement argument rather than a category slogan. Prevention blocks bad claims before they are filed. Detection flags suspicious claims after FNOL, and FRISS, Shift Technology and Verisk are the vendors that do it. Investigation takes a flagged claim and resolves it end to end with an audit-ready record, and until recently the only thing occupying that layer was a manual SIU team. Each layer has a native unit of measurement. Detection is measured in recall and false-positive rate. Handling is measured in cost and time per claim. Investigation is measured in claims resolved, dollars of loss avoided, and whether the finding is defensible. Importing one layer's unit into another layer's procurement is the error, and it is an easy error to make, because the handling metric is the one that arrives already formatted for a slide.

Notice what this does not say. It does not say a handling metric is a bad metric. For claims handling it is the correct metric, and Shift computing it over claims processing is Shift doing the right thing. It says that a metric carries its denominator with it, and that when the denominator changes, the number stops meaning what the buyer thinks it means. An efficiency gain measured over handled claims tells a fraud-investigation buyer almost nothing about the flagged claims that are never handled.

### A general principle about new-technology measurement

The gap between what a technology can do and the statistics used to describe it is a documented general problem. In [NBER Working Paper 24001](https://www.nber.org/papers/w24001), published in November 2017, Brynjolfsson, Rock and Syverson set out four explanations for the clash between AI capability and measured productivity: "false hopes, mismeasurement, redistribution, and implementation lags." They note that "going forward, national statistics could fail to measure the full benefits of the new technologies and some may even have the wrong sign." Their subject is national productivity statistics, not vendor efficiency metrics, and the paper is a working paper rather than a peer-reviewed publication. The transferable part is only the principle: when a technology changes what is possible rather than only how fast the existing work goes, a metric inherited from the prior regime can miss the change, or point the wrong way.

## The outcome column the industry already named

Impact Ratio is the US insurance industry's own outcome metric for special investigations: SIU referrals closed with impact, divided by all SIU referrals. The Coalition Against Insurance Fraud defines it, 57% of carriers capture it, and 43% do not. It is the closest thing this market has to a standard outcome number.

The Coalition and PwC surveyed 93 insurance companies in July and August 2023 for [The Keys to Unlocking SIUs Future Success](https://insurancefraud.org/wp-content/uploads/Coalition-Keys-Study.pdf), 45 of them Coalition members, 41 from APCIA and four from NAMIC. It was the largest insurer participation in any Coalition study to date. The study opens by asking the measurement question directly: "Every insurer strives for leading claims and investigation performance, but what does that really mean? Which standards are being used to even start evaluating key anti-fraud metrics such as overall detection, conversion, acceptance rates and false positive ratios?" That was published two years before an autonomy ladder existed to answer it in the wrong units.

> Among respondents, 57% capture Impact Ratio. Of survey participants, 3% report an Impact Ratio of 71-100% while 22% of companies report their Impact Ratio is between 41-70% and an additional 32% report a 40% impact or less. Whereas 43% of respondents do not capture Impact Ratio.
>
> - Coalition Against Insurance Fraud and PwC, "The Keys to Unlocking SIUs Future Success," 2023, p15

| Impact Ratio reported by US insurers (Coalition Against Insurance Fraud and PwC, 2023, 93 companies) | Value | Share |
| --- | --- | --- |
| Do not capture Impact Ratio at all | 43% | 100% |
| Report an Impact Ratio of 40% or less | 32% | 74% |
| Report an Impact Ratio of 41-70% | 22% | 51% |
| Report an Impact Ratio of 71-100% | 3% | 7% |

Sit with that distribution for a second. Fewer than one company in thirty reports an Impact Ratio above 70%. Nearly a third of the companies that do measure it report 40% or less, meaning most of their referrals close without an identifiable impact. And the largest single group, 43%, does not compute the ratio at all. An industry in that condition has no outcome number ready to hand when a vendor arrives with an efficiency number. The efficiency number wins the scorecard by walkover, because it is the only quantity on the table that both sides of the procurement can agree on.

It gets harder. Even among the carriers that do evaluate economic impact, there is no shared definition of the dollars. The same study found 82% of respondents evaluate economic impact, but the basis splits five ways: 38% use the mitigated amount, 34% use the estimated or actual claim value, 6% use the claim reserve at the time of referral, 2% use file counts and 20% use some other basis. Five denominators for one word. A vendor cannot be held to an outcome number that the buyer has not defined, and a market that has not converged on a definition defaults to the quantity that is already standardised, which is cost and time per claim.

There is a structural reason the cost language is native and the yield language is foreign. The same study found only 12% of carriers treat the SIU as a company profit center. 35% expense SIU operating costs by line of business, 33% leave them unallocated, and 17% allocate them to the sponsoring business unit. Seven carriers in eight book the SIU as a cost. When a function is booked as a cost, the metric that gets asked for is cost per unit, and the metric that would require a yield definition never gets built. The accounting treatment selects the measurement, and the measurement then selects the vendor.

The same bias shows up in what SIUs count day to day: 89% count referrals for suspected fraud and 70% count referrals accepted, and the outcome-side factors sit lower down the same list. The full evaluation-factor ranking and the post-acceptance cycle-time distribution are laid out in [Verisk Fraud Discovery unified five layers of the fraud stack](/blog/verisk-fraud-discovery-vs-autonomous-investigation/), which reads the same study from the case-management side rather than the autonomy side.

This is where the coverage number belongs. Hesper's internal benchmarks put manual flagged-claim coverage at roughly 25%: of the claims a detection layer flags, about a quarter receive a full investigation and the rest are paid, denied without full work, or queued until they age out. An investigation layer that runs 15+ investigation phases in parallel on every flagged claim takes that to 100%, in hours rather than weeks against a 14+ day manual baseline. All four of those figures are Hesper internal benchmarks and should be read as ours, not as industry statistics. The reason they belong in this section is that no efficiency percentage, however honestly measured, contains a coverage term. Coverage decides how much of a flagged book is ever resolved, and it lives in the outcome column.

## Where an efficiency number and a coverage number diverge

Coverage is the share of flagged claims that receives a full investigation rather than a disposition. It does not appear in any efficiency formula, because efficiency is computed over work already being done. Two carriers can post an identical efficiency gain and resolve very different shares of the same flagged book.

### The coverage-adjusted arithmetic

Run the calculation a finance reviewer would run, with both inputs named. Take Shift's published L4 figure, a 50% efficiency gain, and apply it to a carrier whose flagged-claim investigation coverage is roughly 25%, which is a Hesper internal benchmark rather than a published industry statistic. The gain lands on the quarter of the flagged book that was already being investigated. The other three quarters are unchanged, because they were never in the denominator. Coverage-adjusted, the efficiency gain touches one flagged claim in four. That is a derivation combining a Shift figure with a Hesper benchmark. It is not a Shift claim, not a measured result, and not a criticism of the 50% figure, which is doing exactly what an efficiency figure is supposed to do.

The finance question that follows is the one that decides the business case: what is the unit cost of the work that is not being done. Hesper's internal benchmarks put a manual investigation at roughly $2,500 per case against roughly $150 per case for an agent-run one, with an investigator carrying 200+ cases and completing roughly 10 investigations a month, against 800+ per investigator per month when the agent does the execution and the investigator does the judgement. Those are our numbers. The point for a model is not the ratio between them but what the first number explains: at $2,500 and ten a month, coverage is capacity-bound, and capacity is bought in salaries, one headcount approval at a time.

The argument that does not work here is headcount displacement, and it should not be reached for. An efficiency gain applied to an existing caseload produces an expense saving a CFO can model quickly. An SIU that resolves four times as many flagged claims produces a loss-cost effect that is larger and harder to estimate, and sits on a different line in a different place in the P&L. The harder one loses to the easier one in a spreadsheet unless somebody insists on putting it there. The internal version of that model, the scorecard a carrier builds for its own investigation AI rather than the number a vendor publishes, is laid out in [how carriers should measure AI ROI in fraud investigation](/blog/measuring-fraud-investigation-ai-roi/).

The Claims VP version of the same question is basis points. An efficiency gain on the handling of investigated claims lands in loss adjustment expense, which is real money and a comparatively small line. Coverage lands in indemnity, which is the large line. If the flagged book is around 10% of claims and three quarters of that book is never fully worked, the indemnity exposure sitting in the unworked portion is not a rounding error against an expense saving. The board-facing risk is asymmetric too. Nobody writes a board memo about an efficiency gain that came in at 40% instead of 50%. Board memos get written when a documented fraud case is paid after a vendor was bought to prevent exactly that.

| Dimension | Throughput / handling side | Outcome / investigation side |
| --- | --- | --- |
| What it measures | Cost and time per unit of work handled | Dollars of loss avoided and findings that survive review |
| Shift's ARISE columns | Efficiency gain: 10 / 20 / 30 / 50 / 80% | Indemnity impact: none stated / 1 / 1 / 2 / 3% |
| Shift Claims product reporting | 30% faster claims handling; 60% overall automation rate | 3% lower claims losses |
| Sedgwick figure via Insurance Journal | 80% faster processing times for some carriers on low-severity claims | No paired outcome figure published in the article |
| What US SIUs count most | Referrals sent and referrals accepted lead the evaluation-factor ranking (CAIF/PwC) | Impact Ratio: 57% capture it, 43% do not (CAIF/PwC) |
| Distribution of the outcome ratio | Not applicable | 3% report 71-100%; 22% report 41-70%; 32% report 40% or less |
| How the SIU is booked | 12% of carriers treat the SIU as a company profit center; the rest book it as a cost | A yield metric is only native where the unit is booked as a profit center |
| Coverage | Not contained in any efficiency metric; efficiency is a ratio over work already being done | ~25% of flagged claims investigated manually vs 100% with an investigation layer (Hesper internal benchmark) |

### Why the queue grows faster than the capacity

The coverage gap is not static, and the direction it moves in is against the carrier. The Coalition and PwC found 47% of insurers report referral rates of 3% or less, against an estimate of roughly 10% of claims being potentially fraudulent that NICB, the Insurance Information Institute, the Coalition, the NAIC and IASIU converge on. The Coalition separately puts US insurance fraud at [$308.6 billion a year](https://insurancefraud.org/fraud-stats/) and finds fraud in about 10% of property-casualty losses. Better detection closes the gap between 3% and 10% by producing more referrals, which is the correct thing for a detection layer to do, and which enlarges the queue the investigation layer has to clear.

Staffing does not respond on the same curve. 53% of carriers set SIU staffing by total referrals sent, and another 12% by referrals assigned per investigator, so two thirds size the unit off queue volume. That is why throughput framing feels native inside an SIU: the unit is sized by the queue, so the queue is what gets managed. It is also why the queue never clears. Sizing off volume means capacity chases the queue at whatever ratio last year's budget allowed, and detection improvements arrive faster than headcount approvals do.

The supply side has a hard ceiling underneath all of it. There are 39,500 private detectives and investigators in the United States, which is the number the outsourcing thesis runs into directly; the capacity argument and the arithmetic behind it are in [outsourced insurance investigation services vs autonomous AI](/blog/outsourced-investigation-services-vs-autonomous-ai/). Coverage cannot be closed by buying more hours, because the hours are priced by an occupation that is not growing at the rate the queue is. That is the structural reason coverage belongs in a procurement scorecard rather than in a staffing plan.

## The measurement gap is the industry's own stated problem

The AI ROI measurement gap in insurance is a problem carriers name themselves. The Coalition's 2024 technology study lists "Lack of cost/benefit analysis (ROI)" among the challenges carriers report with anti-fraud technology, and an AM Best survey reported in May 2026 found only 13% of carriers very confident in their ability to develop an AI ROI measurement.

Take the Coalition study first, with its provenance stated plainly. The [2024 State of Insurance Fraud Technology Study](https://insurancefraud.org/wp-content/uploads/SHIFT-State-of-Insurance-Fraud-Technology-Study.pdf) is the sixth edition, published December 10, 2024, fielded between October 3 and November 14 across 35 carriers in P&C and life and disability, with 69% of respondents from SIU. The deck is sponsored by Shift Technology and hosted by the Coalition. A Shift-sponsored study is being used here to support a point Shift has no interest in making, which is worth flagging rather than hiding.

The study's benefits and challenges artifacts are word clouds, where type size encodes relative prominence rather than a percentage, so they should be read as ranking with no numbers attached. On the benefits side the most prominent items are increased speed of detection, more referrals and higher quality referrals. "Increased mitigation of losses determined to be fraudulent after investigation" appears at lower prominence. That is the same asymmetry as the ARISE table, drawn from the buyers rather than from a vendor: the throughput benefits are front of mind, the outcome benefit is present but smaller.

The challenges cloud names the problem in the industry's own words. Two items sit in it verbatim: "Lack of cost/benefit analysis (ROI)" and "SIU cannot handle the volume of potentially fraudulent claims." Those are the two halves of this post, written by the buyers, in a deck published two years ago. The carriers know they cannot compute the return, and they know the unit cannot absorb the volume. An efficiency percentage answers neither one.

The AM Best data puts a number on the confidence gap. [Insurance Journal reported in May 2026](https://www.insurancejournal.com/news/international/2026/05/21/870833.htm) on an AM Best survey of approximately 150 rated carriers and MGAs fielded in November 2025, which found that "Only 13% of survey respondents felt very confident in their organization's ability to accurately develop an AI ROI measurement." The same survey found roughly one in five describing their AI implementation as already at an advanced stage, and 53% describing themselves as cautious pacesetters rather than first movers. A market where 87% are not very confident they can measure AI ROI is a market that will accept whatever measurement a vendor hands it, and will score the RFP on that.

Efficiency-only reporting is not a Shift habit. It is the category default, and Shift is one of the few that departs from it. [Insurance Journal reported in March 2026](https://www.insurancejournal.com/news/national/2026/03/13/861869.htm) that using AI to handle low-severity claims "has led to 80% faster processing times for some carriers," a figure attributed in the piece to a Sedgwick report. No paired outcome figure appears beside it in the article. That is the shape of most published claims-AI results: a throughput number, precisely stated, with no column next to it for what happened to the money.

Compliance has its own version of the same gap, and it is worth one line. In the Coalition and PwC study, reporting cases to the state department of insurance under mandatory reporting sits at the very bottom of the evaluation-factor list, below every volume metric on it. A measurement culture built almost entirely on volume produces a program that is hard to describe to a regulator in outcome terms, which is a problem that surfaces at antifraud-plan filing time rather than at procurement time.

## Three vendor postures on publishing an outcome number

Vendors take one of three postures on outcome reporting. Shift publishes two columns, an efficiency figure and an indemnity figure, in the same table. FRISS publishes throughput-shaped figures with no paired outcome column. Verisk's Fraud Discovery announcement published no accuracy, throughput, recovery or efficiency metric at all.

Ranked on transparency rather than on product, Shift comes first of the three and by a clear margin. ARISE separates efficiency from indemnity and puts both on the page. Shift's product reporting does the same thing: the Shift Claims launch paired 30% faster claims handling and a 60% overall automation rate with 3% lower claims losses and greater than 99% accuracy in claims assessment. Whatever else is true, a buyer reading Shift's material can find an outcome figure without asking. Those four numbers are unpacked in [Shift Claims agentic AI vs Hesper](/blog/shift-claims-agentic-ai-vs-hesper/) and are not re-argued here. The only point that matters for this post is that the outcome column exists.

[FRISS](https://www.friss.com/) publishes 75% fewer false positives and 90% of honest claims fast-tracked. Both are throughput-shaped. A false-positive reduction is a statement about how much work does not have to be done, and a fast-track rate is a statement about how quickly clean claims clear. Neither is dishonest, both are operationally meaningful, and neither is an outcome column. Verisk's Fraud Discovery release, examined in the [Fraud Discovery post](/blog/verisk-fraud-discovery-vs-autonomous-investigation/), publishes no accuracy, throughput, recovery or efficiency metric of any kind, which is the third and weakest posture for a buyer: nothing to score at all.

| Vendor | Published throughput figure | Published outcome figure | Posture |
| --- | --- | --- | --- |
| Shift Technology | ARISE efficiency gain 10 / 20 / 30 / 50 / 80%; Shift Claims 30% faster handling, 60% automation rate | ARISE indemnity impact none stated / 1 / 1 / 2 / 3%; Shift Claims 3% lower claims losses | Two columns. The most complete of the three. |
| FRISS | 75% fewer false positives; 90% of honest claims fast-tracked | None published | One column, throughput-shaped |
| Verisk (Fraud Discovery) | None published | None published | No metric to score |
| Hesper AI | No efficiency-gain percentage published, by choice | ~25% to 100% flagged-claim coverage (Hesper internal benchmark) | Coverage column, no efficiency column |

Hesper's own posture belongs in that table, including its gap. Hesper publishes a coverage number, roughly 25% to 100% of flagged claims, labelled as a Hesper internal benchmark, and does not publish an efficiency-gain percentage. That is deliberate, and it is the same argument running in the other direction. An efficiency percentage computed over a 25% denominator is precisely the artifact this post argues a buyer should not price on, so publishing one would contradict the post. What a buyer should ask us for is exactly what they should ask everyone for: what happens to the three quarters of the flagged book that currently sits outside the denominator.

None of this makes the vendors interchangeable or adversarial. Hesper AI sits downstream of detection. Complementary to FRISS, Shift Technology, and Verisk - not a replacement. A carrier running a detection layer that uncovers fraud at scale has more flagged claims arriving at an SIU that can work a quarter of them, which is a better problem than the alternative and is still a problem. The buyer's-guide view of how the layers fit together, and which vendor occupies which one, is in the [AI fraud platforms compared buyer's guide](/blog/ai-fraud-platforms-compared-2026-pillar/).

## The metric set to put in the RFP

Eight metrics convert an autonomy claim into something a carrier can price: investigated coverage rate, Impact Ratio, economic impact per closed referral, the mitigation and denial split, post-acceptance cycle time, referral return rate with reason codes, defensibility rate, and coverage-adjusted efficiency. Five carry published industry benchmarks, one carries a regulatory anchor, two are ours.

| Metric | What it measures | Benchmark to hold the vendor against | Source of the benchmark |
| --- | --- | --- | --- |
| Investigated coverage rate | Share of accepted referrals that receive a full investigation rather than a disposition | ~25% manual baseline; 100% with an investigation layer | Hesper internal benchmark |
| Impact Ratio | SIU referrals closed with impact, divided by all SIU referrals | 43% of carriers do not capture it; only 3% report 71-100% | Coalition Against Insurance Fraud and PwC, 2023, p15 |
| Economic impact per closed referral | Dollars of loss avoided per referral closed | 48% of SIUs use Total Economic Impact per closed referral, and the market has no common definition of the dollars | Coalition Against Insurance Fraud and PwC, 2023, p11 and p15 |
| Mitigation and denial split | Referred claims mitigated, referred claims denied, claims closed without payment | Tracked by a minority of SIUs and ranked below referral counts in the industry's own evaluation-factor list | Coalition Against Insurance Fraud and PwC, 2023, p11 |
| Post-acceptance cycle time | Elapsed time from referral acceptance to case closure | 64% of SIUs capture it; 36% do not capture it at all, and most of those who do report a band measured in weeks | Coalition Against Insurance Fraud and PwC, 2023, p17 |
| Referral return rate and reason codes | Share of referrals returned or cancelled before investigation, split by cause | 75% of carriers return 20% or less; of returns, 76% did not contain sufficient elements of suspected fraud and 5% were analytics false positives | Coalition Against Insurance Fraud and PwC, 2023, p16 |
| Defensibility rate | Share of findings that survive DOI review, an examination under oath, or litigation with the decision chain intact | No industry benchmark is published. Propose the metric and require a sample file rather than a percentage | NAIC Model Act 680; California 10 CCR 2698.36 |
| Coverage-adjusted efficiency | The vendor's stated efficiency gain multiplied by the share of the flagged book actually investigated | At a ~25% coverage baseline, a 50% efficiency gain touches one flagged claim in four | Derivation. Inputs: Shift's published L4 figure and a Hesper internal benchmark |

> **The one question to put at the top of the RFP**
>
> Ask every vendor the same thing: what is the denominator of your efficiency percentage, and what happens to the claims outside it? An honest answer names a scope, a share of the flagged book, and a disposition for the remainder. A vendor that answers with a larger percentage has not answered the question.

Then ask the companion question, and ask it of everyone including us: show me your outcome column. Shift can answer it, because ARISE carries an indemnity-impact column running 1 to 3% and Shift Claims reports 3% lower claims losses. A vendor that cannot produce an outcome column at all is in a weaker position than Shift, not a stronger one. Absence of a number is not modesty and it is not caution. It is an unscoreable line in an evaluation, and it should be treated that way.

Metric seven, defensibility rate, is the one with no industry benchmark, and it should go in the RFP anyway because the regulatory requirement exists whether or not a benchmark does. NAIC Model Act 680, adopted in 48 states, requires insurers to file antifraud plans describing how suspected fraud is detected and investigated. California's 10 CCR 2698.36 requires documented decisions on the claims an SIU handles. The question for a vendor is narrow: when a state DOI examiner pulls one of these files, can the decision chain be reconstructed from sources, reasoning and timestamps, or does it terminate in a score with nothing behind it. Do not accept a percentage in answer to that question. Ask for a file.

The reason to run the whole set is not to catch a vendor out. It is that the thing a carrier buys, when it buys a fraud-investigation capability, is the second column: claims resolved, dollars of loss avoided, findings that hold up. Efficiency is how you get there and it is not the thing itself. Make every flagged claim investigable, then measure what came of it.

## Key takeaways

- Shift's ARISE ladder, as published, attaches efficiency gains of 10, 20, 30, 50 and 80% to its five autonomy levels, and the June 3, 2026 report states the figures are drawn from production deployments rather than aspiration.
- The same table carries a second column: indemnity impact runs from no stated figure at L1 to 3% at L5, so across the ladder the efficiency figure multiplies eightfold while the outcome figure moves three percentage points.
- Shift places its own production capability at L4, whose paired figures are 50% efficiency and 2% indemnity impact; the 80% figure belongs to L5, the top of the taxonomy rather than a product claim.
- An efficiency gain is a ratio over work already in scope, so it cannot describe investigated-claim coverage, which is the variable that decides how much of a flagged book is ever resolved; Hesper's internal benchmark puts manual coverage at roughly 25%.
- The industry already named the outcome metric, the Coalition calls it Impact Ratio, and 43% of carriers do not capture it, which is why an efficiency percentage wins an RFP scorecard by default.

### What is the ARISE framework's efficiency gain at each level?

Shift Technology's ARISE framework attaches an efficiency gain to each of its five autonomy levels: 10% at L1 Answers, 20% at L2 Recommends, 30% at L3 Initiates, 50% at L4 Solves and 80% at L5 Exceeds. The report published on June 3, 2026 sets those figures out in a table alongside a second column, indemnity impact, which runs from no stated figure at L1 to 1% at L2 and L3, 2% at L4 and 3% at L5. The report states the metrics are drawn from Shift's current production deployments rather than being aspirational targets. Shift places its own production capability at L4, so the 80% figure describes the top of the taxonomy rather than a shipped product.

### Does an AI efficiency gain percentage tell you how much fraud you will recover?

No. An efficiency gain is a ratio computed over the work already in scope, which is cost or time per claim handled. It says nothing about how many suspicious claims get resolved, how many findings survive review, or how many dollars of loss are avoided. Shift's own ARISE table keeps the two apart: at Level 4 the efficiency figure is 50% and the paired indemnity-impact figure is 2%. Across the whole ladder, efficiency multiplies eightfold while the indemnity figure moves three percentage points. If you are buying a fraud-investigation capability, the second column is the one you are purchasing, and an efficiency percentage on its own will not tell you what it will be.

### What is Impact Ratio in SIU measurement?

Impact Ratio is the US insurance industry's own outcome metric for special investigations. The Coalition Against Insurance Fraud defines it as SIU referrals closed with impact divided by all SIU referrals, and describes it as giving insight into the proportion of claims that are settled, denied, compromised or withdrawn. In the Coalition's 2023 study with PwC, which surveyed 93 insurance companies, 57% of respondents captured Impact Ratio and 43% did not. Of survey participants, 3% reported an Impact Ratio of 71-100%, 22% reported 41-70% and 32% reported 40% or less. It is the closest thing the market has to a standard outcome number for investigation work, and most carriers are not computing it.

### How should a carrier measure ROI on an AI claims autonomy investment?

Score the outcome column, not only the throughput column. Ask the vendor for the denominator of its efficiency percentage and what happens to the claims outside it, then hold the answer against eight metrics: investigated coverage rate, Impact Ratio, economic impact per closed referral, the mitigation and denial split, post-acceptance cycle time, referral return rate with reason codes, defensibility rate, and coverage-adjusted efficiency. Five of those have published industry benchmarks in the Coalition and PwC 2023 study. The measurement problem is real and carriers say so: an AM Best survey of around 150 rated carriers and MGAs, reported by Insurance Journal in May 2026, found only 13% felt very confident in their ability to develop an AI ROI measurement.

### What does a 50% efficiency gain actually measure in claims processing?

It measures the reduction in cost or time per claim handled, computed over the claims already being handled. Shift's Level 4 definition describes an agent that acts end to end at 99%+ accuracy by applying contractual, regulatory and insurer-specific logic consistently at scale, which is deterministic work where the correct answer is knowable in advance. Cost per claim handled is the right metric for that class of work. What the figure does not measure is how many claims enter the process, how many suspicious claims get resolved, or what the carrier ends up paying out. In Shift's own table the 50% efficiency gain at Level 4 is paired with a 2% indemnity impact, and those two numbers answer different questions.

### Can a carrier hit an efficiency target while investigating the same number of claims?

Yes, and that is the structural risk. Efficiency is a ratio over work already being done, so improving it compresses the cost of the existing caseload without enlarging it. Hesper's internal benchmarks put manual flagged-claim coverage at roughly 25%, meaning three quarters of the flagged book is paid, denied without full work, or queued indefinitely. A 50% efficiency gain applied to that quarter leaves the other three quarters exactly where they were. The Coalition and PwC found that 53% of carriers set SIU staffing by total referrals sent and another 12% by referrals assigned per investigator, so the queue grows as detection improves while the labour per case is what actually caps coverage.

### Do insurance AI vendors publish outcome metrics or only efficiency metrics?

It varies, and there is no market convention. Shift Technology publishes both: the ARISE table carries an efficiency column and an indemnity-impact column, and its Shift Claims launch reported 3% lower claims losses alongside 30% faster handling and a 60% automation rate. FRISS publishes throughput-shaped figures on its site, 75% fewer false positives and 90% of honest claims fast-tracked, without a paired outcome column. Verisk's Fraud Discovery announcement in September 2026 published no accuracy, throughput, recovery or efficiency metric at all. Of the three postures, publishing both columns is the most useful to a buyer, so ask every vendor for its outcome column and treat a missing one as the weaker position rather than the modest one.

### What questions should be in an RFP for an AI claims investigation vendor?

Lead with the denominator question: what scope of work is your efficiency percentage computed over, and what happens to the claims outside it. Then ask for the outcome column, meaning Impact Ratio or its equivalent, economic impact per closed referral, and the mitigation and denial split. Ask for post-acceptance cycle time, which the Coalition and PwC found 36% of SIUs do not capture at all. Ask what share of accepted referrals receives a full investigation rather than a disposition. Finally ask for the defensibility posture: whether the output reconstructs the reasoning well enough for a state DOI review under NAIC Model Act 680 and California 10 CCR 2698.36, and ask to see a sample file rather than a percentage.
