A decision log tells an examiner what your system decided. It does not tell them what the claim had been checked for before the decision was taken, and the second record is the one that carries weight two years later. Drew Young, founder of Veriscopic, named the gap in September 2026 by splitting AI accountability into three layers: system evidence, governance evidence and decision evidence. No US property and casualty carrier had to wait for AI to be asked for the third layer.
Young's framing question is the useful artefact: "Why was the insurer entitled to rely on that decision when it was made?" Set it against NAIC Model Regulation 902, the model unfair property and casualty claims settlement practices rule. Section 4.B requires "detailed documentation" in each claim file "in order to permit reconstruction of the insurer's activities relative to each claim." Activities, not conclusions. Section 3.G then defines "Investigation" as "all activities of an insurer directly or indirectly related to the determination of liabilities", which includes the ones that found nothing. Pennsylvania wrote its own version of the test in December 1978.
The operational distinction is between logging a decision and recording an investigation. Five Sigma published the clearest public specification of the first on 17 September 2026: six fields that reconstruct one decision. They are the right six, and a carrier that cannot produce them has a problem. What none of them carries is the set of checks that ran on the claim and came back clean, which is the half of the file an examiner uses to decide whether the investigation happened.
That half has a cost attached, which is why it goes missing. A cleared check produces nothing the adjuster needs and one more line to type, so under caseload pressure it is the first thing to stop being written down. The decision survives in the file and the search that justified it does not. The adjacent question, which claims decisions stay with a person at all, is mapped in which claims decisions stay human.
Insurance decisions have long afterlives
Answer
What is decision evidence, and how is it different from an audit trail?
Decision evidence is the record that shows why an organisation was entitled to rely on a decision at the moment it was made. Drew Young separates it from system evidence, meaning what the model or workflow did, and governance evidence, meaning what controls and approvals were in place. An audit trail is system evidence.
The framework comes from a piece Young published in Insurance Edge on 22 September 2026. He stacks three questions. System evidence: "what did the model, workflow or agent do?" Governance evidence: "what controls, approvals, policies or oversight mechanisms were in place?" Decision evidence: "why was the organisation entitled to rely on this decision in this specific context?" The first two are what most AI assurance work produces. The third is what gets tested when somebody comes back.
Young is also proposing a standard, which is worth naming plainly. Veriscopic promotes VES, described in its own materials as "an interoperable decision-evidence layer for consequential insurance decisions" that does not "replace runtime records, audit logs or governance evidence" but connects them to the business decision they support. Version 1.2 went final on 21 September 2026. Nothing here endorses it and Hesper does not implement it. The three-layer framing stands on its own, which is why it is worth crediting and then arguing with.
The companion piece he published on 1 September 2026 is the stronger one for claims. Its point is that the evidence behind a decision may no longer exist together, "in the state in which the decision occurred", because it ends up "distributed across claims systems, underwriting platforms, emails, documents, model logs, data stores and people's memories." That is a contested claim file three years on: the adjuster has left, the vendor report is in an inbox, the index search is a screenshot, and the reasoning is nowhere.
He also names the chain that makes this an MGA problem and not only a carrier problem: "A delegated underwriter makes a judgement. A carrier assumes the risk. A reinsurer assumes part of it. Years later, another organisation may need to understand the basis on which the original decision occurred." For a head of claims operating under delegated authority, that chain is an annual event, not a thought experiment. It is the claims audit a capacity partner runs on a sample of files, and the question in that audit is not what the decision was. It is what the file shows was looked at.
Joe Kim, chief executive of the conversational-AI vendor Druid AI, told Insurance Edge in a separate opinion piece on 22 September 2026 that "an agent's actions should be observable and auditable, with organisations able to measure whether it continues to operate within its intended mandate" - a general software-governance point from outside insurance rather than a claims one. Observability of an action is still system evidence, and an input to decision evidence rather than a substitute for it.
Six fields describe a decision. An investigation has more.
Answer
What should a claim decision record contain?
Five Sigma's published record carries six fields: the decision and timestamp, the inputs available, the actor with version, the delegated authority, the override path and the override history. Those are the right six for logging a decision. None of them records which checks ran on the claim and came back clean.
Five Sigma's Tirtza Bensoussan published the specification on 17 September 2026 and defines the artefact precisely: "The claim-file artifact that reconstructs one decision: inputs, actor, model version, authority, and override history." The examiner framing in the same post is correct, and it is the sentence most AI governance documents never reach: "An examiner will ask you to reconstruct a specific decision on a specific claim on a specific date, and show who or what made it, on what basis, under whose authority, and who could have overridden it." That is an accurate description of a market conduct interrogatory.
Read that definition again and notice its scope. One decision. The six fields are complete with respect to that object, and a carrier that can produce them on demand is ahead of most of the market. The argument is not that the list is wrong. It is that the object is smaller than the thing being examined, because what an examiner tests is the investigation that preceded the decision.
Make it concrete. A property claim is denied on a late-reporting ground. The decision log records the denial, the timestamp, the adjuster, the authority level, the documents in the file and the fact that nobody overrode it. Three years later the policyholder's counsel asks what else the carrier considered. Was the prior-loss history pulled. Were the other addresses associated with the insured run. Was the weather data for the loss date checked against the reported cause. Was the vacancy exclusion tested and found not to apply, or was it never read. Every one of those is an activity inside the Model 902 definition, and not one of them is a field in the log.
That gap closes only when the artefact is produced by the work rather than written up after it. Hesper runs 15+ investigation phases in parallel on a claim instead of in a queue - a Hesper internal benchmark, not an industry figure - and what it leaves behind is the output of every phase, not a summary of the ones that mattered. On a file where three phases moved the outcome and twelve cleared, the record shows the twelve that came back clean alongside the three that did not, with the same timestamp discipline on both. Running the phases in parallel is what makes that affordable. At manual throughput, the clean phases are the first thing that stops being written down.
The claim-file rules already ask for the activities
Answer
Do US claim-file rules require documenting the investigation or just the decision?
The investigation. NAIC Model Regulation 902 requires documentation sufficient to permit reconstruction of the insurer's activities relative to each claim. California adds a second test on top of event reconstruction: the licensee's actions must be determinable from the file. New York and Pennsylvania use the same reconstruction construction.
Four texts run a reconstruction test on the claim file, and three of them are older than any claims AI on the market. None of them uses the phrase decision evidence. All of them describe it.
Model 902 repays reading past Section 4.B. Section 3.E defines "Documentation" to include "all pertinent communications, transactions, notes, work papers, claim forms, bills and explanation of benefits forms relative to the claim" - artefacts, not conclusions. Section 4.C requires each relevant document to be "noted as to date received, date processed or date mailed." A rule that cares about processing dates is a rule about sequence, and sequence matters only if somebody intends to check what was known when.
The definitional move in Section 3.G does the heavy lifting. "Investigation" means "all activities of an insurer directly or indirectly related to the determination of liabilities under coverages afforded by an insurance policy or insurance contract." Three features of that sentence matter. It is plural. It is unordered. And it draws no distinction between activities that changed the outcome and activities that did not. A record organised around the outcome keeps a subset of what the definition names, and the subset it drops is the larger one.
Fifteen illustrative activities behind one liability determination. The three that moved the outcome are what a decision record keeps. The twelve that cleared are the negative space a reconstruction test asks about.
California wrote the second test explicitly. 10 CCR 2695.3(a) requires files to contain documents, notes and work papers "in such detail that pertinent events and the dates of the events can be reconstructed and the licensee's actions pertaining to the claim can be determined." Two tests joined by an and: the events, and what the insurer did about them. A file can pass the first and fail the second, which is the exact shape of the gap. The SIU-side documentation layer California stacks on top of that is walked through in the 10 CCR 2698 compliance guide.
What a market conduct examiner actually opens
Answer
What does a market conduct examiner test on claim files?
A random sample of files, measured against standards one at a time, against a tolerance. The NAIC examination standards summary lists eleven claims standards. Two land on this subject: Standard 2, timely investigations are conducted, and Standard 5, claim files are adequately documented. They carry separate samples.
The NAIC Market Regulation Handbook examination standards summary sets out eleven claims standards. Standard 2: "Timely investigations are conducted." Standard 5: "Claim files are adequately documented." Standard 9: "Denied and closed without payment claims are handled in accordance with policy provisions and state law." Three separate tests on the same file, and the first two decide whether an AI-era claims operation is defensible.
That separation is what a decision log cannot survive. It answers Standard 5 for one line item: a record exists, it is dated, it names the actor. It does not speak to Standard 2 at all, because nothing in a decision log describes an investigation. An investigation record answers both from one artefact. That is a different architecture, not a bigger log.
Examinations run on a tolerance rather than on perfection. A Kansas Insurance Department report of examination of the Liberty Mutual Group, covering the period to 31 December 2004, puts it plainly: "A tolerance standard of 7% is used for claim procedures and 10% is used for all other procedures." That report is twenty-two years old and tolerances vary by state, so read the figure as illustrative of the method. The structural detail that has not changed sits in the same report: Standard 2 and Standard 5 are tested separately, each with its own sample.
The clearest itemised public example of what a sample returns is the California Department of Insurance targeted examination of the California FAIR Plan Association, covering claims closed between 1 January 2017 and 18 March 2021 and adopted on 25 May 2022. Examiners randomly selected 259 claim files and cited 418 violations. The distribution is the interesting part.
Thirty-eight citations for failing to "conduct and diligently pursue a thorough, fair and objective investigation". Four for failing to maintain documents in enough detail that events and dates could be reconstructed. Thirteen more for failing to document the justification for an adjustment on account of betterment, depreciation or salvage. The file did not fail on what was not written down. It failed on what was not checked. The review period ends in March 2021, so this is pre-AI claims handling, which is the point: the failure mode predates the technology and the technology inherits it.
Penalties attach. The Florida Office of Insurance Regulation's insurer compliance report for October 2025 lists seven property and casualty consent orders in the third quarter of that year, each described as violations resulting from a market conduct examination related to Hurricane Ian claims-handling operations, or Ian and Idalia. Fines run from $50,000 to $400,000 and add up to $1,975,000 across the seven; the report prints no total, so that sum is arithmetic. The orders are not described as documentation findings and should not be read as such. What they establish is that claims-handling examinations carry money.
The reason carriers struggle to produce an investigation record is throughput, not ignorance. A manual SIU investigation runs 14+ days per case against caseloads of 200+ cases per investigator, both Hesper internal benchmarks, so roughly 25% of flagged claims get a full workup and the rest get a disposition. Under that constraint the negative findings go unwritten first. An investigator who clears a prior-loss search in seconds does not stop to record that it returned nothing. Why most flagged claims never get investigated is the same arithmetic from the capacity side.
The AI rulebook covers the model. The claim file covers the claim.
Answer
Does the NAIC AI model bulletin require insurers to document individual claim decisions?
Not directly. Section 4 of the bulletin lists what a department may request: the written AI Systems Program, model inventories, validation and drift documentation, and third-party contracts. Those are program-level and model-level records. The claim-level obligation comes from the unfair claims settlement practices rules in each state.
The NAIC adoption map, status as of 31 August 2026, records 26 jurisdictions that have adopted the Model Bulletin on the Use of Artificial Intelligence Systems by Insurers, adopted by the NAIC on 4 December 2023, plus four more - California, Colorado, New York and Texas - carrying insurance-specific AI regulation or guidance of their own. Thirty jurisdictions is a real compliance surface, and a programme built against it is worth having.
The bulletin says twice that it is not the source of the claim-level obligation. Section 3 states that "compliance with these standards is required regardless of the tools and methods Insurers use to make such decisions", and that decisions resulting from AI use are "subject to the Department's examination to determine that the reliance on AI Systems are compliant with all applicable existing legal standards." The NAIC's own issue brief of March 2026 is blunter: AI "does not alter insurers' legal obligations. Existing state insurance laws apply regardless of whether decisions are made by humans, algorithms, or third-party vendors."
What Section 4 of the AI bulletin actually asks for
The bulletin's Section 4 list covers the written AI Systems Program, model inventories and descriptions, data provenance and lineage, documentation of validation, testing and auditing including model drift, and third-party due diligence and contracts. Every item is program-level or model-level. Nothing in Section 4 asks for the file on one claim. A carrier that satisfies Section 4 has built system evidence and governance evidence and can still hold no decision evidence at claim level.
Two of the four jurisdictions with their own rules show how uneven the reach is. Colorado's amended Regulation 10-1-1, 3 CCR 702-10, effective 13 November 2023 with amendments effective 15 October 2025, requires a documented governance framework with an up-to-date inventory under version control plus documented testing, monitoring and annual review. Its applicability runs to individually issued life, private passenger automobile and health benefit plans, and the statutory definition of insurance practice it incorporates reaches "claims management in the transaction of insurance". New York runs the other way: DFS Circular Letter No. 7 of 11 July 2024 says at paragraph 8 that it "is not intended to address phases of the insurance product lifecycle other than underwriting and pricing." A programme built to CL7 was built for the front of the policy.
The EU AI Act gets read into this gap and mostly does not fill it. Article 12(1) requires high-risk systems to "technically allow for the automatic recording of events (logs) over the lifetime of the system", and Article 26(6) makes deployers keep those logs for at least six months. The obligation is on logs, which is system evidence. And Annex III point 5(c) scopes insurance high-risk use to "risk assessment and pricing in relation to natural persons in the case of life and health insurance". Claims handling is not on that list.
The one instrument drafted at claim level did not pass. An NCOIL model act, summarised in the minutes of the NAIC Big Data and AI Working Group meeting of 29 September 2025, would have barred AI as the sole basis for claim denials or adjustments and required insurers to "maintain documentation on the basis for the denial, including information provided by the algorithm or AI system". NCOIL paused it in March 2026, with committee chair Erik Dilan telling Repairer Driven News on 18 March 2026 that it would be "very difficult to get the consensus necessary to successfully adopt that model". The same minutes describe an AI Systems Evaluation Tool of "five optional exhibits that can be incorporated into market conduct or financial examinations", with Pennsylvania "already using versions of these tools in financial exams and analyses" - examination tooling moving faster than legislation, and taken apart for claims models in the NAIC AI risk evaluation piece. New Jersey's A5494 is the live state version of the narrower question, and it regulates who signs rather than what is in the file.
Negative findings are evidence, and audit practice says so
Answer
Does a claim file have to document what was ruled out?
Yes, if the file is meant to show the investigation happened. Audit practice is explicit about the equivalent duty: PCAOB Auditing Standard No. 3 requires documenting the procedures performed, identifying the items inspected, and retaining information that contradicts the final conclusion. Missing documentation casts doubt that the work was done.
Claims is not the only discipline that had to decide what a defensible file looks like. Auditing wrote it down. PCAOB Auditing Standard No. 3, now AS 1215, requires at paragraph 6 that the auditor "document the procedures performed, evidence obtained, and conclusions reached" and that the documentation "clearly demonstrate that the work was in fact performed", to a standard where an experienced auditor with no previous connection can understand the nature, timing, extent and results of the procedures. Procedures, evidence, conclusions, in that order, with the conclusion last.
Two paragraphs go further than anything in the claims rulebook. Paragraph 10 requires documentation of procedures involving inspection of documents to "include identification of the items inspected". Paragraph 8 requires documentation to include, beyond what supports the final conclusion, "information the auditor has identified relating to significant findings or issues that is inconsistent with or contradicts the auditor's final conclusions." Contradictory evidence stays in the file by rule. Appendix A paragraph A10 closes the loop: where documentation does not exist for a procedure on a significant matter, it "casts doubt as to whether the necessary work was done."
The California claims rules say the same thing in a different register. 10 CCR 2695.7(d) requires every insurer to "conduct and diligently pursue a thorough, fair and objective investigation" - a statement about the search, not the answer, and the only way to evidence a search is to record what it covered. 2695.7(b)(1) then requires a denial to state "all bases for such rejection or denial and the factual and legal bases for each reason given". All bases, each with its factual support, which presupposes a file that distinguishes grounds tested and held from grounds tested and discarded.
The practical difference is production cost. A decision log is cheap because it writes down an event that already happened. An investigation record is expensive at human throughput, because every cleared check is another minute of typing that produces no benefit to the adjuster doing it. That is the honest reason the industry ended up with logs. It is a capacity problem wearing a documentation costume.
Running the investigation as software changes that arithmetic. When the checks execute as agent actions, the record of each is a by-product rather than an extra task, and a cleared check costs the same to preserve as one that hits. Coverage of flagged claims moves from 25% to 100%, both Hesper internal benchmarks, because the binding constraint was never the method, it was the hours, and the file arrives in hours rather than weeks. What that file has to contain is specified in the defensibility standard for investigation AI, with the architecture underneath it in the guide to autonomous AI claims investigation.
The file you hand a regulator, a reinsurer, or a jury
Answer
Who reads a claim file after the claim closes?
Three audiences with one test between them. A market conduct examiner asks whether the investigation was conducted and documented. A capacity partner's claims audit asks what the delegated decision rested on. A court asks whether the record is trustworthy enough to admit. All three are reconstruction questions.
Admissibility is the strictest version of Young's question, and the federal rules state it as a property of the record rather than of the conclusion. Federal Rule of Evidence 803(6) admits a business record made at or near the time by someone with knowledge, kept in the course of a regularly conducted activity, where making the record "was a regular practice of that activity", unless "the method or circumstances of preparation indicate a lack of trustworthiness". Every clause is about process discipline, not about being right.
Rule 901(b)(9) allows authentication by "evidence describing a process or system and showing that it produces an accurate result". Rules 902(13) and 902(14), both added on 1 December 2017, make a record generated by an electronic process self-authenticating on a qualified person's certification. A file assembled by hand under deadline pressure and a file assembled by an agent that logs every tool call face the same trustworthiness test, and only one carries the timestamps by construction. The courtroom half of this argument is worked through in what survives a courtroom.
On the bad-faith side, the California line of cases holds that an insurer cannot reasonably and in good faith deny payment without thoroughly investigating the foundation for the denial; Jordan v. Allstate Insurance Co., 148 Cal.App.4th 1062 (2007), is one published application. Read that next to Standard 2 and the overlap is close to total.
The capacity-partner version is the one that catches MGAs. A delegated-authority audit runs on a sample of files a year or more after the decisions in them, conducted by someone with no access to the people who made them. What the auditor sees is what the file holds, and a file of conclusions reads as a file of assertions.
Hesper AI is an AI claims resolution platform: agents take a claim from first notice to final recovery, and the evidence is a by-product of doing the work rather than a layer that reads the file afterwards. Clean claims resolve straight through; suspicious claims get an investigation-grade workup. The part that is hard to copy is that the same discipline runs the whole path. Coverage analysis produces a cited position naming the endorsements it tested and rejected. The investigation writes down every phase, including the ones that cleared. Settlement valuation carries the comparables it used and the ones it discarded. The recovery screen records the files where no recoverable third party was found. Evidence behind every decision, not only behind the suspicious ones. Where each sits on the claim path is mapped in the guide to automating the claims lifecycle, and the recovery end in AI-driven subrogation.
This describes obligations, not compliance
Nothing here asserts that Hesper or any other system satisfies NAIC Model Regulation 902, 10 CCR 2695.3, 11 NYCRR 216.11, 31 Pa. Code 146.3 or any state's unfair claims settlement practices act. Those obligations sit on the insurer, and whether a given file meets them is a judgement for the carrier's own counsel on the facts of that file. The assertable part is narrower: these are the tests the texts describe, and this is what an investigation record holds that a decision log does not.
The reason this is tractable now and was not in 1978 is unglamorous. Sampling-era investigation depth was a capacity constraint, not a policy choice. Nobody decided that cleared checks should go unrecorded; there were 200+ cases per investigator and 14+ days of work in each real one, so the record compressed to the findings that moved the file. When the investigation runs on every file instead of a sample, the record of what was ruled out stops being an extra artefact a carrier has to fund and becomes a by-product of the work. That is what makes Young's third layer buildable at claim level, and why the shortest route to decision evidence is producing the investigation that generates it.
Key takeaways
- Decision evidence, in Drew Young's three-layer framing, is the record of why an insurer was entitled to rely on a decision when it was made, distinct from the system evidence and governance evidence most AI assurance work produces.
- The obligation is not new: NAIC Model Regulation 902 Section 4.B requires documentation sufficient to permit reconstruction of the insurer's activities relative to each claim, and Pennsylvania has run a version of the same test since December 1978.
- Five Sigma's six decision-record fields are the right six for logging a decision, and none of them carries the checks that ran on the claim and came back clean.
- NAIC market conduct examinations test timely investigations and adequate file documentation as two separate standards with separate samples, so a decision log answers one while an investigation record answers both.
- In the California Department of Insurance examination of the California FAIR Plan Association, 259 sampled files produced 38 citations for a deficient investigation against four for a deficient file, which puts the exposure in what was not checked.