Design Partner ProgramWe're accepting applications for the next cohort of design partners in finance, insurance, healthcare, and HR. Apply now →

meilynx

Reference · Insurance

NAIC AI Systems Evaluation Tool: The AI Risk Evaluation Supplement Explained

The questionnaire state regulators are building to examine an insurer's AI, piloted in 12 states in 2026

Last reviewed October 11, 2026

At a glance

Current name
AI Risk Evaluation Supplement, renamed from AI Systems Evaluation Tool by July 2026
Status
In draft. Version 7.0 is planned for adoption at the NAIC Fall National Meeting, 14 to 17 November 2026

The AI Systems Evaluation Tool is a set of questionnaires that state insurance regulators can use to find out how an insurer uses AI and whether its governance of that use is working. The NAIC's Big Data and Artificial Intelligence (H) Working Group drafts it, under a charge to build tools that help regulators identify and assess the financial and consumer risks of AI Systems on an ongoing basis.Supplement v5.0 · Intent By its 22 July 2026 meeting the Working Group had renamed it the AI Risk Evaluation Supplement, in response to stakeholder feedback on the name and to reduce confusion about what the document is and is not.Working Group, Summer 2026 · 22 July minutes

Twelve states piloted it from March to September 2026, in market conduct exams, financial exams and financial analysis.Pilot summary The current public draft is version 5.0. The Working Group plans one more draft for comment and then a version 7.0 for adoption at the NAIC Fall National Meeting, held 14 to 17 November 2026 in Dallas.Working Group, 8 Oct 2026 · 31 August minutesNAIC Fall National Meeting As of 11 October 2026 it remains a draft.

A supplement to the examination handbooks

The exhibits add to the review procedures examiners already follow in market conduct, financial analysis and financial examination work, and the NAIC's existing resources stay authoritative. Requests made under the supplement are coordinated through the Market Regulation Handbook, the Financial Condition Examiners Handbook and the Financial Analysis Handbook, and the handbooks' guidance decides which insurers receive an inquiry.Supplement v5.0 · Intent

Every exhibit is optional, and the instructions on each tell regulators to cut it down for a limited-scope exam. The supplement suggests starting with Exhibit A and stopping there when an insurer's use of AI is limited or low in inherent risk. Its own example is a targeted claims exam that returns to the Market Regulation Handbook once Exhibit A shows the insurer's AI falls outside the exam's scope. The answers feed the regulator's view of the insurer's inherent risk and shape the nature, timing and extent of the work that follows.Supplement v5.0 · Instructions

Responses are held confidential under the requesting state's authority, and a regulator cites its examination or other authority when it asks.Supplement v5.0 · Confidentiality During the pilot, Pennsylvania said it would administer the tool through financial analysis and Iowa through its financial examinations, each to preserve confidentiality.Working Group, 1 Jun 2026 · 24 March minutes

Twelve states piloted it in 2026

The pilot ran in California, Colorado, Connecticut, Florida, Iowa, Louisiana, Maryland, Pennsylvania, Rhode Island, Vermont, Virginia and Wisconsin, from March to September 2026. Its purpose was to learn whether the tool helps insurers explain their AI governance and helps regulators understand it, and to inform long-term recommendations for market conduct and financial risk assessment. Each state focused on domestic insurers across property and casualty, life and health, could adapt the questions to its own needs, and was to spend more time on high-risk AI systems than on low-risk back-office ones.Pilot summary

Each pilot state's domestic regulator chose the insurers. By late March most pilot states had selected one to ten insurers, two had sent inquiries to more than ten, and property and casualty and life insurers outnumbered health insurers.Working Group, 1 Jun 2026 · 24 March minutes States delivered the tool as a formal examination, as a survey or data call, or as a hybrid of the two.Working Group, 1 Jun 2026 · Pilot update

Before the pilot began, the Working Group's chair said he believed participation would not be voluntary for the insurers selected, and that pilot states would coordinate so that an insurer does not receive inquiries from several states.Working Group, 9 Feb 2026 The Working Group's schedule allowed that not every state pilot would finish by 30 September 2026, and what states learn continues to feed the drafts through adoption.Working Group, Summer 2026 · Pilot timeline

Version 7.0 is planned for adoption in November 2026

The tool was first exposed for public comment on 7 July 2025, for 30 days.Working Group, 16 Jul 2025 Version 5.0 followed on 31 August 2026, with written comments due 29 September.Working Group page The Working Group released it as a clean document without tracked changes, because the tables had been restructured too heavily for a readable redline.Working Group, 8 Oct 2026 · 31 August minutes

The schedule set out on 31 August runs in three steps: a public meeting in early October to hear comments on version 5.0, then a version 6.0 with a 14-day comment period and two further public meetings, then a version 7.0 presented for adoption at the Fall National Meeting.Working Group, 8 Oct 2026 · 31 August minutes The Working Group scheduled a public meeting on 8 October and its next on 2 November 2026. As of 11 October, its page carried version 5.0 and the comments received on it, and no version 6.0.Working Group page The Fall National Meeting runs 14 to 17 November 2026 in Dallas.NAIC Fall National Meeting

In February the Working Group described the goal as adoption at the Fall National Meeting, with states then using the tool on a voluntary basis in 2027.Working Group, 9 Feb 2026

Insurers in the 12 pilot states received drafts of these questions in 2026, and every version is published on the Working Group's page. Preparing for the request does not depend on the adoption vote.

From a count of AI use down to the data behind each model

The supplement suggests using Exhibit A first and moving to Exhibits B to D only where the answers warrant it.Supplement v5.0 · Instructions Version 5.0 asks an insurer for the following.

  • A count of AI use (Exhibit A). How many AI Systems are in use and how many went live in a period the regulator sets, and how many AI models have a direct consumer impact or a material financial impact, each split into generalized linear models, generative or agentic AI models, and other AI and machine learning models. The rows run across operations from marketing, quoting, underwriting and ratemaking to claims, customer service, utilization management, fraud, investments, reserves, catastrophe triage and reinsurance. An insurer with an existing AI Systems inventory may offer it in place of the table.Supplement v5.0 · Exhibit A
  • The records behind the count (Exhibit A). The insurer's materiality calculation, its risk assessment process for AI Systems, and a model inventory that gives each model's use, use case, program area, inherent risk level, and consumer and financial impact.Supplement v5.0 · Exhibit A, Part Two
  • The AIS Program (Exhibit B). As a narrative or a checklist. The checklist asks the insurer to name the document and page that cover each process, among them unfair trade practices, adverse consumer outcomes, consumer data, enterprise risk management and the ORSA, financial reporting, employee training, vendor procurement, consumer complaints, consumer disclosure, materiality and oversight of vendor-built AI. The narrative asks about board reporting, independent validation, explainability, how the insurer assesses the autonomy and reversibility of its AI Systems, and, for uses with direct consumer impact, human review, error handling and any staffing reductions.Supplement v5.0 · Exhibit B
  • Each high-risk model (Exhibit C). For each model in production that the insurer's own risk criteria rate high: name and version, use case, model type, implementation date, whether it was built internally or by a third party (with the vendor's name), whether the insurer can modify it, its risk classification, risks and limitations, whether it automates, augments or supports a decision, how it was validated before deployment and is monitored now, the date of its last test, its effect on the financial statements, how it is reviewed for legal compliance, and any regulatory action taken over it.Supplement v5.0 · Exhibit C
  • The data (Exhibit D). For a list of data types running from aerial imagery and consumer risk scores to age, gender, ethnicity or race, medical information and voice analysis: how the model uses each one, and whether it comes from inside the insurer or from a named third-party vendor. A regulator may follow up by asking for a data dictionary.Supplement v5.0 · Exhibit D

It turns the Model Bulletin's examination list into fixed questions

The NAIC Model Bulletin on the Use of Artificial Intelligence Systems by Insurers, adopted by the NAIC on 4 December 2023, tells insurers what a department may request in an investigation or market conduct action: the written AIS Program and evidence of its adoption, inventories of the predictive models and AI Systems behind decisions that can produce adverse consumer outcomes, data lineage and bias analysis, validation and drift testing, and the diligence, contracts and audits behind third-party AI.Model Bulletin · Section 4

The supplement asks for the same material in a standard form, and gives Exhibit B over to the AIS Program itself. It defines its terms by the bulletin where it can.Supplement v5.0 · Definitions Version 5.0 also takes its definition of an AIS Program from the bulletin's language.Summary of changes

It reaches further in two directions. The bulletin frames its list around investigations and market conduct actions.Model Bulletin · Section 4 The supplement is written for financial analysts and financial examiners as well, so Exhibit B asks how AI risk enters enterprise risk management, the ORSA and financial reporting.Supplement v5.0 · Exhibit B It measures materiality the way the Financial Condition Examiners Handbook does.Supplement v5.0 · Materiality It also treats generative and agentic AI as a model type of their own and defines agentic AI.Supplement v5.0 · Definitions

Exhibit B says its questions about an AIS Program are “not intended to be interpreted as creating new requirements.”Supplement v5.0 · Exhibit B The expectations come from the bulletin and each state's insurance laws, and the supplement is how an examiner checks them.

What to have ready before the request arrives

Each exhibit asks for something an insurer either keeps already or has to reconstruct against a deadline. These are the records that answer it.

  • An inventory of AI Systems and AI models by operation area, with each model's type (generalized linear, generative or agentic, other) and flags for direct consumer impact and material financial impact, so the Exhibit A counts can be reproduced from it.Supplement v5.0 · Exhibit A
  • A documented materiality threshold with the reasoning for it, and a risk assessment that rates each model's inherent risk before controls. A regulator may set the threshold or ask for the insurer's own.Supplement v5.0 · Exhibit A
  • The written AIS Program with its adoption date, review frequency and the role of the board or management, indexed so each checklist item points to a document and page.Supplement v5.0 · Exhibit B
  • For each generative or agentic AI use that affects consumers: whether a person reviews its outputs, how exceptions and errors are found and resolved, and how a decision the model recommends or makes can be explained.Supplement v5.0 · Exhibit B
  • For each high-risk model: the version in production, its risk classification and limitations, the validation done before deployment, the monitoring since (drift, accuracy, unfair discrimination) and the date of the last test.Supplement v5.0 · Exhibit C
  • For third-party AI and data: the vendor behind each model and data source, the standards applied when procuring them, whether the insurer can modify the model, and how vendor-built systems are governed, monitored and tested.Supplement v5.0 · Exhibits B to D
  • Consumer-impact controls: how complaints that stem from an AI System are identified, tracked and addressed, and how consumers are told AI Systems are in use.Supplement v5.0 · Exhibit B
  • A record of what has already been filed. If the same information went to this or another state's department, the response can say so, and the regulator may accept the earlier submission while it is still current.Supplement v5.0 · Instructions

Control mapping

What a reviewer expects to be able to see.

ObligationWhat the system must doEvidence a reviewer expects
AI use counts (Exhibit A)Report AI Systems and AI models by operation area, model type and impactAn inventory that reproduces the counts, with model type and impact flags per model
Model inventory (Exhibit A)List each model with its use case, program area, inherent risk and impactThe model inventory, the materiality rationale and the risk assessment method
AIS Program (Exhibit B)Show a written program, its adoption and how it operatesProgram document with adoption date and review cycle, board reporting, and a page reference per checklist item
High-risk models (Exhibit C)Document validation, monitoring and legal review for each high-risk modelVersion history, pre-deployment validation, dated test results for drift, accuracy and unfair discrimination
Model data (Exhibit D)Identify each data element a model uses and where it comes fromData inventory by model with internal or vendor source, and a data dictionary
Third-party AIGovern, monitor and test AI Systems and data supplied by vendorsProcurement standards, vendor names per model and data source, oversight and test records

Key dates

Timeline, drawn to scale

From the Model Bulletin to an evaluation supplement

CompletedUpcomingPlanned adoption
2024202520264 DEC 2023Model Bulletin adopted7 JUL 2025Tool first exposed24 MAR 2026Pilot under way22 JUL 2026Renamed31 AUG 2026Version 5.0 exposed14 NOV 2026Adoption planned
Dates from the NAIC Big Data and Artificial Intelligence (H) Working Group, its meeting minutes and the NAIC meeting calendar. The November date is the Working Group's plan; the supplement had not been adopted as of the review date. Status as of October 11, 2026.
  • 4 December 2023NAIC adopts the Model Bulletin on the Use of Artificial Intelligence Systems by Insurers.
  • 7 July 2025First draft of the AI Systems Evaluation Tool exposed for a 30-day comment period ending 6 August 2025.
  • 24 March 2026Spring National Meeting: the Working Group reports the pilot under way in 12 states, having begun earlier in March.
  • 22 July 2026The Working Group reports the tool renamed the AI Risk Evaluation Supplement.
  • 31 August 2026Version 5.0 exposed for a 30-day comment period ending 29 September 2026.
  • 30 September 2026Scheduled end of the pilot period and of the survey of pilot insurers; some state pilots may run later.
  • 8 October 2026Public meeting to hear comments on version 5.0.
  • 2 November 2026Next scheduled public meeting of the Working Group on the supplement.
  • 14 November 2026NAIC Fall National Meeting opens in Dallas (14 to 17 November), where version 7.0 is planned to be presented for adoption.

Sources cited

Common gaps

Where an insurer answering the supplement is most likely to come up short.

  • Counts no one can reproduce. Exhibit A asks for counts by operation area and model type, and its second part asks for the model inventory behind them. An inventory assembled by survey gives a different number each time it is run, and the two answers will not match.
  • Generative AI left off the inventory. Exhibit A gives generative and agentic AI models their own columns, and Exhibit B asks whether a person is in the loop for uses with direct consumer impact. An assistant that summarizes a claim file for an adjuster augments a claims decision.
  • A program without page numbers. The checklist asks for the document name and page that cover each process. A process the AIS Program does not address shows up as an empty line in the response.
  • Vendor models with no vendor detail. Exhibit C asks who built each model and whether the insurer can change it, and Exhibit D asks which vendor supplies each third-party data element. Asked whether vendor names could be optional because of confidentiality agreements, a Working Group vice chair said that is for the examination staff and the insurer to discuss.Working Group, 9 Feb 2026
  • No materiality threshold of its own. A regulator may set the threshold or ask the insurer to state its own and explain it. An insurer without a documented threshold answers on the regulator's terms.
  • Waiting for the adoption vote. Adoption is planned for November 2026. The questions are public now, and insurers in the 12 pilot states received drafts of them this year.

Last reviewed October 11, 2026. This reference summarises publicly available regulatory guidance and is provided for general information. It is not legal advice. Obligations depend on an institution's charter, registration status, size, and activities. Verify against the primary sources cited above and consult counsel before relying on any summary here.

Regulatory updates

When a regulator changes what an AI examination asks for, hear about it first.

Short notes on SR 26-2, NYDFS 500, FINRA, the NAIC bulletin, the EU AI Act, and the employment-AI statutes, plus what we ship. A few emails a month.