Reference · Insurance
NAIC AI Systems Evaluation Tool: The AI Risk Evaluation Supplement Explained
The questionnaire state regulators are building to examine an insurer's AI, piloted in 12 states in 2026
At a glance
- Current name
- AI Risk Evaluation Supplement, renamed from AI Systems Evaluation Tool by July 2026
- Status
- In draft. Version 7.0 is planned for adoption at the NAIC Fall National Meeting, 14 to 17 November 2026
The AI Systems Evaluation Tool is a set of questionnaires that state insurance regulators can use to find out how an insurer uses AI and whether its governance of that use is working. The NAIC's Big Data and Artificial Intelligence (H) Working Group drafts it, under a charge to build tools that help regulators identify and assess the financial and consumer risks of AI Systems on an ongoing basis.Supplement v5.0 · Intent By its 22 July 2026 meeting the Working Group had renamed it the AI Risk Evaluation Supplement, in response to stakeholder feedback on the name and to reduce confusion about what the document is and is not.Working Group, Summer 2026 · 22 July minutes
Twelve states piloted it from March to September 2026, in market conduct exams, financial exams and financial analysis.Pilot summary The current public draft is version 5.0. The Working Group plans one more draft for comment and then a version 7.0 for adoption at the NAIC Fall National Meeting, held 14 to 17 November 2026 in Dallas.Working Group, 8 Oct 2026 · 31 August minutesNAIC Fall National Meeting As of 11 October 2026 it remains a draft.
A supplement to the examination handbooks
The exhibits add to the review procedures examiners already follow in market conduct, financial analysis and financial examination work, and the NAIC's existing resources stay authoritative. Requests made under the supplement are coordinated through the Market Regulation Handbook, the Financial Condition Examiners Handbook and the Financial Analysis Handbook, and the handbooks' guidance decides which insurers receive an inquiry.Supplement v5.0 · Intent
Every exhibit is optional, and the instructions on each tell regulators to cut it down for a limited-scope exam. The supplement suggests starting with Exhibit A and stopping there when an insurer's use of AI is limited or low in inherent risk. Its own example is a targeted claims exam that returns to the Market Regulation Handbook once Exhibit A shows the insurer's AI falls outside the exam's scope. The answers feed the regulator's view of the insurer's inherent risk and shape the nature, timing and extent of the work that follows.Supplement v5.0 · Instructions
Responses are held confidential under the requesting state's authority, and a regulator cites its examination or other authority when it asks.Supplement v5.0 · Confidentiality During the pilot, Pennsylvania said it would administer the tool through financial analysis and Iowa through its financial examinations, each to preserve confidentiality.Working Group, 1 Jun 2026 · 24 March minutes
Twelve states piloted it in 2026
The pilot ran in California, Colorado, Connecticut, Florida, Iowa, Louisiana, Maryland, Pennsylvania, Rhode Island, Vermont, Virginia and Wisconsin, from March to September 2026. Its purpose was to learn whether the tool helps insurers explain their AI governance and helps regulators understand it, and to inform long-term recommendations for market conduct and financial risk assessment. Each state focused on domestic insurers across property and casualty, life and health, could adapt the questions to its own needs, and was to spend more time on high-risk AI systems than on low-risk back-office ones.Pilot summary
Each pilot state's domestic regulator chose the insurers. By late March most pilot states had selected one to ten insurers, two had sent inquiries to more than ten, and property and casualty and life insurers outnumbered health insurers.Working Group, 1 Jun 2026 · 24 March minutes States delivered the tool as a formal examination, as a survey or data call, or as a hybrid of the two.Working Group, 1 Jun 2026 · Pilot update
Before the pilot began, the Working Group's chair said he believed participation would not be voluntary for the insurers selected, and that pilot states would coordinate so that an insurer does not receive inquiries from several states.Working Group, 9 Feb 2026 The Working Group's schedule allowed that not every state pilot would finish by 30 September 2026, and what states learn continues to feed the drafts through adoption.Working Group, Summer 2026 · Pilot timeline
Version 7.0 is planned for adoption in November 2026
The tool was first exposed for public comment on 7 July 2025, for 30 days.Working Group, 16 Jul 2025 Version 5.0 followed on 31 August 2026, with written comments due 29 September.Working Group page The Working Group released it as a clean document without tracked changes, because the tables had been restructured too heavily for a readable redline.Working Group, 8 Oct 2026 · 31 August minutes
The schedule set out on 31 August runs in three steps: a public meeting in early October to hear comments on version 5.0, then a version 6.0 with a 14-day comment period and two further public meetings, then a version 7.0 presented for adoption at the Fall National Meeting.Working Group, 8 Oct 2026 · 31 August minutes The Working Group scheduled a public meeting on 8 October and its next on 2 November 2026. As of 11 October, its page carried version 5.0 and the comments received on it, and no version 6.0.Working Group page The Fall National Meeting runs 14 to 17 November 2026 in Dallas.NAIC Fall National Meeting
In February the Working Group described the goal as adoption at the Fall National Meeting, with states then using the tool on a voluntary basis in 2027.Working Group, 9 Feb 2026
From a count of AI use down to the data behind each model
The supplement suggests using Exhibit A first and moving to Exhibits B to D only where the answers warrant it.Supplement v5.0 · Instructions Version 5.0 asks an insurer for the following.
- A count of AI use (Exhibit A). How many AI Systems are in use and how many went live in a period the regulator sets, and how many AI models have a direct consumer impact or a material financial impact, each split into generalized linear models, generative or agentic AI models, and other AI and machine learning models. The rows run across operations from marketing, quoting, underwriting and ratemaking to claims, customer service, utilization management, fraud, investments, reserves, catastrophe triage and reinsurance. An insurer with an existing AI Systems inventory may offer it in place of the table.Supplement v5.0 · Exhibit A
- The records behind the count (Exhibit A). The insurer's materiality calculation, its risk assessment process for AI Systems, and a model inventory that gives each model's use, use case, program area, inherent risk level, and consumer and financial impact.Supplement v5.0 · Exhibit A, Part Two
- The AIS Program (Exhibit B). As a narrative or a checklist. The checklist asks the insurer to name the document and page that cover each process, among them unfair trade practices, adverse consumer outcomes, consumer data, enterprise risk management and the ORSA, financial reporting, employee training, vendor procurement, consumer complaints, consumer disclosure, materiality and oversight of vendor-built AI. The narrative asks about board reporting, independent validation, explainability, how the insurer assesses the autonomy and reversibility of its AI Systems, and, for uses with direct consumer impact, human review, error handling and any staffing reductions.Supplement v5.0 · Exhibit B
- Each high-risk model (Exhibit C). For each model in production that the insurer's own risk criteria rate high: name and version, use case, model type, implementation date, whether it was built internally or by a third party (with the vendor's name), whether the insurer can modify it, its risk classification, risks and limitations, whether it automates, augments or supports a decision, how it was validated before deployment and is monitored now, the date of its last test, its effect on the financial statements, how it is reviewed for legal compliance, and any regulatory action taken over it.Supplement v5.0 · Exhibit C
- The data (Exhibit D). For a list of data types running from aerial imagery and consumer risk scores to age, gender, ethnicity or race, medical information and voice analysis: how the model uses each one, and whether it comes from inside the insurer or from a named third-party vendor. A regulator may follow up by asking for a data dictionary.Supplement v5.0 · Exhibit D
It turns the Model Bulletin's examination list into fixed questions
The NAIC Model Bulletin on the Use of Artificial Intelligence Systems by Insurers, adopted by the NAIC on 4 December 2023, tells insurers what a department may request in an investigation or market conduct action: the written AIS Program and evidence of its adoption, inventories of the predictive models and AI Systems behind decisions that can produce adverse consumer outcomes, data lineage and bias analysis, validation and drift testing, and the diligence, contracts and audits behind third-party AI.Model Bulletin · Section 4
The supplement asks for the same material in a standard form, and gives Exhibit B over to the AIS Program itself. It defines its terms by the bulletin where it can.Supplement v5.0 · Definitions Version 5.0 also takes its definition of an AIS Program from the bulletin's language.Summary of changes
It reaches further in two directions. The bulletin frames its list around investigations and market conduct actions.Model Bulletin · Section 4 The supplement is written for financial analysts and financial examiners as well, so Exhibit B asks how AI risk enters enterprise risk management, the ORSA and financial reporting.Supplement v5.0 · Exhibit B It measures materiality the way the Financial Condition Examiners Handbook does.Supplement v5.0 · Materiality It also treats generative and agentic AI as a model type of their own and defines agentic AI.Supplement v5.0 · Definitions
What to have ready before the request arrives
Each exhibit asks for something an insurer either keeps already or has to reconstruct against a deadline. These are the records that answer it.
- An inventory of AI Systems and AI models by operation area, with each model's type (generalized linear, generative or agentic, other) and flags for direct consumer impact and material financial impact, so the Exhibit A counts can be reproduced from it.Supplement v5.0 · Exhibit A
- A documented materiality threshold with the reasoning for it, and a risk assessment that rates each model's inherent risk before controls. A regulator may set the threshold or ask for the insurer's own.Supplement v5.0 · Exhibit A
- The written AIS Program with its adoption date, review frequency and the role of the board or management, indexed so each checklist item points to a document and page.Supplement v5.0 · Exhibit B
- For each generative or agentic AI use that affects consumers: whether a person reviews its outputs, how exceptions and errors are found and resolved, and how a decision the model recommends or makes can be explained.Supplement v5.0 · Exhibit B
- For each high-risk model: the version in production, its risk classification and limitations, the validation done before deployment, the monitoring since (drift, accuracy, unfair discrimination) and the date of the last test.Supplement v5.0 · Exhibit C
- For third-party AI and data: the vendor behind each model and data source, the standards applied when procuring them, whether the insurer can modify the model, and how vendor-built systems are governed, monitored and tested.Supplement v5.0 · Exhibits B to D
- Consumer-impact controls: how complaints that stem from an AI System are identified, tracked and addressed, and how consumers are told AI Systems are in use.Supplement v5.0 · Exhibit B
- A record of what has already been filed. If the same information went to this or another state's department, the response can say so, and the regulator may accept the earlier submission while it is still current.Supplement v5.0 · Instructions
Control mapping
What a reviewer expects to be able to see.
| Obligation | What the system must do | Evidence a reviewer expects |
|---|---|---|
| AI use counts (Exhibit A) | Report AI Systems and AI models by operation area, model type and impact | An inventory that reproduces the counts, with model type and impact flags per model |
| Model inventory (Exhibit A) | List each model with its use case, program area, inherent risk and impact | The model inventory, the materiality rationale and the risk assessment method |
| AIS Program (Exhibit B) | Show a written program, its adoption and how it operates | Program document with adoption date and review cycle, board reporting, and a page reference per checklist item |
| High-risk models (Exhibit C) | Document validation, monitoring and legal review for each high-risk model | Version history, pre-deployment validation, dated test results for drift, accuracy and unfair discrimination |
| Model data (Exhibit D) | Identify each data element a model uses and where it comes from | Data inventory by model with internal or vendor source, and a data dictionary |
| Third-party AI | Govern, monitor and test AI Systems and data supplied by vendors | Procurement standards, vendor names per model and data source, oversight and test records |
Key dates
Timeline, drawn to scale
From the Model Bulletin to an evaluation supplement
- 4 December 2023NAIC adopts the Model Bulletin on the Use of Artificial Intelligence Systems by Insurers.
- 7 July 2025First draft of the AI Systems Evaluation Tool exposed for a 30-day comment period ending 6 August 2025.
- 24 March 2026Spring National Meeting: the Working Group reports the pilot under way in 12 states, having begun earlier in March.
- 22 July 2026The Working Group reports the tool renamed the AI Risk Evaluation Supplement.
- 31 August 2026Version 5.0 exposed for a 30-day comment period ending 29 September 2026.
- 30 September 2026Scheduled end of the pilot period and of the survey of pilot insurers; some state pilots may run later.
- 8 October 2026Public meeting to hear comments on version 5.0.
- 2 November 2026Next scheduled public meeting of the Working Group on the supplement.
- 14 November 2026NAIC Fall National Meeting opens in Dallas (14 to 17 November), where version 7.0 is planned to be presented for adoption.
Sources cited
- ×23NAIC, Artificial Intelligence (AI) Risk Evaluation Supplement, version 5.0 exposure draft (31 August 2026, Word document)
- ×2NAIC Big Data and Artificial Intelligence (H) Working Group, 2026 Summer National Meeting materials: minutes of 22 July 2026 and pilot timeline (PDF)
- ×2NAIC, AI Systems Evaluation Tool Pilot: Pilot Project Background (PDF)
- ×3NAIC Big Data and Artificial Intelligence (H) Working Group, agenda and materials for 8 October 2026, with draft minutes of 31 August 2026 (PDF)
- ×2NAIC, 2026 Fall National Meeting
- ×3NAIC Big Data and Artificial Intelligence (H) Working Group, materials for 1 June 2026: minutes of 24 March 2026 and pilot update (PDF)
- ×3NAIC Big Data and Artificial Intelligence (H) Working Group, minutes of 9 February 2026 (PDF)
- ×1NAIC Big Data and Artificial Intelligence (H) Working Group, minutes of 16 July 2025 (PDF)
- ×2NAIC Big Data and Artificial Intelligence (H) Working Group: exposure drafts and meeting schedule
- ×2NAIC Model Bulletin: Use of Artificial Intelligence Systems by Insurers (adopted text, PDF)
- ×1NAIC staff, AI Risk Evaluation Supplement: Summary of Changes from version 4.0 (PDF)
Common gaps
Where an insurer answering the supplement is most likely to come up short.
- Counts no one can reproduce. Exhibit A asks for counts by operation area and model type, and its second part asks for the model inventory behind them. An inventory assembled by survey gives a different number each time it is run, and the two answers will not match.
- Generative AI left off the inventory. Exhibit A gives generative and agentic AI models their own columns, and Exhibit B asks whether a person is in the loop for uses with direct consumer impact. An assistant that summarizes a claim file for an adjuster augments a claims decision.
- A program without page numbers. The checklist asks for the document name and page that cover each process. A process the AIS Program does not address shows up as an empty line in the response.
- Vendor models with no vendor detail. Exhibit C asks who built each model and whether the insurer can change it, and Exhibit D asks which vendor supplies each third-party data element. Asked whether vendor names could be optional because of confidentiality agreements, a Working Group vice chair said that is for the examination staff and the insurer to discuss.Working Group, 9 Feb 2026
- No materiality threshold of its own. A regulator may set the threshold or ask the insurer to state its own and explain it. An insurer without a documented threshold answers on the regulator's terms.
- Waiting for the adoption vote. Adoption is planned for November 2026. The questions are public now, and insurers in the 12 pilot states received drafts of them this year.
Related
Last reviewed October 11, 2026. This reference summarises publicly available regulatory guidance and is provided for general information. It is not legal advice. Obligations depend on an institution's charter, registration status, size, and activities. Verify against the primary sources cited above and consult counsel before relying on any summary here.