CHAI Is Not a Regulator. It Is a Library.
The assurance labs failed. The guidance survived. Here is how health leaders use what is left.
View all my published articles
In October 2025, the Secretary of Health and Human Services wrote that we must not let CHAI build a “regulatory cartel.” Four months later, the assurance labs program failed. Many health leaders stopped watching. That was a mistake. CHAI lost the fight to become a gatekeeper. It won a different fight, and few people noticed. It is now the best free library of health AI evaluation methods that your team can use this quarter.
Executive summary
The federal position is clear. CHAI has no regulatory authority. Senior officials said so in public. The plan for a national network of assurance labs did not survive. Reporting in February 2026 showed that the plan left CHAI members and health system leaders in confusion.
Two things did survive. The first is the partnership with The Joint Commission, announced in June 2025, which is building AI playbooks and a certification program on an accreditation platform. The second is the work group output. In July 2026, the Ambient AI Work Group released an implementation playbook with 42 best practices and a public testing and evaluation framework.
The second thing shows the model that works. The framework gives literature-backed metrics with named studies, lifecycle stages, and action thresholds. It does not give pass or fail limits. It tells you to set your own baseline and to calibrate to your own patients. This is guidance, not gatekeeping, and guidance is what health leaders asked for.
Use it that way. Put the retention and model-training questions into your next request for proposal. Ask each vendor for primary evidence, not an attestation. Keep per-encounter records that show which controls were active. Watch the certification program, because accreditation is the path that still has force.
What the fight was about
The Coalition for Health AI started in 2022. It grew to about 3,000 member organizations. Its main plan was a national network of assurance labs. These labs would test health AI products before hospitals put them into service. The previous administration supported the plan, in part because the FDA did not have the staff to do this work.
The current administration rejected it. On October 8, 2025, HHS Secretary Robert F. Kennedy Jr. posted his warning about a regulatory cartel. He praised an opinion piece by Deputy Secretary Jim O’Neill and FDA Commissioner Marty Makary. O’Neill told Politico that the department does not support the coalition, and that the group holds no regulatory power.
The criticism was not new. In June 2024, when CHAI opened a draft framework for public comment, some lawmakers said that large member companies could use the rules to block smaller competitors.
CHAI CEO Brian Anderson answered in a letter to members. He described an open process. He said that members rarely agree on everything, and that the group makes room for disagreement. He also said that the strategy matches current federal AI policy.
Then the labs plan collapsed. In February 2026, Fierce Healthcare reported that the assurance lab network never took shape. Affiliates and hospital executives told the publication that the effort created confusion instead of clarity.
That is the story most executives know. It is also where most executives stopped reading.
What survived
Two things survived, and both matter more than the fight.
The first is the accreditation path. In June 2025, CHAI and The Joint Commission announced a partnership to co-develop AI playbooks, tools, and a certification program. The program sits on The Joint Commission’s standards platform. The two groups said the guidance would reach more than 80 percent of healthcare organizations in the United States. A voluntary framework asks. An accreditation standard tells. If you track only one thing from this article, track this program.
The second is the work group library. While the policy fight continued, the work groups kept publishing. That output is free, public, and versioned on GitHub. Your team can read it today.
What good guidance looks like
On July 15, 2026, the Ambient AI Work Group released two documents. Oregon Health and Science University, Suki, Nabla, and Infinitus led the group. Pediatric practices, community health centers, integrated delivery networks, and academic medical centers all took part.
The first document is an implementation playbook. It gives 42 best practices across five lifecycle stages: procurement, pre-deployment, pilot, deployment, and monitoring. Most of the content is not about the model. It is about the operational problems that stop projects. Consent and recording rules change from state to state. Pediatric and behavioral health visits make those rules harder. Vendors store patient audio for periods that they do not always disclose. Per-seat licences ignore part-time clinicians and trainees. Staff use unapproved tools when the approved tool is slow. Vendors update models quietly, and performance drifts.
The best practices answer those problems in plain terms. Treat ambient documentation as a recording that state law governs. Build a workflow that lets a patient withdraw consent in the middle of a visit. Fund the tool as shared infrastructure, not as a productivity tax on each clinician. Measure success through clinician wellbeing and patient experience, not through adoption counts.
The second document is a testing and evaluation framework. It groups metrics under five principles: usefulness, fairness, safety, privacy, and business value. Each metric card names the responsible AI principle, the lifecycle phase, the persona who owns it, the published studies behind it, a benchmark, and, for the highest-risk metrics, an action threshold and a required response.
The numbers are specific. A documentation time reduction of about 15 percent per appointment. A drop of at least 0.44 points in work exhaustion on the Stanford Professional Fulfillment Index, with a number needed to treat of 1.68 clinicians. A speech recognition error gap of 0.10 or more between demographic groups, which triggers remediation before further use. A weekly increase of 5.8 percent in relative value units. A target of zero known clinically significant errors at sign-off, with clinician review on 100 percent of notes.
Here is the part that makes the framework useful. The authors state that these values are reference points from the cited literature. They are not universal pass or fail limits. Performance changes with setting, specialty, workflow, and patient population. Organizations must set a local baseline and revisit it as the case mix changes.
That is the difference between a library and a gate. A gate gives you a score. A library gives you a method, the evidence behind it, and the responsibility to think.
Four moves you can make this quarter
Put data questions in the procurement document. Ask where the vendor stores session audio and transcripts. Ask how long it keeps them. Ask whether your clinical data trains the vendor’s models. The framework notes that practice varies widely across the market. Audio retention ranges from immediate deletion to about 90 days. Transcript retention ranges from seven days to no deletion at all. Some vendors do not disclose the answer.
Ask for primary evidence, not attestation. A vendor statement is not proof. Ask for the configuration export and the contract language that set the retention period. Add a clause that forbids model training on your clinical data without your written approval.
Keep per-encounter control evidence. For each recorded visit, keep a record of the consent scope, the clinician sign-off, the retention setting, and the model version in use. The framework ties this directly to active litigation over ambient documentation and to shrinking AI insurance cover. If you cannot reconstruct which controls were active on a given day, you cannot defend the encounter.
Measure people, not adoption. Seat counts prove that you bought something. Burnout scores, note quality, and patient experience prove that it worked.
Where the library is thin
The published use cases now cover ambient AI, sepsis risk prediction, discharge summarization, clinical decision support, EHR information retrieval, general health advice chatbots, mental health, clinical trials, agentic systems, and prior authorization criteria matching.
Pharmacy is not on that list. Neither is much of the operational core of a hospital. If your programme sits outside the published set, the metric card structure still transfers, but the numbers do not. Build your own cards. Cite your own literature. Set your own thresholds.
One more caution. Two ambient AI vendors helped lead the group that defined how ambient AI should be evaluated. The benchmarks come from independent studies in JAMA Network Open, NEJM AI, and Mayo Clinic Proceedings, so the evidence stands on its own. The question of who writes the rules for a category is still fair to ask, and it is the same question that started the cartel fight.
The bottom line
CHAI tried to be a regulator and failed. It became a library and succeeded. Your team does not need CHAI’s permission to deploy AI. Your team needs its methods, its citations, and its questions. Take them. Adapt them to your patients. Then watch the certification programme, because that is where voluntary guidance turns into something your surveyors will ask about.
Disclosure: my company builds a governance layer for healthcare AI agents, so I have a commercial interest in this subject.
Sources:
Becker’s Health IT and Healthcare IT News, October 2025.
Fierce Healthcare, February 2026.
The Joint Commission and CHAI joint announcement, June 2025.
Paul J. Swider is CEO and Chief AI Officer at RealActivity, a Microsoft Partner specializing in mission-critical AI for healthcare systems. He has 30+ years in healthcare technology, has trained over 3,000 engineers across GE, IDX, and Microsoft, and is the founder of BOSHUG, the Boston Healthcare Cloud & AI Community spanning 50+ countries.


