A healthcare AI governance framework is a standing decision process with four working parts: how a tool gets proposed and approved, what evidence the approvers require before saying yes, who owns the tool after it goes live, and what specifically triggers pausing or retiring it. That is the whole framework. Everything else is documentation around those four decisions.
It is worth being blunt about why this matters, because the failure mode is not usually a hospital with no policy. It is a hospital with an excellent policy and no owner. The document exists, the committee met twice, and there is a sepsis model quietly running in production that nobody has evaluated in eighteen months because the person who championed it moved to a different health system. That is a governance failure, and no amount of additional policy language fixes it.
Part One: Intake, So Shadow AI Has Somewhere To Go
Every AI tool in the building arrives through one of three doors: procurement buys it, the EHR vendor ships it in an update, or a clinician starts using a consumer tool on their own. Governance frameworks are usually built for the first door and are then surprised by the other two. The EHR door matters because a model can appear in your environment without a purchase order, as part of a version upgrade. The clinician door matters because it is where protected health information leaves the building fastest.
A workable intake process is short enough that people actually use it. If the form takes forty minutes, staff will route around it, and you will have less visibility than before you built the process. The intake step only needs to answer three things: what does this tool do, what data does it touch, and does it influence a clinical decision. Those three answers determine how much scrutiny follows.
Part Two: The Evidence Bar, Scaled to Risk
Not every tool deserves the same review. A model that predicts patient deterioration and a tool that drafts appointment reminder text should not go through the same committee process, and if they do, the committee will be too slow for the reminders and too fast for the deterioration model. Tiering by clinical influence is what keeps the process usable.
- Influences a clinical decision: requires validation evidence on a population resembling yours, a named clinical owner, documented FDA status, and a monitoring plan before go-live
- Touches protected health information but not clinical decisions: requires a Business Associate Agreement, a data flow review, and a retention and model-training answer in writing
- Neither: requires intake registration so it is on the inventory, and little else
The validation question is the one most often waved through. A vendor will show you performance metrics. The question that matters is what population those metrics came from. A readmission or sepsis model validated on an academic medical center in one region can behave meaningfully differently on a rural population with a different case mix, different baseline acuity, and different documentation habits. Asking for the validation population is not an insult to the vendor. It is the single most useful question in the entire review.
Part Three: Named Ownership After Go-Live
This is the part that separates a real framework from a published one. Every approved tool needs a named individual, not a department, who is accountable for it after deployment. Departments do not notice model drift. People do, and only when it is their job.
The owner is responsible for a small, recurring set of questions: is the model still performing the way it did at approval, has the vendor pushed an update that changed behavior, are clinicians actually using the output or silently ignoring it, and has anything about our patient population shifted. That last one is easy to miss. A model does not have to change for its performance to degrade. The population underneath it can change instead.
Part Four: The Off Switch
Almost no published governance framework specifies what triggers turning a tool off, which is strange, because that is the only part of governance that constitutes actual control. Approval without a decommission process is not governance. It is procurement with extra meetings.
Triggers worth writing down in advance: performance falls below a stated threshold, the vendor changes the model materially without notice, an adverse event is linked to the output, the clinical owner leaves and is not replaced within a set window, or the tool has simply stopped being used. Writing these down before deployment matters because the moment you need them is the moment it is hardest to make the decision on the merits.
Where the External Frameworks Fit
NIST's AI Risk Management Framework and the work coming out of the Coalition for Health AI are useful reference structures, and borrowing their vocabulary saves argument time. But they are voluntary scaffolding, not compliance obligations, and adopting one does not answer any of the four questions above for your specific building. Use them for structure. Do not mistake having cited them for having governed anything.
The regulatory floor underneath all of this is separate and non-optional. Which specific rules apply depends on what the tool does, and that is worth understanding before the committee convenes rather than during it.
The Working Checklist
- Keep one inventory covering purchased tools, EHR-supplied models, local builds, and approved consumer AI use
- Assign a risk tier based on clinical influence, patient data, scale, and ability to cause harm
- Record the evidence, regulatory classification, data flow, owner, monitoring metric, review date, and stop condition
- Review material vendor and model changes before they reach production
- Retire tools that lose an owner, fall below a threshold, create an adverse event, or are no longer used
Use the compliance map to classify obligations and the responsible AI tests to set the evidence and monitoring bar. Joint Commission's certification framework and CHAI's governance playbooks provide useful external structure for the same working system.