The AI Governance Gap
Most healthcare AI governance programs are built around a single category of risk: the tools the organization deliberately purchased. However, when organizations take a hard look at where AI is actually operating inside their environment, the tools they knowingly procured are typically the smallest piece of the picture.
Limited visibility creates regulatory, operational, and data privacy risks. Organizations cannot effectively validate, monitor, or govern AI systems they have not identified. Effective governance starts by auditing all AI tools currently in use, including hidden or unauthorized tools, then establishing risk-based oversight.
AI Governance Review: Where to Begin
Most teams center their governance discussions on complying with federal regulations, such as the new HIPAA security rule, or meeting the intent of the latest quarterly HITRUST update. The best practices within these guidelines, including encryption, multi-factor authentication, and incident response testing, are all worth doing. But national laws and frameworks move slowly, while state laws are changing rapidly.
More than a dozen states now have statutes specifically governing AI’s use in prior authorization and utilization review, and organizations need to track those. The requirements are fairly consistent across states. AI can assist with a review, but a qualified human has to make the final determination. Decisions must reflect the individual patient’s circumstances rather than aggregate or population-level data. And organizations have to disclose when AI played a role.
State regulations aren’t especially controversial. The harder question is knowing what you have to apply the regulations to.
Inventory: Understanding Where AI Actually Lives
In a typical environment, AI shows up in four areas, and governance programs are usually built to monitor only the first area.
1. Procured on purpose
Most AI governance programs focus on oversight of purchased tools. The ones where someone ran a security review, negotiated a BAA, and put a name on file. These tools get monitored, tested, and revisited.
2. Embedded AI
This is AI that arrived inside a system you already own. Your EHR shipped a summarization feature. The payer platform added a coding assistant. The ticketing system turned on a suggestion engine in a minor release. No one procured the tool. No one read the release notes. It’s live and undocumented.
3. Built in house
These tools or fine-tuned models are built by internal teams on internal data. These are usually well-intentioned but poorly documented. They live in infrastructure originally stood up for proof of concept and never formalized.
4. Personal accounts
This is often the largest category and almost never on the company’s asset inventory. These include shadow AI tools used by employees on personal accounts. When organizations actually go looking, they often find two to three times more shadow AI solutions than what leadership expected.

Tier Your Findings Based on Risk
After you’ve found tools across the organization in all four categories, tier them by risk not technology. Evaluate what each tool’s decisions affect. Ask, “Are the outcomes reversible?” or “What happens downstream if the tool is wrong?” A triage or coverage determination tool warrants more security than a note-taking assistant.

Two Key Considerations for Testing
A vendor’s published performance numbers came from someone else’s patient population and someone else’s documentation practices. They tell you little about how the tool will behave on your data. Validate performance against internal data in testing before a tool influences any real decision. Define “acceptable performance” during procurement, not after go-live, when everyone involved has an incentive to grade the tool favorably.
Recognize that the model validated during procurement may not be the model running six months later. Vendors retrain, tune, and swap underlying models on their own schedules. If your contract doesn’t require notices of model changes, your organization may not know about them. Set thresholds, assign ownership to someone who understands the workflow and the infrastructure, and keep the governance committee focused on intake.
Governing Employee Use
Governance programs need to account for personal use, given that it’s often the biggest category of AI. Banning AI tools rarely works. When employees need help, they move to personal devices, which remove organizational visibility entirely.
A better approach is to give employees an approved option backed by real contract terms, then build the acceptable use policy around data classification rather than a list of approved products. A policy defining which categories of data can and can’t leave the environment will hold up through the next release cycle.
Also, product tiers often matter more than employees realize. Consumer and enterprise versions of the same product often carry different data retention and handling terms, which frequently determine whether a BAA is even possible. A free version of a tool that an organization approved at an enterprise tier is outside the agreement.
Employees need to know and understand AI policies clearly. Role-specific training with examples pulled from the work people do is essential. For example, telling someone not to put PHI into a chatbot doesn’t register with the person pasting in a scheduling exception, or a case summary they believe is already de-identified.
Humans must verify the accuracy of AI-generated output, especially anything entering a patient chart. Prompts and outputs are records. They can contain ePHI, they may be discoverable, and they shouldn’t be sitting in a vendor’s logs outside the organization’s retention policy.
Pay Attention to the Environment Underneath
Most organizations already have an access review process. It usually doesn’t know AI agents exist.
An agent or service account with write access to claims, scheduling, or clinical systems is a privileged identity, and it belongs in both the access review and decommissioning cycles. Many of these accounts were created for a pilot and never turned off. A useful starting question is to ask who owns each one.
Non-production environments deserve a second look as well. PHI sitting in test and development environments is a familiar problem, but AI development makes it worse: the data doesn’t just sit in a table anymore. It gets embedded, indexed, cached, and copied into places your data map doesn’t account for.
Retrieval systems deserve the closest scrutiny of all. A model with access to clinical notes will surface what’s in them, potentially including sensitive information that shouldn’t be surfaced. DLP rules written for file shares won’t catch this exposure. Neither will egress monitoring, because monitoring what leaves the organization through a prompt is fundamentally different than monitoring what leaves through an attachment. Most organizations have only built the latter capability.
Governance Starts with an Inventory
To build a strong AI governance program, start with a comprehensive AI inventory to better understand usage across the organization. From there, you can gauge risks and develop policies to address them.
This can be a big lift for some organizations. AArete works with health plans to close this kind of visibility gap. We help clients turn AI governance into an operating discipline, spanning access reviews, monitoring, and employee policy design, so AI adoption strengthens performance without introducing risk the organization can’t see.
Meet the Author

Robert Vitelli
Director, Cybersecurity