Quick Summary

No individual large language model is compliant with HIPAA by itself; compliance applies to the entire system. Each vendor that touches PHI must have a business associate agreement (BAA) covering the specific service you use, and your design must restrict, record, and protect PHI at all stages.

  • A BAA applies only to certain products and endpoints and not to a vendor's whole catalog.
  • A production generative AI system includes numerous PHI hops beyond the model itself, including telephony, embeddings, vector stores, logs, alerts, and review queues.
  • The HIPAA Security Rule is fully enforced, and the Office for Civil Rights's enforcement focus is risk analysis. Your AI components should therefore be included in that analysis.
  • The most effective design control is to minimize what the model sees.

Note that Cypherox develops AI systems for healthcare clients and therefore has an interest in this matter. The following content is intended as engineering advice and not as legal advice. You should consult your privacy officer or legal advisor regarding your decisions.

What "HIPAA-Compliant Generative AI" Actually Means

HIPAA applies to organizations, not to software; it covers covered entities (such as providers, health plans, and clearinghouses) as well as the business associates who act on their behalf regarding PHI. A model cannot be compliant, but a deployment can.

In practice, a generative AI system complies with HIPAA if three conditions are met. Each organization that creates, receives, stores, or transmits PHI must have a BAA covering the specific service that is being used. The safeguards required by the Security Rule, such as risk analysis, access control, audit controls, integrity, and transmission security, also apply to the AI components.

The third condition is that the limits on the use and disclosure set out in the Privacy Rule, including the minimum necessary standard, determine where the data flows. When a vendor refers to its model as 'HIPAA-compliant,' it generally means something more limited in that it will enter into a BAA for some of its products.

The Rules in Force as of September 2026

Do not act in accordance with the rule currently in place, but with the one that is under discussion. In the Fall 2026 Unified Agenda, HHS transferred the proposed amendments to the HIPAA Security Rule to its Long-Term Actions agenda and set July 2027 as the expected time for taking final action. The present Security Rule is still fully enforceable, and the OCR continues to expect that covered entities and business associates will put in place reasonable and appropriate safeguards.

Enforcement focuses on a single requirement. On April 24, 2026, the OCR reached settlement agreements with four regulated entities following separate investigations into ransomware incidents that impacted more than 427,000 people and involved payments totaling $1,165,000. These cases are part of the OCR's Risk Analysis Initiative, whose common feature is that the entities had not conducted an accurate and complete assessment of the risks to the confidentiality, integrity, and availability of ePHI.

The breach figures show how serious the situation is. By June 2026, the HHS breach database showed that in 2025 there had been 772 healthcare breaches affecting 500 or more individuals, exposing the personal health information (PHI) of 139,721,832 people and making that year the worst in history for large-scale breaches. The biggest incident in 2025 took place at a business associate, Conduent, and involved the compromise of the personal health information of more than 62 million Americans.

Each AI vendor you include becomes another business associate in that chain. If your risk assessment omits your LLM endpoint, embedding service, or vector database, then that omission is the very gap that OCR is always referring to.

Where PHI Moves in a Generative AI System

Most teams focus on a single stage: the prompt sent to the model. A production system involves many more stages. The simplest way to identify them is by tracing a real build.

Cypherox has developed an AI-powered healthcare voicebot designed for continuous patient monitoring. The architecture that has been published makes use of Twilio for making outbound calls and for voice streaming, employs Kore.ai for recognizing intent, extracting entities, and detecting sentiment, includes a React and Node.js dashboard, uses PHP for scheduling and generating reports, and stores patient records, call logs, and health readings in MySQL. The system automatically saves the call recordings and transcripts. When emergency alerts occur, the system sends them to doctors via push notifications and email, along with the patient's name, the breached threshold, and the exact response delivered.

The system carries out its clinical responsibilities. Weekly follow-up coverage has increased from about 30% when done manually to 100% automation, and the provider saves over 200 staff hours each month. Now let's count the PHI hops: the phone contact, the interactive AI platform, the database, the stored recordings and transcripts, the dashboard, and the alert emails.

All of them fall within the scope of compliance. This alert email is a good example: it is clinically actionable because it includes the patient's name and the exact reading, and because of these details, the email method stays within your HIPAA boundaries.

Here is the basic outline for any generative AI system.

Component

What PHI it sees

What to confirm

Telephony and voice

Caller audio, phone numbers, recordings

Service is on the vendor's HIPAA-eligible list; where recordings are stored

Speech-to-text and conversational AI

Transcripts, extracted entities such as vitals and medications

BAA in place; retention period; whether your data trains vendor models

LLM inference API

Prompts and completions

BAA covers the specific endpoint and features you call

Embedding model

Text chunks sent for vectorizing

BAA if any chunk contains PHI

Vector database

Embeddings plus stored source text

BAA; per-document access control

Orchestration and agent tools

Tool inputs, outputs, and agent memory

Least-privilege permissions; what gets persisted

Logs, traces, and eval datasets

Full prompts and responses

BAA with observability vendor; retention; redaction

Notifications and alerts

Names, readings, clinical flags

Channel security; minimum content per message

Human review queues

Flagged conversations

Role-based access; audit trail

The rows teams tend to overlook are generally the final four; although they don't show up in the architecture diagram, they contain the same data as the model call.

What a BAA Covers, and What It Doesn't

Begin with the definition as set out by HHS: a cloud service provider that processes or stores ePHI on behalf of a covered entity is considered to be a business associate even if it has only encrypted data and does not possess the decryption key. Storing only encrypted blobs does not exempt the vendor from the scope.

Omitting a BAA has consequences; for example, HHS cites a case in which OCR reached a resolution agreement and drew up a corrective action plan with a covered entity that stored the ePHI of more than 3,000 individuals on a cloud server without a BAA.

A more frequent error is to think that a BAA applies to all the products a vendor offers. OpenAI makes this clear in its documentation. According to its Business Associate Addendum, BAA-eligible endpoints may process PHI. Still, Web Search with live internet access is not HIPAA-eligible and therefore not covered by a BAA. The rule about zero data retention is a separate requirement: it affects how long OpenAI keeps data, but on its own it does not make OpenAI a business associate, nor does it authorize the processing of PHI.

Telephony operates in the same manner. Twilio requires either the Security Edition or the Enterprise Edition to sign a BAA, and customers handling PHI should limit themselves to the products Twilio lists as HIPAA-eligible. Moreover, Twilio makes it clear that signing a BAA alone does not make an application HIPAA-compliant; it must adhere to its architecture guidelines.

The model can be covered on one route and not on another. The Gemini API available via AI Studio does not have a BAA, even though the same model family is offered under a BAA through Vertex AI. As a practical matter, your BAA review should include the endpoints, features, and regions, not the names of the vendors.

Six Architecture Decisions That Keep PHI Under Control

Six Architecture Decisions That Keep PHI Under Control

1. Send the model only what the task needs

The Privacy Rule's minimum necessary standard requires you to restrict PHI to what is necessary for the purpose at hand. Exceptions exist, for example, when disclosing information to a provider for treatment. In most cases, AI use falls within operations, and prompts assembled "just in case" breach this rule.

A consistent practice is to replace direct identifiers with tokens before the prompt leaves your environment, then reassociate the response within your own systems. The model can reason about "Patient 7F3A, systolic 142, missed two doses" just as well as a patient with a name. Only your database knows who 7F3A is.

The voicebot illustrates a similar principle: doctors require medical accuracy and trend data, while caregivers need simple summaries. Both groups get information from the same database through two specially designed portal views. This is what "minimum necessary" should look like in software.

2. Know when de-identification actually works

Data that has been properly de-identified is not considered PHI. Under the Safe Harbor method, a covered entity must eliminate all 18 listed identifiers and must not know that the remaining information could identify an individual. HHS also notes that a code or record identifier derived from PHI usually must be removed under Safe Harbor.

Free text constitutes PHI and should be treated as such under a BAA unless an expert determination says otherwise. Clinical notes and call transcripts can contain identifiers in many places, for example, "my daughter drove me to the clinic on Elm Street after my work at the plant." Pattern-based redaction fails to catch phrasing like this.

De-identification is worthwhile for evaluation datasets and analytics. HHS says a cloud provider that receives and holds only properly de-identified information is not considered a business associate.

3. Choose the hosting path deliberately

Three realistic options exist, and each shifts responsibility differently.

PathWho signs the BAAWhat you must verifyBest fit
Direct model APIThe model vendorWhich endpoints and features the BAA coversTeams standardized on one vendor's models
Cloud AI platformYour cloud providerThat the specific model and region are on the eligible listTeams already running PHI workloads in that cloud
Self-hosted open modelYour cloud or infrastructure provider onlyPatching, access control, and evaluation, all owned by youStrict data-residency needs or high, steady volume

Although the cloud path can combine several BAAs into one, check the small print. Because AWS's reference lists of HIPAA-eligible services explicitly exclude Amazon Bedrock, make sure your specific model appears on the eligible list.

Location is also important. Although HIPAA allows a cloud provider to store ePHI outside the United States pursuant to a BAA, OCR notes that offshore storage risks can vary and must be considered in your risk analysis. Regardless of your data residency decision, put it in writing.

4. Treat retrieval and memory as PHI stores

With retrieval-augmented generation, the source material stays close to its embeddings. If the source material comes from charts or transcripts, your vector database falls within the scope of an ePHI system. It must meet the same encryption, access control, backup, and disposal requirements as your EHR integration.

The harder issue is permissions. Before any chunks reach the model, filter them according to what the requesting user may see. It is not access control to ask the model not to reveal records the user cannot access.

The same applies to agent memory and conversation history. Decide how long to retain each, then delete it on a set schedule. When building agents that invoke tools in various clinical systems, limit each tool to the minimum amount of data it requires. This is where careful AI agent development begins to pay off.

5. Log for audit without building a second PHI leak

The Audit Controls required by the HIPAA Security Rule must record and examine activity in systems that contain ePHI. For generative AI, the audit log should show who made the request, which records were retrieved, which model version was used, what output it generated, and what action was taken thereafter.

The problem is that prompt logs contain PHI. Once you store them, your observability tools become a business associate. For example, Datadog's LLM Observability documentation advises HIPAA customers to connect only an OpenAI account subject to a BAA and configured for zero data retention.

A better approach is to split the logs. Ordinary application logs contain only references, such as record IDs, hashes, model versions, and latency. The full prompt and response content is stored in a restricted store with its own retention policy and access review.

6. Defend against injection and keep humans at clinical risk

Generative AI has introduced attack types that traditional healthcare software could not address. Prompt injection has again taken the number one position in the OWASP Top 10 for LLM Applications, and sensitive information disclosure has moved from sixth to second in the 2025 list. In healthcare, indirect injection usually occurs through documents the system processes, such as referral letters, faxed records, or patient emails.

The defensive measures are architectural in nature. Give each tool only the most limited permission possible. Obtain human approval for any high-impact actions, and check outputs for PHI before they leave the system.

Clinical escalation requires the same discipline. In the voicebot, each node has a fallback sequence: re-prompt, then rephrase, then flag the case for human review. An individual critical reading triggers an alert right away, as does a pattern of borderline readings over several calls.

HIPAA is not the only regulation concerning AI-based patient risk assessment; the rule in section 1557 at 45 CFR 92.210 requires covered entities to make reasonable efforts to identify and reduce discrimination resulting from patient care decision support tools, and compliance is due by May 1, 2025. This section is listed in the eCFR as up to date as of September 2026. Whenever your model is used to triage or score patients, record the input variables it uses and the reasons for doing so.

Pre-Launch Checklist for HIPAA-Compliant Generative AI

Pre-Launch Checklist for HIPAA-Compliant Generative AI
  1. Each of the components in the PHI map is assigned a named owner and has a documented data flow.
  2. All vendors with access to PHI have a written BAA that covers the specific services, endpoints, and regions used.
  3. Features outside the scope of the BAA, such as live web search, are turned off in the code, not just by policy.
  4. Whenever the task permits it, direct identifiers are tokenized before being included in the prompts.
  5. Before the model looks at them, the results are filtered according to user permissions.
  6. Audit records include details such as the requester's name, the records retrieved, the model version, the output, and the action, with the complete PHI stored separately.
  7. Testing for prompt injection includes not just chat input but also the documents and messages that the system takes in.
  8. When clinical escalation procedures lead to a human, the risk-scoring inputs are recorded for review under Section 1557.
  9. Your current Security Rule risk analysis includes the AI components.

Frequently Asked Questions

A BAA does not include consumer ChatGPT and must never have access to PHI. OpenAI provides access to PHI via its API on eligible endpoints under a BAA, as well as through sales-managed options such as ChatGPT for Healthcare. However, ChatGPT Business is not eligible.
Not at all. A BAA covers only the vendor's responsibilities for the specific services it lists. Your organization remains liable for risk analysis, access control, the minimum necessary design, audit logging, and configuration. Twilio's documentation makes this point plainly clear: a BAA by itself is not enough to make an application compliant.
So, if either one stores or handles PHI, and embedding services receive raw text chunks. In contrast, most vector databases store the main text alongside the embeddings, so both companies would be considered business associates in such cases.
Yes, provided that it is correctly de-identified using either the Safe Harbor method or expert determination. The Department of Health and Human Services says that a provider with only de-identified data is not considered a business associate. Since free-text notes and transcripts generally do not satisfy the Safe Harbor criteria by simple redaction, they should be carefully verified.
On the contrary, the overhaul suggested in January 2025 is still only a proposed rule, and the final decision is now expected in July 2027. The current Security Rule remains fully enforced, and the Office of Civil Rights's current priority is addressing risk analysis failures.
HIPAA permits it when a BAA is in place, and other rules are met. OCR notes that offshore processing can carry different risks, which your risk analysis must address. State laws and client contracts may add tighter data residency requirements.

Conclusion

HIPAA-compliant generative AI is not a product you buy. It is a set of design decisions you make and document. The model call is one hop among many, and the gaps usually sit in the parts nobody drew on the whiteboard: the embedding service, the prompt logs, the alert email, and the review queue.

Start with the PHI map. Put a scoped BAA behind every component that sees PHI, and minimize what the model receives. Log in a way that proves what transpired without creating a second copy of the chart.

Teams that do this work before launch get something beyond compliance: a system they can explain to an auditor, a clinician, and a patient. If you are planning a healthcare AI build, Cypherox's healthcare software team and generative AI development practice can help you map it before you write the first prompt. Talk to our engineers about your architecture.

Vipinraj Nair

About the Author

Vipinraj Nair LinkedIn

Founder & CEO

Vipinraj Nair is the Founder and CEO of Cypherox Technologies, which he started in 2015. He leads the company's work across custom software, web and mobile development, and AI solutions for startups, SMEs, and enterprises worldwide. He writes on technology trends, custom development, and how businesses put emerging tech to practical use.