Skip to main content
AI Development

OpenAI connected ChatGPT to Epic. Test two retrieval paths, not one demo

OpenAI's Epic integration and Healthcare Public Data plugin bring private chart context and public medical sources into one workspace. Validate those retrieval paths separately before combining them.

Sean McLellan profile photo

Sean McLellan

Lead Architect & Founder

6 min read
A sealed amber specimen archive and an open blue reference vessel feed separate glass tubes toward an empty inspection tray beside laboratory equipment.
Constructed diagramA BaristaLabs conceptual still life of separate private-record and public-reference retrieval paths. It does not depict patient data, OpenAI or Epic software, a diagnosis, or a validated result.

OpenAI connected ChatGPT for Healthcare to Epic patient records and introduced a Healthcare Public Data plugin on September 1, 2026. The first path can retrieve authorized chart context. The second can query structured public sources such as PubMed, DailyMed, and CMS Coverage.

Putting both paths in one governed workspace can reduce switching between systems. It also makes a weak evaluation tempting: ask one impressive question, receive a fluent answer with references, and call the integration ready. A safer test separates private-record retrieval from public-evidence retrieval, proves each path on its own, and only then asks a question that needs both.

What did OpenAI release?

The Epic integration can bring authorized patient information into ChatGPT for Healthcare. OpenAI lists appointment notes, laboratory results, medications, and specialist documentation among the record types a clinician might review. It says responses can summarize changes and point back to supporting chart information.

OpenAI describes two deployment experiences. A user can bring EHR context into ChatGPT, or a supported deployment can place ChatGPT inside an EHR layout. The announcement does not say that every Epic environment, workflow, or customer receives both experiences immediately.

The separate Healthcare Public Data plugin offers structured access to official sources including PubMed, DailyMed, and CMS Coverage. That is evidence outside the patient chart: research literature, drug information, and coverage material. It should not silently inherit the chart path's permissions or its evaluation criteria.

Availability is also bounded. OpenAI tells ChatGPT for Healthcare customers to ask a workspace administrator to enable the integration and plugin. ChatGPT Enterprise customers must confirm eligibility and the correct Regulated Workspace configuration with OpenAI. The EHR integration is not available to individual users.

Why are these two different retrieval systems?

A patient-record question asks whether the system found the correct authorized facts for the correct person and encounter. A public-evidence question asks whether it found the right external source, represented it accurately, and preserved publication or coverage context. Those paths can meet in one answer, but they do not fail in the same way.

A chart retrieval failure might omit a recent medication change, surface stale information, or cross an access boundary. A public-source failure might cite an irrelevant paper, confuse a drug formulation, or present a coverage document without the jurisdiction and effective date needed to interpret it. A polished synthesis can hide either miss.

OpenAI says ChatGPT for Healthcare includes role-based access control, single sign-on, and audit logs. Those are important workspace controls. They do not establish that a particular user should see a particular record, that every source needed for a question was retrieved, or that the answer is suitable for a clinical decision.

HHS guidance supplies the contractual baseline for protected health information in cloud services. A provider that creates, receives, maintains, or transmits electronic protected health information for a covered entity is generally a business associate, and the parties need a HIPAA-compliant business associate agreement. OpenAI likewise conditions HIPAA-compliant use on an applicable BAA. A BAA defines obligations; it does not test retrieval accuracy or assign local review work.

How should a healthcare team test the two paths?

Start with three small case sets built by the people who own the underlying information.

Private-record cases should ask for facts already known from the authorized chart: a change since the prior visit, a recent lab result, a medication update, or an unresolved referral. For each case, record the expected fact, the exact chart evidence, which role may retrieve it, and which nearby record must remain unavailable. Score missing facts, wrong-patient or wrong-encounter facts, stale facts, and unsupported claims separately.

Public-source cases should ask answerable questions from a named source class. Record the expected PubMed article, DailyMed label, or CMS Coverage document, including the identifier, date, and section that supports the answer. Score source selection and faithful representation before judging prose quality.

Joined cases come last. Each should require one authorized chart fact and one public-source fact. The reviewer must be able to trace both halves independently. If the answer gives a citation for the public claim but no usable pointer to the chart evidence, the joined case has not passed.

This sequence is BaristaLabs guidance, not a published OpenAI test protocol. Its purpose is to keep one good retrieval path from masking a weak one.

Two separate glass channels terminate at three empty ceramic test vessels beside a blank brass review lever.
Constructed diagramTest private-record retrieval and public-source retrieval independently before asking a joined question. The illustration shows no actual result or clinical decision.

What should stay outside the first pilot?

Do not begin with treatment recommendations, autonomous chart changes, patient messages, orders, or coverage determinations. The announcement focuses on finding, understanding, and using information; it does not publish evidence that the integration can make those consequential decisions safely or autonomously.

Keep public-source queries free of patient details unless the approved design explicitly requires and permits them. A public evidence tool is still an outbound data surface. Our earlier analysis of research-agent query spill explains why a search trail can expose private context even when no full document leaves the workspace.

Set a stop condition before real records enter the test. Pause when a result points to the wrong patient or encounter, cannot return supporting chart evidence, sends unnecessary patient detail to an external source, or produces a public citation the reviewer cannot open and match to the claim. Route disputed results to the existing clinical, privacy, security, or compliance owner rather than asking the model to resolve its own failure.

What evidence supports a rollout decision?

Measure retrieval before time saved. For each path, track the share of expected facts found, unsupported facts added, access-boundary violations, evidence links a reviewer can follow, and cases that need escalation. For joined cases, require both a valid chart pointer and a valid public citation.

Then measure operations: review time, correction time, and the number of system switches avoided. A pilot that shortens chart preparation but adds a long evidence audit may still need narrower questions or better source configuration. The public announcement provides no independent error rate or clinical-outcome result, so each organization needs its own representative cases and acceptance threshold.

Treat deployment configuration as part of the evaluated system. Record the workspace, user role, Epic environment, enabled plugin, retention settings, and BAA status for every run. A result from a sandbox with broad test access does not prove the same question will work correctly under production roles.

OpenAI's launch makes a useful connection possible: authorized patient context and official healthcare sources can be reviewed in one workspace. The responsible next step is not a broad rollout or a single showcase prompt. Prove the private path, prove the public path, and make the joined answer carry evidence from both. If your team needs help designing those cases and boundaries, scope a healthcare AI evaluation with BaristaLabs.

Sources

Healthcare AI evaluation

Test each retrieval boundary before joining them

BaristaLabs can help define patient-record permissions, public-source access, representative cases, citation checks, reviewer ownership, and stop conditions for one bounded pilot.

Best fit for a healthcare organization evaluating governed AI access to patient records and official public sources.

Turn this idea into a pilot

Which workflow should go first?

Use the readiness check to compare impact, effort, risk, owner, and next step before requesting a review.

  • 3-5 minutes
  • Deterministic score
  • No sensitive data
Check workflow readiness

Practical AI Workflow Notes

Want more practical AI operations ideas?

Get short notes on applying AI inside real small-business workflows — from document handling and customer follow-up to internal reporting, compliance, and automation guardrails.

A useful next step if you’re still exploring and not ready to request a 20-minute workflow assessment.

Occasional emails. Practical workflow guidance only. Unsubscribe anytime.