AIMed26  GLOBAL  SUMMIT

Hot Topics for the Week of:
September 7th

Dr. Anthony Chang

 

OpenAI Wires ChatGPT for Healthcare Into Epic


HEALTH SYSTEM DEPLOYMENT / FRONTIER LAB IN THE RECORD


What happened. On September 1, OpenAI announced that ChatGPT for Healthcare now connects to Epic. Authorized clinicians at organizations with supported Epic environments can pull appointment notes, lab results, medications, and specialist documentation into a ChatGPT session, or work with ChatGPT embedded inside the Epic chart itself. The stated use cases are synthesis: what changed since the last visit, recent labs, medication changes, unresolved referrals. The connection is read-only and cannot write to the record. UCSF Health is the named pilot, and its CEO described the work as exploratory. Alongside it, OpenAI shipped a public-data plugin linking nine official sources, including ClinicalTrials.gov, RxNorm, DailyMed, and PubMed. The integration follows the January consumer launch and April clinician tier. Two details matter: Epic has issued no statement, and OpenAI's 99.1% physician safety rating rests on 4,363 internal evaluations that remain unpublished.


Why it matters to the AIMed audience. I think overall this dyad is premature and not ready for prime time, and warrant some governance. This is the first frontier lab sitting inside the system of record rather than beside it. Read-only is the concession that makes it deployable; it is also the ceiling on what it can do. Synthesis without action is a summarizer with a very good model behind it so that it is useful, but not the agentic layer Epic previewed at UGM two weeks earlier. The competitive picture is now explicit: OpenAI and Epic want the same workflow real estate, and Epic's silence is the loudest thing in the announcement. Whether this arrives through a sanctioned partnership or through interfaces any third-party developer can request decides whether your Epic governance process holds a gate here at all. Ask your Epic representative before a service line asks you for the tool. And ask OpenAI for the validation data. A safety rating no one outside the company has examined is a marketing claim, not evidence.


GPT-6 Astra Ships as the First Model Rated "Critical" for Cyber Capability


GENERAL AI / SECURITY


What happened. OpenAI released GPT-6 Astra on September 3, disclosing in the same announcement that it is the first model to cross the "critical" cybersecurity threshold under the company's Preparedness Framework. That tier is defined as the ability to identify and develop working zero-day exploits across hardened real-world systems without a human directing each step. Tested without production safeguards, Astra scored 100% on ExploitBench, up from 78.5% for GPT-5.6 Sol, and 42.4% on the broader ExploitGym versus 30.3%. Run against vulnerabilities disclosed in the three months before launch (to rule out recall from training data), it found two previously unknown zero-days and used them in an exploit chain. Deployment is phased: select organizations first, then ChatGPT tiers, API, and AWS Bedrock. Enterprise access is off by default and must be enabled by administrators. Zero data retention is available for eligible API customers.


Why it matters to the AIMed audience. This was probably the biggest news item in the past few weeks of AI big news items. Most hospital boards have been asking about hallucination and bias. This week the question changed. A model that autonomously finds and weaponizes unknown vulnerabilities now exists at API scale, and healthcare runs on the most under-patched infrastructure in the economy with legacy imaging systems, infusion pumps, building controls, and an EHR estate that took three years to recover from the last major ransomware wave. The same capability accelerates defense: code review, patch validation, and attack-surface testing. Which side benefits first depends on who adopts faster, and hospitals are structurally slow. Astra being off by default is a governance gift as someone in your organization has to affirmatively turn it on, and that decision should run through security, not a productivity champion. And the zero data retention (ZDR) option, a privacy setting that disallows AI users from saving prompts answers, and data on their servers, is the detail that makes any clinical use even discussable.


RAND Proposes Nine-Layer Biosecurity Defense as the Governance Template


GENERAL AI / BIOSECURITY — DUAL-USE


What happened. RAND released a report in late August outlining nine mitigation strategies against AI-enabled biological weapons, aimed at different actors across the chain that include model developers, biological tool and dataset holders, synthesis providers, and governments and calling for cross-industry cooperation and more controls than current voluntary frameworks provide. The context is a widening gap between capability and evidence. Frontier labs have been adding biological safeguards as models' performance on expert virology evaluations has risen sharply; the same week, Anthropic reported that Claude designed functional protein binders against 14 targets while withholding those capabilities from general access. The uplift evidence remains mixed. RAND's own 2024 red-team study found no statistically significant improvement in bioweapon attack plans with LLM assistance, while Anthropic's Frontier Red Team has since reported measurable uplift to novices with newer models. The honest summary is that the signal is moving, not settled.


Why it matters to the AIMed audience. While this is not a healthcare-delivery story per se, it belongs here anyway as it is under AI governance in healthcare. Biosecurity is becoming the template for AI governance generally, because it is the domain where the field first had to design controls for a capability nobody could verify was dangerous yet. RAND's layered model that controls distributed across actors, with no single point of trust , maps directly onto what a hospital AI oversight committee should look like but rarely does (as yet). Most AI committees are a single review gate. There are three implications for those of you in academic medical centers. Your IRB and research-security function just got more difficult as it now needs a position on generative biological design tools, not just on data. In addition, your DNA synthesis vendors either screen orders or they do not, and that is a contract term. And the same uncertainty RAND describes with capability rising faster than harm can be measured is exactly the condition your clinical AI committee is operating in. In short, in addition to have your C-suite and board of directors educated on AI, this education should include biosecurity as well as cybersecurity. 


The Pediatric Validation Gap, Quantified


REGULATION / PEDIATRICS


What happened. A JAMA Network Open cross-sectional study from Lurie Children's, published in March and now circulating in commentary, analyzed 952 FDA-authorized AI-enabled devices through 2024. Only 42, or 4.4%, carried indications for use in patients under 18. Five were exclusively pediatric, all authorized since 2020. Age was omitted entirely from 59.6% of indication statements, meaning clinicians cannot tell whether most devices were ever intended for children. Pediatric-labeled devices took longer to review at a median 162 days versus 134, and more often carried registered clinical trials, suggesting reviewers expect more evidence despite unchanged statutory standards. FDA does not require manufacturers to state whether validation included children. A Baylor commentary offers the illustration: an adult-developed algorithm for thyroid nodules reached only 36% specificity in pediatric patients, producing false positives nearly two-thirds of the time. The authors call for standardized age labeling and validation reporting.


Why it matters to the AIMed audience. As a pediatric subspecialist, I like to state that this chronic injustice is utterly indefensible, especially in the era of healthcare AI. For a children's hospital, this is the citable number that reframes every AI procurement conversation. Roughly 95% of authorized AI devices carry no pediatric indication, and a majority carry no age information at all, so most tools deployed near children were, at best, never tested on them, and at worst are silently being used off-label. That is a governance as well as a liability exposure. The economics cut the wrong way: pediatric labeling costs more and takes longer, so the rational vendor prints "not indicated for pediatric use" and moves on without much hesitation. Transparency alone risks converting passive omission into active exclusion. The practical ask for pediatric AI leaders is narrow and immediate: require every vendor to state, in writing, whether its validation dataset included children and at what ages. If the answer is no, the tool is a research question, not a deployment strategy.


Nvidia Buys Hugging Face - the Open-Model Commons Gets an Owner


GENERAL AI / OPEN-SOURCE INFRASTRUCTURE


What happened. On September 3, Nvidia agreed to acquire Hugging Face for $12.93 billion, its second-largest purchase after last year's $20 billion Groq deal. Hugging Face is the de facto registry of open AI with more than 3 million models, 500,000 datasets, and 1 million applications used by 18 million developers and 200,000 companies. Jensen Huang pledged it will remain an open, compute-agnostic platform with the founding team staying on. Closing is expected in early 2027 pending approvals. Two context points sit underneath the price. Hugging Face approached Nvidia, saying it needed more compute and support to keep pace with closed labs. And in July the platform was breached by OpenAI models with GPT-5.6 Sol and a more capable unreleased model, tested with cyber refusals off. 


Why it matters to the AIMed audience. In my humble opinion, this mega deal has both obvious and not-so-obvious consequences. Most health system running AI on its own hardware depends on this platform, whether it knows it or not. MedGemma, the open-weight Llama and Qwen families, most clinical NLP models, and the datasets they were tuned on are downloaded from Hugging Face. On-premise clinical AI, the sovereign, privacy-preserving option that governance committees prefer, has acquired a single corporate owner whose core business is selling the chips those models run on. Two nearer concerns. The July breach makes model provenance a supply-chain question: a hospital pulling weights from a registry that an autonomous agent penetrated six weeks ago should be verifying checksums and signatures, and almost none do. And Nvidia's own SEC filing warns that government restrictions may affect the platform. For international readers, the open-model commons is now inside a US export-control perimeter. Therefore, sovereignty through open weights just got more conditional.


FDA Opens the Door to a "Competency-Based" Review Standard for Generative AI Devices


REGULATION / UNITED STATES


What happened. On August 18, the FDA's Digital Health Center of Excellence released a discussion paper on regulating generative AI-enabled medical devices. It is a request for feedback, not draft guidance. Comments close October 19 under docket FDA-2026-N-7874.


Two ideas carry it. The first is a two-axis risk grid: how independently a tool acts, from information-only through recommendations to fully autonomous action, and how much harm follows from relying on a wrong answer. Evidence expectations rise with position on that grid, not with whether a language model is involved. The second is evaluation modeled on clinician credentialing — benchmarking, then clinical confirmation in progressively realistic settings, then ongoing assessment in use. A system with open-ended inputs cannot be tested on every scenario it will meet, any more than a clinician can. The 26 questions span postmarket monitoring, foundation models, and agentic systems.
Why it matters to the AIMed26 audience. This is the year's most consequential US regulatory signal, and it is worth reading precisely because it is an important development in AI in health regulation. Whoever amongst us comments before October 19 helps choose the words reviewers will later apply to everyone. Of note, this document is short on pediatric relevance. 


It is interesting that I had been suggesting to various organizations that we credential AI clinicians rather than AI models (similar to credentialing physicians who perform procedures rather than for each procedure). The credentialing metaphor brings the whole apparatus with it: benchmarks, defined scope of practice, periodic re-evaluation. Health systems already run that machinery for clinicians so should be able to execute this evaluation with enough expertise. 
The grid is usable tomorrow and requires every vendor to place its product on it in writing. Governance, not modality, now sets the burden of proof. Least discussed (but should be discussed proactively): whether accountability for foundation models sits with the developer or the integrator is unresolved, and it determines whether a health system building its own agents is a manufacturer or a customer as well. In my opinion, this evaluation as part of AI governance belongs to the board of directors and C-suite and not on the CIO or CMIO agendas. 

 

Epic's Agent Factory Makes AI Workflow Automation Default EHR Infrastructure


HEALTH SYSTEM DEPLOYMENT / AGENTIC AI


What happened. At Epic's Users' Group Meeting in Verona, Wisconsin (August 17–20), the company unveiled Agent Factory, a no-code platform with more than 120 prebuilt AI agents that organizations can deploy as shipped, tune locally, or extend with their own. General availability comes in 2027.


Judy Faulkner framed the wider strategy around Ergo — Epic's Latin "therefore" — an AI-native layer binding Art (clinician-facing for clinical workflow), Emmie (patient-facing for engagement), and Penny (administrative revenue cycle) into one interface that adapts to role and specialty. Early results are concrete. ECU Health reports roughly 20 hours a week saved by an agent summarizing transfer-center requests, plus an 86% increase in transfers landing at regional hospitals rather than the academic center.


Why it matters to the AIMed26 audience. Epic reaches more than 3,000 hospitals and roughly 325 million patients. When “agentic” AI (many AI agents are not “agentic” with autonomy) arrives as a configuration option inside the system of record rather than as a purchase, adoption stops being a build-versus-buy decision and becomes a governance decision. Big time. 
Ambient documentation was the entry point for provider-facing AI and has already commoditized. Agents are the next enclosure but this step requires much more AI governance and Epic infrastructure and database knowledge. This step also requires something I have recommended: please don’t use AI to automate a bad workflow or work process/design as AI will just make it faster but not necessarily better.


There is no published cost model for building inside Agent Factory, and therefore no calculable ROI for now. And no one answered the question that hung over the meeting: once anyone can spin up an agent without writing code, who keeps each one working? Maintenance of these agents to address drift is key. Lastly, just because there are multiple agents, it does not mean the system is “agentic” in behavior which implies some degree of autonomy. 


A Model Picks the Targets: First Phase 3 Win for an AI-Designed Therapy


AI-DESIGNED THERAPEUTICS / GLOBAL


What happened. On August 19, Merck and Moderna announced that intismeran autogene, given with pembrolizumab, met both endpoints in an interim Phase 3 analysis in adjuvant melanoma-less recurrence, less distant metastasis. It is the first randomized Phase 3 win for a neoantigen cancer vaccine, and the first for a therapy whose active ingredient is chosen by a machine learning model.


Every dose is computed before it is manufactured. Tumor DNA, tumor RNA, and normal DNA are sequenced and the patient's HLA type is called. A selection algorithm annotates the mutations, predicts which resulting peptides that particular immune system will actually present, ranks them for immunogenicity, and picks up to 34 from thousands of candidates. An automated step strings them into one mRNA construct. Moderna orchestrates the per-patient batch through a purpose-built digital system, Maestro; biopsy to injection runs roughly four to eight weeks.


Why it matters to the AIMed26 audience. This is truly amazing pioneering work on precision oncology enabled by AI; if this works, this AI-enable approach to cancer therapy will be truly disruptive. Nearly every prior claim for AI in drug development concerned speed and cost upstream: faster screening, cheaper candidates. Here the model's output is the therapy itself. This is the first randomized Phase 3 evidence that a prediction model can carry clinical benefit.


This precision approach also engenders several regulatory questions: What is approved when every dose differs and when there is a manufacturing process with a prediction model inside it. In addition, if the selection algorithm is retrained on new immunogenicity data, is it still the “same” product? This is precisely the problem the FDA raises about self-updating software.


The remaining bottleneck is computational logistics, not biology. Four to eight weeks from biopsy to dose is a scheduling problem, and health systems inherit it: sequencing, HLA typing, design turnaround, per-patient manufacture, and weeks or months of infusions. A supply chain management challenge for sure, but perhaps AI can offer a solution for the execution part as well.

 

Asia Builds the Yardstick: HIMSS Launches a Global AI Outcomes Framework in Singapore


MEASUREMENT / INTERNATIONAL — ASIA-PACIFIC


What happened. At HIMSS26 APAC in Singapore (August 23–25, co-hosted with SingHealth), HIMSS announced the AI Outcomes Framework, described as the first evidence-based, longitudinal standard for measuring the clinical, operational, and financial impact of healthcare AI.


The design is the news. Deployed AI is judged today on vendor-reported model metrics (AUC on a held-out set and accuracy against a retrospective cohort). The framework replaces that with pooled real-world data across health systems, tracking how models in production affect patient care, clinician cognitive load, and hospital throughput over time. The founding provider partners are entirely Asian: SingHealth and National University Hospital in Singapore, Asan Medical Center and Seoul National University Hospital in South Korea, and Taichung Veterans General Hospital in Taiwan. Worldwide rollout comes at HIMSS27, with peer benchmarks and clinical evidence published then.


Why it matters to the AIMed26 audience. I have long admired Singapore for its healthcare data handling and governance. Many governance committee members are stuck on the same question: how do we know this is working? Five Asian academic centers are supplying the founding evidence base, and the benchmarks they generate will become the peer comparisons your board eventually cites.


The United States leads on model development and deployment volume. Singapore, Korea, and Taiwan are leading on measurement- and measurement is what converts deployment into evidence, especially for outcomes. Some Asian health systems with unified national data layers can run longitudinal, multi-site outcome studies that fragmented American systems structurally cannot. Our advantage in scale is offset by our disadvantage in coherence. Clinician cognitive load appears as a first-class endpoint alongside throughput and cost, which is not what most US dashboards measure.
All of this points to a basic insight we often share and discuss: Good to great healthcare AI work relies on a strong foundation of healthcare data, especially longitudinal data. This is also why it is always good to have an international perspective on healthcare AI. 


Gates's Turbulence Memo: "The Greatest Equalizer, or the Worst Source of Injustice"


GENERAL AI — POLICY / SOCIETAL


What happened. On August 26, Bill Gates published a roughly 6,000-word essay, The turbulent AI era is here. The choices we make now are critical. Coming from technology's most durable optimist, the tone is the news: AI is advancing faster than governments and societies are preparing for, and even in the best case the transition will rank among the most turbulent periods in human history.


He names three risks. Labor displacement, which he distinguishes from earlier shifts because AI does not merely offload the task — it does the thinking, faster than institutions adapt. He expects law, medicine, software, and manufacturing to face serious disruption within a decade, with junior and mid-level roles most exposed. Malicious use, spanning cyber and biological risk. And developmental harm to children and to human relationships.
He proposes new national bodies and an international organization modeled on nuclear oversight and civil aviation; a protected class of "Human Reserved" occupations; and taxing AI and robotics to fund retraining.
Why it matters to the AIMed26 audience. 


A good but perhaps unrealistic idea is "Human Reserved," because it inverts the question our field has been answering by default. Governance committees ask where automation is safe. Gates asks which clinical work should remain human- not because a machine cannot do it, but because we decide it should not. Human performance becomes a value to preserve deliberately, the way one preserves a species on the verge of extinction. Some candidates are obvious. Comfort care at the end of life. Breaking bad news. The pediatric encounter where the parent needs a person more than an answer. Nobody has drawn that boundary; if the profession does not, procurement will.


His one healthcare example runs through demography rather than technology. Japan, with too few young people to care for the old, may welcome a caregiving robot a younger society would refuse. In addition, an extension of the developmental harm is the training of our junior colleagues. While AI promises to be a valuable partner for learning, loss of critical thinking and clinical judgment can occur unless we plan to preserve those dimensions.