Inside OpenAI’s public-first bet on medical AI
This article summarizes an episode of Cognitive Revolution’s video series featuring Karan Singhal, head of Health AI at OpenAI.

Photo credit: Shutterstock
OpenAI is positioning its AI models to assist with medical diagnostics and research. Karan Singhal, the company’s head of health AI, argues that the fastest path to adoption is not training on private patient records, but creating a publicly accessible AI tool first.
Singhal believes that avoiding private data initially allows the company to build a reliable public utility. By refining the system through mass usage, the team can establish a safety baseline before integrating the system into complex clinical environments.
Executing the initial roadmap
Releasing a widely accessible model also creates liability risks. To mitigate this, OpenAI invested heavily in safety research and protocol design before exposing the system to general users.
Singhal describes the strategy as operating in three phases, starting with foundational safety research. He notes that they are “now in this phase of adoption, and over 230 million people a week are using [the platform] for various health and wellness-related queries.”
Building for compliance
While the public platform serves millions, general-purpose chatbots cannot meet the strict regulatory standards required in healthcare institutions. To solve this problem and move beyond general wellness queries, OpenAI developed a separate product designed specifically to function within clinical environments.
Singhal notes this tool is “purpose-built for the workflows of health professionals,” ensuring it meets HIPAA (Health Insurance Portability and Accountability Act) compliance where standard models do not. He breaks this foundation down into two phases:
- Specialization adapts the technology to fit strict hospital rules.
- Integration helps medical staff adopt the system into their daily routines.
Evaluating baseline reliability
Even a medical AI that provides correct answers most of the time is still dangerous if it fails unpredictably. Developers must actively search for the model’s worst potential response to ensure stability under pressure.
Singhal points to a metric called “worst of N,” where the team measures the poorest performance across multiple samples to gauge consistency.
Additionally, the team uses a framework called HealthBench to measure “about 49,000 different axes on which model performance could differ,” including tone and factual accuracy.
Bypassing traditional gatekeepers
Once these models are proven reliable, they can help resolve the common struggle patients face when trying to describe symptoms to busy doctors. This communication gap often creates an imbalance where valid concerns are unfortunately ignored or misdiagnosed.
Singhal explains that AI allows patients to “integrate knowledge and information across both their previous history and the latest medical evidence.” This data access provides a level of ownership that enables patients to advocate for better care.
Forcing open clinical trials
Stay updated on the go with our mobile app.
Get latest insights with smoother, more personalized experience through TIA mobile app.






