OpenEvidence has introduced a family of four medical AI models, including a research model that the company says is the first AI system to achieve a perfect score on MedQA, a widely used medical AI benchmark.
The new model, Darwin, scored 100% on MedQA and also recorded results of 72.8% on MedXpertQA, 82.7% on HealthBench Professional and 87.2% on NOHARM. OpenEvidence reported that these results placed Darwin ahead of other models evaluated on the benchmarks, including Claude Fable 5 and Gemini 3.7.
Announced on September 3, the model family is named after figures associated with the history of medicine. Three models — Osler, Sackett and Snow — are production tools available free of charge to verified clinicians, while Darwin remains in research preview.
Osler has become OpenEvidence's default model, replacing the system that previously powered the platform. Designed for speed, it produces answers in approximately five seconds. Sackett is intended for clinical questions requiring closer assessment of the strength and weight of available evidence, with responses taking around 30 seconds.
Snow succeeds the company's Deep Consult feature. It conducts a broader review of medical literature over approximately five minutes before generating a report, supporting questions requiring more extensive evidence synthesis.
Darwin is available only through an application process. OpenEvidence attributed the restricted access to potential dual-use risks associated with advanced capabilities in areas including virology, immunology and genetics. Current users include institutional partners such as the National Organization for Rare Disorders and accredited academic researchers.
The launch comes as OpenEvidence continues expanding its role in clinician-focused AI. The company was valued at $12 billion as of January 2026 and has positioned its platform around providing clinicians with evidence-based medical information and decision support.
OpenEvidence said it intends to progressively incorporate capabilities developed through Darwin into its production models. That process will depend on safeguards being tested and validated alongside institutional and research partners, allowing capabilities developed in the research environment to move into clinician-facing applications as safety requirements are addressed.
Click here for the original news story.