Healthcare AI Safety Gap · Don't break patients
    Peer-reviewed / official12 Jun 2026·US

    Purpose-built clinical AI tools lost to general chatbots in head-to-head test

    Nature Medicine (NYU Langone), published 12 Jun but only breaking into mainstream health-tech debate via a 29 Jul STAT deep-dive: OpenEvidence and UpToDate Expert AI underperformed 3 frontier LLMs on all 3 benchmarks tested.

    3 of 3
    evaluations where general frontier models beat dedicated clinical AI tools, incl. 1,800 clinician-reviewed real-world queries

    What the research found

    NYU Langone researchers benchmarked two clinical AI products used by hundreds of thousands of US physicians, OpenEvidence and UpToDate Expert AI, against three general-purpose frontier models (GPT-5.2, Gemini 3.1 Pro, Claude Opus 4.6) on medical-knowledge questions, clinician-alignment items, and 100 real de-identified physician queries blindly rated by 12 US clinicians (1,800 annotations). The general-purpose models won on every measure; the clinical-specific tools scored only about as well as Google's AI Overview search feature. Published mid-June, the paper only became a flashpoint in health-AI circles after a STAT deep dive on 29 July, which described the online reaction as unlike anything the author had seen a single paper trigger.

    Why it matters for providers

    Hospitals and clinicians have been choosing 'clinical' AI products partly on the assumption that specialisation means extra safety — this independent benchmark says that assumption needs testing, not taking on faith, before a tool reaches a patient chart.

    Original source
    Nature Medicine (NYU Langone Health / NYU Grossman School of Medicine)
    Read the full report ↗
    Peer-reviewed or official source — the most reliable tier.