What the research found
NYU Langone researchers benchmarked two clinical AI products used by hundreds of thousands of US physicians, OpenEvidence and UpToDate Expert AI, against three general-purpose frontier models (GPT-5.2, Gemini 3.1 Pro, Claude Opus 4.6) on medical-knowledge questions, clinician-alignment items, and 100 real de-identified physician queries blindly rated by 12 US clinicians (1,800 annotations). The general-purpose models won on every measure; the clinical-specific tools scored only about as well as Google's AI Overview search feature. Published mid-June, the paper only became a flashpoint in health-AI circles after a STAT deep dive on 29 July, which described the online reaction as unlike anything the author had seen a single paper trigger.
Hospitals and clinicians have been choosing 'clinical' AI products partly on the assumption that specialisation means extra safety — this independent benchmark says that assumption needs testing, not taking on faith, before a tool reaches a patient chart.