Healthcare AI Safety Gap · Don't break patients
    Peer-reviewed / official1 Sep 2026·US

    Cancer trials formally reported a quarter of the adverse events sitting in their own notes

    JAMA Network Open: across four randomised immunotherapy trials, expert reviewers found 200 adverse events in the encounter notes — only 47 had been formally reported. A large language model reading the same notes caught 137 of them.

    24%
    of the adverse events reviewers found in trial encounter notes had actually been reported by the trials

    What the research found

    Researchers compared a large language model against a consensus of human expert reviewers at identifying adverse events in encounter notes from four randomised immunotherapy trials. Of the 200 adverse events reviewer consensus identified in those notes, only 47 (24%) appeared in the trials' formal reporting, while the LLM independently captured 137 (69%), scoring a mean encounter F1 of 0.76 (95% CI, 0.70–0.82) against reviewer consensus — comparable to individual human reviewers. The comparison cut both ways: trial reporting captured clinically serious events the reviewers missed, including atrial fibrillation and cytokine release syndrome, while several of the extra events reviewers surfaced were minor, such as fatigue and urinary frequency.

    Why it matters for providers

    The finding isn't that the model is clever — it's that the safety record everyone already trusted was incomplete before any AI touched it, which makes an LLM reading the same notes a credible second pair of eyes rather than a replacement for the first.

    Original source
    JAMA Network Open (University of California, San Francisco & Fred Hutchinson Cancer Center)
    Read the full report ↗
    Peer-reviewed or official source — the most reliable tier.