What the research found
Researchers compared a large language model against a consensus of human expert reviewers at identifying adverse events in encounter notes from four randomised immunotherapy trials. Of the 200 adverse events reviewer consensus identified in those notes, only 47 (24%) appeared in the trials' formal reporting, while the LLM independently captured 137 (69%), scoring a mean encounter F1 of 0.76 (95% CI, 0.70–0.82) against reviewer consensus — comparable to individual human reviewers. The comparison cut both ways: trial reporting captured clinically serious events the reviewers missed, including atrial fibrillation and cytokine release syndrome, while several of the extra events reviewers surfaced were minor, such as fatigue and urinary frequency.
The finding isn't that the model is clever — it's that the safety record everyone already trusted was incomplete before any AI touched it, which makes an LLM reading the same notes a credible second pair of eyes rather than a replacement for the first.