What the research found
The international I3LUNG study built a multimodal, explainable AI model from 2,396 patients with advanced non-small cell lung cancer and tested whether it helped 20 physicians predict how patients would respond to immunotherapy. With AI support, accuracy rose from 0.57 to 0.65 and sensitivity from 0.72 to 0.87, and physicians took up correct AI suggestions 74.5% of the time. However, lung-cancer experts followed the AI's incorrect suggestions 72.2% of the time (non-experts 63.6%), and performance fell in external validation, where AUCs ranged from 0.55 to 0.72.
A tool that improves the average can still make a confident expert more likely to adopt a wrong answer, so keeping a human in the room only protects patients if that human is still checking the AI.