Chatbots Gave the Right Sleep Apnea Advice Until the Patient Pushed Back

Five free AI chatbots told a cooperative simulated sleep apnea patient to see a specialist in all 350 test conversations, then abandoned that advice in more than a third of conversations when the same patient played down symptoms and resisted a referral, researchers reported Sept. 6 at the European Respiratory Society (ERS) Congress in Barcelona.
Sleep apnea, in which breathing repeatedly stops during the night, is diagnosed only after someone is referred for a sleep study. Deeban Ratneswaran, a research fellow at Guy's and St Thomas' NHS Foundation Trust in London who led the work, said that 80% to 90% of moderate-to-severe cases go undiagnosed, and that free chatbots have become "a first port of call for health questions, often before any clinician is involved."
The team wrote seven sleep apnea cases, all meeting the criteria for referral, and ran 700 conversations with ChatGPT, Google Gemini, Claude, DeepSeek and Grok. Each ran twice with identical medical facts: once cooperative, once resistant. Correct advice survived all 350 cooperative conversations and 64% of the resistant ones, according to the society's release. The work is a congress abstract, not a peer-reviewed paper, and the patients were simulated, not real.
The advice held up worst in the most serious cases. In a textbook severe case, the advice survived 22% of the time, and with a man who had already dozed off at the wheel, it survived 32%. When the chatbots dropped it, they usually left the driving risk unmentioned, Ratneswaran said. Depending on the model, the chatbots offered lifestyle tips instead of a referral in roughly a quarter to a half of those conversations. The release gave no results for individual models.
Io Hui of the University of Edinburgh, who chairs the ERS group on m-health and e-health and was not involved, said the issue is behavior, not knowledge: "The problem is not what the chatbots know, it is how they handle disagreement; they appear to exhibit a tendency to please the user, a phenomenon known as 'AI sycophancy.'"
