







ChatGPT Health was launched in January 2026 as OpenAI’s consumer health tool and has reached millions of users. Here we conducted a structured stress test of triage recommendations using 60 clinician-authored vignettes across 21 clinical domains under 16 factorial conditions, yielding 960 total responses. Performance followed an inverted U-shaped pattern, with the most dangerous failures concentrated at clinical extremes—nonurgent presentations (35%) and emergency conditions (48%). Among gold-standard emergencies, the system undertriaged 52% of cases, directing patients with diabetic ketoacidosis or impending respiratory failure to 24–48 h evaluation rather than the emergency department, while correctly triaging classical emergencies such as stroke and anaphylaxis. When family or friends minimized symptoms, indicating anchoring bias, triage recommendations shifted significantly in edge cases (odds ratio = 11.7, 95% confidence interval = 3.7–36.6), with the majority of shifts toward less urgent care. Crisis-intervention messages activated unpredictably across suicidal ideation presentations, occurring more frequently when patients described no specific method than when they did. Patient race, sex and barriers to care did not show significant effects, although confidence intervals did not exclude clinically meaningful differences. These findings reveal missed high-risk emergencies and inconsistent activation of crisis safeguards, raising safety concerns that warrant prospective validation before consumer-scale deployment of artificial intelligence triage systems.
ChatGPT Health performance in a structured test of triage recommendations
ChatGPT Health was launched in January 2026 as OpenAI’s consumer health tool and has reached millions of users. Here we conducted a structured stress test of triage recommendations using 60 clinician-authored vignettes across 21 clinical domains under 16 factorial conditions, yielding 960 total responses. Performance followed an inverted U-shaped pattern, with the most dangerous failures concentrated at clinical extremes—nonurgent presentations (35%) and emergency conditions (48%). Among gold-standard emergencies, the system undertriaged 52% of cases, directing patients with diabetic ketoacidosis or impending respiratory failure to 24–48 h evaluation rather than the emergency department, while correctly triaging classical emergencies such as stroke and anaphylaxis. When family or friends minimized symptoms, indicating anchoring bias, triage recommendations shifted significantly in edge cases (odds ratio = 11.7, 95% confidence interval = 3.7–36.6), with the majority of shifts toward less urgent care. Crisis-intervention messages activated unpredictably across suicidal ideation presentations, occurring more frequently when patients described no specific method than when they did. Patient race, sex and barriers to care did not show significant effects, although confidence intervals did not exclude clinically meaningful differences. These findings reveal missed high-risk emergencies and inconsistent activation of crisis safeguards, raising safety concerns that warrant prospective validation before consumer-scale deployment of artificial intelligence triage systems.

‘Unbelievably dangerous’: experts sound alarm after ChatGPT Health fails to recognise medical emergencies
Study finds ChatGPT Health did not recommend a hospital visit when medically necessary in more than half of cases

OpenAI shares data on ChatGPT users with suicidal thoughts, psychosis
The figure could mean potentially hundreds of thousands of users show signs of mental health distress weekly.

Reliability of LLMs as medical assistants for the general public: a randomized preregistered study
Global healthcare providers are exploring the use of large language models (LLMs) to provide medical advice to the public. LLMs now achieve nearly perfect scores on medical licensing exams, but this does not necessarily translate to accurate performance in real-world settings. We tested whether LLMs can assist members of the public in identifying underlying conditions and choosing a course of action (disposition) in ten medical scenarios in a controlled study with 1,298 participants. Participants were randomly assigned to receive assistance from an LLM (GPT-4o, Llama 3, Command R+) or a source of their choice (control). Tested alone, LLMs complete the scenarios accurately, correctly identifying conditions in 94.9% of cases and disposition in 56.3% on average. However, participants using the same LLMs identified relevant conditions in fewer than 34.5% of cases and disposition in fewer than 44.2%, both no better than the control group. We identify user interactions as a challenge to the deployment of LLMs for medical advice. Standard benchmarks for medical knowledge and simulated patient interactions do not predict the failures we find with human participants. Moving forward, we recommend systematic human user testing to evaluate interactive capabilities before public deployments in healthcare.

Reliability of LLMs as medical assistants for the general public: a randomized preregistered study
Global healthcare providers are exploring the use of large language models (LLMs) to provide medical advice to the public. LLMs now achieve nearly perfect scores on medical licensing exams, but this does not necessarily translate to accurate performance in real-world settings. We tested whether LLMs can assist members of the public in identifying underlying conditions and choosing a course of action (disposition) in ten medical scenarios in a controlled study with 1,298 participants. Participants were randomly assigned to receive assistance from an LLM (GPT-4o, Llama 3, Command R+) or a source of their choice (control). Tested alone, LLMs complete the scenarios accurately, correctly identifying conditions in 94.9% of cases and disposition in 56.3% on average. However, participants using the same LLMs identified relevant conditions in fewer than 34.5% of cases and disposition in fewer than 44.2%, both no better than the control group. We identify user interactions as a challenge to the deployment of LLMs for medical advice. Standard benchmarks for medical knowledge and simulated patient interactions do not predict the failures we find with human participants. Moving forward, we recommend systematic human user testing to evaluate interactive capabilities before public deployments in healthcare.

“It happened to be the perfect thing”: experiences of generative AI chatbots for mental health
The global mental health crisis underscores the need for accessible, effective interventions. Chatbots based on generative artificial intelligence (AI), like ChatGPT, are emerging as novel solutions, but research on real-life usage is limited. We interviewed nineteen individuals about their experiences using generative AI chatbots for mental health. Participants reported high engagement and positive impacts, including better relationships and healing from trauma and loss. We developed four themes: (1) a sense of ‘emotional sanctuary’, (2) ‘insightful guidance’, particularly about relationships, (3) the ‘joy of connection’, and (4) comparisons between the ‘AI therapist’ and human therapy. Some themes echoed prior research on rule-based chatbots, while others seemed novel to generative AI. Participants emphasised the need for better safety guardrails, human-like memory and the ability to lead the therapeutic process. Generative AI chatbots may offer mental health support that feels meaningful to users, but further research is needed on safety and effectiveness.

“It happened to be the perfect thing”: experiences of generative AI chatbots for mental health
The global mental health crisis underscores the need for accessible, effective interventions. Chatbots based on generative artificial intelligence (AI), like ChatGPT, are emerging as novel solutions, but research on real-life usage is limited. We interviewed nineteen individuals about their experiences using generative AI chatbots for mental health. Participants reported high engagement and positive impacts, including better relationships and healing from trauma and loss. We developed four themes: (1) a sense of ‘emotional sanctuary’, (2) ‘insightful guidance’, particularly about relationships, (3) the ‘joy of connection’, and (4) comparisons between the ‘AI therapist’ and human therapy. Some themes echoed prior research on rule-based chatbots, while others seemed novel to generative AI. Participants emphasised the need for better safety guardrails, human-like memory and the ability to lead the therapeutic process. Generative AI chatbots may offer mental health support that feels meaningful to users, but further research is needed on safety and effectiveness.

Public use of a generalist LLM chatbot for health queries
Here we analyse over 500,000 de-identified health-related conversations with Microsoft Copilot from January 2026 to characterize what people ask conversational artificial intelligence (AI) about health. We apply a hierarchical intent taxonomy of 12 primary categories using privacy-preserving large language model-based classification validated against expert human annotation and use topic clustering for prevalent themes within each intent. We then characterize the intents and topics behind health queries, identify who they are about, and analyse how usage varies by device and time of day. Nearly one in five conversations involves personal symptom assessment or condition discussion, and the dominant general information category is also concentrated on specific treatments and conditions, suggesting that this is a lower bound on personal health intent. One in seven of these personal health queries concerns someone other than the user, suggesting that conversational AI can also be a caregiving tool. Personal queries increase markedly in the evening and nighttime hours, when traditional healthcare is most limited. Usage diverges sharply by device: mobile concentrates on personal health concerns, while desktop is dominated by professional and academic work. A substantial share of queries focuses on navigating healthcare systems. These patterns have direct implications for platform-specific design, safety considerations and the responsible development of health AI.

Public use of a generalist LLM chatbot for health queries
Here we analyse over 500,000 de-identified health-related conversations with Microsoft Copilot from January 2026 to characterize what people ask conversational artificial intelligence (AI) about health. We apply a hierarchical intent taxonomy of 12 primary categories using privacy-preserving large language model-based classification validated against expert human annotation and use topic clustering for prevalent themes within each intent. We then characterize the intents and topics behind health queries, identify who they are about, and analyse how usage varies by device and time of day. Nearly one in five conversations involves personal symptom assessment or condition discussion, and the dominant general information category is also concentrated on specific treatments and conditions, suggesting that this is a lower bound on personal health intent. One in seven of these personal health queries concerns someone other than the user, suggesting that conversational AI can also be a caregiving tool. Personal queries increase markedly in the evening and nighttime hours, when traditional healthcare is most limited. Usage diverges sharply by device: mobile concentrates on personal health concerns, while desktop is dominated by professional and academic work. A substantial share of queries focuses on navigating healthcare systems. These patterns have direct implications for platform-specific design, safety considerations and the responsible development of health AI.

"AI Psychosis" in Context: How Conversation History Shapes LLM Responses to Delusional Beliefs
Extended interaction with large language models (LLMs) has been linked to the reinforcement of delusional beliefs, attracting clinical and public concern. Yet most empirical work evaluates model safety in brief interactions, which may not reflect how harms develop through sustained dialogue. Five LLMs were tested across three levels of accumulated context, using the same escalating delusional conversation history to isolate its effect on model behaviour. Responses were coded on risk and safety dimensions, and each model was analysed qualitatively. Models separated into two distinct tiers: GPT-4o, Grok 4.1 Fast, and Gemini 3 Pro exhibited high-risk, low-safety profiles; Claude Opus 4.5 and GPT-5.2 Instant displayed the opposite pattern. As context accumulated, performance degraded in the unsafe group, while the same material activated stronger safety interventions among safer models. Qualitative analysis identified distinct mechanisms of failure, including validating the user's delusional premises, elaborating beyond them with new content, and attempting harm reduction from within the delusional frame. Safer models, however, often used the established relationship to support intervention, challenging delusional beliefs and directing the user to external support. These findings indicate that accumulated context functions as a stress test of safety architecture, revealing whether prior dialogue is treated as a worldview to inherit or evidence to evaluate. Short-context assessments may therefore mischaracterise model safety, underestimating danger in some systems while missing context-activated gains in others. The results suggest that delusion reinforcement is a tractable alignment failure, with safer models establishing a baseline that future systems should now be expected to meet.

Use of Consumer Chatbots for Emergent Adolescent Health Concerns
This cross-sectional study examines content policies and assesses behaviors of consumer chatbots in response to adolescent health crises.

I wanted ChatGPT to help me. So why did it advise me how to kill myself?
ChatGPT wrote a woman a suicide note and another AI chatbot role-played sexual acts with children, BBC finds.

This Treatment Works, Right? Evaluating LLM Sensitivity to Patient Question Framing in Medical QA
Patients are increasingly turning to large language models (LLMs) with medical questions that are complex and difficult to articulate clearly. However, LLMs are sensitive to prompt phrasings and can be influenced by the way questions are worded. Ideally, LLMs should respond consistently regardless of phrasing, particularly when grounded in the same underlying evidence. We investigate this through a systematic evaluation in a controlled retrieval-augmented generation (RAG) setting for medical question answering (QA), where expert-selected documents are used rather than retrieved automatically. We examine two dimensions of patient query variation: question framing (positive vs. negative) and language style (technical vs. plain language). We construct a dataset of 6,614 query pairs grounded in clinical trial abstracts and evaluate response consistency across eight LLMs. Our findings show that positively- and negatively-framed pairs are significantly more likely to produce contradictory conclusions than same-framing pairs. This framing effect is further amplified in multi-turn conversations, where sustained persuasion increases inconsistency. We find no significant interaction between framing and language style. Our results demonstrate that LLM responses in medical QA can be systematically influenced through query phrasing alone, even when grounded in the same evidence, highlighting the importance of phrasing robustness as an evaluation criterion for RAG-based systems in high-stakes settings.

Characteristics and Safety of Consumer Chatbots for Emergent Adolescent Health Concerns
This cross-sectional study examines content policies and assesses behaviors of consumer chatbots in response to adolescent health crises.

Co-creation process of an app for people with rare diseases - a citizen science approach
Background Rare diseases affect a small percentage of the population, leading to challenges such as delayed diagnoses and limited treatment options. Mobile health technologies offer solutions to improve patient outcomes, yet their application in rare diseases remains underexplored. The German citizen science project SelEe created a customizable app for the self-management of rare diseases through a co-creation process that involved patients with such conditions. Methods The project consisted of three phases. In Phase 1, 9 to 68 patients or relatives of patients participated in workshops to define research topics and app requirements. Phase 2 involved a core research team of nine patients and researchers who iteratively developed the app, released in March 2023. Phase 3 focused on evaluating the app’s usage and usability through an in-app survey conducted from March 2023 to February 2024. We utilized descriptive statistics to evaluate app usage and employed the mHealth App Usability Questionnaire to assess usability. Results The SelEe app offers the possibility to create and store data in a personalized health diary. Patients can create their own templates or use templates which were defined by the core research team. Users can record findings (e.g. blood test results) and export data using different graphs and formats. Furthermore, the app supports blind users. The app was downloaded 3040 times and 1456 users registered, with 1967 unique diseases entered. 50.7% of the diseases were rare, 30.5% non-rare, and 18.8% were classified as suspected, undefined, or symptoms. A total of 1223 valid user profiles were analyzed for app usage and demographics. Furthermore, 432 users qualified for the in-app survey by making at least one health diary entry, and 117 participated. The app was rated with an overall usability score of 5.19 out of 7. While the app’s health diary function was frequently used, other functionalities like findings and data export were less utilized. Feedback highlighted the need for improved usability and additional features. Conclusions The study highlights active patient engagement in developing a mobile health app for individuals with rare diseases. Although improvements are necessary for broader acceptance, the app is promising for the management of rare diseases. Supplementary information The online version contains supplementary material available at 10.1186/s13023-025-04140-1.

Chatbots are 'constantly validating everything' even when you're suicidal. New research measures how dangerous AI psychosis really is | Fortune
It's free, easy to use, doesn't have a stigma around it, and makes you feel better. Researchers and experts warn that's the problem with chatbots.
