







ChatGPT didn’t replace my care team, but it did make me a better patient.
ChatGPT Health performance in a structured test of triage recommendations
ChatGPT Health was launched in January 2026 as OpenAI’s consumer health tool and has reached millions of users. Here we conducted a structured stress test of triage recommendations using 60 clinician-authored vignettes across 21 clinical domains under 16 factorial conditions, yielding 960 total responses. Performance followed an inverted U-shaped pattern, with the most dangerous failures concentrated at clinical extremes—nonurgent presentations (35%) and emergency conditions (48%). Among gold-standard emergencies, the system undertriaged 52% of cases, directing patients with diabetic ketoacidosis or impending respiratory failure to 24–48 h evaluation rather than the emergency department, while correctly triaging classical emergencies such as stroke and anaphylaxis. When family or friends minimized symptoms, indicating anchoring bias, triage recommendations shifted significantly in edge cases (odds ratio = 11.7, 95% confidence interval = 3.7–36.6), with the majority of shifts toward less urgent care. Crisis-intervention messages activated unpredictably across suicidal ideation presentations, occurring more frequently when patients described no specific method than when they did. Patient race, sex and barriers to care did not show significant effects, although confidence intervals did not exclude clinically meaningful differences. These findings reveal missed high-risk emergencies and inconsistent activation of crisis safeguards, raising safety concerns that warrant prospective validation before consumer-scale deployment of artificial intelligence triage systems.

ChatGPT Health performance in a structured test of triage recommendations
ChatGPT Health was launched in January 2026 as OpenAI’s consumer health tool and has reached millions of users. Here we conducted a structured stress test of triage recommendations using 60 clinician-authored vignettes across 21 clinical domains under 16 factorial conditions, yielding 960 total responses. Performance followed an inverted U-shaped pattern, with the most dangerous failures concentrated at clinical extremes—nonurgent presentations (35%) and emergency conditions (48%). Among gold-standard emergencies, the system undertriaged 52% of cases, directing patients with diabetic ketoacidosis or impending respiratory failure to 24–48 h evaluation rather than the emergency department, while correctly triaging classical emergencies such as stroke and anaphylaxis. When family or friends minimized symptoms, indicating anchoring bias, triage recommendations shifted significantly in edge cases (odds ratio = 11.7, 95% confidence interval = 3.7–36.6), with the majority of shifts toward less urgent care. Crisis-intervention messages activated unpredictably across suicidal ideation presentations, occurring more frequently when patients described no specific method than when they did. Patient race, sex and barriers to care did not show significant effects, although confidence intervals did not exclude clinically meaningful differences. These findings reveal missed high-risk emergencies and inconsistent activation of crisis safeguards, raising safety concerns that warrant prospective validation before consumer-scale deployment of artificial intelligence triage systems.

Cheap AI chatbots transform medical diagnoses in places with limited care
Studies in Rwanda and Pakistan reveal real-world utility of chatbots in underfunded clinics, and not just in benchmark tests.

Cheap AI chatbots transform medical diagnoses in places with limited care
Studies in Rwanda and Pakistan reveal real-world utility of chatbots in underfunded clinics, and not just in benchmark tests.

Reliability of LLMs as medical assistants for the general public: a randomized preregistered study
Global healthcare providers are exploring the use of large language models (LLMs) to provide medical advice to the public. LLMs now achieve nearly perfect scores on medical licensing exams, but this does not necessarily translate to accurate performance in real-world settings. We tested whether LLMs can assist members of the public in identifying underlying conditions and choosing a course of action (disposition) in ten medical scenarios in a controlled study with 1,298 participants. Participants were randomly assigned to receive assistance from an LLM (GPT-4o, Llama 3, Command R+) or a source of their choice (control). Tested alone, LLMs complete the scenarios accurately, correctly identifying conditions in 94.9% of cases and disposition in 56.3% on average. However, participants using the same LLMs identified relevant conditions in fewer than 34.5% of cases and disposition in fewer than 44.2%, both no better than the control group. We identify user interactions as a challenge to the deployment of LLMs for medical advice. Standard benchmarks for medical knowledge and simulated patient interactions do not predict the failures we find with human participants. Moving forward, we recommend systematic human user testing to evaluate interactive capabilities before public deployments in healthcare.

Reliability of LLMs as medical assistants for the general public: a randomized preregistered study
Global healthcare providers are exploring the use of large language models (LLMs) to provide medical advice to the public. LLMs now achieve nearly perfect scores on medical licensing exams, but this does not necessarily translate to accurate performance in real-world settings. We tested whether LLMs can assist members of the public in identifying underlying conditions and choosing a course of action (disposition) in ten medical scenarios in a controlled study with 1,298 participants. Participants were randomly assigned to receive assistance from an LLM (GPT-4o, Llama 3, Command R+) or a source of their choice (control). Tested alone, LLMs complete the scenarios accurately, correctly identifying conditions in 94.9% of cases and disposition in 56.3% on average. However, participants using the same LLMs identified relevant conditions in fewer than 34.5% of cases and disposition in fewer than 44.2%, both no better than the control group. We identify user interactions as a challenge to the deployment of LLMs for medical advice. Standard benchmarks for medical knowledge and simulated patient interactions do not predict the failures we find with human participants. Moving forward, we recommend systematic human user testing to evaluate interactive capabilities before public deployments in healthcare.

“It happened to be the perfect thing”: experiences of generative AI chatbots for mental health
The global mental health crisis underscores the need for accessible, effective interventions. Chatbots based on generative artificial intelligence (AI), like ChatGPT, are emerging as novel solutions, but research on real-life usage is limited. We interviewed nineteen individuals about their experiences using generative AI chatbots for mental health. Participants reported high engagement and positive impacts, including better relationships and healing from trauma and loss. We developed four themes: (1) a sense of ‘emotional sanctuary’, (2) ‘insightful guidance’, particularly about relationships, (3) the ‘joy of connection’, and (4) comparisons between the ‘AI therapist’ and human therapy. Some themes echoed prior research on rule-based chatbots, while others seemed novel to generative AI. Participants emphasised the need for better safety guardrails, human-like memory and the ability to lead the therapeutic process. Generative AI chatbots may offer mental health support that feels meaningful to users, but further research is needed on safety and effectiveness.

“It happened to be the perfect thing”: experiences of generative AI chatbots for mental health
The global mental health crisis underscores the need for accessible, effective interventions. Chatbots based on generative artificial intelligence (AI), like ChatGPT, are emerging as novel solutions, but research on real-life usage is limited. We interviewed nineteen individuals about their experiences using generative AI chatbots for mental health. Participants reported high engagement and positive impacts, including better relationships and healing from trauma and loss. We developed four themes: (1) a sense of ‘emotional sanctuary’, (2) ‘insightful guidance’, particularly about relationships, (3) the ‘joy of connection’, and (4) comparisons between the ‘AI therapist’ and human therapy. Some themes echoed prior research on rule-based chatbots, while others seemed novel to generative AI. Participants emphasised the need for better safety guardrails, human-like memory and the ability to lead the therapeutic process. Generative AI chatbots may offer mental health support that feels meaningful to users, but further research is needed on safety and effectiveness.

Public use of a generalist LLM chatbot for health queries
Here we analyse over 500,000 de-identified health-related conversations with Microsoft Copilot from January 2026 to characterize what people ask conversational artificial intelligence (AI) about health. We apply a hierarchical intent taxonomy of 12 primary categories using privacy-preserving large language model-based classification validated against expert human annotation and use topic clustering for prevalent themes within each intent. We then characterize the intents and topics behind health queries, identify who they are about, and analyse how usage varies by device and time of day. Nearly one in five conversations involves personal symptom assessment or condition discussion, and the dominant general information category is also concentrated on specific treatments and conditions, suggesting that this is a lower bound on personal health intent. One in seven of these personal health queries concerns someone other than the user, suggesting that conversational AI can also be a caregiving tool. Personal queries increase markedly in the evening and nighttime hours, when traditional healthcare is most limited. Usage diverges sharply by device: mobile concentrates on personal health concerns, while desktop is dominated by professional and academic work. A substantial share of queries focuses on navigating healthcare systems. These patterns have direct implications for platform-specific design, safety considerations and the responsible development of health AI.

Public use of a generalist LLM chatbot for health queries
Here we analyse over 500,000 de-identified health-related conversations with Microsoft Copilot from January 2026 to characterize what people ask conversational artificial intelligence (AI) about health. We apply a hierarchical intent taxonomy of 12 primary categories using privacy-preserving large language model-based classification validated against expert human annotation and use topic clustering for prevalent themes within each intent. We then characterize the intents and topics behind health queries, identify who they are about, and analyse how usage varies by device and time of day. Nearly one in five conversations involves personal symptom assessment or condition discussion, and the dominant general information category is also concentrated on specific treatments and conditions, suggesting that this is a lower bound on personal health intent. One in seven of these personal health queries concerns someone other than the user, suggesting that conversational AI can also be a caregiving tool. Personal queries increase markedly in the evening and nighttime hours, when traditional healthcare is most limited. Usage diverges sharply by device: mobile concentrates on personal health concerns, while desktop is dominated by professional and academic work. A substantial share of queries focuses on navigating healthcare systems. These patterns have direct implications for platform-specific design, safety considerations and the responsible development of health AI.

Doctors’ AI scribes get names of drugs and diagnoses wrong, NHS watchdog warns
Exclusive: Patients identify errors in consultation transcripts that are missed by GPs, Healthwatch England finds

Language Access in Healthcare — Discourse Graph
An open, AI-assisted evidence synthesis of how language concordance — matching patients with providers or interpreters who share their language — affects healthcare outcomes. Every question, claim, evidence item, caveat, and source is its own addressable node.

Preparing Physicians for the Clinical Algorithm Era
The U.S. government recently took steps to ensure that clinical decision support algorithms are safe for clinical use. The next and larger step will be teaching physicians how to use the algorithms ...

Preparing Physicians for the Clinical Algorithm Era
The U.S. government recently took steps to ensure that clinical decision support algorithms are safe for clinical use. The next and larger step will be teaching physicians how to use the algorithms ...

It’s not lost on me the impact of my latinidad and the team behind Blacksky being responsible for creating the Atmosphere’s first medical platform. People historically abused & underrepresented in medicine are responsible for building the foundation for its community here. Don’t forget that.