







This summary was generated using automated tools and was not authored or reviewed by the article's author(s). It is provided to support discovery, help readers assess relevance, and assist readers from adjacent research areas in understanding the work. It is intended to complement the author-supplied abstract, which remains the primary summary of the paper. The full article remains the authoritative version of record. Click here to learn more.
Fact Sheet 3: ME/CFS: Information for Medical Professionals
Fact Sheet 3: ME/CFS: Information for Medical Professionals Published December 2025 Link to pdf: ME/CFS: Information for Medical Professionals.pdf Discussion thread: Fact sheet #3: Information...
Reliability of LLMs as medical assistants for the general public: a randomized preregistered study
Global healthcare providers are exploring the use of large language models (LLMs) to provide medical advice to the public. LLMs now achieve nearly perfect scores on medical licensing exams, but this does not necessarily translate to accurate performance in real-world settings. We tested whether LLMs can assist members of the public in identifying underlying conditions and choosing a course of action (disposition) in ten medical scenarios in a controlled study with 1,298 participants. Participants were randomly assigned to receive assistance from an LLM (GPT-4o, Llama 3, Command R+) or a source of their choice (control). Tested alone, LLMs complete the scenarios accurately, correctly identifying conditions in 94.9% of cases and disposition in 56.3% on average. However, participants using the same LLMs identified relevant conditions in fewer than 34.5% of cases and disposition in fewer than 44.2%, both no better than the control group. We identify user interactions as a challenge to the deployment of LLMs for medical advice. Standard benchmarks for medical knowledge and simulated patient interactions do not predict the failures we find with human participants. Moving forward, we recommend systematic human user testing to evaluate interactive capabilities before public deployments in healthcare.

Reliability of LLMs as medical assistants for the general public: a randomized preregistered study
Global healthcare providers are exploring the use of large language models (LLMs) to provide medical advice to the public. LLMs now achieve nearly perfect scores on medical licensing exams, but this does not necessarily translate to accurate performance in real-world settings. We tested whether LLMs can assist members of the public in identifying underlying conditions and choosing a course of action (disposition) in ten medical scenarios in a controlled study with 1,298 participants. Participants were randomly assigned to receive assistance from an LLM (GPT-4o, Llama 3, Command R+) or a source of their choice (control). Tested alone, LLMs complete the scenarios accurately, correctly identifying conditions in 94.9% of cases and disposition in 56.3% on average. However, participants using the same LLMs identified relevant conditions in fewer than 34.5% of cases and disposition in fewer than 44.2%, both no better than the control group. We identify user interactions as a challenge to the deployment of LLMs for medical advice. Standard benchmarks for medical knowledge and simulated patient interactions do not predict the failures we find with human participants. Moving forward, we recommend systematic human user testing to evaluate interactive capabilities before public deployments in healthcare.

Evidence appraisal: a scoping review, conceptual framework, and research agenda
Abstract Objective Critical appraisal of clinical evidence promises to help prevent, detect, and address flaws related to study importance, ethics, validity, applicability, and reporting. These research issues are of growing concern. The purpose of this scoping review is to survey the current literature on evidence appraisal to develop a conceptual framework and an informatics research agenda. Methods We conducted an iterative literature search of Medline for discussion or research on the critical appraisal of clinical evidence. After title and abstract review, 121 articles were included in the analysis. We performed qualitative thematic analysis to describe the evidence appraisal architecture and its issues and opportunities. From this analysis, we derived a conceptual framework and an informatics research agenda. Results We identified 68 themes in 10 categories. This analysis revealed that the practice of evidence appraisal is quite common but is rarely subjected to documentation, organization, validation, integration, or uptake. This is related to underdeveloped tools, scant incentives, and insufficient acquisition of appraisal data and transformation of the data into usable knowledge. Discussion The gaps in acquiring appraisal data, transforming the data into actionable information and knowledge, and ensuring its dissemination and adoption can be addressed with proven informatics approaches. Conclusions Evidence appraisal faces several challenges, but implementing an informatics research agenda would likely help realize the potential of evidence appraisal for improving the rigor and value of clinical evidence.


Social Uses of Personal Health Information Within PatientsLikeMe, an Online Patient Community: What Can Happen When Patients Have Access to One Another’s Data
Background: This project investigates the ways in which patients respond to the shared use of what is often considered private information: personal health data. There is a growing demand for patient access to personal health records. The predominant model for this record is a repository of all clinically relevant health information kept securely and viewed privately by patients and their health care providers. While this type of record does seem to have beneficial effects for the patient–physician relationship, the complexity and novelty of these data coupled with the lack of research in this area means the utility of personal health information for the primary stakeholders—the patients—is not well documented or understood. Objective: PatientsLikeMe is an online community built to support information exchange between patients. The site provides customized disease-specific outcome and visualization tools to help patients understand and share information about their condition. We begin this paper by describing the components and design of the online community. We then identify and analyze how users of this platform reference personal health information within patient-to-patient dialogues. Methods: Patients diagnosed with amyotrophic lateral sclerosis (ALS) post data on their current treatments, symptoms, and outcomes. These data are displayed graphically within personal health profiles and are reflected in composite community-level symptom and treatment reports. Users review and discuss these data within the Forum, private messaging, and comments posted on each other’s profiles. We analyzed member communications that referenced individual-level personal health data to determine how patient peers use personal health information within patient-to-patient exchanges. Results: Qualitative analysis of a sample of 123 comments (about 2% of the total) posted within the community revealed a variety of commenting and questioning behaviors by patient members. Members referenced data to locate others with particular experiences to answer specific health-related questions, to proffer personally acquired disease-management knowledge to those most likely to benefit from it, and to foster and solidify relationships based on shared concerns. Conclusions: Few studies examine the use of personal health information by patients themselves. This project suggests how patients who choose to explicitly share health data within a community may benefit from the process, helping them engage in dialogues that may inform disease self-management. We recommend that future designs make each patient’s health information as clear as possible, automate matching of people with similar conditions and using similar treatments, and integrate data into online platforms for health conversations.
Preparing Physicians for the Clinical Algorithm Era
The U.S. government recently took steps to ensure that clinical decision support algorithms are safe for clinical use. The next and larger step will be teaching physicians how to use the algorithms ...

Preparing Physicians for the Clinical Algorithm Era
The U.S. government recently took steps to ensure that clinical decision support algorithms are safe for clinical use. The next and larger step will be teaching physicians how to use the algorithms ...

Qualitative research: standards, challenges, and guidelines
Qualitative research methods could help us to improve our understanding of medicine. Rather than thinking of qualitative and quantitative strategies as incompatible, they should be seen as complementary. Although procedures for textual interpretation differ from those of statistical analysis, because of the different type of data used and questions to be answered, the underlying principles are much the same. In this article I propose relevance, validity, and reflexivity as overall standards for qualitative inquiry.

Hallucination by proxy in LLM-assisted differential diagnosis
Current evidence suggests that LLM assistance could augment the diagnostic accuracy of clinicians. However, these systems are black boxes, susceptible to hallucinations, and project a potentially...

Hallucination by proxy in LLM-assisted differential diagnosis
Current evidence suggests that LLM assistance could augment the diagnostic accuracy of clinicians. However, these systems are black boxes, susceptible to hallucinations, and project a potentially...

I had strong priors against LLMs for medicine. There are a lot of doctors in my family and I grew up viewing doctors as careful, skilled professionals. I had plenty of bad medical experiences, but I thought it would be hard to do better. Then an LLM found a cure for my 2 decade chronic condition...
An editorial was published in Nature recently claiming that glam journal publication of LLMs (like DeepSeek-R1 this case) marks a step towards greater transparency, accountability & credibility nature.com/articles/d41586-025-02979-9. I have thoughts ... 1/
https://www.nature.com/articles/d41586-025-02979-9
t.coBy request of a lot of people who were looking for an example of how I do clinical notes for a new appointment with a specialist about extremely complex shit that's involved a lot of referral runaround and/or dildobucket doctors: dropbox.com/scl/fi/6dx9mbztu2i40987mjsw8/…
I read this result as: LLMs do more bullshit citations, name-dropping without engaging.
infoDOCKET
Citing Less Critically: #LLMs Reshape the Rhetoric and Reach of #Scientific #Citation (New Research Article (preprint); via @arxiv.bsky.social) arxiv.org/abs/2609.01432 #scholcomm #citations #libraries #AI #GenAI
I had strong priors against LLMs for medicine. There are a lot of doctors in my family and I grew up viewing doctors as careful, skilled professionals. I had plenty of bad medical experiences, but I thought it would be hard to do better. Then an LLM found a cure for my 2 decade chronic condition...