







An open, AI-assisted evidence synthesis of how language concordance — matching patients with providers or interpreters who share their language — affects healthcare outcomes. Every question, claim, evidence item, caveat, and source is its own addressable node.
General-purpose large language models outperform specialized clinical AI tools on medical benchmarks
Specialized clinical artificial intelligence (AI) tools are entering medical practice despite scarce independent evaluation. We quantitatively evaluate two clinical AI tools, OpenEvidence and UpToDate Expert AI, built on large language models (LLMs) against three frontier LLMs: GPT-5.2, Gemini 3.1 Pro and Claude Opus 4.6. Our evaluation has three stages: (1) 500 MedQA questions testing medical knowledge, (2) 500 HealthBench items measuring alignment with clinicians and (3) the real clinical queries (RCQ) benchmark, built from 100 de-identified queries from physicians to a general-purpose language model in a live clinical environment. For the RCQ benchmark, 12 US clinicians performed randomized, blinded review of model outputs, producing 1,800 model–question annotations. Frontier LLMs outperformed clinical AI tools in all three evaluations. Clinical AI tools performed comparably to auto-enabled Google Search AI Overview on the RCQ. These findings highlight the need for independent, real-world evaluation of AI tools before they enter clinical settings.

Document analysis in health policy research: the READ approach
Abstract. Document analysis is one of the most commonly used and powerful methods in health policy research. While existing qualitative research manuals of

Like a Therapist, But Not: Reddit Narratives of AI in Mental Health Contexts
Large language models (LLMs) are increasingly used for emotional support and mental health-related interactions outside clinical settings, yet little is known about how people evaluate and relate to these systems in everyday use. We analyze 5,126 Reddit posts from 47 mental health communities describing experiential or exploratory use of AI for emotional support or therapy. Grounded in the Technology Acceptance Model and therapeutic alliance theory, we develop a theory-informed annotation framework and apply a hybrid LLM-human pipeline to analyze evaluative language, adoption-related attitudes, and relational alignment at scale. Our results show that engagement is shaped primarily by narrated outcomes, trust, and response quality, rather than emotional bond alone. Positive sentiment is most strongly associated with task and goal alignment, while companionship-oriented use more often involves misaligned alliances and reported risks such as dependence and symptom escalation. Overall, this work demonstrates how theory-grounded constructs can be operationalized in large-scale discourse analysis and highlights the importance of studying how users interpret language technologies in sensitive, real-world contexts.

Your medical provider might be recording your mental health care visits – The Markup
Mental health providers are increasingly using AI technology to record conversations, raising privacy concerns among patients and practitioners.

Doctors’ AI scribes get names of drugs and diagnoses wrong, NHS watchdog warns
Exclusive: Patients identify errors in consultation transcripts that are missed by GPs, Healthwatch England finds

The shrinking landscape of linguistic diversity in the age of large language models
Language is far more than a communication tool; it encodes a wealth of information about a person’s identity, psychological state and social context, providing valuable insights for diverse fields including psychology, marketing and healthcare. Across three studies spanning seven datasets in different domains and over 880,000 texts, we show that the widespread adoption of large language models (LLMs) as writing assistants is linked to declines in linguistic diversity, interfering with the societal and psychological insights language provides. While core content is retained when LLMs polish and rewrite texts, LLMs also homogenize writing styles, reducing writing-complexity variance by a statistically significant 21–50% across datasets and models (P ≤ 0.05), and amplify patterns associated with dominant characteristics while suppressing others, emphasizing conformity over individuality. These trends hold across different LLMs, prompts and contexts, with potential implications for diagnostic processes, personalization efforts, hiring assessments and cultural preservation.

Performance of a large language model on the reasoning tasks of a physician
More than 65 years ago, complex clinical diagnostic reasoning cases were introduced as the gold standard for the evaluation of expert medical computing systems, a standard that has held ever since. In this study, we report the results of a physician ...

The Collective Intelligence Project
We’ve launched an open, collaborative platform to build evaluations that test what matters to you. We empower a global community to create qualitative benchmarks for any domain—from medical chatbots to legal assistance. Just as Wikipedia democratized knowledge, Weval aims to democratize evaluation, ensuring that AI works for, and represents, everyone.

Getting started - Docs
Discourse Graphs are a tool and ecosystem for collaborative knowledge synthesis, enabling researchers to map ideas and arguments in a modular, composable graph format.
Americans Turning to AI to Supplement Healthcare Visits
One in four Americans use AI for health information. Most do so to supplement care, but some are using AI in place of a provider visit when barriers arise.

This Treatment Works, Right? Evaluating LLM Sensitivity to Patient Question Framing in Medical QA
Patients are increasingly turning to large language models (LLMs) with medical questions that are complex and difficult to articulate clearly. However, LLMs are sensitive to prompt phrasings and can be influenced by the way questions are worded. Ideally, LLMs should respond consistently regardless of phrasing, particularly when grounded in the same underlying evidence. We investigate this through a systematic evaluation in a controlled retrieval-augmented generation (RAG) setting for medical question answering (QA), where expert-selected documents are used rather than retrieved automatically. We examine two dimensions of patient query variation: question framing (positive vs. negative) and language style (technical vs. plain language). We construct a dataset of 6,614 query pairs grounded in clinical trial abstracts and evaluate response consistency across eight LLMs. Our findings show that positively- and negatively-framed pairs are significantly more likely to produce contradictory conclusions than same-framing pairs. This framing effect is further amplified in multi-turn conversations, where sustained persuasion increases inconsistency. We find no significant interaction between framing and language style. Our results demonstrate that LLM responses in medical QA can be systematically influenced through query phrasing alone, even when grounded in the same evidence, highlighting the importance of phrasing robustness as an evaluation criterion for RAG-based systems in high-stakes settings.

Co-Writing with Opinionated Language Models Affects Users’ Views
As part of the Digital Library's transition to Open Access, new features for researchers are available in the Premium Edition. Click here to learn more.

Co-Writing with Opinionated Language Models Affects Users’ Views
As part of the Digital Library's transition to Open Access, new features for researchers are available in the Premium Edition. Click here to learn more.

Avoiding Ableist Language: Suggestions for Autism Researchers - Kristen Bottema-Beutel, Steven K. Kapp, Jessica Nina Lester, Noah J. Sasson, Brittany N. Hand, 2021
In this commentary, we describe how language used to communicate about autism within much of autism research can reflect and perpetuate ableist ideologies (i.e....

Cheap AI chatbots transform medical diagnoses in places with limited care
Studies in Rwanda and Pakistan reveal real-world utility of chatbots in underfunded clinics, and not just in benchmark tests.

Cheap AI chatbots transform medical diagnoses in places with limited care
Studies in Rwanda and Pakistan reveal real-world utility of chatbots in underfunded clinics, and not just in benchmark tests.
