







By request of a lot of people who were looking for an example of how I do clinical notes for a new appointment with a specialist about extremely complex shit that's involved a lot of referral runaround and/or dildobucket doctors: dropbox.com/scl/fi/6dx9mbztu2i40987mjsw8/…
Feb 16, 2026 at 10:01 PM
Your doctor’s AI notetaker may be making things up, Ontario audit finds
Made-up therapy referrals, incorrect prescriptions among the common mistakes.

Advancing AMIE towards expert-level audio-visual clinical consultations
Anil Palepu, Senior Research Scientist, and Mike Schaekermann, Research Lead, Google

Note-ifying all the things (with Boris Mann)
Conversation about note-taking, the overwhelm of tools, and future possibilities. (shortlink:…

Fact Sheet 3: ME/CFS: Information for Medical Professionals
Fact Sheet 3: ME/CFS: Information for Medical Professionals Published December 2025 Link to pdf: ME/CFS: Information for Medical Professionals.pdf Discussion thread: Fact sheet #3: Information...
Bluenotes: Community Notes for ATProto
I'll briefly demonstrate Bluenotes, a fork of the Bluesky with Community Notes. Then I'll discuss a proposed "Open Community Notes" standard and lexicon for Community Notes on ATProto, and discuss challenges such as preserving contributor anonymity and defending against manipulation.

Doctors’ AI scribes get names of drugs and diagnoses wrong, NHS watchdog warns
Exclusive: Patients identify errors in consultation transcripts that are missed by GPs, Healthwatch England finds

Towards Expert-level Medical AI for Real-time Video Consultations
Audio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effective assessment of illness through non-verbal cues. While text-based AI has shown promise, it discards essential perceptual dimensions and limits patients who cannot articulate symptoms in writing. Early efforts to extend medical AI to audio-visual interaction have demonstrated feasibility but not reached clinician-level performance. Here, we provide the first demonstration of expert-level AI in real-time clinical video consultations using AMIE (Articulate Medical Intelligence Explorer) in a video configuration. AMIE (Video) is a Gemini-based multi-agent system integrating low-latency dialogue, clinical reasoning, and real-time audio-visual perception. To guide development, we established a taxonomy and automated evaluations for clinical audio-visual cues in telehealth settings. In a randomized Objective Structured Clinical Examination (OSCE) study with 30 primary care physicians (PCPs), 15 patient actors and 100 clinical scenarios, we compared AMIE (Video), its text-only counterpart AMIE (Text), and PCPs consulting via video. Clinical evaluators rated AMIE (Video) on par or better than PCPs in history-taking, diagnosis, management, and physical observation and examination. Patient actors preferred AMIE's approach to assessing and explaining conditions, while PCPs were preferred for rapport and partnership building. In modality ablation, patient actors preferred AMIE (Video)'s interface over text chat for communicative effectiveness, convenience, and feeling understood. Limitations remain in fine anatomical precision, subtle affective nuances, and high-frequency movements. While further research is needed before real-world translation, these results mark an important milestone toward AI systems capable of augmenting care across the sensory complexity of clinical practice.

What have note-taking PKMs accomplished, really?
As somebody who writes a lot, I've always been interested in personal knowledge management systems and note-taking apps. But what do these frameworks and methodologies actually give to researchers and the broader world? Has there been a meaningful increase of wisdom since these became popular? I went to look for an answer.


Your medical provider might be recording your mental health care visits – The Markup
Mental health providers are increasingly using AI technology to record conversations, raising privacy concerns among patients and practitioners.



Reliability of LLMs as medical assistants for the general public: a randomized preregistered study
Global healthcare providers are exploring the use of large language models (LLMs) to provide medical advice to the public. LLMs now achieve nearly perfect scores on medical licensing exams, but this does not necessarily translate to accurate performance in real-world settings. We tested whether LLMs can assist members of the public in identifying underlying conditions and choosing a course of action (disposition) in ten medical scenarios in a controlled study with 1,298 participants. Participants were randomly assigned to receive assistance from an LLM (GPT-4o, Llama 3, Command R+) or a source of their choice (control). Tested alone, LLMs complete the scenarios accurately, correctly identifying conditions in 94.9% of cases and disposition in 56.3% on average. However, participants using the same LLMs identified relevant conditions in fewer than 34.5% of cases and disposition in fewer than 44.2%, both no better than the control group. We identify user interactions as a challenge to the deployment of LLMs for medical advice. Standard benchmarks for medical knowledge and simulated patient interactions do not predict the failures we find with human participants. Moving forward, we recommend systematic human user testing to evaluate interactive capabilities before public deployments in healthcare.

Reliability of LLMs as medical assistants for the general public: a randomized preregistered study
Global healthcare providers are exploring the use of large language models (LLMs) to provide medical advice to the public. LLMs now achieve nearly perfect scores on medical licensing exams, but this does not necessarily translate to accurate performance in real-world settings. We tested whether LLMs can assist members of the public in identifying underlying conditions and choosing a course of action (disposition) in ten medical scenarios in a controlled study with 1,298 participants. Participants were randomly assigned to receive assistance from an LLM (GPT-4o, Llama 3, Command R+) or a source of their choice (control). Tested alone, LLMs complete the scenarios accurately, correctly identifying conditions in 94.9% of cases and disposition in 56.3% on average. However, participants using the same LLMs identified relevant conditions in fewer than 34.5% of cases and disposition in fewer than 44.2%, both no better than the control group. We identify user interactions as a challenge to the deployment of LLMs for medical advice. Standard benchmarks for medical knowledge and simulated patient interactions do not predict the failures we find with human participants. Moving forward, we recommend systematic human user testing to evaluate interactive capabilities before public deployments in healthcare.
