







<span> <p><span>This is the fourth in a series of short reports that help business, education, and policy leaders understand the technical details of working w
Prompting Science Report 4: Playing Pretend: Expert Personas Don't Improve Factual Accuracy
<span> <p><span>This is the fourth in a series of short reports that help business, education, and policy leaders understand the technical details of working w
Prompting Science Report 4: Playing Pretend: Expert Personas Don't Improve Factual Accuracy
<span> <p><span>This is the fourth in a series of short reports that help business, education, and policy leaders understand the technical details of working with AI through rigorous testing. Here, we ask whether assigning personas to models improves performance on difficult objective multiple-choice questions. We study both domain-specific expert personas and low-knowledge personas, evaluating six models on GPQA Diamond (Rein et al. 2024) and MMLU-Pro (Wang et al. 2024), graduate-level questions spanning science, engineering, and law. </span></p> <p><span>We tested three approaches:</span></p> <p><span>• In-Domain Experts: Assigning the model an expert persona (“you are a physics expert”) matched to the problem type (physics problems) had no significant impact on performance (with the exception of the Gemini 2.0 Flash model).</span></p> <p><span> • Off-Domain Experts (Domain-Mismatched): Assigning the model an expert persona (“you are a physics expert”) not matched to the problem type (law problems) resulted in marginal differences.</span></p> <p><span> • Low-Knowledge Personas: We assigned the model negative capability personas (layperson, young child, toddler), which were generally harmful to benchmark accuracy. </span></p> <p><span>Across both benchmarks, persona prompts generally did not improve accuracy relative to a no-persona baseline. Expert personas showed no consistent benefit across models, with few exceptions. Domain-mismatched expert personas sometimes degraded performance. Low-knowledge personas often reduced accuracy. These results are about the accuracy of answers only; personas may serve other purposes (such as altering the tone of outputs), beyond improving factual performance.</span></p></span>
When Nature Calls: The Enshittification of Science and Its Enablers
Proof-of-work papers, policy laundering, and the collapse of self-correction

Avoiding Digital Productivity Traps - Cal Newport
Last week in this newsletter, I summarized some interesting results from a study that analyzed the behavior of 164,000 knowledge workers. It found that introducing ... Read more

Technical Report: Playing Pretend: Expert Personas Don't Improve Factual Accuracy
[vc_row martech_row_background_position=”None” css=”.vc_custom_1742942106856{margin-bottom: 24px !important;}”][vc_column][/vc_column][/vc_row][vc_row martech_row_background_position=”None” css=”.vc_custom_1750970721780{margin-bottom: 40px !important;}”][vc_column width=”5/6″ css=”.vc_custom_1750970738029{margin-bottom: 40px !important;}”][vc_column_text css=”.vc_custom_1764963741477{margin-bottom: 1em !important;}”]This study investigates whether persona prompting improves AI performance on challenging academic benchmarks. We find that despite widespread adoption, assigning expert personas (e.g., “You are a world-class physics expert”) does not reliably improve accuracy. Domain-mismatched…Read More

Howtown
AT Protocol: explanation for non-techies? - Debbie's Blatherings - by Debbie Ridpath Ohi
Working on an explanation without tech jargon, for fellow creatives
Cory Doctorow: The people who tell you ‘AI is changing everything’ are lying
It has become impossible to tell managers mesmerised by artificial intelligence that the tools are not, in fact, helpful. So employees just play along with the fiction to keep their jobs, writes our tech columnist

Push-button science
Technological advances change not only what we can learn as scientists, but also how science is conducted. Here we explore how automation and outsourcing are affecting the act of doing science.
I work, I think? - Annotated
How AI may quietly dismantle the feedback loop that turns inexperienced people into competent ones, and why my work matters to me.
AI, peer review and the human activity of science
When researchers cede their scientific judgement to machines, we lose something important.

Everything is miscellaneous : the power of the new digital disorder
Includes bibliographical references (p. [235]-257) and index; Prologue : information in space -- The new order of order -- Alphabetization and its discontents -- The geography of knowledge -- Lumps and splits -- The laws of the jungle -- Smart leaves -- Social knowing -- What nothing says -- Messiness as a virtue -- The work of knowledge -- Coda : misc; Philosopher Weinberger shows how the digital revolution is radically changing the way we make sense of our lives. Human beings constantly collect, label, and organize data--but today, the shift from the physical to the digital is mixing, burning, and ripping our lives apart. In the past, everything had its one place--the physical world demanded it--but now everything has its places: multiple categories, multiple shelves. Everything is suddenly miscellaneous. Weinberger charts the new principles of digital order that are remaking business, education, politics, science, and culture. He examines how Rand McNally decides what information not to include in a physical map (and why Google Earth is winning that battle), how Staples stores emulate online shopping to increase sales, why your children's teachers will stop having them memorize facts, and how the shift to digital music stands as the model for the future.--From publisher description; From A to Z, Everything Is Miscellaneous will completely reshape the way you think - and what you know - about the world. Includes information on alphabetical order, Amaxon.com, animals, Aristotle, authority, Bettmann Archive, blogs (weblogs), books, broadcasting, British Broadcasting Corporation (BBC), business, card catalog, categories and categorization, clusters, companies, Colon Classification, conversation, Melvil Dewey, Dewey Decimal Classification system, Encyclopaedia Britannica, encyclopedia, essentialism, experts, faceted classification system, first order of order, Flickr.com, Google, Great Books of the Western World, ancient Greeks, health and medical information, identifiers, index, inventory tracking, knowledge, labels, leaf and leaves, libraries, Library of Congress, links, Carolus Linnaeus, lumping and splitting, maps and mapping, marketing, meaning, metadata, multiple listing services (MLS), names of people, neutrality or neutral point of view, New York Public Library, Online Computer Library Center (OCLC), order and organization, people, physical space, everything having place, Plato, race, S.R. Ranganathan, Eleanor Rosch, Joshua Schacter, science, second order of order, simplicity, social constructivism, social knowledge, social networks, sorting, species, standardization, tags, taxonomies, third order of roder, topical categorization, tree, Uniform Product Code (UPC), users, Jimmy Wales, web, Wikipedia, etc

How Google and AI Nearly Made a Seasoned Reporter Spiral — ProPublica
I thought I had missed something major in my reporting. Turns out I had stumbled into an AI-fueled feedback loop that involved a real LLC’s fictional website and a search engine that’s thrusting unreliable answers on users.

Anthropomorphism Is Breaking Our Ability to Judge AI
Tech Policy Press fellow James Ball asks, how should we interact with a technology designed to ‘speak’ with us on what appear to be human terms?

Illusions of Understanding in the Sciences
Scientists seek to understand the causes of observed phenomena. Beliefs that they have succeeded are based on understanding that is rarely or possibly never complete, and varies in depth and quality. Most often scientists believe they understand more than they do, making their belief an illusion. This illusion then persists in explanations scientists provide in print, in talks, or in discussions. The illusion that a scientist has a valid and complete explanation tends to be magnified when the data are well described by mathematical and computer simulation models due to the precision of such models and their ability to predict well; prediction does not imply causality, but gives the illusion that it does. The first part of this essay supports the case for the universality of partial and incomplete levels of understanding by showing the difficulty of reaching a deep level of understanding for even a simple analysis and model that most scientists use and believe they understand: linear regression. The second part highlights some implications of the existence of many levels of understanding and explanation, and their use by scientists for design, testing, analysis, and theory development. It discusses the way that deduction and induction depend on the levels of understanding and the implications of the illusion that a scientist’s understanding is deep. It makes a case that the many incomplete levels of understanding affect, often unwittingly, the ways scientists design experiments, test theories, comprehend, communicate, and teach.

1/ Lots of thoughts about the new White House science and tech policy report, meanwhile just a few quick notes (I only skimmed it so far but plan to read through it)
Science: A New Golden Age
www.whitehouse.gov