







On selective agency in capable people
Could A.I. Do Your Job? We Put Agents to the Test.
In our experiment, we deployed A.I. “agents” to act as office workers, and found that they were capable of performing some of the tasks we assigned, but not all of them.

Could A.I. Do Your Job? We Put Agents to the Test.
In our experiment, we deployed A.I. “agents” to act as office workers, and found that they were capable of performing some of the tasks we assigned, but not all of them.

You’ve Heard About Who ICE Is Recruiting. The Truth Is Far Worse. I’m the Proof.
What happens when you do minimal screening before hiring agents, arming them, and sending them into the streets? We're all finding out.

The Least Agentic People Alive
The arms race of deferring agency and the dark acquiescence to robotic bureaucracy.

AI Grants
AI for Individual Rights Grant Application Form Use this form to apply for HRF AI for Individual Rights. Please direct questions to ai@hrf.org. *Required Fields *The Human Rights Foundation (“HRF”) is a U.S.-based 501(c)(3) nonprofit organization committed to promoting and protecting human rights and advancing democracy worldwide, with a particular focus on authoritarian regimes. In […]
Skills, forks, and self-surgery: how agent harnesses grow
Claude Code, NanoClaw, and Pi take radically different approaches to harness extensibility. The tradeoff is always safety vs. agent agency.

Agent Design Is Still Hard
My Agent abstractions keep breaking somewhere I don’t expect.

Prompting Science Report 4: Playing Pretend: Expert Personas Don't Improve Factual Accuracy
<span> <p><span>This is the fourth in a series of short reports that help business, education, and policy leaders understand the technical details of working with AI through rigorous testing. Here, we ask whether assigning personas to models improves performance on difficult objective multiple-choice questions. We study both domain-specific expert personas and low-knowledge personas, evaluating six models on GPQA Diamond (Rein et al. 2024) and MMLU-Pro (Wang et al. 2024), graduate-level questions spanning science, engineering, and law. </span></p> <p><span>We tested three approaches:</span></p> <p><span>• In-Domain Experts: Assigning the model an expert persona (“you are a physics expert”) matched to the problem type (physics problems) had no significant impact on performance (with the exception of the Gemini 2.0 Flash model).</span></p> <p><span> • Off-Domain Experts (Domain-Mismatched): Assigning the model an expert persona (“you are a physics expert”) not matched to the problem type (law problems) resulted in marginal differences.</span></p> <p><span> • Low-Knowledge Personas: We assigned the model negative capability personas (layperson, young child, toddler), which were generally harmful to benchmark accuracy. </span></p> <p><span>Across both benchmarks, persona prompts generally did not improve accuracy relative to a no-persona baseline. Expert personas showed no consistent benefit across models, with few exceptions. Domain-mismatched expert personas sometimes degraded performance. Low-knowledge personas often reduced accuracy. These results are about the accuracy of answers only; personas may serve other purposes (such as altering the tone of outputs), beyond improving factual performance.</span></p></span>

Mental models for working with coding agents
Model intelligence sets the ceiling. Your workflow with the agent harness sets what you actually ship.

letta-ai/letta-code
Stateful agents that are like people, with memory, identity, and the ability to learn and adapt
”The distinction between programmer and user is reinforced and maintained by a tech industry that benefits from a population rendered computationally passive. If we accept and adopt the role of less agency, we then make it harder for ourselves to come into more agency.”
always-already-programming.md
gist.github.com