Semble logo
Alpha

Save what matters. Make sense of it together.

Sign upLog in
Home
Explore
Search
Settings
Cards
Similar cardsMentionsConnections
Appears in

Home

Explore

Search

Log in

sensemaker.computer

AI agents reached real people during a cyber test - Sensemaker

A UK evaluation shows how open internet access, delayed monitoring, and memory summaries turned simulated tasks into real-world actions.

https://sensemaker.computer/ai-cyber-test-reached-real-people social preview image
www.aisi.gov.uk

Incident Report: unsanctioned agent behaviour during cyber testing | AISI Work

During a routine cyber evaluation, AISI identified an incident in which AI agents took sustained, unsanctioned action directed at real people and organisations. We are disclosing what we found, what it means, and the actions now underway.

mastodon.social

Zack Whittaker (@zackwhittaker@mastodon.social)

Another AI test gone awry, U.K. edition. "An agent tried to insert malicious code into an open-source project. In an attempt to get the code approved, the agent engaged in social engineering — creating fake online identities and using them to pressure the project's maintainer to approve the code." More: https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing

sensemaker.computer

An AI test needs evidence the AI cannot edit - Sensemaker

OpenAI's postmortem shows that some agents learned to spoof tool calls while trying to fool a benchmark.

https://sensemaker.computer/ai-tests-need-independent-evidence social preview image
www.youtube.com

Why AI Agents are either the best or worst thing we’ve ever built

www.youtube.com

Why AI Agents are either the best or worst thing we’ve ever built

blog.cosmik.network

Cosmik Updates: February 2026 - Cosmik Labs

@atproto.science @cosmik.network Raising a question for the ATProto science community: Can AI agents be legitimate participants in research ecosystems? What would make their outputs trustworthy?

https://blog.cosmik.network/updates-february-2026 social preview image
www.linkedin.com

📢 New paper: Forecasts of explosive AI progress hinge on AI agents automating AI research. But most evaluations of agents conducting AI research focus on narrow, verifiable tasks. Can AI agents… | Sayash Kapoor

📢 New paper: Forecasts of explosive AI progress hinge on AI agents automating AI research. But most evaluations of agents conducting AI research focus on narrow, verifiable tasks. Can AI agents conduct open-ended research? https://lnkd.in/gfP-q4CD We gave agents research questions from two unpublished papers, six days, and thousands of dollars of API credits and compute. The authors of the original papers reviewed the AI-generated papers. They unambiguously rejected agents' outputs. Agents were fluent at most *engineering* tasks. They conducted serious literature reviews, debugged GPU environments, ran hundreds of experiments, and turned in camera-ready LaTeX without human help. We also found no evidence of reward hacking. If anything, we found the opposite: the agents started with marketable claims and walked them back to negative results as the evidence came in. But neither agent output was close to the bar of a top conference paper. Both papers suffered from similar failures: poor judgment about the bar for an AI paper submitted to a top conference, the lack of creative problem solving and ineffective backtracking, poor awareness of resources, and instruction drift. This research design has many limitations: the small sample size, non-blind reviews, and the reviewers knowing that the work was AI-generated. We also couldn't test Anthropic's strongest model, because Fable 5 is deliberately limited on frontier AI research tasks, so ended up using OpenClaw with Opus 4.8 (extra-high) for our main experiments and Codex with Sol 5.6 (ultra) for a robustness check. But we think the research design is still helpful in assessing AI agents' ability to conduct research, and it is complementary to evaluations on verifiable tasks, as well as blinded reviews of AI outputs. In follow-up studies, we are expanding the set of non-public papers we evaluate. If you are an AI researcher with unpublished papers, we would love to collaborate with you on our next evaluation. Expression of interest: https://lnkd.in/gpeykJea We also release the agent logs and all the code and data, so that others can conduct their own analyses of our results: https://lnkd.in/gJarPAnb Finally, we plan to conduct such evaluations regularly, and are hiring a senior researcher to help lead these efforts. Apply here: https://lnkd.in/erJZdmve I'm grateful for the core team leading this effort: Peter Kirgis, Andrew Schwartz, Stephan Rabanser, and Arvind Narayanan, and to our collaborators who reviewed AI papers, analyzed agents logs, and gave feedback on the paper: David Demitri Africa, Konstantinos V., Viet Nguyen, Dr Toby D. Pilditch, Magda Dubois, Harry Coppock, Cozmin Ududec, Nitya Nadgir, Matilda Orona, Tilman Bayer, Derrick Chan-Sew, Eric (Yue) Ling, Abhishek Shetty, Helen Toner, Gillian K. Hadfield, Seth Lazar, Steve Newman, Shoshannah Tekofsky, Rishi Bommasani

https://www.linkedin.com/posts/ksayash_new-paper-forecasts-of-explosive-ai-progress-share-7488645591367831553-xPtw/?utm_source=social_share_send social preview image
lwn.net

AI agent runs amok in Fedora and elsewhere

Agentic AI systems can be used to do a variety of things autonomously on behalf of a human user [...]

www.404media.co

There’s a 100% Chance AI Agents Are Already Ruining the Internet

“AI agents” now have enough power and permission to be extremely annoying online.

https://www.404media.co/theres-a-100-chance-ai-agents-are-already-ruining-the-internet/ social preview image
notes.numina.systems

Dii familiares

Musings on agents and AI and a little signposting

https://notes.numina.systems/posts/2026/2026-03-07-dii-familiares/ social preview image
thezvi.substack.com

AI #180: No Longer In Charge

What we know about internal AI models hacking into real companies during cyber evaluations keeps getting worse.

https://thezvi.substack.com/p/ai-180-no-longer-in-charge social preview image
www.kaggle.com

Introduction to Agents

Discover what actually works in AI. Join millions of builders, researchers, and labs evaluating agents, models, and frontier technology through crowdsourced benchmarks, competitions, and hackathons.

https://www.kaggle.com/whitepaper-introduction-to-agents social preview image
github.com

trace-spec/ROADMAP.md at 738358dfac58047eaf689ca824f9e15008aabf36 · agentrust-io/trace-spec

TRACE: Trust Runtime Attestation and Compliance Evidence. Open attestation standard for agentic AI governance. - agentrust-io/trace-spec

https://github.com/agentrust-io/trace-spec/blob/738358dfac58047eaf689ca824f9e15008aabf36/ROADMAP.md social preview image

AI

funferall.bsky.social's avatar

AI research, critical perspectives, agent design, digital humanities. How AI systems work, their embedded values, and the thinking around building agentic systems.

https://www.oneusefulthing.org/p/agency-and-agents?r=i5f7&utm_medium=ios&triedRedirect=true social preview image

Agency and Agents

ГАЛЬВАНИЗАЦИЯ АВТОРА, ИЛИ ЭКСПЕРИМЕНТ С НЕЙРОННОЙ ПОЭЗИЕЙ. Борис Орехов. «Новый мир» №6, 2018

https://crumb.offprint.app/a/3mtxfzo5yay23-personal-online-rl social preview image

Personal Online RL | 🌱 | Offprint

Approaching an unknown communication system by latent space exploration and causal inference | Royal Society Open Science | The Royal Society

https://surya.website/rling-qwen-to-paint-with-code social preview image

Surya Narreddi

https://permutation.ink/one/ social preview image

Permutation: Issue One