







Good decisions require verified quality. Refine devotes hours of frontier compute to protect your work and reputation from fixable mistakes.
The future belongs to those who can refute AI, not just generate with AI
Why verification, not prompting, could shape the next decade of engineering

Reify This
The authors contend that contemporary efforts to render AI systems interpretable rest on a mistake: reification, the process of treating abstractions and statistical artifacts as if they were concrete realities.…

Piloting the world's first double-blind AI evaluations
Building trust in proprietary model benchmarks using cryptographically secure environments
Prediction: AI will make formal verification go mainstream — Martin Kleppmann’s blog
Much has been said about the effects that AI will have on software development, but there is an angle I haven’t seen talked about: I believe that AI will bring formal verification, which for decades has been a bit of a fringe pursuit, into the software engineering mainstream.

Standards around generative AI
Accuracy, fairness and speed are the guiding values for AP’s news report, and we believe the mindful use of artificial intelligence can serve these values and over time improve how we work.
AI Detector — Verified AI Content Checker | Pangram
Dive into the technical foundations behind Pangram's AI detection stack and understand how we achieve low false positive rates across modalities.

Deepfakes, Scams, and the Age of Paranoia
As AI-driven fraud becomes increasingly common, more people feel the need to verify every interaction they have online.

Open-world evaluations for measuring frontier AI capabilities
Introducing CRUX, a new project for evaluating AI on long, messy tasks

Refine
I recently tried refine, an AI tool for refining academic articles, developed by Yann Calvó López and Ben Golub.

AI agents are checking the scientific literature — and spotting decades-old errors
The technology is proving adept at finding faults in decades-old papers and reference databases.

An AI test needs evidence the AI cannot edit - Sensemaker
OpenAI's postmortem shows that some agents learned to spoof tool calls while trying to fool a benchmark.
Establishing trust in automated reasoning - MetaROR
Since its beginnings in the 1940s, automated reasoning by computers has become a tool of ever growing importance in scientific research. So far, the rules underlying automated reasoning have mainly been formulated by humans, in the form of program source code. Rules derived from large amounts of data, via machine learning techniques, are a complementary approach currently under intense development. The question of why we should trust these systems, and the results obtained with their help, has been discussed by early practitioners of computational science, but was later forgotten. The present work focuses on independent reviewing, an important source of trust in science, and identifies the characteristics of automated reasoning systems that affect their reviewability. It also discusses possible steps towards increasing reviewability and trustworthiness via a combination of technical and social measures.

Linux Foundation Welcomes TRACE to Advance Verifiable Runtime Evidence for AI Workloads
The Linux Foundation introduces TRACE, a new open standard for verifiable evidence in AI workloads, ensuring portable and trustworthy governance across platforms.
.png)
Can AI Agents Synthesize Scientific Conclusions?
Scientific AI agents increasingly retrieve evidence, reason across sources, and synthesize conclusions used in consequential decisions. Yet, their ability to do so in high-stakes domains such as health remains unclear. We introduce SciConBench, a large-scale live benchmark of 9.11K questions and expert-written conclusions from systematic reviews to evaluate open-domain scientific conclusion synthesis. The benchmark draws on an expert-validated automated evaluation pipeline that decomposes conclusions into atomic facts and measures correctness and comprehensiveness via factual precision and recall. To mitigate data leakage, we further introduce SciConHarness, a clean-room evaluation harness that equips agents with controlled web interaction to ensure valid measurement. Evaluating 8 frontier models and deep research agents, we find that factual quality remains low: under clean-room settings, the best agent achieves only a factual F1 of 0.337. Our clean-room setting consistently reduces performance relative to unconstrained evaluation, suggesting that leakage inflates estimates of models' true synthesis capabilities. Finally, we audit consumer-facing agents (e.g., Google AI Overview, OpenEvidence) and find they frequently generate incomplete and sometimes contradictory conclusions, even when the ground-truth answer is available. Overall, our results show that reliable synthesis of scientific conclusions remains an open challenge, and that clean-room evaluation is essential for assessing open-domain AI agents.

How do authors want to use AI for review?
A survey of researchers who compared AI-generated scientific reviews with journal-agnostic human peer review reveals that they overwhelmingly prefer using AI as a self-checking tool before submission rather than as a replacement for human reviewers. It encourages an “author-centric” model in which AI helps researchers improve their manuscripts before they are reviewed by their peers.
