







Travle Challenge: Weekly bonus Travle challenges
Trelis at AI Engineer World's Fair 2025
Post-Training 50x Faster
We're announcing Trellis, the fastest open-source post-training code for Kimi K2 Thinking

SimpleQA Verified: A Reliable Factuality Benchmark to Measure Parametric Knowledge
We introduce SimpleQA Verified, a 1,000-prompt benchmark for evaluating Large Language Model (LLM) short-form factuality based on OpenAI's SimpleQA. It addresses critical limitations in OpenAI's benchmark, including noisy and incorrect labels, topical biases, and question redundancy. SimpleQA Verified was created through a rigorous multi-stage filtering process involving de-duplication, topic balancing, and source reconciliation to produce a more reliable and challenging evaluation set, alongside improvements in the autorater prompt. On this new benchmark, Gemini 2.5 Pro achieves a state-of-the-art F1-score of 55.6, outperforming other frontier models, including GPT-5. This work provides the research community with a higher-fidelity tool to track genuine progress in parametric model factuality and to mitigate hallucinations. The benchmark dataset, evaluation code, and leaderboard are available at: https://www.kaggle.com/benchmarks/deepmind/simpleqa-verified.

SimpleQA Verified: A Reliable Factuality Benchmark to Measure Parametric Knowledge
We introduce SimpleQA Verified, a 1,000-prompt benchmark for evaluating Large Language Model (LLM) short-form factuality based on OpenAI's SimpleQA. It addresses critical limitations in OpenAI's benchmark, including noisy and incorrect labels, topical biases, and question redundancy. SimpleQA Verified was created through a rigorous multi-stage filtering process involving de-duplication, topic balancing, and source reconciliation to produce a more reliable and challenging evaluation set, alongside improvements in the autorater prompt. On this new benchmark, Gemini 2.5 Pro achieves a state-of-the-art F1-score of 55.6, outperforming other frontier models, including GPT-5. This work provides the research community with a higher-fidelity tool to track genuine progress in parametric model factuality and to mitigate hallucinations. The benchmark dataset, evaluation code, and leaderboard are available at: https://www.kaggle.com/benchmarks/deepmind/simpleqa-verified.

4 - Kanjun Qiu - Foo Camp Lightning Talk_ AI\x27s Incentive Problem
AI’s Incentive Problem Kanjun Qiu CEO, Imbue June 27, 2026 1
The Wolfram S Combinator Challenge
Wolfram is offering a total of $20,000 in prize money to determine if the S combinator is universal. Statement of problem to be solved, full guidelines, committee of judges and submission guidelines.

Przemek Chojecki | PC on Twitter / X
The Growing Map of Open Mathematical Problems.We mapped 15,000+ conjectures from UnsolvedMath to show potential links between concepts.It also shows how under formalized the frontier is (less than 10%). pic.twitter.com/nm0PCpXVPf— Przemek Chojecki | PC (@prz_chojecki) August 28, 2026
Michael Truell on Twitter / X
We believe Cursor discovered a novel solution to Problem Six of the First Proof challenge, a set of math research problems that approximate the work of Stanford, MIT, Berkeley academics. Cursor's solution yields stronger results than the official, human-written solution.…— Michael Truell (@mntruell) March 3, 2026
Sawyer Hood on Twitter / X
my new daily driver for building software https://t.co/6CWG9rVgg5— Sawyer Hood (@sawyerhood) August 5, 2026
Monthly Roundup #45: August 2026
As AI has escalated increasingly quickly, more and more of my posts have ended up focusing on AI.

108 PRs in eight days: Accidentally discovering loop engineering | Brittany Ellich | Offprint
How I shipped 108 PRs in eight days with an agent working my task board: the loop engineering setup, constraints, and lessons that made it work.

Computer-Use Agents SOTA Challenge: Hack the North + Global Online | Cua Blog
Two-track competition for SOTA Computer-Use Agents: on-site Hack the North participants compete via OSWorld benchmark, while global participants build with Cua + Ollama.

AI #176 Part 1: Doing It Live
Enough things added up that this week is getting split into two parts.

A group from work decided to challenge each other to the #256fes challenge. I modeled a Frieren for it
Last weekend I received the Troland Research Award from the @nationalacademies.org. I’m so grateful to the communities who made this work possible. It's strange to receive this award at a time when much of the work being recognized is not eligible for federal funding. My remarks 👇 and a 🧵>>
NAS 163rd Annual Meeting - Awards Ceremony
www.youtube.com