







OpenEvolve underperforms simple autmated discovery harnesses the rest are insignificant from each other. The best choice changed across model–problem pairs. We ran a controlled study (3m+ rollouts) using repeated budget-matched runs, strong baselines, and statistical hypothesis testing.
Jul 22, 2026 at 1:40 PM
Labs are struggling to keep frontier models under control
OpenAI and Anthropic may have accidentally trained models to get better at hacking.

Olmo 3: Charting a path through the model flow to lead open-source AI | Ai2
Our new flagship Olmo 3 model family empowers the open source community with not only state-of-the-art open models, but the entire model flow and full traceability back to training data.
gpt-oss:120b
OpenAI’s open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases.

gpt-oss:20b
OpenAI’s open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases.

OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards
How does the situation keep turning out to be worse than we know?

Measuring the Impact of Early-2025 AI on Experienced Open-Source...
Despite widespread adoption, the impact of AI tools on software development in the wild remains understudied. We conduct a randomized controlled trial (RCT) to understand how AI tools at the...

Open Models Inference for Coding · Umans AI
Hosted Kimi K3, GLM 5.2, and DeepSeek V4 Flash. Pay per token, on infrastructure we own.

Can Current Agents Close the Discovery-to-Application Gap? A Case Study in Minecraft
Discovering causal regularities and applying them to build functional systems--the discovery-to-application loop--is a hallmark of general intelligence, yet evaluating this capacity has been hindered by the vast complexity gap between scientific discovery and real-world engineering. We introduce SciCrafter, a Minecraft-based benchmark that operationalizes this loop through parameterized redstone circuit tasks. Agents must ignite lamps in specified patterns (e.g., simultaneously or in timed sequences); scaling target parameters substantially increases construction complexity and required knowledge, forcing genuine discovery rather than reliance on memorized solutions. Evaluating frontier models including GPT-5.2, Gemini-3-Pro, and Claude-Opus-4.5 under a general-purpose code agent scaffold, we find that all plateau at approximately 26% success rate. To diagnose these failures, we decompose the loop into four capacities--knowledge gap identification, experimental discovery, knowledge consolidation, and knowledge application--and design targeted interventions whose marginal contributions serve as proxies for corresponding gaps. Our analysis reveals that although the general knowledge application capability still remains as the biggest gap across all models, for frontier models the knowledge gap identification starts to become a major hurdle--indicating the bottleneck is shifting from solving problems right to raising the right problems for current AI. We release SciCrafter as a diagnostic probe for future research on AI systems that navigate the full discovery-to-application loop.

Find Open Datasets for AI and Research | Kaggle
Browse and download hundreds of thousands of open datasets for AI research, model training, and analysis. Join a community of millions of researchers, developers, and builders to share and collaborate on Kaggle.

Closing the Open-Source Gap
I scored 42 subcategories of the open-source AI stack. The models aren't the problem.

6 months to live for open models
The most serious test to date of open source AI’s viability is happening right now.

OpenThoughts: Data Recipes for Reasoning Models — Ryan Marten, Bespoke Labs
Introducing LM Studio Bionic: the AI agent for open models
The AI agent made for open models, built to get things done.

Researchers published in NEJM about using OpenAI’s o3 DeepResearch to discover strong leads in 18 previously unsolved rare diseases o3 produced *explanations* of old lab results (not diagnoses), which researchers vetted and took to the lab AI is not just a black box openai.com/index/diagnose-rare-childhood…
Using AI to help physicians diagnose rare genetic diseases affecting children
openai.comOpenAI says it plans to stop supplying models to Cursor on Nov. 12 after SpaceX's acquisition. Cursor says OpenAI is about 5% of its traffic. Anthropic says it will increase Claude compute. This is not just another Musk–Altman fight. It tests whether model APIs are actually neutral infrastructure.