







Here are 58 words of prompts to GPT-5.6 Pro that got the model to discover that the long-standing Dinitz-Garg-Goemans conjecture is false. Increasingly, prompt crafting is over-rated, ask for what you want. (Which itself can be a hard problem)
Jul 22, 2026 at 6:59 PM
Mathematics with large language models as provers and verifiers
During 2024 and 2025 the discussion about the theorem-proving capabilities of large language models started reporting interesting success stories, mostly to do with difficult exercises (such as problems from the International Mathematical Olympiad), but also with conjectures [Feldman & Karbasi, arXiv:2509.18383v1] formulated for the purpose of verifying whether the artificial intelligence could prove it. In this paper we report a theorem proving feat achieved by ChatGPT by using a protocol involving different prover and verifier instances of the gpt-5 model working collaboratively. To make sure that the produced proofs do not suffer from hallucinations, the final proof is formally verified by the lean proof assistant, and the conformance of premises and conclusion of the lean code is verified by a human. Our methodology is by no means complete or exact. It was nonetheless able to solve five out of six 2025 IMO problems, and close about a third of the sixty-six number theory conjectures in [Cohen, Journal of Integer Sequences, 2025].

Mathematics with large language models as provers and verifiers
During 2024 and 2025 the discussion about the theorem-proving capabilities of large language models started reporting interesting success stories, mostly to do with difficult exercises (such as problems from the International Mathematical Olympiad), but also with conjectures [Feldman & Karbasi, arXiv:2509.18383v1] formulated for the purpose of verifying whether the artificial intelligence could prove it. In this paper we report a theorem proving feat achieved by ChatGPT by using a protocol involving different prover and verifier instances of the gpt-5 model working collaboratively. To make sure that the produced proofs do not suffer from hallucinations, the final proof is formally verified by the lean proof assistant, and the conformance of premises and conclusion of the lean code is verified by a human. Our methodology is by no means complete or exact. It was nonetheless able to solve five out of six 2025 IMO problems, and close about a third of the sixty-six number theory conjectures in [Cohen, Journal of Integer Sequences, 2025].

Mo on Twitter / X
Deciphering OpenAI's misleading claim that GPT-5.2 made a new discovery in theoretical physics. https://t.co/Lysx647aKr pic.twitter.com/xNj08evmXx— Mo (@atmoio) February 14, 2026
clem 🤗 on Twitter / X
The main breakthrough of GPT-5 was to route your messages between a couple of different models to give you the best, cheapest & fastest answer possible.This is cool but imagine if you could do this not only for a couple of models but hundreds of them, big and small, fast and… pic.twitter.com/ww2ApYTN3D— clem 🤗 (@ClementDelangue) October 17, 2025

A recent experience with ChatGPT 5.5 Pro
We are all having to keep revising upwards our assessments of the mathematical capabilities of large language models. I have just made a fairly large revision as a result of ChatGPT 5.5 Pro, to whi…

Nathan Lambert on Twitter / X
Reasoning model reports I recommend reading:2025-01-22 - DeepSeek R1 - https://t.co/FanOPm8R472025-01-22 - Kimi 1.5 - https://t.co/NN8Nr1E2xi2025-03-31 - Open-Reasoner-Zero - https://t.co/H5ycSmzMHk2025-04-10 - Seed-Thinking 1.5 - https://t.co/t1ytgV1ZZK2025-04-30 - Phi-4…— Nathan Lambert (@natolambert) January 2, 2026
Model Release Notes | OpenAI Help Center
We’re beginning the rollout of GPT-5.6 Sol in ChatGPT, our flagship reasoning model for complex work across coding, research, science, cybersecurity, computer use, and design.GPT-5.6 Sol is rolling out to eligible paid ChatGPT plans. Free, Go, and logged-out users are not included. Availability may vary during rollout, and managed-workspace access can depend on administrator settings. Availability for other GPT-5.6 family models varies by product and plan; check the model picker or current rate card in the product you use.


Small Models Have Arrived
For the past few weeks, I've been playing with gpt-5.6-luna. It is shockingly capable, fast, and smart. I regularly see it do ~100 tps, and rip around my codebase, email, and knowledge base.
François Chollet on Twitter / X
GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game.In fact,…— François Chollet (@fchollet) September 3, 2026