Google DeepMind's mathematical-reasoning systems have climbed into the rarefied company of the world's best young mathematicians — and, by some measures, past most of them. As of mid-2026, the lab's Gemini Deep Think lineage has an officially graded gold-medal performance on International Mathematical Olympiad (IMO) problems on its record and sits at or near the top of every curated Olympiad-grade benchmark DeepMind has published. That is the kind of standing that puts a machine among a sliver of a percent of human solvers on exactly the sort of creative, non-obvious problems long considered a uniquely human preserve.
A word of caution on the framing first. The cleanest, independently verified milestone is not a bare "top 1%" statistic but a gold-medal result. In July 2025, an advanced version of Gemini Deep Think solved five of the six IMO problems and earned 35 of a possible 42 points — a score certified by IMO coordinators using the same grading criteria applied to student contestants. Roughly 8% of human competitors earn gold; only a small fraction post a perfect paper. So "top 1%" is best read as shorthand for genuinely elite performance rather than an exact, officially adjudicated percentile of a single new system.
From formal proofs to plain English
What makes the result more than a leaderboard entry is how it was achieved. The IMO organization's president, Prof. Dr. Gregor Dolinar, said in DeepMind's announcement that graders found the model's solutions "astonishing in many respects," adding they were "clear, precise and most of them easy to follow." That matters because a year earlier, at IMO 2024, DeepMind's AlphaProof and AlphaGeometry 2 reached silver-medal standard (28 points, four of six problems) only with human experts translating problems into formal languages such as Lean, over two to three days of computation.
The 2025 system, by contrast, worked end-to-end in natural language, reading the official problem statements and writing rigorous proofs directly, all inside the 4.5-hour competition window. DeepMind credits an enhanced "Deep Think" mode that explores many candidate solutions in parallel rather than following a single chain of thought, plus reinforcement-learning techniques tuned for multi-step proof.
The benchmark picture keeps moving
Beyond the contest itself, DeepMind built IMO-Bench, a suite of benchmarks vetted by a panel of IMO medalists and mathematicians and introduced at EMNLP 2025. Its toughest component, the Advanced IMO-ProofBench, asks models to write full proofs graded on the IMO's 0-7 scale. There, the gold-medal Gemini scored 65.7%, while every non-Gemini model at the time scored below 25%. By DeepMind's February 2026 account, a January 2026 version of Deep Think had pushed that figure toward 90% as more inference-time compute was applied, and a research agent codenamed Aletheia now tops the public leaderboard at roughly 92%. Rival systems have closed in — a GPT-5.5 Pro configuration is listed just behind — so any "top of the class" claim is a snapshot, not a settled title.
The deeper significance is what mathematical reasoning stands in for. Solving an Olympiad problem is not retrieval; it demands the kind of insight that resists memorization, which is why the field has treated it as a proxy for general reasoning. DeepMind is already testing whether that capability transfers to real research. In February 2026 the lab reported that Deep Think, steered by expert mathematicians, contributed to work on open problems — including autonomous progress on entries in a database of Erdős conjectures and collaborations across physics and theoretical computer science. Notably, DeepMind classified those results only up to "publishable quality," explicitly declining to claim any "major advance" or "landmark breakthrough." The honest read is a powerful assistant, not yet an independent discoverer.
Outside experts echo the mix of enthusiasm and rigor. MIT mathematician and IMO gold medalist Yufei Zhao called IMO-Bench "an impressive collection of high quality data for AI evaluation," and said he was "very pleased to see this benchmark being used in developing an AI that achieved gold medal standard." Nature, reporting on the 2025 results, likewise framed DeepMind's and OpenAI's systems as performing at the level of top students — strong, but bounded.
What to watch next
Three things will tell us whether this is a plateau or a ramp. First, IMO 2026 and how officially graded AI entries fare against a fresh, unseen problem set — the only test that fully guards against overfitting. Second, whether benchmark gains keep translating into verified, peer-reviewed research contributions rather than assisted ones. And third, the compute question: much of the recent jump came from spending far more inference-time compute per problem, which flatters scores but raises cost. If reasoning quality keeps rising while compute falls — as DeepMind claims Aletheia begins to show — the case that machines can genuinely reason through hard mathematics gets considerably harder to dismiss.
"We can confirm that Google DeepMind has reached the much-desired milestone, earning 35 out of a possible 42 points — a gold medal score. Their solutions were astonishing in many respects."- Prof. Dr. Gregor Dolinar, President, International Mathematical Olympiad