Nvidia published research paper arXiv:2609.10712 and released code, checkpoints, data, and benchmark suite
Nvidia releases math AI behind 30-point Olympiad gold
Open blueprints aim to widen access to elite reasoning tools, but heavy computing demands keep full reproduction out of reach for smaller labs.
In a nutshell
Nvidia made public the complete blueprint for an artificial intelligence system that scored a gold-medal-level 30 out of 42 points at the 2026 International Mathematical Olympiad using written natural language instead of formal coding languages. The company released weights, code, and a 200-problem benchmark, but the steep computing hardware needed to train and test the model leaves independent verification out of reach for most academic labs.
Highlights
- Nvidia's Nemotron 3 Ultra scored 30 out of 42 points, clearing the 29-point gold medal threshold at the 2026 International Mathematical Olympiad.
- The reasoning pipeline generates, critiques, and refines proofs entirely in human language rather than specialized logic code.
- The released Nemotron-IMO-Bench provides researchers with 200 new Olympiad problems to test automated reasoning.
- Training the system required a cluster of 512 GB200 graphics processors.
2026 IMO Competition Score Thresholds
pointsThe findingNvidia's system cleared the gold medal standard by one point.
- How the score achieved by Nvidia's AI system compares with the gold medal threshold and the total possible points at the competition.
- Cutoff — Minimum score required to earn a gold medal at the 2026 International Mathematical Olympiad
- Nemotron — Score achieved by Nvidia's Nemotron 3 Ultra system using natural-language proofs
- Maximum — Highest possible total score across all competition problems
- Shows automated natural-language reasoning can reach top-tier student competition levels, though hardware costs continue to constrain independent verification.
From the Editor’s Diary
Open-sourcing complex model architectures does not guarantee democratic access when verifying the results requires supercomputing infrastructure.
Who's involved
Nvidia
U.S. designer of computing hardware and artificial intelligence systems
goal → Proving advanced multi-step reasoning can work without rigid coding provers
Ivan Moshkov
Lead researcher and paper author at Nvidia
goal → Distributing a public, step-by-step training system for elite mathematics
Open-Source AI Research Community
Independent machine learning laboratories and academic scholars
goal → Testing the natural-language verification method on broader scientific tasks
International Mathematical Olympiad
Global annual mathematics championship for high school students
goal → Serving as the benchmark for testing machine problem-solving
In short
Nvidia has opened up the design of an artificial intelligence system that earned a gold medal score in top-tier high school mathematics, moving automated reasoning into plain language.
The release makes it likely that research laboratories will try to adapt the conversational verification method to complex scientific benchmarks.
Whether that transition succeeds remains uncertain, because testing and training the setup requires massive computing power that few outside industry leaders possess.
How it unfolded
Nvidia Publishes Code and Paper for Math System
On Sept. 9, 2026, an Nvidia research team led by Ivan Moshkov released a technical paper detailing its Olympiad system alongside code on GitHub and model files on the Hugging Face repository.
Release of Model Checkpoints and Benchmarks
Industry outlets quickly highlighted the public distribution of the NeMo-Skills testing pipeline and NeMo-RL reinforcement learning recipes, which train software through trial and feedback.
Reproduction Hurdles Surface Across Research Community
Scrutiny turned toward implementation barriers, as independent researchers found that matching the 512-processor training scale and heavy test-time compute exceeded typical academic budgets.
Where things stand
Nvidia has fully published the paper, training data, fine-tuning recipes, test scripts, weights, and evaluation dataset. Human graders awarded the system 30 out of 42 points under official competition scoring rules, without external internet access or symbolic proof engines.
Full independent reproduction of the training run remains unconfirmed because the initial stage required 512 GB200 graphics processors alongside substantial power to generate test-time solutions. Research groups are now examining whether this iterative language critique transfers to other complex scientific disciplines.