Starting from Nemotron 3 Ultra, the authors train two specialist checkpoints using supervised fine-tuning and reinforcement learning. An iterative search generates, verifies and refines candidate proofs, and a separate high-compute stage selects each final submission.

The system scored 30 out of 42 points at IMO 2026, reaching the gold-medal threshold, and operates entirely in natural language, with no formal prover, external tools or internet access.

The release covers the two post-trained checkpoints, the training data, the training and inference code, the submitted solutions and Nemotron-IMO-Bench, a new benchmark of 200 novel olympiad-level problems.