News

NVIDIA's open Nemotron models reach gold level at IMO and IOI 2026, and beat the top human coder

NVIDIA's Nemotron 3 systems scored 30 of 42 at IMO 2026 and 535.4 of 600 at IOI 2026, above the best human. Checkpoints, data and code are on Hugging Face.

djcroman news card: NVIDIA's open Nemotron models reach gold level at IMO and IOI 2026

Two of the hardest competitions for young people on the planet test two very different skills. The International Mathematical Olympiad (IMO) asks for rigorous proofs written in plain language. The International Olympiad in Informatics (IOI) asks for algorithms that pass hidden test cases under strict time and submission limits. On 7 October 2026, NVIDIA reported that one family of its own open models, Nemotron 3, reached gold medal level at both events this year. At the IOI it even scored more points than the best human contestant.

The headline numbers are strong, but the more interesting part is what NVIDIA did next. Instead of keeping the method secret, it published the fine tuned checkpoints, the training data, the inference code, the submitted proofs and a new benchmark on Hugging Face. Coming on the same day that OpenAI dropped hundreds of AI written math manuscripts on GitHub, it adds to a clear trend: elite level math and coding is no longer something only a few closed labs can show off.

Bar chart of the 2026 olympiad results: at IOI 2026 Nemotron 3 Ultra CC scored 535.4 of 600 points, above the top human contestant at 498.27 and the gold threshold at 361.12; at IMO 2026 the Nemotron 3 Ultra system scored 30 of 42 points, above the gold threshold of 29

What NVIDIA announced

The announcement is a post on NVIDIA's Hugging Face blog titled "One Model Family, Two Gold Level Results", written by Aleksander Ficek, Igor Gitman, Sean Narenthiran, Mehrzad Samadi and Somshubra Majumdar. It summarises two research papers that appeared on arXiv in September: one on competitive programming and one on olympiad mathematics.

The results in short:

  • IOI 2026: a system called Nemotron 3 Ultra CC (the CC stands for competitive coding) scored 535.4 out of 600 points. The gold medal threshold was 361.12 and the top human contestant scored 498.27.
  • IMO 2026: a system built from three Nemotron 3 Ultra checkpoints scored 30 out of 42 points, with full marks on four of the six problems. The official gold threshold was 29.

Both systems started from the same base, the general Nemotron 3 model family, and were specialised with standard post training methods. NVIDIA's point is that you do not need to build a new foundation model for every challenge. You take a capable base, feed it carefully chosen problems and high quality reasoning, and pair it with a smart loop that generates, checks and improves answers.

The fine print: what counts and what does not

Before anyone writes "AI wins the olympiad", it is worth reading the conditions closely, because NVIDIA is careful about them itself.

The IOI result came from a live run during IOI 2026, under the same time limits, internet restrictions and submission limits as the human contestants. But it was an unofficial and unsupervised benchmark. The AI was not a registered participant and does not appear in the official ranking. So it is fair to say the system outscored the top human on the same problem set, but not that it "won" the IOI. According to the paper, this is the first time an AI system has scored higher than the best human contestant on an IOI problem set.

The IMO result is a bit stronger on the verification side. The proofs the system submitted were graded by official IMO graders, so the 30 points are not a self assessment. The system also worked entirely in natural language, with no formal proof checker such as Lean, no calculator style tools and no internet access. A score of 30 sits just one point above the gold line, so this is gold medal level, not a perfect paper.

That nuance matters. Olympiad scores are now a marketing battleground for AI labs, and "gold" can mean very different things depending on who grades the work and under which rules. NVIDIA spells out both conditions, which makes its claim easier to trust and easier to compare.

How the coding model got to gold

For competitive programming, NVIDIA curated 22,000 problems and generated synthetic reasoning traces, worked examples of how to think through a task step by step. With that data it trained two specialists:

  • Nemotron 3 Nano CC, a mixture of experts model with 30 billion total parameters, of which 3 billion are active for each token. It received supervised fine tuning (SFT) and reinforcement learning (RL).
  • Nemotron 3 Ultra CC, with 550 billion total parameters and 55 billion active. It received only SFT.

The team tested both on last year's IOI 2025 problems, and the progression shows how much each step adds. Nano started at 130 points before post training, reached 280 after SFT and 291 after RL. Then came GenCorrect, NVIDIA's test time strategy that generates many candidate solutions, runs them against tests and refines them over several feedback rounds. With GenCorrect, the small Nano model jumped to 468 points, above the IOI 2025 gold threshold of 438.3. The big Ultra model reached 502 with the same strategy.

Bar chart of IOI 2025 scores out of 600: Nano CC before training 130, after SFT 280, after SFT and RL 291, with GenCorrect 468; gpt-oss-120b with GenCluster 446.75; Ultra CC with GenCorrect 502; a green line marks the gold threshold at 438.3

Two lessons stand out. First, most of the gain came from supervised fine tuning, with RL adding a smaller but consistent improvement on top. Second, scale changes the recipe. For the stronger Ultra model, a single epoch of SFT was already enough to beat the fully trained Nano model on IOI, ICPC problems and LiveCodeBench Pro. That finding shaped the competition version of Ultra CC that then scored 535.4 at IOI 2026.

The jump from 291 to 468 points for the same Nano model is also a reminder of how much "thinking time" matters. The model did not get smarter between those two numbers. It simply got more attempts, better feedback and a system that kept the best ideas. In October 2025, NVIDIA had already shown a similar effect with its GenCluster method, which pushed OpenAI's open weight gpt-oss-120b to 446.75 points on IOI 2025. The new work combines that kind of search with models that were trained for the job.

How the math system learned to prove, check and revise

Olympiad mathematics is a different beast. There are no hidden test cases that tell you whether an answer is right. A proof is either complete and rigorous or it is not, and judging that is itself a hard task. NVIDIA's math team therefore trained the model not only to write proofs but also to criticise them.

Starting from Nemotron 3 Ultra, the team built two specialists. The SFT checkpoint learned from 414,890 quality filtered examples covering 15,818 unique proof problems. Those examples covered four jobs: generating a proof, refining it, verifying it and even meta verification, which means judging whether a verification was correct. The RL checkpoint was trained on 9,597 proof problems that were chosen because they sat right at the edge of what the model could already do.

Both specialists beat the general model in development tests, and they turned out to be good at different things. The SFT checkpoint was strongest in the first search round, while the RL checkpoint gave the best overall result as a single model. So the final system used all three together: the general model and both specialists.

Five step diagram of the IMO pipeline: generate candidate proofs with three Nemotron 3 Ultra checkpoints, score each attempt, write critiques of the weak steps, refine the best attempts over several rounds, and let a separate high compute stage select the proof that gets submitted

For every IMO problem the system generated candidate proofs, scored them, wrote critiques and refined the most promising attempts. A separate stage with a large compute budget then chose the final submission. NVIDIA stresses that neither fine tuning alone nor brute force sampling alone produced the medal. In its words, the results came from co designing the model, the data and the inference loop. Using two complementary checkpoints was worth more than drawing more samples from one of them.

Everything is on Hugging Face

This is where the announcement differs from many earlier olympiad headlines. When Google DeepMind and OpenAI reported gold medal level performance at IMO 2025, they did so with closed systems that outsiders could not download or inspect. NVIDIA has released most of the ingredients for its results:

  • The Nemotron Labs IMO 2026 collection with the two math checkpoints (Nemotron 3 Labs Ultra Math SFT and RL), the general Nemotron 3 Ultra 550B model, and both training datasets.
  • Nemotron IMO Bench, a new benchmark with 200 olympiad level problems.
  • The NeMo Skills repository on GitHub with the IMO inference pipeline, the prompts, the submitted proofs and a quickstart to reproduce the run, plus the IOI evaluation and inference pipeline.
  • The Nemotron 3 Ultra CC coding model on Hugging Face, published in a compressed NVFP4 format.
Overview tiles of the released training material: 414,890 SFT examples for olympiad proofs, 15,818 unique proof problems, 9,597 proof problems for RL, 22,000 curated coding problems, 200 new problems in Nemotron IMO Bench, and two math checkpoints plus a 550B coding model

"Open" has limits here. These are very large models. A 550 billion parameter mixture of experts model is not something you run on a gaming PC, even if only 55 billion parameters are active at a time, and the competition runs used substantial compute on top. NVIDIA admits the training and inference runs were "substantial". The models also come under NVIDIA's own model license, so anyone planning commercial use should read the model card first. Still, researchers and universities now have a complete, documented recipe that they can study, rerun on smaller models or attack with their own critiques.

Why this matters

For NVIDIA, this is a statement about its software, not only its chips. The company sells the hardware almost every AI lab trains on, but it also wants Nemotron to be the default open base for companies that build their own specialists. Showing that one base model can be bent into a world class coder and a world class mathematician is exactly the pitch: bring your domain data, use our recipe, get an expert model.

For the AI race, the results narrow the gap between open and closed systems on the hardest public reasoning tests. A year ago, gold medal claims at these competitions came almost exclusively from closed labs. Now an open model family has matched that level in mathematics and gone beyond the best human score in informatics, with the method published alongside it.

For people who care about math and education, the story has two sides. Tools that can write and check olympiad proofs could become powerful tutors and research assistants. At the same time, the debate around OpenAI's math release today shows that mathematicians want receipts: complete proofs, independent checks and honest reporting of failures. NVIDIA's decision to publish the submitted proofs and have the IMO work graded by official graders is a good example of how such claims should be made.

What to watch next

The open questions are practical ones. Can smaller labs reproduce these scores with the released recipe, or does it only work at NVIDIA's compute scale? How well do the specialists perform on problems that look nothing like competition tasks, such as real research mathematics or messy production code? And will other labs follow with the same level of openness, or keep their olympiad systems closed?

For now, the takeaway is simple. An open model family reached gold medal level in both mathematics and informatics in 2026, the scores come with clear conditions, and the training data and code are there for anyone who wants to check the work.

Sources

Source: huggingface.co

Newsletter

The AI news that matters, in your inbox.