Fine-Tuning Nemotron For Gold-Level Performance In Two Math Competitions
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Fine-Tuning Nemotron For Gold-Level Performance In Two Math Competitions on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get tech for your team delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Hugging Face says specialized systems built from its Nemotron 3 family scored 30 of 42 at the 2026 International Mathematical Olympiad and 535.4 of 600 in an International Olympiad in Informatics run. IMO proofs received official grading; the IOI result was unofficial and excluded from the competition ranking.

Hugging Face says two specialized systems based on its Nemotron 3 model family reached gold-level scores in 2026 international mathematics and programming competitions, as detailed in the original analysis. The IMO system earned 30 of 42 points on written proofs graded by official graders, while a separate IOI system scored 535.4 of 600 in an unofficial run that did not count toward the event’s official ranking.

The two results came from distinct systems and evaluation settings. At the International Olympiad in Informatics (IOI), Hugging Face used a competition-specific version of Nemotron-3-Ultra-CC, post-trained with supervised fine-tuning and paired with GenCorrect, a process that generates candidate code, evaluates it and iteratively refines solutions. The company says the system operated prospectively under the contest’s time, internet-access and submission constraints. Its score was above the reported 361.12-point gold threshold and the top human score of 498.27, but the run was unsupervised and unofficial.

For the International Mathematical Olympiad (IMO), Hugging Face combined its general Nemotron 3 Ultra model with supervised fine-tuning and reinforcement-learning checkpoints. The system generated candidate written proofs, scored and critiqued them, and revised selected attempts. The company reports that official graders awarded the submissions 30 points, above the stated 29-point gold threshold, with full credit on four of six problems. Hugging Face says this system used no formal prover, external tools or internet access.

The training pipelines also differed. The IOI project drew on 22,000 programming problems and synthetic reasoning traces. For IMO, Hugging Face reports that its supervised fine-tuning data contained 414,890 quality-filtered examples from 15,818 proof problems; its reinforcement-learning model was trained on 9,597 problems selected near the model’s capability frontier. These figures and performance claims come from Hugging Face’s account.

At a glance
reportWhen: Results reported for the 2026 competiti…
The developmentHugging Face reported that fine-tuned Nemotron systems achieved scores above the stated gold thresholds in mathematics and programming competitions, with different levels of official validation.
At a glance
reportWhen: Reported after the 2026 competitions
The developmentHugging Face reported that systems fine-tuned from Nemotron 3 scored above the gold thresholds at IOI 2026 and IMO 2026.

Two Tasks, Different Evidence

The report offers evidence that specialized fine-tuning paired with repeated checking and revision can produce strong results on tightly scored tasks. But the two scores should not be treated as equivalent proof of performance: the IMO submissions received official grading, while the IOI score came from a company-reported run outside the official ranking.

The distinction matters to readers comparing AI systems with human contest results. The IMO score reflects official assessment of the submitted proofs in that competition. The IOI result is a promising benchmark claim under reported contest-like constraints, not an official medal or placement. Neither score by itself establishes how well the models handle broader mathematical research, unfamiliar proof styles or everyday programming work.

Hugging Face presents the projects as evidence that a shared base model can be adapted into specialists for different disciplines, rather than requiring a separate foundation model for each. That is the company’s interpretation of its results; independent replication would help establish how broadly the approach works.

Amazon

AI programming problem solver

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From IOI Experiments to Proofs

Hugging Face describes the 2026 work as an extension of its experiments at the 2025 IOI, where it tested post-training and additional computation at inference time. The company reported that a Nemotron-3-Nano-CC system rose from 130 points before post-training to 280 after supervised fine-tuning and 291 after reinforcement learning. With GenCorrect, it reached 468, above the reported 438.3 gold threshold for that year; an Ultra-CC version scored 502 using the same test-time strategy.

The newer projects applied related methods to two different kinds of problems. IOI tasks require executable programs that pass hidden tests, while IMO tasks require rigorous written mathematical arguments. Hugging Face says its IMO work found complementary strengths across supervised fine-tuning and reinforcement-learning checkpoints, prompting the team to combine them with the general model. The reported results therefore reflect both model training and procedures for generating, checking and revising answers.

“Success at both points to something broader.”

— Hugging Face

Independent Checks Still Needed

The IOI result was not part of the official ranking. The supplied account describes a prospective, unsupervised run under competition constraints but does not explain how the run was audited or report independent replication. Its score should not be described as an official IOI medal.

Both results are reported by the organization that developed the systems. The account does not establish how performance would transfer to other competitions, unseen proof styles or practical workloads, nor does it settle how data selection and compute budgets affected the scores. The IMO grading is official for the submitted proofs, but that does not independently validate every broader claim about the method.

Hugging Face says it is releasing IMO checkpoints, datasets and a 200-problem benchmark, but the source material does not provide a complete description of repository details or a release timetable. It also does not say whether the IOI evaluation will receive outside scrutiny.

Checkpoints and Benchmarks

Hugging Face says its Nemotron Labs IMO 2026 collection includes supervised fine-tuning and reinforcement-learning checkpoints, both training datasets, and Nemotron-IMO-Bench, a benchmark of 200 olympiad-level problems. The company also points to a paper describing the IMO training and generate-verify-refine system, as well as a NeMo-Skills repository.

Those materials could allow researchers to inspect the training approach and compare results on additional problems. The next useful evidence would include independent tests of the published benchmark and external scrutiny or replication of the IOI run. Hugging Face has not provided a timetable for further releases in the supplied material.

Key Questions

Did Nemotron officially win gold at both competitions?

No. Hugging Face reports that the IMO proofs received official grading and scored above the stated gold threshold. The IOI score came from an unofficial run and was excluded from the official ranking.

What were the reported scores?

The IMO system scored 30 of 42 points, with full credit on four of six problems. The IOI system scored 535.4 of 600 in the company-reported unofficial run.

How did the systems generate their answers?

The IOI system paired supervised fine-tuning with GenCorrect to generate, evaluate and refine code. The IMO system combined a general model with supervised fine-tuning and reinforcement-learning checkpoints, then generated, critiqued and revised candidate proofs.

Are the results independently verified?

The IMO submissions were graded by official competition graders, according to Hugging Face. The supplied account does not report independent replication of either system’s overall results, and it gives limited information about external auditing of the unofficial IOI run.

Primary source: Hugging Face · via ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

OpenAI’s Approach To Fast-Tracking AI Breakthroughs

OpenAI has published an internal account titled ‘Research acceleration,’ highlighting how AI tools may speed up research processes, but full details are pending.

The NextBigFuture Insight: XAI Grok 4.6 Offers Frontier AI At 85% Reduced Cost

xAI’s Grok 4.6 reportedly offers near-frontier AI performance with 85% reduced costs, though key details and verification are still pending.

Discover The Power Of AI In Imagine Image 2.0 By X.ai

xAI releases Grok Imagine Image 2.0, enhancing image editing, multi-reference support, and text handling, available via Grok’s platform.

A Deep Dive Into Multi-Vector (Late Interaction) Sentence Transformer Models For AI

Exploring Sentence Transformers v6.0’s new MultiVectorEncoder for ColBERT-style late interaction retrieval, supporting text and visual document search.