AceReason-Nemotron 1.1
Zihan Liu, Zhuoling Yang, Yang Chen, et al.
AceReason-Nemotron 1.1 scales SFT data and uses stage-wise RL with tuned temperatures to achieve state-of-the-art reasoning in 7B models.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Zihan Liu, Zhuoling Yang, Yang Chen, et al.
AceReason-Nemotron 1.1 scales SFT data and uses stage-wise RL with tuned temperatures to achieve state-of-the-art reasoning in 7B models.
Unknown
Large-scale reinforcement learning boosts reasoning in small/mid-sized models by training on math then code prompts.
Unknown
Nemotron CrossThink uses reinforcement learning with multi-domain, verifiable data to improve LLM reasoning across diverse tasks beyond math.
Unknown
OpenMath Nemotron introduces a series of mathematical reasoning models trained on 540K problems and 3.2M solutions, achieving top performance in the AIMO-2 competition.
Unknown
An open-source family of heterogeneous reasoning models (Nano, Super, Ultra) with dynamic reasoning toggle, trained via NAS, distillation, and RL.
Unknown
A bilingual Hindi-English language model built by continuously pre-training Nemotron-Mini 4B on 400B real and synthetic tokens.
Unknown
Nvidia's Nemotron-4 340B models and reward model enable synthetic data generation for training smaller language models, with over 98% of alignment data being synthetic.
Unknown
Nvidia's Nemotron-4 15B is a multilingual language model trained on 8 trillion tokens, achieving strong performance across English, code, and multilingual tasks.