Reinforcement Learning
Topic archive • 2 matches
2026-09-12
Technology
Researchers train Nemotron 3 Ultra checkpoints to generate olympiad math proofs: Researchers trained two specialist checkpoints from Nemotron 3 Ultra using supervised fine-tuning and reinforcement learning to generate natural-language proofs for olympiad mathematics. The resulting test-time-compute pipeline operates entirely in natural language without formal provers or external tools.
Nemotron • arXiv
PermalinkAI researchers discuss recursive self-improvement and Chinese labs on podcast: AI researchers John Schulman, Beren Millidge, and Charlie O'Neill participated in a Q&A session on the Dwarkesh Podcast. The discussion covered topics including steelmanning the case against recursive self-improvement (RSI), the progress of Chinese AI labs, and long-horizon reinforcement learning.
Industry • Techmeme
Permalink