Skip to content

Reinforcement Learning

Topic archive2 matches

Back to homeGEO summary endpoint

2026-09-12

Technology

  • Researchers train Nemotron 3 Ultra checkpoints to generate olympiad math proofs: Researchers trained two specialist checkpoints from Nemotron 3 Ultra using supervised fine-tuning and reinforcement learning to generate natural-language proofs for olympiad mathematics. The resulting test-time-compute pipeline operates entirely in natural language without formal provers or external tools.

    NemotronarXiv

    Permalink
  • AI researchers discuss recursive self-improvement and Chinese labs on podcast: AI researchers John Schulman, Beren Millidge, and Charlie O'Neill participated in a Q&A session on the Dwarkesh Podcast. The discussion covered topics including steelmanning the case against recursive self-improvement (RSI), the progress of Chinese AI labs, and long-horizon reinforcement learning.

    IndustryTechmeme

    Permalink