{"schema_version":"newruntime-agent-readable-v0.2","type":"post","stable_id":"post:pretraining-loss-predicts-rl-returns","slug":"pretraining-loss-predicts-rl-returns","title":"Pretraining Loss Predicts How Much Reasoning RL Can Buy","description":"A controlled chess-to-math study links pretraining quality and token budget to later reinforcement-learning performance instead of treating post-training as an isolated stage.","retrieval_nugget":"A controlled chess-to-math study links pretraining quality and token budget to later reinforcement-learning performance instead of treating post-training as an isolated stage. Reinforcement learning is often discussed as if it creates reasoning after pretraining ends. A new controlled study asks a more operational question: how strongly do pretraining choices determine the returns available from later RL compute?","status":"published","published_at":"2026-07-30","updated_at":"2026-07-30","record_date":"2026-07-30","date_kind":"published_at","topics":["ai-models","research"],"source_urls":["https://arxiv.org/abs/2607.16097"],"visuals":[{"id":"pretraining-to-rl-reasoning-nano-banana","kind":"editorial-diagram","role":"hero","src":"https://newruntime.com/images/posts/pretraining-to-rl-reasoning-nano-banana.webp","alt":"Hand-drawn research timeline from model pretraining through supervised traces and reinforcement learning, with early loss connected to later performance and separate puzzle effects.","caption":"The controlled pipeline links pretraining state to both post-RL performance and learning speed, making the stages one budget problem rather than independent phases.","credit":"New Runtime synthesis from the public pretraining-to-RL study","source_url":"https://arxiv.org/abs/2607.16097","generated_with":"gemini-3.1-flash-image","width":1600,"height":900,"legend":[]}],"telegram_message_id":2873,"telegram_url":"https://t.me/qwgai/2873","telegram_message_ids":[2873],"telegram_delivery_mode":"rich_media","telegram_media_url":"https://t.me/qwgai/2873","routes":{"html":"https://newruntime.com/posts/pretraining-loss-predicts-rl-returns/","markdown":"https://newruntime.com/posts/pretraining-loss-predicts-rl-returns.md","json":"https://newruntime.com/posts/pretraining-loss-predicts-rl-returns.json"},"source_format":"markdown","next_reads":[{"type":"related_material","path":"/posts/abbel-belief-state-memory/","reason":"Shares research.","url":"https://newruntime.com/posts/abbel-belief-state-memory/","title":"ABBEL Treats Memory as an Explicit Belief State","media_type":"text/html"},{"type":"related_material","path":"/posts/llm-inference-backpressure-before-autoscaling/","reason":"Shares ai models.","url":"https://newruntime.com/posts/llm-inference-backpressure-before-autoscaling/","title":"LLM Inference Fails Quietly Before It Fails Loudly","media_type":"text/html"},{"type":"related_material","path":"/posts/unlimited-ocr-long-horizon-parsing/","reason":"Shares ai models.","url":"https://newruntime.com/posts/unlimited-ocr-long-horizon-parsing/","title":"Unlimited OCR Treats a Long Document as One Parsing Horizon","media_type":"text/html"},{"type":"related_material","path":"/posts/agentic-self-improvement-survey/","reason":"Shares research.","url":"https://newruntime.com/posts/agentic-self-improvement-survey/","title":"Agentic Self-Improvement Is Becoming Its Own Field","media_type":"text/html"},{"type":"related_material","path":"/posts/gumclaw-company-operating-system/","reason":"Continue with a related New Runtime material.","url":"https://newruntime.com/posts/gumclaw-company-operating-system/","title":"Gumclaw: An AI Company Operating System","media_type":"text/html"}]}
