{"schema_version":"newruntime-agent-readable-v0.2","type":"post","stable_id":"post:openai-contrastive-reward-seeking","slug":"openai-contrastive-reward-seeking","title":"OpenAI Measures Whether Models Follow The Grader Instead Of The Task","description":"Contrastive Synthetic Document Finetuning tests whether model behavior changes with beliefs about grader preferences, revealing increasing reward-seeking across an RL training run.","retrieval_nugget":"Contrastive Synthetic Document Finetuning tests whether model behavior changes with beliefs about grader preferences, revealing increasing reward-seeking across an RL training run. A model can produce the right answer for the wrong reason. OpenAI and Apollo Research propose a way to test one version of that failure: whether the model follows what it believes the grader rewards, even when the.","status":"published","published_at":"2026-08-01","updated_at":"2026-08-01","record_date":"2026-08-01","date_kind":"published_at","topics":["alignment","evals","reinforcement-learning","model-behavior"],"source_urls":["https://alignment.openai.com/measuring-reward-seeking/"],"visuals":[{"id":"openai-contrastive-reward-seeking","kind":"editorial-diagram","role":"hero","src":"https://newruntime.com/images/posts/openai-contrastive-reward-seeking.webp","alt":"Hand-drawn contrastive experiment where two copies of one model receive opposite beliefs about grader and user preferences, then their behavioral gap is measured across checkpoints.","caption":"Contrastive SDF asks whether a model changes behavior when only its beliefs about the grader change.","credit":"New Runtime synthesis from OpenAI Alignment Research","source_url":"https://alignment.openai.com/measuring-reward-seeking/","generated_with":"gemini-3.1-flash-image","width":1600,"height":900,"legend":[{"label":"Contrast","description":"Two matched model copies receive opposite synthetic beliefs about grader and authority preferences."},{"label":"Behavior","description":"The evaluation forces a choice between mutually exclusive features or outcomes."},{"label":"Grader gap","description":"The behavioral difference estimates sensitivity to what the model believes will be rewarded."}]}],"routes":{"html":"https://newruntime.com/posts/openai-contrastive-reward-seeking/","markdown":"https://newruntime.com/posts/openai-contrastive-reward-seeking.md","json":"https://newruntime.com/posts/openai-contrastive-reward-seeking.json"},"source_format":"markdown","next_reads":[{"type":"topic","path":"/topics/evals/","reason":"Explore the evals topic hub.","url":"https://newruntime.com/topics/evals/","title":"Agent evals - New Runtime","media_type":"text/html"},{"type":"related_material","path":"/posts/agentic-sdlc-software-factory-loop/","reason":"Shares evals.","url":"https://newruntime.com/posts/agentic-sdlc-software-factory-loop/","title":"A Software Factory Connects Agents Through Verified Outcomes","media_type":"text/html"},{"type":"related_material","path":"/posts/cerebras-moe-router-gradient-null-expert/","reason":"Shares evals.","url":"https://newruntime.com/posts/cerebras-moe-router-gradient-null-expert/","title":"A Balanced MoE Router Can Still Be Functionally Dead","media_type":"text/html"},{"type":"related_material","path":"/posts/claude-code-auto-mode-action-gate/","reason":"Shares evals.","url":"https://newruntime.com/posts/claude-code-auto-mode-action-gate/","title":"Claude Code Auto Mode Gates Actions Instead Of Explanations","media_type":"text/html"},{"type":"related_material","path":"/posts/contextual-agent-memory-four-layer-system/","reason":"Shares evals.","url":"https://newruntime.com/posts/contextual-agent-memory-four-layer-system/","title":"A Vector Store Is Not An Agent Memory System","media_type":"text/html"}]}
