{"schema_version":"newruntime-agent-readable-v0.2","type":"post","stable_id":"post:cerebras-moe-router-gradient-null-expert","slug":"cerebras-moe-router-gradient-null-expert","title":"A Balanced MoE Router Can Still Be Functionally Dead","description":"Cerebras demonstrates how top-k normalization can erase the cross-entropy gradient to an MoE router even while expert utilization appears perfectly balanced.","retrieval_nugget":"Cerebras demonstrates how top-k normalization can erase the cross-entropy gradient to an MoE router even while expert utilization appears perfectly balanced. Expert utilization is a necessary MoE metric, but it can conceal a router that is not learning from the task. Cerebras demonstrates the failure on a 124-million-parameter GPT-2-style model with four experts trained on FineWeb.","status":"published","published_at":"2026-08-03","updated_at":"2026-08-03","record_date":"2026-08-03","date_kind":"published_at","topics":["open-models","model-architecture","evals","training"],"source_urls":["https://www.cerebras.ai/blog/moe-guide-debug"],"visuals":[{"id":"cerebras-moe-router-gradient-null-expert","kind":"editorial-diagram","role":"hero","src":"https://newruntime.com/images/posts/cerebras-moe-router-gradient-null-expert.webp","alt":"Hand-drawn contrast showing a top-one mixture-of-experts router with balanced traffic but no task gradient, followed by a phantom comparison expert that restores a learning signal to the router.","caption":"Load balance measures where tokens went; gradient flow reveals whether the router learned which expert was useful.","credit":"New Runtime synthesis from Cerebras","source_url":"https://www.cerebras.ai/blog/moe-guide-debug","generated_with":"gemini-3.1-flash-image","width":1600,"height":900,"legend":[{"label":"Balanced traffic","description":"Four experts can each receive roughly one quarter of tokens while specialization remains absent."},{"label":"Lost gradient","description":"With top-one routing, normalizing one selected gate divides it by itself and removes relative weight."},{"label":"Contrast option","description":"A null expert provides a comparison without adding an active expert computation."},{"label":"Recovered learning","description":"The task-loss gradient reaches the router and restores the expected quality gain."}]}],"routes":{"html":"https://newruntime.com/posts/cerebras-moe-router-gradient-null-expert/","markdown":"https://newruntime.com/posts/cerebras-moe-router-gradient-null-expert.md","json":"https://newruntime.com/posts/cerebras-moe-router-gradient-null-expert.json"},"source_format":"markdown","next_reads":[{"type":"topic","path":"/topics/evals/","reason":"Explore the evals topic hub.","url":"https://newruntime.com/topics/evals/","title":"Agent evals - New Runtime","media_type":"text/html"},{"type":"topic","path":"/topics/open-models/","reason":"Explore the open models topic hub.","url":"https://newruntime.com/topics/open-models/","title":"Open Models - New Runtime","media_type":"text/html"},{"type":"related_material","path":"/posts/fireworks-lora-fullft-three-test-protocol/","reason":"Shares evals and open models.","url":"https://newruntime.com/posts/fireworks-lora-fullft-three-test-protocol/","title":"Run Three Tests Before Replacing LoRA With Full Fine-Tuning","media_type":"text/html"},{"type":"related_material","path":"/posts/agentic-sdlc-software-factory-loop/","reason":"Shares evals.","url":"https://newruntime.com/posts/agentic-sdlc-software-factory-loop/","title":"A Software Factory Connects Agents Through Verified Outcomes","media_type":"text/html"},{"type":"related_material","path":"/posts/claude-code-auto-mode-action-gate/","reason":"Shares evals.","url":"https://newruntime.com/posts/claude-code-auto-mode-action-gate/","title":"Claude Code Auto Mode Gates Actions Instead Of Explanations","media_type":"text/html"}]}
