{"schema_version":"newruntime-agent-readable-v0.2","type":"post","stable_id":"post:dropbox-dspy-relevance-judge","slug":"dropbox-dspy-relevance-judge","title":"Dropbox Uses DSPy To Move A Relevance Judge To Cheaper Models","description":"Dropbox Dash optimized a human-calibrated relevance judge with DSPy, reducing disagreement, shortening model migration, and adding structural-output reliability to the objective.","retrieval_nugget":"Dropbox Dash optimized a human-calibrated relevance judge with DSPy, reducing disagreement, shortening model migration, and adding structural-output reliability to the objective. Dropbox Dash uses an LLM judge to score query-document relevance across ranking, training-data generation, and offline evaluation. Moving that judge to a cheaper model exposed a familiar problem: a prompt tuned for one model did not transfer cleanly to.","status":"published","published_at":"2026-08-01","updated_at":"2026-08-01","record_date":"2026-08-01","date_kind":"published_at","topics":["dspy","evals","llm-judges","model-economics"],"source_urls":["https://dropbox.tech/machine-learning/optimizing-dropbox-dash-relevance-judge-with-dspy","https://github.com/stanfordnlp/dspy"],"visuals":[{"id":"dropbox-dspy-relevance-judge","kind":"editorial-diagram","role":"hero","src":"https://newruntime.com/images/posts/dropbox-dspy-relevance-judge.webp","alt":"Hand-drawn optimization loop where human relevance labels and JSON validity feed DSPy, which adapts a judge prompt for cheaper target models under regression gates.","caption":"Dropbox optimizes judge behavior against human agreement and parseable output, then constrains changes when production stability matters.","credit":"New Runtime synthesis from Dropbox.Tech","source_url":"https://dropbox.tech/machine-learning/optimizing-dropbox-dash-relevance-judge-with-dspy","generated_with":"gemini-3.1-flash-image","width":1600,"height":900,"legend":[{"label":"Objective","description":"Human 1-to-5 ratings, explanations, and malformed-output penalties define success."},{"label":"Optimize","description":"DSPy analyzes disagreement and searches for prompt changes tailored to the target model."},{"label":"Constrain","description":"Guardrails prevent example copying, rating-scale drift, and unsafe full-prompt rewrites."}]}],"telegram_message_id":2910,"telegram_url":"https://t.me/qwgai/2910","telegram_message_ids":[2910,2911],"telegram_delivery_mode":"text_then_media","telegram_media_url":"https://t.me/qwgai/2911","routes":{"html":"https://newruntime.com/posts/dropbox-dspy-relevance-judge/","markdown":"https://newruntime.com/posts/dropbox-dspy-relevance-judge.md","json":"https://newruntime.com/posts/dropbox-dspy-relevance-judge.json"},"source_format":"markdown","next_reads":[{"type":"topic","path":"/topics/evals/","reason":"Explore the evals topic hub.","url":"https://newruntime.com/topics/evals/","title":"Agent evals - New Runtime","media_type":"text/html"},{"type":"related_material","path":"/posts/agentic-sdlc-software-factory-loop/","reason":"Shares evals.","url":"https://newruntime.com/posts/agentic-sdlc-software-factory-loop/","title":"A Software Factory Connects Agents Through Verified Outcomes","media_type":"text/html"},{"type":"related_material","path":"/posts/cerebras-moe-router-gradient-null-expert/","reason":"Shares evals.","url":"https://newruntime.com/posts/cerebras-moe-router-gradient-null-expert/","title":"A Balanced MoE Router Can Still Be Functionally Dead","media_type":"text/html"},{"type":"related_material","path":"/posts/claude-code-auto-mode-action-gate/","reason":"Shares evals.","url":"https://newruntime.com/posts/claude-code-auto-mode-action-gate/","title":"Claude Code Auto Mode Gates Actions Instead Of Explanations","media_type":"text/html"},{"type":"related_material","path":"/posts/contextual-agent-memory-four-layer-system/","reason":"Shares evals.","url":"https://newruntime.com/posts/contextual-agent-memory-four-layer-system/","title":"A Vector Store Is Not An Agent Memory System","media_type":"text/html"}]}
