{"schema_version":"newruntime-agent-readable-v0.2","type":"post","stable_id":"post:ramp-adaptive-llm-routing","slug":"ramp-adaptive-llm-routing","title":"Ramp Teaches Its Gateway To Route By Failure, Latency, And Cost","description":"Ramp's internal LLM gateway uses failure-aware online learning to reorder model and service-tier candidates, cutting spend without relaxing request deadlines.","retrieval_nugget":"Ramp's internal LLM gateway uses failure-aware online learning to reorder model and service-tier candidates, cutting spend without relaxing request deadlines. Ramp's internal LLM gateway processes trillions of tokens per day. That scale gave the company a useful optimization surface: the gateway can learn which provider, model, and service tier is most likely to satisfy a particular request right now.","status":"published","published_at":"2026-08-01","updated_at":"2026-08-01","record_date":"2026-08-01","date_kind":"published_at","topics":["inference","agent-economics","model-routing","reliability"],"source_urls":["https://builders.ramp.com/post/thompson-sampling-model-routing"],"visuals":[{"id":"ramp-adaptive-llm-routing","kind":"editorial-diagram","role":"hero","src":"https://newruntime.com/images/posts/ramp-adaptive-llm-routing.webp","alt":"Hand-drawn gateway flow where caller preferences enter a failure filter and latency sampler, then produce a cost-aware ordered route with fallbacks.","caption":"Ramp turns model routing into a live control loop over provider failure, request deadlines, and price.","credit":"New Runtime synthesis from Ramp Builders","source_url":"https://builders.ramp.com/post/thompson-sampling-model-routing","generated_with":"gemini-3.1-flash-image","width":1600,"height":900,"legend":[{"label":"Caller intent","description":"The application supplies an ordered list of acceptable models and a latency deadline."},{"label":"Live evidence","description":"Provider failures and latency distributions are updated from recent traffic."},{"label":"Adaptive route","description":"The gateway reorders candidates by bad-outcome probability and relative cost, then falls back when needed."}]}],"telegram_message_id":2882,"telegram_url":"https://t.me/qwgai/2882","telegram_message_ids":[2882,2883],"telegram_delivery_mode":"text_then_media","telegram_media_url":"https://t.me/qwgai/2883","routes":{"html":"https://newruntime.com/posts/ramp-adaptive-llm-routing/","markdown":"https://newruntime.com/posts/ramp-adaptive-llm-routing.md","json":"https://newruntime.com/posts/ramp-adaptive-llm-routing.json"},"source_format":"markdown","next_reads":[{"type":"topic","path":"/topics/agent-economics/","reason":"Explore the agent economics topic hub.","url":"https://newruntime.com/topics/agent-economics/","title":"Agent economics - New Runtime","media_type":"text/html"},{"type":"topic","path":"/topics/inference/","reason":"Explore the inference topic hub.","url":"https://newruntime.com/topics/inference/","title":"Inference - New Runtime","media_type":"text/html"},{"type":"related_material","path":"/posts/llm-costs-operational-risk-controls/","reason":"Shares agent economics and reliability.","url":"https://newruntime.com/posts/llm-costs-operational-risk-controls/","title":"LLM Costs Need Prevention, Detection, And Mitigation","media_type":"text/html"},{"type":"related_material","path":"/posts/openai-gpt-5-6-efficiency-stack/","reason":"Shares agent economics and inference.","url":"https://newruntime.com/posts/openai-gpt-5-6-efficiency-stack/","title":"OpenAI Shows Efficiency Is a Full-Stack Agent Problem","media_type":"text/html"},{"type":"related_material","path":"/posts/chatgpt-agent-loop-efficiency-stack/","reason":"Shares inference.","url":"https://newruntime.com/posts/chatgpt-agent-loop-efficiency-stack/","title":"ChatGPT Cuts Repeated Work Across The Agent Stack","media_type":"text/html"}]}
