{"schema_version":"newruntime-topic-hub-v0.2","type":"topic_hub","stable_id":"topic_hub:inference","slug":"inference","title":"Inference - New Runtime","description":"A New Runtime topic hub collecting signals, patterns, field notes, and public sources about inference.","retrieval_nugget":"A New Runtime topic hub collecting signals, patterns, field notes, and public sources about inference. Inference is tracked here as an evidence-linked topic, not as a static glossary entry. The page connects raw observations to pattern hypotheses, longer analysis, and public sources. Use it as the canonical landing page before drilling into individual records.","answer":["Inference is tracked here as an evidence-linked topic, not as a static glossary entry.","The page connects raw observations to pattern hypotheses, longer analysis, and public sources.","Use it as the canonical landing page before drilling into individual records."],"search_intents":["inference","inference AI agents","inference software"],"status":"featured","last_updated":"2026-08-03","record_date":"2026-08-03","date_kind":"last_updated","counts":{"total":37,"signals":31,"patterns":0,"posts":6,"atlas":0,"sources":63},"routes":{"html":"https://newruntime.com/topics/inference/","markdown":"https://newruntime.com/topics/inference.md","json":"https://newruntime.com/topics/inference.json"},"source_urls":["https://ai.google.dev/gemini-api/docs/gemini-3","https://aiechoes.substack.com/p/building-a-biomedical-graphrag-when","https://anthropic.com/news/claude-sonnet-4-6","https://app.klingai.com/global/release-notes/vaxrndo66h?type=dialog","https://arxiv.org/pdf/2501.07572","https://arxiv.org/pdf/2501.14249","https://blog.bytebytego.com/p/how-chatgpt-optimizes-its-agent-loop","https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4","https://blog.google/technology/google-labs/opal-expansion-160","https://blog.nilenso.com/blog/2025/11/04/a-short-lesson-in-simpler-prompts","https://builders.ramp.com/post/thompson-sampling-model-routing","https://cdn.openai.com/pdf/5e10f4ab-d6f7-442e-9508-59515c65e35d/browsecomp.pdf"],"top_sources":["https://ai.google.dev/gemini-api/docs/gemini-3","https://aiechoes.substack.com/p/building-a-biomedical-graphrag-when","https://anthropic.com/news/claude-sonnet-4-6","https://app.klingai.com/global/release-notes/vaxrndo66h?type=dialog","https://arxiv.org/pdf/2501.07572","https://arxiv.org/pdf/2501.14249","https://blog.bytebytego.com/p/how-chatgpt-optimizes-its-agent-loop","https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4","https://blog.google/technology/google-labs/opal-expansion-160","https://blog.nilenso.com/blog/2025/11/04/a-short-lesson-in-simpler-prompts","https://builders.ramp.com/post/thompson-sampling-model-routing","https://cdn.openai.com/pdf/5e10f4ab-d6f7-442e-9508-59515c65e35d/browsecomp.pdf"],"evidence_records":[{"kind":"Field Note","stable_id":"post:chatgpt-agent-loop-efficiency-stack","slug":"chatgpt-agent-loop-efficiency-stack","title":"ChatGPT Cuts Repeated Work Across The Agent Stack","date":"2026-08-03","record_date":"2026-08-03","date_kind":"updated_or_published_at","source_count":1,"url":"https://newruntime.com/posts/chatgpt-agent-loop-efficiency-stack/"},{"kind":"Field Note","stable_id":"post:fireworks-lora-fullft-three-test-protocol","slug":"fireworks-lora-fullft-three-test-protocol","title":"Run Three Tests Before Replacing LoRA With Full Fine-Tuning","date":"2026-08-03","record_date":"2026-08-03","date_kind":"updated_or_published_at","source_count":1,"url":"https://newruntime.com/posts/fireworks-lora-fullft-three-test-protocol/"},{"kind":"Field Note","stable_id":"post:openai-gpt-5-6-price-performance-frontier","slug":"openai-gpt-5-6-price-performance-frontier","title":"GPT-5.6 Turns Efficiency Work Into API Economics","date":"2026-08-01","record_date":"2026-08-01","date_kind":"updated_or_published_at","source_count":1,"url":"https://newruntime.com/posts/openai-gpt-5-6-price-performance-frontier/"},{"kind":"Field Note","stable_id":"post:ramp-adaptive-llm-routing","slug":"ramp-adaptive-llm-routing","title":"Ramp Teaches Its Gateway To Route By Failure, Latency, And Cost","date":"2026-08-01","record_date":"2026-08-01","date_kind":"updated_or_published_at","source_count":1,"url":"https://newruntime.com/posts/ramp-adaptive-llm-routing/"},{"kind":"Field Note","stable_id":"post:vercel-ai-gateway-operating-budget","slug":"vercel-ai-gateway-operating-budget","title":"Vercel AI Gateway Adds Runtime Budget Controls","date":"2026-08-01","record_date":"2026-08-01","date_kind":"updated_or_published_at","source_count":4,"url":"https://newruntime.com/posts/vercel-ai-gateway-operating-budget/"},{"kind":"Field Note","stable_id":"post:openai-gpt-5-6-efficiency-stack","slug":"openai-gpt-5-6-efficiency-stack","title":"OpenAI Shows Efficiency Is a Full-Stack Agent Problem","date":"2026-07-30","record_date":"2026-07-30","date_kind":"updated_or_published_at","source_count":3,"url":"https://newruntime.com/posts/openai-gpt-5-6-efficiency-stack/"},{"kind":"Raw Signal","stable_id":"signal:colibri-runs-a-744b-model-in-25-gb-of-memory","slug":"colibri-runs-a-744b-model-in-25-gb-of-memory","title":"Colibri Runs a 744B Model in 25 GB of Memory","date":"2026-07-15","record_date":"2026-07-15","date_kind":"observed_at","source_count":2,"url":"https://newruntime.com/signals/colibri-runs-a-744b-model-in-25-gb-of-memory/"},{"kind":"Raw Signal","stable_id":"signal:inference-gateway-makes-model-switching-operational","slug":"inference-gateway-makes-model-switching-operational","title":"An inference gateway makes model switching operational","date":"2026-07-02","record_date":"2026-07-02","date_kind":"observed_at","source_count":1,"url":"https://newruntime.com/signals/inference-gateway-makes-model-switching-operational/"},{"kind":"Raw Signal","stable_id":"signal:sakana-fugu-presents-multiple-agents-as-one-model-endpoint","slug":"sakana-fugu-presents-multiple-agents-as-one-model-endpoint","title":"Sakana Fugu Presents Multiple Agents as One Model Endpoint","date":"2026-06-25","record_date":"2026-06-25","date_kind":"observed_at","source_count":1,"url":"https://newruntime.com/signals/sakana-fugu-presents-multiple-agents-as-one-model-endpoint/"},{"kind":"Raw Signal","stable_id":"signal:github-yichuan-w-leann-routable-model-components","slug":"github-yichuan-w-leann-routable-model-components","title":"GitHub / yichuan-w/LEANN: Routable Model Components","date":"2026-05-22","record_date":"2026-05-22","date_kind":"observed_at","source_count":1,"url":"https://newruntime.com/signals/github-yichuan-w-leann-routable-model-components/"},{"kind":"Raw Signal","stable_id":"signal:github-chenglou-pretext-executable-design-context","slug":"github-chenglou-pretext-executable-design-context","title":"GitHub / chenglou/pretext: Executable Design Context","date":"2026-04-30","record_date":"2026-04-30","date_kind":"observed_at","source_count":3,"url":"https://newruntime.com/signals/github-chenglou-pretext-executable-design-context/"},{"kind":"Raw Signal","stable_id":"signal:openai-routable-model-components","slug":"openai-routable-model-components","title":"OpenAI: Routable Model Components","date":"2026-04-24","record_date":"2026-04-24","date_kind":"observed_at","source_count":1,"url":"https://newruntime.com/signals/openai-routable-model-components/"}],"patterns":[],"field_notes":[{"kind":"Field Note","stable_id":"post:chatgpt-agent-loop-efficiency-stack","slug":"chatgpt-agent-loop-efficiency-stack","title":"ChatGPT Cuts Repeated Work Across The Agent Stack","description":"ByteByteGo's OpenAI engineering walkthrough connects persistent sessions, stable prompt prefixes, deferred tools, delta tokenization, cache-aware routing, and split inference.","date":"2026-08-03","record_date":"2026-08-03","date_kind":"updated_or_published_at","topics":["agent-harness","context-engineering","inference","performance"],"source_count":1,"url":"https://newruntime.com/posts/chatgpt-agent-loop-efficiency-stack/"},{"kind":"Field Note","stable_id":"post:fireworks-lora-fullft-three-test-protocol","slug":"fireworks-lora-fullft-three-test-protocol","title":"Run Three Tests Before Replacing LoRA With Full Fine-Tuning","description":"Fireworks shows how data coverage, optimization, and adapter capacity can create or close an apparent quality gap between LoRA and full fine-tuning.","date":"2026-08-03","record_date":"2026-08-03","date_kind":"updated_or_published_at","topics":["evals","inference","open-models","post-training"],"source_count":1,"url":"https://newruntime.com/posts/fireworks-lora-fullft-three-test-protocol/"},{"kind":"Field Note","stable_id":"post:openai-gpt-5-6-price-performance-frontier","slug":"openai-gpt-5-6-price-performance-frontier","title":"GPT-5.6 Turns Efficiency Work Into API Economics","description":"OpenAI turned GPT-5.6 serving and kernel efficiency gains into lower Luna and Terra prices, plus a faster Sol mode for latency-sensitive API workloads.","date":"2026-08-01","record_date":"2026-08-01","date_kind":"updated_or_published_at","topics":["coding-agents","enterprise-ai","inference","model-economics"],"source_count":1,"url":"https://newruntime.com/posts/openai-gpt-5-6-price-performance-frontier/"},{"kind":"Field Note","stable_id":"post:ramp-adaptive-llm-routing","slug":"ramp-adaptive-llm-routing","title":"Ramp Teaches Its Gateway To Route By Failure, Latency, And Cost","description":"Ramp's internal LLM gateway uses failure-aware online learning to reorder model and service-tier candidates, cutting spend without relaxing request deadlines.","date":"2026-08-01","record_date":"2026-08-01","date_kind":"updated_or_published_at","topics":["agent-economics","inference","model-routing","reliability"],"source_count":1,"url":"https://newruntime.com/posts/ramp-adaptive-llm-routing/"},{"kind":"Field Note","stable_id":"post:vercel-ai-gateway-operating-budget","slug":"vercel-ai-gateway-operating-budget","title":"Vercel AI Gateway Adds Runtime Budget Controls","description":"Vercel's July 31 AI Gateway releases combine team and project spend budgets, unified fast mode, Laguna S 2.1 capacity, and updated MCP support into a practical inference control layer.","date":"2026-08-01","record_date":"2026-08-01","date_kind":"updated_or_published_at","topics":["agents","api-design","developer-tools","inference"],"source_count":4,"url":"https://newruntime.com/posts/vercel-ai-gateway-operating-budget/"},{"kind":"Field Note","stable_id":"post:openai-gpt-5-6-efficiency-stack","slug":"openai-gpt-5-6-efficiency-stack","title":"OpenAI Shows Efficiency Is a Full-Stack Agent Problem","description":"OpenAI's GPT-5.6 efficiency write-up connects model training, inference optimization, and the Codex/ChatGPT Work harness into one compounding cost-performance loop.","date":"2026-07-30","record_date":"2026-07-30","date_kind":"updated_or_published_at","topics":["agent-economics","agent-harness","context-engineering","inference","models"],"source_count":3,"url":"https://newruntime.com/posts/openai-gpt-5-6-efficiency-stack/"}],"raw_signals":[{"kind":"Raw Signal","stable_id":"signal:colibri-runs-a-744b-model-in-25-gb-of-memory","slug":"colibri-runs-a-744b-model-in-25-gb-of-memory","title":"Colibri Runs a 744B Model in 25 GB of Memory","description":"An experimental C inference engine uses aggressive storage and memory techniques to make a very large GLM model runnable without a GPU.","date":"2026-07-15","record_date":"2026-07-15","date_kind":"observed_at","topics":["efficiency","inference","local-models"],"source_count":2,"metric":"notable","url":"https://newruntime.com/signals/colibri-runs-a-744b-model-in-25-gb-of-memory/"},{"kind":"Raw Signal","stable_id":"signal:inference-gateway-makes-model-switching-operational","slug":"inference-gateway-makes-model-switching-operational","title":"An inference gateway makes model switching operational","description":"Inference.net exposes model routing through a stable gateway so clients can change providers without rewriting every integration.","date":"2026-07-02","record_date":"2026-07-02","date_kind":"observed_at","topics":["agent-infrastructure","inference","model-routing"],"source_count":1,"metric":"notable","url":"https://newruntime.com/signals/inference-gateway-makes-model-switching-operational/"},{"kind":"Raw Signal","stable_id":"signal:sakana-fugu-presents-multiple-agents-as-one-model-endpoint","slug":"sakana-fugu-presents-multiple-agents-as-one-model-endpoint","title":"Sakana Fugu Presents Multiple Agents as One Model Endpoint","description":"Fugu hides a multi-agent debate and synthesis system behind a model-like interface, making orchestration an implementation detail of inference.","date":"2026-06-25","record_date":"2026-06-25","date_kind":"observed_at","topics":["inference","model-routing","multi-agent"],"source_count":1,"metric":"notable","url":"https://newruntime.com/signals/sakana-fugu-presents-multiple-agents-as-one-model-endpoint/"},{"kind":"Raw Signal","stable_id":"signal:github-yichuan-w-leann-routable-model-components","slug":"github-yichuan-w-leann-routable-model-components","title":"GitHub / yichuan-w/LEANN: Routable Model Components","description":"The archive captures GitHub / yichuan-w/LEANN as a dated public record from GitHub / yichuan-w/LEANN. It documents models becoming replaceable or specialized components inside a more durable runtime and is retained as supporting evidence for the routable model components trend.","date":"2026-05-22","record_date":"2026-05-22","date_kind":"observed_at","topics":["agent-memory","agent-protocols","context-engineering","inference","local-models","model-routing","retrieval"],"source_count":1,"metric":"notable","url":"https://newruntime.com/signals/github-yichuan-w-leann-routable-model-components/"},{"kind":"Raw Signal","stable_id":"signal:github-chenglou-pretext-executable-design-context","slug":"github-chenglou-pretext-executable-design-context","title":"GitHub / chenglou/pretext: Executable Design Context","description":"The archive captures GitHub / chenglou/pretext as a dated public record from GitHub / chenglou/pretext. It documents design systems becoming machine-readable context, constraints, and review loops and is retained as supporting evidence for the executable design context trend.","date":"2026-04-30","record_date":"2026-04-30","date_kind":"observed_at","topics":["agent-context","agent-memory","design-systems","frontend","inference","local-models","model-routing"],"source_count":3,"metric":"notable","url":"https://newruntime.com/signals/github-chenglou-pretext-executable-design-context/"},{"kind":"Raw Signal","stable_id":"signal:openai-routable-model-components","slug":"openai-routable-model-components","title":"OpenAI: Routable Model Components","description":"The archive captures OpenAI as a dated public record from OpenAI. It documents models becoming replaceable or specialized components inside a more durable runtime and is retained as supporting evidence for the routable model components trend.","date":"2026-04-24","record_date":"2026-04-24","date_kind":"observed_at","topics":["agent-interfaces","agent-runtime","agent-tools","agents","inference","local-models","model-routing"],"source_count":1,"metric":"notable","url":"https://newruntime.com/signals/openai-routable-model-components/"},{"kind":"Raw Signal","stable_id":"signal:google-developers-tools-routable-model-components","slug":"google-developers-tools-routable-model-components","title":"Google / Developers Tools: Routable Model Components","description":"The archive captures Google / Developers Tools as a dated public record from Google / Developers Tools. It documents models becoming replaceable or specialized components inside a more durable runtime and is retained as supporting evidence for the routable model components trend.","date":"2026-04-04","record_date":"2026-04-04","date_kind":"observed_at","topics":["agent-harness","coding-agents","google","inference","local-models","model-routing","skills"],"source_count":2,"metric":"notable","url":"https://newruntime.com/signals/google-developers-tools-routable-model-components/"},{"kind":"Raw Signal","stable_id":"signal:google-developers-bring-state-of-art-agentic-skills-to-ai-native-operating-models","slug":"google-developers-bring-state-of-art-agentic-skills-to-ai-native-operating-models","title":"Google Developers / Bring State Of Art Agentic Skills To: Ai-Native Operating Models","description":"The archive captures Google Developers / Bring State Of Art Agentic Skills To as a dated public record from Google Developers / Bring State Of Art Agentic Skills To. It documents AI adoption shifting jobs, coordination, review, and organizational capacity and is retained as supporting evidence for the AI-native operating models trend.","date":"2026-04-03","record_date":"2026-04-03","date_kind":"observed_at","topics":["ai-adoption","future-of-work","google","inference","local-models","model-routing","org-design"],"source_count":1,"metric":"notable","url":"https://newruntime.com/signals/google-developers-bring-state-of-art-agentic-skills-to-ai-native-operating-models/"},{"kind":"Raw Signal","stable_id":"signal:anthropic-claude-sonnet-routable-model-components","slug":"anthropic-claude-sonnet-routable-model-components","title":"Anthropic / Claude Sonnet: Routable Model Components","description":"The archive captures Anthropic / Claude Sonnet as a dated public record from Anthropic / Claude Sonnet. It documents models becoming replaceable or specialized components inside a more durable runtime and is retained as supporting evidence for the routable model components trend.","date":"2026-02-18","record_date":"2026-02-18","date_kind":"observed_at","topics":["agent-memory","agent-protocols","inference","interoperability","local-models","mcp","model-routing"],"source_count":2,"metric":"notable","url":"https://newruntime.com/signals/anthropic-claude-sonnet-routable-model-components/"},{"kind":"Raw Signal","stable_id":"signal:github-cloudflare-moltworker-agent-security-runtime-boundaries","slug":"github-cloudflare-moltworker-agent-security-runtime-boundaries","title":"GitHub / cloudflare/moltworker: Agent Security Runtime Boundaries","description":"The archive captures GitHub / cloudflare/moltworker as a dated public record from GitHub / cloudflare/moltworker. It documents agent security expanding from prompt policy into memory, tools, sandboxes, and execution boundaries and is retained as branch-opening evidence for the agent security runtime boundaries trend.","date":"2026-02-04","record_date":"2026-02-04","date_kind":"observed_at","topics":["agent-security","agents","inference","local-models","model-routing","prompt-injection","sandbox"],"source_count":3,"metric":"structural","url":"https://newruntime.com/signals/github-cloudflare-moltworker-agent-security-runtime-boundaries/"},{"kind":"Raw Signal","stable_id":"signal:github-tobi-qmd-routable-model-components","slug":"github-tobi-qmd-routable-model-components","title":"GitHub / tobi/qmd: Routable Model Components","description":"The archive captures GitHub / tobi/qmd as a dated public record from GitHub / tobi/qmd. It documents models becoming replaceable or specialized components inside a more durable runtime and is retained as pressure-testing evidence for the routable model components trend.","date":"2026-02-04","record_date":"2026-02-04","date_kind":"observed_at","topics":["agent-memory","agent-protocols","context-engineering","inference","local-models","model-routing","retrieval"],"source_count":1,"metric":"notable","url":"https://newruntime.com/signals/github-tobi-qmd-routable-model-components/"},{"kind":"Raw Signal","stable_id":"signal:simonwillison-routable-model-components","slug":"simonwillison-routable-model-components","title":"Simonwillison: Routable Model Components","description":"The archive captures Simonwillison as a dated public record from Simonwillison. It documents models becoming replaceable or specialized components inside a more durable runtime and is retained as supporting evidence for the routable model components trend.","date":"2026-02-03","record_date":"2026-02-03","date_kind":"observed_at","topics":["agent-harness","agents","coding-agents","inference","local-models","model-routing","skills"],"source_count":4,"metric":"notable","url":"https://newruntime.com/signals/simonwillison-routable-model-components/"},{"kind":"Raw Signal","stable_id":"signal:github-openclaw-openclaw-routable-model-components","slug":"github-openclaw-openclaw-routable-model-components","title":"GitHub / openclaw/openclaw: Routable Model Components","description":"The archive captures GitHub / openclaw/openclaw as a dated public record from GitHub / openclaw/openclaw. It documents models becoming replaceable or specialized components inside a more durable runtime and is retained as pressure-testing evidence for the routable model components trend.","date":"2026-02-02","record_date":"2026-02-02","date_kind":"observed_at","topics":["agent-interfaces","agent-runtime","agent-tools","agents","inference","local-models","model-routing"],"source_count":4,"metric":"notable","url":"https://newruntime.com/signals/github-openclaw-openclaw-routable-model-components/"},{"kind":"Raw Signal","stable_id":"signal:google-agent-security-runtime-boundaries","slug":"google-agent-security-runtime-boundaries","title":"Google: Agent Security Runtime Boundaries","description":"The archive captures X source as a dated public record from X source. It documents agent security expanding from prompt policy into memory, tools, sandboxes, and execution boundaries and is retained as branch-opening evidence for the agent security runtime boundaries trend.","date":"2026-01-31","record_date":"2026-01-31","date_kind":"observed_at","topics":["agent-security","agents","inference","local-models","model-routing","prompt-injection","sandbox"],"source_count":1,"metric":"structural","url":"https://newruntime.com/signals/google-agent-security-runtime-boundaries/"},{"kind":"Raw Signal","stable_id":"signal:docs-agent-security-runtime-boundaries","slug":"docs-agent-security-runtime-boundaries","title":"Docs: Agent Security Runtime Boundaries","description":"The archive captures Docs as a dated public record from Docs. It documents agent security expanding from prompt policy into memory, tools, sandboxes, and execution boundaries and is retained as branch-opening evidence for the agent security runtime boundaries trend.","date":"2026-01-28","record_date":"2026-01-28","date_kind":"observed_at","topics":["agent-security","agents","inference","local-models","model-routing","prompt-injection","sandbox"],"source_count":1,"metric":"structural","url":"https://newruntime.com/signals/docs-agent-security-runtime-boundaries/"},{"kind":"Raw Signal","stable_id":"signal:github-camel-ai-camel-routable-model-components","slug":"github-camel-ai-camel-routable-model-components","title":"GitHub / camel-ai/camel: Routable Model Components","description":"The archive captures GitHub / camel-ai/camel as a dated public record from GitHub / camel-ai/camel. It documents models becoming replaceable or specialized components inside a more durable runtime and is retained as supporting evidence for the routable model components trend.","date":"2026-01-20","record_date":"2026-01-20","date_kind":"observed_at","topics":["agent-protocols","ai-adoption","future-of-work","inference","local-models","model-routing","org-design"],"source_count":2,"metric":"notable","url":"https://newruntime.com/signals/github-camel-ai-camel-routable-model-components/"},{"kind":"Raw Signal","stable_id":"signal:github-owner-repo-routable-model-components","slug":"github-owner-repo-routable-model-components","title":"GitHub / owner/repo: Routable Model Components","description":"The archive captures GitHub / owner/repo as a dated public record from GitHub / owner/repo. It documents models becoming replaceable or specialized components inside a more durable runtime and is retained as supporting evidence for the routable model components trend.","date":"2026-01-15","record_date":"2026-01-15","date_kind":"observed_at","topics":["agent-harness","agent-interfaces","computer-use","generative-ui","inference","local-models","model-routing"],"source_count":2,"metric":"notable","url":"https://newruntime.com/signals/github-owner-repo-routable-model-components/"},{"kind":"Raw Signal","stable_id":"signal:interconnects-plots-that-explain-state-of-routable-model-components","slug":"interconnects-plots-that-explain-state-of-routable-model-components","title":"Interconnects / Plots That Explain State Of: Routable Model Components","description":"The archive captures Interconnects / Plots That Explain State Of as a dated public record from Interconnects / Plots That Explain State Of. It documents models becoming replaceable or specialized components inside a more durable runtime and is retained as supporting evidence for the routable model components trend.","date":"2026-01-13","record_date":"2026-01-13","date_kind":"observed_at","topics":["agents","ai-adoption","future-of-work","inference","local-models","model-routing","org-design"],"source_count":1,"metric":"notable","url":"https://newruntime.com/signals/interconnects-plots-that-explain-state-of-routable-model-components/"},{"kind":"Raw Signal","stable_id":"signal:google-research-titans-miras-helping-ai-have-long-term-governed-context-and-memory","slug":"google-research-titans-miras-helping-ai-have-long-term-governed-context-and-memory","title":"Google Research / Titans Miras Helping Ai Have Long Term: Governed Context And Memory","description":"The archive captures Google Research / Titans Miras Helping Ai Have Long Term as a dated public record from Google Research / Titans Miras Helping Ai Have Long Term. It documents context and memory becoming governed infrastructure with update and provenance loops and is retained as supporting evidence for the governed context and memory trend.","date":"2025-12-20","record_date":"2025-12-20","date_kind":"observed_at","topics":["agent-memory","agents","context-engineering","inference","local-models","model-routing","retrieval"],"source_count":1,"metric":"notable","url":"https://newruntime.com/signals/google-research-titans-miras-helping-ai-have-long-term-governed-context-and-memory/"},{"kind":"Raw Signal","stable_id":"signal:github-mixedbread-ai-mgrep-routable-model-components","slug":"github-mixedbread-ai-mgrep-routable-model-components","title":"GitHub / mixedbread-ai/mgrep: Routable Model Components","description":"The archive captures GitHub / mixedbread-ai/mgrep as a dated public record from GitHub / mixedbread-ai/mgrep. It documents models becoming replaceable or specialized components inside a more durable runtime and is retained as supporting evidence for the routable model components trend.","date":"2025-12-19","record_date":"2025-12-19","date_kind":"observed_at","topics":["agent-harness","agent-memory","coding-agents","inference","local-models","model-routing","skills"],"source_count":1,"metric":"notable","url":"https://newruntime.com/signals/github-mixedbread-ai-mgrep-routable-model-components/"},{"kind":"Raw Signal","stable_id":"signal:app-release-notes-generative-media-infrastructure","slug":"app-release-notes-generative-media-infrastructure","title":"App / Release Notes: Generative Media Infrastructure","description":"The archive captures App / Release Notes as a dated public record from App / Release Notes. It documents image, video, audio, and multimodal generation becoming application infrastructure and is retained as branch-opening evidence for the generative media infrastructure trend.","date":"2025-12-03","record_date":"2025-12-03","date_kind":"observed_at","topics":["generative-media","inference","local-models","media-infrastructure","model-routing","multimodal","new-models"],"source_count":2,"metric":"structural","url":"https://newruntime.com/signals/app-release-notes-generative-media-infrastructure/"},{"kind":"Raw Signal","stable_id":"signal:hugging-face-deepseek-v3-routable-model-components","slug":"hugging-face-deepseek-v3-routable-model-components","title":"Hugging Face / Deepseek V3: Routable Model Components","description":"The archive captures Hugging Face / Deepseek V3 as a dated public record from Hugging Face / Deepseek V3. It documents models becoming replaceable or specialized components inside a more durable runtime and is retained as supporting evidence for the routable model components trend.","date":"2025-12-03","record_date":"2025-12-03","date_kind":"observed_at","topics":["agent-protocols","agents","inference","interoperability","local-models","mcp","model-routing"],"source_count":2,"metric":"notable","url":"https://newruntime.com/signals/hugging-face-deepseek-v3-routable-model-components/"},{"kind":"Raw Signal","stable_id":"signal:aiechoes-building-biomedical-graphrag-when-governed-context-and-memory","slug":"aiechoes-building-biomedical-graphrag-when-governed-context-and-memory","title":"Aiechoes / Building Biomedical Graphrag When: Governed Context And Memory","description":"The archive captures Aiechoes / Building Biomedical Graphrag When as a dated public record from Aiechoes / Building Biomedical Graphrag When. It documents context and memory becoming governed infrastructure with update and provenance loops and is retained as pressure-testing evidence for the governed context and memory trend.","date":"2025-11-26","record_date":"2025-11-26","date_kind":"observed_at","topics":["agent-memory","context-engineering","inference","local-models","model-routing","rag","retrieval"],"source_count":1,"metric":"notable","url":"https://newruntime.com/signals/aiechoes-building-biomedical-graphrag-when-governed-context-and-memory/"},{"kind":"Raw Signal","stable_id":"signal:google-ai-gemini-api-routable-model-components","slug":"google-ai-gemini-api-routable-model-components","title":"Google AI / Gemini Api: Routable Model Components","description":"The archive captures Google AI / Gemini Api as a dated public record from Google AI / Gemini Api. It documents models becoming replaceable or specialized components inside a more durable runtime and is retained as supporting evidence for the routable model components trend.","date":"2025-11-22","record_date":"2025-11-22","date_kind":"observed_at","topics":["agent-harness","agent-memory","coding-agents","inference","local-models","model-routing","skills"],"source_count":1,"metric":"notable","url":"https://newruntime.com/signals/google-ai-gemini-api-routable-model-components/"}],"atlas_records":[],"next_reads":[{"type":"related_material","path":"/posts/chatgpt-agent-loop-efficiency-stack/","reason":"Continue through the Inference topic.","url":"https://newruntime.com/posts/chatgpt-agent-loop-efficiency-stack/","title":"ChatGPT Cuts Repeated Work Across The Agent Stack","media_type":"text/html"},{"type":"related_material","path":"/posts/fireworks-lora-fullft-three-test-protocol/","reason":"Continue through the Inference topic.","url":"https://newruntime.com/posts/fireworks-lora-fullft-three-test-protocol/","title":"Run Three Tests Before Replacing LoRA With Full Fine-Tuning","media_type":"text/html"},{"type":"related_material","path":"/posts/openai-gpt-5-6-price-performance-frontier/","reason":"Continue through the Inference topic.","url":"https://newruntime.com/posts/openai-gpt-5-6-price-performance-frontier/","title":"GPT-5.6 Turns Efficiency Work Into API Economics","media_type":"text/html"},{"type":"related_material","path":"/posts/ramp-adaptive-llm-routing/","reason":"Continue through the Inference topic.","url":"https://newruntime.com/posts/ramp-adaptive-llm-routing/","title":"Ramp Teaches Its Gateway To Route By Failure, Latency, And Cost","media_type":"text/html"},{"type":"related_material","path":"/posts/vercel-ai-gateway-operating-budget/","reason":"Continue through the Inference topic.","url":"https://newruntime.com/posts/vercel-ai-gateway-operating-budget/","title":"Vercel AI Gateway Adds Runtime Budget Controls","media_type":"text/html"}]}
