Rented intelligence -> owned learning loop

Organizations are moving the durable learning asset out of one model provider and into user-owned traces, evaluations, corrections, and promotion rules.

Tectonic shift

Before: Prompts, corrections, memory, and successful behavior accumulate inside one model provider After: A user-owned evaluation and trace layer improves and compares replaceable models Current stage: emerging This New Runtime record is an evidence-linked retrieval unit. Use its canonical page, machine-readable representations, dates, scope, and public source URLs to verify the claim before reusing it.

Source ledger

Publishable sources attached to this record.

6 public sources
#SourceRolePublic status
1devblogs.microsoft.comarticleprimary receiptsource_urls
2learn.microsoft.comsourcesupporting receiptsource_urls
3docs.cloud.google.comdocssupporting receiptsource_urls
4opentelemetry.iosourcesupporting receiptsource_urls
5openai.comsourcesupporting receiptsource_urls
6privacy.claude.comsourcesupporting receiptsource_urls

Organizations are beginning to treat their accumulated evaluation and correction data as a durable product asset. Models can change while the criteria for good work, representative tasks, failure cases, accepted outputs, and promotion history continue to compound.

What is changing?

A portable harness is only the enabling layer. The more consequential change is ownership of the improvement loop around it.

That loop includes:

  • factual traces of model inputs, outputs, tool calls, and outcomes;
  • owner corrections, accepted results, retries, and escalation decisions;
  • versioned tasks, rubrics, protected cases, and cost constraints;
  • prompts, skills, tool contracts, and routing policies tested against the same evaluation set;
  • an audit trail showing why one candidate configuration or model was promoted.

When those artifacts live outside a model product, active models can be compared on the organization’s real work rather than on a public leaderboard. Changing the model no longer means discarding the evidence used to improve the system.

Evidence

Microsoft Foundry now describes a closed learning loop with a swappable model, traces for every run, organization-defined rubrics, and optimization across instructions, skills, tools, and model choice. Its Agent Optimizer evaluates candidates against the same task set and ranks them by outcome and token cost.

Google’s agent evaluation documentation treats a trace as a factual, immutable record containing model inputs, responses, and tool calls, then uses those traces as the basis for scoring and iterative optimization. OpenTelemetry’s GenAI conventions make model, token, prompt, completion, tool call, and tool-result telemetry portable enough to observe across different agent stacks.

Boundary

This shift should not be justified by claiming that every provider trains on all customer data. OpenAI and Anthropic state that their business and API products do not use customer inputs or outputs for model training by default.

The broader dependency risk remains: prompts, memories, corrections, eval datasets, and successful workflow history can still become trapped inside one provider’s product even when that provider does not train on the data.

Counter-evidence

Provider-native optimizers can exploit undocumented model behavior and proprietary runtime features that a portable learning layer cannot reproduce. Model changes also alter planning, tool use, latency, and safety, so switching is never free.

Revision trigger

Revise this shift if cross-provider replay remains too lossy for meaningful comparison, or if provider-native learning systems consistently beat user-owned traces and evaluations on reliability, total cost, and migration risk.

Discovery graph / next reads

Continue through New Runtime

Open the graph
  1. 01related materialAgent Harness Optimization Is Becoming an Outer-Loop DisciplineField Note documenting this shift.
  2. 02related materialHermes Agent Makes the Learning Loop Part of the RuntimeField Note documenting this shift.
  3. 03related materialThe Coding Harness Is Becoming Independent From the ModelField Note documenting this shift.
  4. 04topicAgent runtime - New RuntimeExplore the agent runtime topic hub.
  5. 05topicAgent evals - New RuntimeExplore the evals topic hub.

These links are also published in this page’s JSON twin and as typed edges in DiscoveryGraph v1.

Who read this page?Machine requests, hidden until opened

Loading the privacy-safe route aggregate…

Open the JSON contract