---
type: "post"
stable_id: "post:the-bitter-lesson-of-tool-calling-tests-general-learning-against-hand-built-schemas"
slug: "the-bitter-lesson-of-tool-calling-tests-general-learning-against-hand-built-schemas"
title: "The Bitter Lesson of Tool Calling Tests General Learning Against Hand-Built Schemas"
description: "A new paper asks whether general learning principles can outperform increasingly hand-engineered tool-calling schemes."
retrieval_nugget: "The useful question is where structure belongs: in a fixed schema, the training signal, the runtime harness, or an evaluation loop. Results should be read within the paper's tasks and tool environment rather than as a universal rule."
published_at: "2026-08-14"
updated_at: "2026-08-15"
record_date: "2026-08-14"
date_kind: "discovered_at"
topics: ["tool-calling","research","evals"]
entities: ["The Bitter Lesson of Tool Calling"]
source_urls: ["https://arxiv.org/abs/2608.06370"]
source_format: "research_paper"
editorial_timing: {"lane":"regular_hourly","scheduled_at":"2026-08-19T17:00:00+03:00","real_news_delta":"owner-selected verified story"}
origin: {"basket_id":"5854f7b2-5954-4d2a-8997-81596a49da74","basket_revision":1,"target_kind":"story_cluster","target_id":"97104458-3d8f-4378-9bad-01ab380528ed","owner_selection":"18","route":"openclaw"}
visual_decision: {"outcome":"text_only","status":"not_applicable","reason_code":"concise_text_sufficient","owner_reviewed":true,"reviewed_by":"owner-and-codex"}
schema_version: "newruntime-agent-readable-v0.2"
status: "published"
visuals: []
routes: {"html":"https://newruntime.com/posts/the-bitter-lesson-of-tool-calling-tests-general-learning-against-hand-built-schemas/","markdown":"https://newruntime.com/posts/the-bitter-lesson-of-tool-calling-tests-general-learning-against-hand-built-schemas.md","json":"https://newruntime.com/posts/the-bitter-lesson-of-tool-calling-tests-general-learning-against-hand-built-schemas.json"}
---

# The Bitter Lesson of Tool Calling Tests General Learning Against Hand-Built Schemas

## Retrieval answer

The useful question is where structure belongs: in a fixed schema, the training signal, the runtime harness, or an evaluation loop. Results should be read within the paper's tasks and tool environment rather than as a universal rule.

A new paper asks whether general learning principles can outperform increasingly hand-engineered tool-calling schemes.

New Runtime reading: The useful question is where structure belongs: in a fixed schema, the training signal, the runtime harness, or an evaluation loop. Results should be read within the paper's tasks and tool environment rather than as a universal rule.

Evidence boundary: this item uses the listed public sources and keeps vendor, author, or reporter claims attributed. The queued page is an editorial synthesis, not an independent validation of every reported metric.
