Normalized Telegram record

Agent skills need behavioral evals, not prose review

Benchmarks show that expert-authored skills can help while self-generated skills can underperform a no-skill baseline.

Signal contract

Benchmarks show that expert-authored skills can help while self-generated skills can underperform a no-skill baseline. A skill is executable behavior packaged as files, so its quality has to be measured across tasks, models, and harnesses instead of judged by how polished the instructions look. Novelty: structural. Verification: source-inspected. This New Runtime record is an evidence-linked retrieval unit.

Source ledger

Publishable sources attached to this record.

5 public sources
#SourceRolePublic status
1arxiv.orgpaperprimary receiptsource_urls
2arxiv.orgpapersupporting receiptsource_urls
3arxiv.orgpapersupporting receiptsource_urls
4github.comreposupporting receiptsource_urls
5platform.claude.comdocssupporting receiptsource_urls

Observation

Benchmarks show that expert-authored skills can help while self-generated skills can underperform a no-skill baseline.

Why it matters

A skill is executable behavior packaged as files, so its quality has to be measured across tasks, models, and harnesses instead of judged by how polished the instructions look.

Entities

SkillsBench, Claude

Provenance

This public record is an English normalization of QWG AI Telegram message 2707. The complete original-language post remains the canonical raw message.

Open the original Telegram record

Discovery graph / next reads

Continue through New Runtime

Open the graph
  1. 01topicContext engineering - New RuntimeExplore the context engineering topic hub.
  2. 02topicAgent evals - New RuntimeExplore the evals topic hub.
  3. 03related materialBlog / Validating Llm As Judge Systems Under Rating: Harnesses And Portable SkillsShares agent skills and context engineering.
  4. 04related materialOpenAI: A benchmark score reflects the model as well as the harness and settings used to run it.Shares context engineering and evals.
  5. 05related materialOpenAI Developers: We're introducing two new transcription models in the API: • GPT-Live-Transcribe: built for lo...Shares context engineering and evals.

These links are also published in this page’s JSON twin and as typed edges in DiscoveryGraph v1.

Who read this page?Machine requests, hidden until opened

Loading the privacy-safe route aggregate…

Open the JSON contract