Normalized Telegram record

Agent skills need behavioral evals, not prose review

Benchmarks show that expert-authored skills can help while self-generated skills can underperform a no-skill baseline.

Signal contract

  • Benchmarks show that expert-authored skills can help while self-generated skills can underperform a no-skill baseline.
  • A skill is executable behavior packaged as files, so its quality has to be measured across tasks, models, and harnesses instead of judged by how polished the instructions look.
  • Novelty: structural. Verification: source-inspected.

Source ledger

Publishable sources attached to this record.

5 public sources
#SourceRolePublic status
1arxiv.orgpaperprimary receiptsource_urls
2arxiv.orgpapersupporting receiptsource_urls
3arxiv.orgpapersupporting receiptsource_urls
4github.comreposupporting receiptsource_urls
5platform.claude.comdocssupporting receiptsource_urls

Observation

Benchmarks show that expert-authored skills can help while self-generated skills can underperform a no-skill baseline.

Why it matters

A skill is executable behavior packaged as files, so its quality has to be measured across tasks, models, and harnesses instead of judged by how polished the instructions look.

Entities

SkillsBench, Claude

Provenance

This public record is an English normalization of QWG AI Telegram message 2707. The complete original-language post remains the canonical raw message.

Open the original Telegram record