Normalized Telegram record
Agent skills need behavioral evals, not prose review
Benchmarks show that expert-authored skills can help while self-generated skills can underperform a no-skill baseline.
Signal contract
- Benchmarks show that expert-authored skills can help while self-generated skills can underperform a no-skill baseline.
- A skill is executable behavior packaged as files, so its quality has to be measured across tasks, models, and harnesses instead of judged by how polished the instructions look.
- Novelty: structural. Verification: source-inspected.
Source ledger
Publishable sources attached to this record.
| # | Source | Role | Public status |
|---|---|---|---|
| 1 | arxiv.orgpaper | primary receipt | source_urls |
| 2 | arxiv.orgpaper | supporting receipt | source_urls |
| 3 | arxiv.orgpaper | supporting receipt | source_urls |
| 4 | github.comrepo | supporting receipt | source_urls |
| 5 | platform.claude.comdocs | supporting receipt | source_urls |
Observation
Benchmarks show that expert-authored skills can help while self-generated skills can underperform a no-skill baseline.
Why it matters
A skill is executable behavior packaged as files, so its quality has to be measured across tasks, models, and harnesses instead of judged by how polished the instructions look.
Entities
SkillsBench, Claude
Provenance
This public record is an English normalization of QWG AI Telegram message 2707. The complete original-language post remains the canonical raw message.