Learning track 06
Evaluation & Improvement
Measure quality, observe production behavior, validate judges, and turn failures into a data flywheel.
- Guides
- 6
- Sections
- 41
- Cards
- 179
Suggested sequence
Follow the track or enter where the problem begins.
Every guide is self-contained. The sequence only preserves prerequisite order when one topic depends on another.
- 01Vol. 02
LLM Observability & Evaluations
How to understand what is happening inside the pipeline - and what exactly is broken. Tracing, metrics, automatic evaluations, testing prompts and secure deployment of changes.
- Sections
- 8
- Cards
- 35
- 02Vol. 12
Evaluation Process & LLM-as-Judge
How to build the process of evaluating LLM systems from scratch: minimum viable eval, guardrails vs evaluators, why a binary assessment is better than a scale of 1–5, negative scenarios, LLM-as-judge with correlation to people, detection of degradation after release, synthetic data - when it helps and when it harms.
- Sections
- 7
- Cards
- 23
- 03Vol. 27
The Production Observability & Evaluation Loop
This volume does not repeat the general review of evals. It shows the operational cycle of the team: how incident cases, datasets, offline evals, human review, canary gates and rollout/rollback decisions are born from production traces and user signals. It’s not “what evals are,” it’s “how the team lives with evals and observability every day.”
- Sections
- 7
- Cards
- 35
- 04Vol. 29
The Data Flywheel: From Production Cases to Release Gates
This volume is about the main engine of a high-quality LLM team: how production traces and user pain turn into curated cases, label taxonomy, human review queue, regression set, canary set and decisions about what to fix prompt, what to retriever, and what to train. So it's not "where to get the dataset," but how the team grows it out of their own work and turns the data into a system improvement cycle.
- Sections
- 7
- Cards
- 35
- 05Vol. 32
LLM-as-Judge Meta-Evaluation
A separate volume is not about eval in general, but about the validation of the judge itself: agreement with people, metrics for binary and ordinal verdicts, resistance to repeated runs, confidence calibration, bias audits, pairwise ranking and minimum reporting standard. Agreement → classifier view → stability → calibration → bias → reporting.
- Sections
- 7
- Cards
- 25
- 06Vol. 33
Core Metrics for LLM Pipelines
One volume about the most useful metrics for LLM systems: from average, variance and standard deviation to accuracy, precision, recall, F1, MRR, MAP, nDCG@k, calibration, pass@k, latency and cost per success. Each card answers four questions: what it measures, how to calculate it, how to read, and where to apply it in the pipeline.
- Sections
- 5
- Cards
- 26