Learning track 06

Evaluation & Improvement

Measure quality, observe production behavior, validate judges, and turn failures into a data flywheel.

Guides
6
Sections
41
Cards
179

Suggested sequence

Follow the track or enter where the problem begins.

Every guide is self-contained. The sequence only preserves prerequisite order when one topic depends on another.

  1. 01
    Vol. 02

    LLM Observability & Evaluations

    How to understand what is happening inside the pipeline - and what exactly is broken. Tracing, metrics, automatic evaluations, testing prompts and secure deployment of changes.

    Sections
    8
    Cards
    35
    Open guide
  2. 02
    Vol. 12

    Evaluation Process & LLM-as-Judge

    How to build the process of evaluating LLM systems from scratch: minimum viable eval, guardrails vs evaluators, why a binary assessment is better than a scale of 1–5, negative scenarios, LLM-as-judge with correlation to people, detection of degradation after release, synthetic data - when it helps and when it harms.

    Sections
    7
    Cards
    23
    Open guide
  3. 03
    Vol. 27

    The Production Observability & Evaluation Loop

    This volume does not repeat the general review of evals. It shows the operational cycle of the team: how incident cases, datasets, offline evals, human review, canary gates and rollout/rollback decisions are born from production traces and user signals. It’s not “what evals are,” it’s “how the team lives with evals and observability every day.”

    Sections
    7
    Cards
    35
    Open guide
  4. 04
    Vol. 29

    The Data Flywheel: From Production Cases to Release Gates

    This volume is about the main engine of a high-quality LLM team: how production traces and user pain turn into curated cases, label taxonomy, human review queue, regression set, canary set and decisions about what to fix prompt, what to retriever, and what to train. So it's not "where to get the dataset," but how the team grows it out of their own work and turns the data into a system improvement cycle.

    Sections
    7
    Cards
    35
    Open guide
  5. 05
    Vol. 32

    LLM-as-Judge Meta-Evaluation

    A separate volume is not about eval in general, but about the validation of the judge itself: agreement with people, metrics for binary and ordinal verdicts, resistance to repeated runs, confidence calibration, bias audits, pairwise ranking and minimum reporting standard. Agreement → classifier view → stability → calibration → bias → reporting.

    Sections
    7
    Cards
    25
    Open guide
  6. 06
    Vol. 33

    Core Metrics for LLM Pipelines

    One volume about the most useful metrics for LLM systems: from average, variance and standard deviation to accuracy, precision, recall, F1, MRR, MAP, nDCG@k, calibration, pass@k, latency and cost per success. Each card answers four questions: what it measures, how to calculate it, how to read, and where to apply it in the pipeline.

    Sections
    5
    Cards
    26
    Open guide
Track complete?Keep a live route into the next relevant area.