The six pages here lay out the case for evals, the maturity ladder you climb, the reasons LLM evaluation does not look like classical ML eval, and where to start when nothing is labeled. Read them in order if you are new. Skim if you are already shipping.
Most teams arrive thinking the gap is tooling. The gap is usually conceptual: a half-built mental model of what evals are for, and which rung of the ladder pays off this quarter. The pages below fix that before you touch a framework. When you get to guardrails, the condensed reference is guardrails versus evals architecture.
Chapters:
After this section, read error analysis next, or pick a role-based track at /start.