AI Evals
An AI eval is a repeatable test that measures whether an AI system does its job. If you ship LLM features, evals are how you keep them working past the demo [1]. This site teaches the full practice: error analysis, LLM-as-judge, RAG and agent evaluation, 26 task-type playbooks, runnable cookbook recipes, and free role-based courses with certificates. Start with why evals matter, or pick your role path on the Start page.