AI Evals
The evaluation papers worth your time, curated into ten themes: LLM-as-judge, benchmarks, RAG, agents, safety, and more, each with why it matters in practice.