Outcomes
- Translate product quality into clear evaluation dimensions and measurable rubrics
- Build representative test sets, golden datasets, and human-review workflows
- Use LLM-as-judge responsibly with calibration, agreement checks, and failure analysis
- Create an eval operating cadence that improves models, prompts, agents, and product decisions
Who it's for
AI product managers, engineers, researchers, and quality leaders who need a practical system for measuring whether generative-AI experiences are reliable, useful, and improving.
Syllabus
Module 1Evaluation strategy — define quality dimensions, risks, and decision thresholds
Module 2Dataset design — representative cases, edge cases, golden sets, and sampling
Module 3Rubrics and human judgment — turn expert review into consistent labels and scores
Module 4Automated evaluation — LLM-as-judge patterns, calibration, and agreement checks
Module 5Failure analysis — slice results, diagnose regressions, and prioritize fixes
Module 6Eval operations — dashboards, release gates, monitoring, and continuous improvement
Capstone
An end-to-end evaluation system for a real AI feature: rubric, dataset, judge prompts, baseline results, and an improvement plan.
First-cohort curriculum draft — details can be refined before launch.