All courses
Live cohortPM

AI Evals

Design rigorous evaluation systems that turn expert judgment into measurable, repeatable AI quality signals.

Next cohortFeb 22 – Mar 13, 2027
Length3 weeks
Live instruction3 × 2.5-hour classes
CohortFirst cohort
3 weeksFrom “looks good” to measurable quality
6 modulesThe complete evaluation lifecycle
7.5 hours liveBuild and pressure-test the system
1 real featureEvaluated end to end
Why this matters now

“It feels better” is not a product metric.

AI teams can change a prompt, swap a model, add context, or redesign an agent in hours. But without a rigorous evaluation system, they cannot tell whether the change improved the product—or simply moved the failures somewhere less visible.

Traditional software testing asks whether the same input produces the expected output. Generative AI requires a harder question: is this behavior useful, reliable, safe, and good enough for this user and this decision?

Answering that requires representative test cases, explicit quality dimensions, consistent expert judgment, calibrated automated graders, failure analysis, and release thresholds the whole team trusts.

Without those signals, teams debate anecdotes, optimize for demos, and discover regressions through angry users. With them, product, engineering, research, and leadership can make the same decision from the same evidence.

Evals are not the final QA step. They are the feedback system that makes every model, prompt, agent, and product decision improve.
What you'll learn

Everything this certification puts in your hands.

01

Define the evaluation strategy

Translate a vague quality goal into explicit product decisions.

  • Define usefulness, correctness, safety, tone, and task success
  • Map risks by user, workflow, and failure severity
  • Set thresholds for launch, rollback, and human review
02

Build representative datasets

Test the experience your users will actually encounter.

  • Balance common cases, long-tail cases, and adversarial edges
  • Create golden sets with trusted reference judgments
  • Use production sampling without overfitting to easy examples
03

Turn judgment into rubrics

Make expert review consistent enough to become a signal.

  • Write dimensions with observable scoring anchors
  • Design annotation and adjudication workflows
  • Measure reviewer agreement and resolve ambiguity
04

Use LLM-as-judge responsibly

Automate evaluation without blindly trusting another model.

  • Design judge prompts and reference-aware scoring
  • Calibrate automated judges against human experts
  • Track bias, variance, agreement, and judge drift
05

Diagnose failures, not averages

Turn a score into a prioritized product improvement.

  • Slice results by user, task, language, risk, and behavior
  • Cluster recurring failure modes and regressions
  • Connect failures to prompts, retrieval, models, tools, or UX
06

Operationalize continuous evaluation

Make quality part of every release—not an occasional audit.

  • Create dashboards, release gates, and ownership
  • Combine offline evals with online monitoring
  • Run a cadence that improves products, prompts, models, and agents
The capstone

Build an end-to-end evaluation system for a real AI feature.

Leave with more than a score. Produce the quality definition, representative dataset, calibrated judge, baseline results, failure diagnosis, and improvement plan your team needs to make a release decision.

Your proof: run the system on a real feature, establish its baseline, identify the highest-impact failure cluster, and show exactly what should change next.
DefineQuality dimensions, risks, and thresholds
SampleRepresentative cases, edges, and golden data
JudgeRubrics, human labels, and calibrated automation
MeasureBaseline scores, slices, and agreement
DiagnoseFailure clusters and regression causes
DecideRelease gate and prioritized improvement plan
Curriculum

Week by week.

Dates are the published cohort schedule.

Week 1Evaluation Strategy + Dataset DesignFebruary 22–28+
  • Define quality dimensions, risks, and decision thresholds
  • Map failure severity and critical user workflows
  • Build representative cases, edge cases, and a golden set
  • Design a sampling strategy that evolves with production
Build: Evaluation plan + first test dataset
Week 2Rubrics, Human Judgment + Automated EvaluationMarch 1–7+
  • Turn expert judgment into observable scoring anchors
  • Create labeling, review, and adjudication workflows
  • Design and calibrate LLM-as-judge prompts
  • Measure agreement, variance, bias, and judge reliability
Build: Rubric + calibrated judge workflow
Week 3Failure Analysis + Eval OperationsMarch 8–13+
  • Run the baseline and slice results by meaningful dimensions
  • Cluster failures and trace likely causes
  • Define dashboards, release gates, monitoring, and ownership
  • Present the improvement plan and next evaluation cycle
Final outcome: Working eval system + improvement plan
Faculty

Taught by the people doing the work.

Reah Miyara

Senior Director, Google Models

Reah leads evaluation work on Google models. She brings the operating perspective required to turn ambiguous human judgment into quality signals teams can trust across product development and release decisions.

The course focuses on the difficult parts that simple benchmark tutorials skip: deciding what “good” means, building representative evidence, calibrating automated judges, resolving disagreement, diagnosing regressions, and making evaluation part of the product cadence.

What's included

Everything that comes with your seat.

Three 2.5-hour classes

Build each layer live with direct instruction and review.

Evaluation strategy

Define quality, risk, metrics, and decision thresholds.

Dataset and rubric

Create representative cases and consistent scoring criteria.

Calibrated judge prompts

Compare automated evaluation with human expert judgment.

Baseline and failure analysis

Run the system, slice results, and prioritize fixes.

Eval operating plan

Establish dashboards, release gates, monitoring, and cadence.

Questions

Is everything taught live?

The Fellowship is built around live, cohort-based learning. You learn directly from the named practitioners, ask questions, complete applied work, and receive the latest insights from people working at the frontier.

Advanced Product Management is included as an on-demand program for strengthening timeless product fundamentals.

How much time will I need each week?

Most cohorts are designed for busy working professionals and require only a few hours per week. The exact commitment varies by program. You'll get the most value by attending live, completing the exercises, and making time for your capstone.

What if I miss a live session?

Attend live whenever possible because the discussions, feedback, and accountability are important parts of the experience.

However, missing an occasional session won't ruin your progress. You can also retake eligible cohorts during your active membership.

Can I retake a cohort?

Yes. You can retake eligible cohorts at no additional cost while your membership remains active. This allows you to experience the updated curriculum, refresh your knowledge, and apply what you learn to a new project as AI evolves.

Can my employer pay for it?

Yes. Many professionals use their learning and development budgets to cover the Fellowship.

We provide a manager-approval email template explaining the practical value to your company. Corporate cohort options are also available for teams. If your company needs an invoice, email support@productfaculty.com and we'll sort it.