The outcome

Compare a baseline and candidate with evidence strong enough to support a documented deployment decision.

Step by step

A workflow you can repeat.

  1. 01

    Define the decision, user segments, critical tasks, prohibited failures, quality dimensions, latency and cost limits, evaluator owners, thresholds, and rollback rule.

  2. 02

    Build a versioned dataset from authorized synthetic, expert, and production-derived cases, removing secrets and personal data and preserving difficult counterexamples.

  3. 03

    Add deterministic code checks, structured human rubrics, and calibrated model judges, recording judge model, prompt, repetitions, known bias, and disagreement policy.

  4. 04

    Run baseline and candidate with identical data, concurrency, caching, and metadata, then inspect item-level regressions, confidence, variance, failure clusters, latency, and cost.

  5. 05

    Require owner sign-off on thresholds and exceptions, attach the experiment to the release record, canary the candidate, and feed verified production failures back into the dataset.

Working standard

What good use looks like.

  • Version data and evaluators together.
  • Inspect failures beyond averages.
  • Calibrate model judges against humans.

Official references

Check the current product documentation.