Skip to main content
Use an evaluator comparison when deciding whether a different model, question formulation, rubric, or escalation policy is a better checker for a specific task. Do not confuse it with an application Experiment, which compares the system producing the response. The current measurement surface is available. These steps describe the protocol and records to inspect; they do not promise a one-click replacement for every evaluator workflow.

Define what the comparison can establish

Record the task, error types, protected populations, source of reference labels, and decision you intend to make. State whether the candidate must improve failure detection, reduce cost within an agreed quality margin, or satisfy a particular absolute requirement. Keep three claims separate: “matches the baseline,” “matches independently verified references,” and “gives the same answer when repeated.” They measure agreement, correctness, and repeatability respectively.

Freeze the comparable inputs

Pin the baseline validation run and its measurement revision. For an evaluator comparison, both arms should judge the same application input, response, evidence construction, reference labels, and environment. EvalGate’s persisted comparison resolves a pinned baseline revision and observation identity. The evaluator-comparison projection binds the target output into input identity, rather than treating two different responses as the same judging task. The comparison descriptor includes baselineValidationRunId, purpose, taskId, minimumReleaseSupport, recallFloor, unsafeAutoPassCeiling, and an optional requiredApprovalStage. This is a description of fields, not a request to hide server-owned qualification or execution metadata inside a sampling manifest.

Keep development, calibration, and qualification separate

Use development cases to change the questions. Use calibration evidence to choose operating parameters. Use an independently reviewed or independently verifiable sealed holdout for qualification where required by the policy. Do not tune on the holdout and then reuse it as independent confirmation. Repeated runs of one case measure stability; they do not create additional independent cases. Preserve the true grouping unit when several rows come from one document or conversation.

Inspect the full cohort

Review missing arms, unplanned cases, identity mismatches, incomplete execution, unresolved references, and exclusions before interpreting headline metrics. The canonical builder rejects repeated declared independent units and inconsistent pairing. Then inspect detection recall, unsafe automatic passes, false blocks, review referrals, and uncertainty. The current qualification builder applies recall and unsafe-pass requirements; do not assume a displayed false-block metric is also a configured release requirement.

Read economics at the right boundary

The persisted comparison currently combines baseline and candidate provider accounting into experiment-level evidence. It does not, by that fact alone, demonstrate separate per-arm cost or latency improvements. Keep experiment spend distinct from candidate operating cost and complete workflow latency.

Before changing the active evaluator

Check that the final composed outcome agrees with the persisted comparison case, especially when using fallback or stability runs. An unresolved escalation must not masquerade as failure; an early passing child must not override a later terminal failure. A shadow comparison is not activation authorization. Finish the applicable qualification and action-approval workflow before changing a protected consumer. See Qualification and authorization and Read evaluation evidence.