Skip to main content
EvalGate answers several different questions. They should not collapse into one score: A successful request, a valid typed answer, a passing case, a qualified evaluator, and an authorized release are different outcomes.

Start with the object being changed

When you change the application prompt, model, tools, or retrieval pipeline, compare application behavior while keeping the evaluator stable where possible. The Experiments workflow documents that comparison. When you change the evaluator, keep the application responses and admissible reference evidence fixed. Otherwise, a better response can be mistaken for a better evaluator. Use the evaluator comparison guide for that distinction. When you change an operating threshold without executing the corresponding real-world action, describe the result as a policy simulation. A simulated block is not evidence that a production action was prevented.

Choose a checker that matches the requirement

Use a deterministic rule for a requirement that can be checked exactly: an allowed value, a schema, or an arithmetic invariant. Use a generative judge when a reviewed rubric requires an explanation or other generated text. A decision judge such as Jev returns bounded typed observations instead of a written critique. The provider is one part of the evaluator. The questions, acceptance policy, input construction, model identity, and aggregation also affect what its result means.

Understand the four similar-sounding records

A scorer version defines a reusable check. An evaluator release identifies a measurement instrument and its protocol. A validation run measures that instrument using selected evidence. An application experiment compares application variants. They can reference each other, but they are not interchangeable names for a run. Likewise, the Calibration control plane’s approved score mapping supports score comparability. It does not by itself demonstrate task-specific failure detection or grant deployment permission.

Read uncertainty as information

An evaluator can return an observation without an accepted pass/fail interpretation. A run can execute completely while still being inconclusive. A comparison can be statistically promising but lack the independent evidence required for qualification. Before acting, identify the original case, the final composed outcome, any unresolved attempts, the measurement population, and the policy being applied. Never treat passed: false alone as proof of a behavioral violation: it can also accompany non-authoritative outcomes.

Availability boundary

EvalGate’s feature inventory distinguishes controlled-beta workflows from experimental source capabilities and unverified release paths. This conceptual map does not change those labels. A feature appearing in the current source does not establish availability in every installed package or deployment. Continue with Decision judges, Qualification and authorization, and Reading evaluation evidence.