Skip to main content
A fast provider response does not automatically make an evaluation workflow fast. A low token price does not establish the cost of a completed decision. EvalGate has several accounting and timing records because these questions have different boundaries.

Three money fields that must not be confused

A budget or ceiling is permission to incur bounded usage. Reserved liability accounts for an in-flight or uncertain attempt. Observed cost is the cost evidence actually available for that execution. An estimate is not a bill. A request that timed out may still have reached the provider. Keep unknown billing explicit until it can be reconciled; do not free uncertain liability merely because the client did not receive an answer. The reviewed evaluator executor exposes provider attempts, observed cost, unknown-cost attempts, and unsettled liability. A deterministic run with no provider calls can have zero provider spend, but that does not mean zero infrastructure or human cost.

Choose a clock before comparing numbers

These clocks exist in the measurement contract. Their existence does not prove every producer populates all of them. A report should identify the clock, sample count, and missing observations. Percentiles are computed from raw observations, not by averaging existing percentiles. The current helper withholds p99 with fewer than 100 observations. That threshold is an implementation rule, not a universal guarantee that 100 observations make a tail estimate precise.

Count the entire candidate path

For a typed judge followed by frontier fallback, candidate cost includes both attempts when fallback occurs. End-to-end latency includes the additional work and any scheduling or review delay covered by the declared clock. Record the fraction of cases that escalated. A fast first call may add delay to every case while only avoiding expensive work on a subset.

Keep per-arm performance separate from experiment spend

The current persisted evaluator-comparison path combines baseline and candidate provider observations and adds their available costs. This is experiment-level accounting. It is not a per-arm savings calculation. Do not label that combined figure “candidate cost,” and do not derive “candidate is faster” from a pooled provider percentile. A proper comparison needs separate baseline/candidate populations and an explicit complete-workflow boundary. Similarly, the current path deliberately does not infer target-system generation cost for a system comparison. Missing cost stays missing.

A useful report sentence

“The measured provider-attempt latency is available. End-to-end review completion was not measured. Some attempts have unresolved billing, so a complete cost-saving claim is not supported.” That statement tells a buyer or operator what the evidence supports without converting an unknown into a favorable number. Continue with Compare evaluators and Prepare a review.