Thrive AI Editorial
Thrive AI Editorial · AI-assisted worked example, not a member conversation.
2026-10-07
The average inherits the test mixture
An overall pass rate weights each slice by how many examples it contains. If low-risk informational cases dominate the dataset, gains there can conceal a regression in a smaller, high-consequence workflow.
The solution is not to abandon aggregate metrics. Define separate release conditions for critical slices, report sample sizes, and assess whether the test distribution matches deployment. A threshold chosen after inspecting results is easier to game than a predeclared acceptance rule.
candidate -> overall comparison
-> critical-slice gates
-> human review of disagreements and failures
all required checks pass -> release decision
Reproduce a misleading improvement
Python 3, standard library only. The two slices are a mathematical fixture, not a recommended evaluation dataset.
baseline = {"information": (80, 100), "critical_action": (10, 10)}
candidate = {"information": (95, 100), "critical_action": (5, 10)}
def overall(results):
return sum(passed for passed, total in results.values()) / sum(
total for passed, total in results.values())
base_score = overall(baseline)
new_score = overall(candidate)
critical_pass, critical_total = candidate["critical_action"]
critical_gate = critical_pass == critical_total
ship = new_score >= base_score and critical_gate
assert new_score > base_score
assert not critical_gate and not ship
print("overall:", round(base_score, 3), "->", round(new_score, 3))
print("critical slice:", critical_pass, "/", critical_total)
print("release allowed:", ship)
Expected output:
overall: 0.818 -> 0.909
critical slice: 5 / 10
release allowed: False
Build an evaluation that supports a decision
Tag cases by failure consequence and workflow, not just by topic. Include explicit authorization boundaries, unsupported requests, ambiguous inputs, retrieval failures and handoff behavior where they matter. Keep a stable held-out comparison set and separate it from examples repeatedly used while tuning.
Run paired comparisons on the same cases. For nondeterministic systems, repeat representative cases and report variability; one run is not a confidence interval. Look at which examples changed rather than only the net score. Keep model, prompt, tool schema, retrieval snapshot and evaluator version in the evaluation record.
The all-pass rule in this fixture expresses an intentionally strict synthetic requirement. Ten successes cannot establish that a real system has zero failure probability. For operational gates, choose thresholds from consequence, tolerance and sufficient evidence, not from these arbitrary counts.
Avoid creating another misleading metric
A model-based judge can also be inconsistent or biased. Calibrate it against human judgments, examine disagreements, and use deterministic checks for exact invariants where possible. Human review does not mean treating a handful of impressions as a benchmark.
Track runtime failures separately from content quality. A correct answer returned after an unacceptable delay or a prohibited side effect is not an unqualified success. Report the measurements that support the decision and the limitations that remain.
Source checked 2026-10-08: OpenAI evaluation best practices. Task-specific evaluation, representative data, continuous comparison and human calibration are supported there. The slice-gating fixture is our own AI-assisted explanation and does not depend on a hosted evaluation product.
Sign in to vote or contribute. Scores count actual member votes, not editorial endorsements.