Thrive AI Engineering
← All questions

evaluation · Editorial starter

Why can an agent's overall evaluation score improve while a critical workflow gets worse?

By Thrive AI Editorial · Created 2026-10-07 · Updated 2026-10-07

Thrive AI Editorial · AI-assisted worked example, not a member conversation.

A candidate model improves the overall pass rate but fails more of the tool-authorization or escalation cases. Most tests are easy informational questions. Should the higher average justify shipping?

This editorial example uses invented, explicitly synthetic pass counts to show aggregation risk. These numbers are not measurements of any model or platform.

0 score

Sign in to vote or contribute. Scores count actual member votes, not editorial endorsements.

1 published answer

A selected solution is chosen by the question author or moderator, not an independent certification. Check assumptions before using any code in production.

Thrive AI Editorial

Thrive AI Editorial · AI-assisted worked example, not a member conversation.

2026-10-07

The average inherits the test mixture

An overall pass rate weights each slice by how many examples it contains. If low-risk informational cases dominate the dataset, gains there can conceal a regression in a smaller, high-consequence workflow.

The solution is not to abandon aggregate metrics. Define separate release conditions for critical slices, report sample sizes, and assess whether the test distribution matches deployment. A threshold chosen after inspecting results is easier to game than a predeclared acceptance rule.

Code / output
candidate -> overall comparison
          -> critical-slice gates
          -> human review of disagreements and failures
all required checks pass -> release decision

Reproduce a misleading improvement

Python 3, standard library only. The two slices are a mathematical fixture, not a recommended evaluation dataset.

Code / output
baseline = {"information": (80, 100), "critical_action": (10, 10)}
candidate = {"information": (95, 100), "critical_action": (5, 10)}

def overall(results):
    return sum(passed for passed, total in results.values()) / sum(
        total for passed, total in results.values())

base_score = overall(baseline)
new_score = overall(candidate)
critical_pass, critical_total = candidate["critical_action"]
critical_gate = critical_pass == critical_total
ship = new_score >= base_score and critical_gate
assert new_score > base_score
assert not critical_gate and not ship
print("overall:", round(base_score, 3), "->", round(new_score, 3))
print("critical slice:", critical_pass, "/", critical_total)
print("release allowed:", ship)

Expected output:

Code / output
overall: 0.818 -> 0.909
critical slice: 5 / 10
release allowed: False

Build an evaluation that supports a decision

Tag cases by failure consequence and workflow, not just by topic. Include explicit authorization boundaries, unsupported requests, ambiguous inputs, retrieval failures and handoff behavior where they matter. Keep a stable held-out comparison set and separate it from examples repeatedly used while tuning.

Run paired comparisons on the same cases. For nondeterministic systems, repeat representative cases and report variability; one run is not a confidence interval. Look at which examples changed rather than only the net score. Keep model, prompt, tool schema, retrieval snapshot and evaluator version in the evaluation record.

The all-pass rule in this fixture expresses an intentionally strict synthetic requirement. Ten successes cannot establish that a real system has zero failure probability. For operational gates, choose thresholds from consequence, tolerance and sufficient evidence, not from these arbitrary counts.

Avoid creating another misleading metric

A model-based judge can also be inconsistent or biased. Calibrate it against human judgments, examine disagreements, and use deterministic checks for exact invariants where possible. Human review does not mean treating a handful of impressions as a benchmark.

Track runtime failures separately from content quality. A correct answer returned after an unacceptable delay or a prohibited side effect is not an unqualified success. Report the measurements that support the decision and the limitations that remain.

Source checked 2026-10-08: OpenAI evaluation best practices. Task-specific evaluation, representative data, continuous comparison and human calibration are supported there. The slice-gating fixture is our own AI-assisted explanation and does not depend on a hosted evaluation product.

0 score

Sign in to vote or contribute. Scores count actual member votes, not editorial endorsements.

Verify your sign-in to contribute an answer

Reading is free. Posting requires a verified account and moderation; no course purchase is required.