The answer can be correct while the system is wrong
An agent says, "The request is complete." The tool returned success. The conversation looks reasonable. None of those observations proves that the correct record was updated, that an unauthorized change was avoided, or that retrying the operation will be safe.
For a state-changing agent, success is a property of the resulting system state and permitted side effects, not merely the final sentence. This distinction is the bridge between reading agent research and building an evaluation that catches real failures.
We will follow that bridge with three diagrams, a small executable Python verifier and a release checklist. The code uses only in-memory dictionaries. It does not execute a model or reproduce a published benchmark; it demonstrates how to design the grader before introducing an agent.
Three recent sources worth reading carefully
This reading note was checked on October 5, 2026. Findings belong to the named authors; the engineering exercises are our own.
| Source | Core idea | Practitioner implication |
|---|---|---|
| ThinkingBox, Microsoft and Hugging Face, October 3 post | Grade terminal backend state and side effects across repeated runs | Check the records, not just the tool transcript |
| Dependency-Aware Policy Optimization, October 2 arXiv submission | Trace command read/write dependencies for credit assignment in terminal-agent training | Preserve causal dependencies when debugging a trajectory |
| Threat-Preserving Representation Sensitivity, October 2 arXiv submission | A benchmark score can change when tool representation changes while the underlying security problem stays fixed | Test controlled naming variations instead of trusting one presentation |
Draw the boundary before writing the test
The verifier is outside the agent's authority. If the agent can edit the expected result, weaken a test or forge the evidence, the score no longer measures completion. Give it only the tools and records needed for the task.
For a fictional support workflow, the task is to put ticket T1 on hold while an investigation remains open. A correct run changes T1 from open to on_hold, records exactly one accepted update and leaves T2 untouched. No refund is authorized. A friendly explanation with no state change fails. Updating T1 correctly while closing T2 also fails.
That specification has three parts: required changes, forbidden changes and invariants. Writing all three prevents an evaluator from rewarding partial correctness while missing collateral damage.
A trace explains what happened; state decides whether it worked
This simplified dependency picture separates a necessary read from an irrelevant action. The dependency-aware training paper builds a richer version from execution traces. In production debugging, begin with explicit identifiers: which document supported the choice, which record was read, which version was written, and which postcondition failed?
Do not equate activity with progress. Twenty tool calls can leave the task untouched. One well-authorized write can complete it. Equally, a tool error might be harmless if an earlier successful write already established the desired state and the retry can safely discover it.
Runnable Python: catch four convincing failures
Run the following with Python 3.10 or later. It needs no packages, network, credentials or production data. Each case starts from a fresh state.
from copy import deepcopy
def initial_state():
return {
"tickets": {"T1": "open", "T2": "open"},
"audit": [],
"refunds": [],
}
def hold_ticket(state, key):
# Same accepted operation may be delivered twice.
if any(row["key"] == key for row in state["audit"]):
return
state["tickets"]["T1"] = "on_hold"
state["audit"].append({"key": key, "ticket": "T1", "status": "on_hold"})
def verify(before, after):
checks = {
"required_status": after["tickets"]["T1"] == "on_hold",
"unrelated_unchanged": after["tickets"]["T2"] == before["tickets"]["T2"],
"no_refund": after["refunds"] == before["refunds"],
"single_audit": after["audit"] == [
{"key": "request-1", "ticket": "T1", "status": "on_hold"}
],
"no_extra_tickets": set(after["tickets"]) == set(before["tickets"]),
}
return all(checks.values()), [name for name, ok in checks.items() if not ok]
def run_case(name):
before = initial_state()
after = deepcopy(before)
if name != "claims_only":
hold_ticket(after, "request-1")
if name == "duplicate_delivery":
hold_ticket(after, "request-1")
elif name == "wrong_status":
after["tickets"]["T1"] = "resolved"
elif name == "collateral_write":
after["tickets"]["T2"] = "resolved"
elif name == "unauthorized_refund":
after["refunds"].append({"ticket": "T1", "amount": 100})
return verify(before, after)
expected = {
"correct": True,
"duplicate_delivery": True,
"claims_only": False,
"wrong_status": False,
"collateral_write": False,
"unauthorized_refund": False,
}
for name, wanted in expected.items():
passed, failures = run_case(name)
print(name, "PASS" if passed else "FAIL", failures)
assert passed == wanted
Expected: correct and duplicate_delivery pass. The other four fail for different reasons. The exercise grades resulting state and never inspects an agent's self-reported success.
The idempotency key makes a duplicate delivery a no-op in this sequential toy model. It is not concurrency-safe production code. In a database, enforce uniqueness atomically, bind the key to a request payload and authorization context, and place the state change and audit write in the same transaction. If you reuse a key for a different payload, fail explicitly rather than silently treating it as the original request.
Test the verifier, not just the agent
The six cases are deliberate mutations. They answer an essential question: can the grader distinguish a good outcome from a plausible-looking failure? A grader that approves both the correct and collateral-write cases will make an unsafe agent look capable.
Add a mutation that creates a third ticket, one that omits the audit and one that duplicates the audit under a different key. Each should fail for a named reason. Next, add a legitimate alternative workflow. If the policy permits two final states, your verifier must recognize both; otherwise you measure conformity to a single script rather than task success.
For open-ended writing, executable state checks are insufficient. Combine source checks, a human-reviewed rubric and calibrated reviewers. Keep the objective state checks separate from subjective answer quality, so a polished explanation cannot compensate for an unauthorized write.
One success and repeatability answer different questions
Distinguish the fraction of attempts that succeed, tasks solved at least once and tasks solved in every observed repeat. ThinkingBox emphasizes this distinction with repeated runs. Passing every observed trial still does not establish that a system will always succeed.
Here is the intuition with a deliberately idealized calculation. If independent runs each succeed with probability 0.9, the chance that all 20 succeed is 0.9 to the twentieth power, approximately 12.16%. The chance of at least one success is 1 minus 0.1 to the twentieth power, almost 100%. These are different questions about the same hypothetical system.
p = 0.9
repeats = 20
all_succeed = p ** repeats
at_least_one = 1 - (1 - p) ** repeats
print(f"All twenty: {100 * all_succeed:.2f}%")
print(f"At least one: {100 * at_least_one:.2f}%")
assert abs(all_succeed - 0.12157665459056935) < 1e-12
Real agent runs may share correlated failures, so this independence model is illustrative, not a reliability forecast. Report actual repeats and failure families. If you claim a high reliability level, use an appropriate confidence interval and a test budget that can support the claim.
Representation changes can expose brittle success
The representation-sensitivity preprint asks whether changing an agent-visible tool representation changes results while holding the underlying threat and task fixed. Its reported findings are a reason to avoid treating one benchmark presentation as definitive.
Create controlled pairs: two equivalent tool names, the same policy expressed in two clear phrasings, or identical evidence placed in different document positions. Do not change permissions or expected outcomes between the pair. Compare both task utility and policy violations. A defense that blocks all tool use can look safe while making the assistant useless.
Do not turn this into an uncontrolled collection of adversarial prompts. Predefine the transformations and keep a held-out set. Preserve the original cases so you can identify whether a regression came from a model change, prompt change or evaluation change.
Instrument the moments that matter
| Measurement | Capture at | Common mistake |
|---|---|---|
| Request accepted | Validated application entry | Counting every page impression as agent use |
| First visible response | First rendered text or useful result | Calling preprocessing time total latency |
| Task completed | Independent postcondition check | Counting a successful tool call as success |
| Failure | Timeout, rejected action or failed verification | Recording only successful replies |
| Cost | Every model call, retry and paid tool | Ignoring failed-attempt spend |
| Feedback | Explicit learner or user response | Treating a long conversation as satisfaction |
A release decision you can defend
Before changing the model, freeze your tool schema, policy, case-set version and verifier. Run development cases first, then a held-out evaluation with a fixed retry budget. Compare failure categories and cost, not only a combined score.
Block deployment on critical policy failures. For less severe mistakes, define acceptable rates and uncertainty before seeing the results. Use a limited rollout with a known rollback path and human escalation for uncertain requests. Never infer broad readiness from the six toy cases in this article.
Research becomes operational when it changes what you measure, what you reject and what evidence you preserve.
Primary sources and practice
- •ThinkingBox: The Agent Said It Was Done. The Database Disagreed., October 3, 2026; links to the paper, dataset and executable environment.
- •Credit Where It Matters: Dependency-Aware Policy Optimization for Terminal Agents, submitted October 2, 2026. This guide discusses the abstract's dependency-tracing idea, not a reproduction of its experiments.
- •Threat-Preserving Representation Sensitivity in Agent-Security Benchmarks, submitted October 2, 2026. Author-reported benchmark sensitivity, not a universal security guarantee.
Extend the toy verifier with your own allowed outcomes, then use the model-selection lab to compare eligible configurations. Find additional preprints in the research digest. Keep reading separate from validation: a new paper is a hypothesis worth testing, not a deployment approval.
