The Agent Said It Was Done. The Database Disagreed.
A new benchmark from Microsoft and Hugging Face grades AI agents on what actually happened to the data, not on how confidently the agent described what it did.
A customer contacts support about a $745 kitchen appliance. The package has been stuck in a courier exception in Nashville for 15 days. The AI agent assigned to the ticket makes nine tool calls, reads the refund policy correctly, and closes the ticket as resolved. According to the benchmark researchers who studied this scenario, the required end state was "on hold." The agent set it to "solved." The customer's actual question went unanswered.
The failing check involves a single field.
ThinkingBox-Bench, a joint release from Microsoft and Hugging Face reported on October 3, 2026, uses this scenario to illustrate what it is measuring. It measures the state of the database when the agent stopped working, not the agent's reasoning chain or the grammar of its final reply.
How Most Benchmarks Grade Agents
Most agent evaluations ask a version of the same question: did the agent say the right thing? The model gets a task, produces a response, and a judge, sometimes human, sometimes another model, scores the quality of that output. This works well enough for question-answering and single-turn tasks. It runs into trouble when the agent's job is to change something.
An agent operating inside an enterprise system is not primarily a text generator. It is a process executor. It reads records, updates fields, triggers workflows, and modifies state across systems that other processes and people depend on. In that context, a fluent summary of what the agent intended to do is not evidence that it did anything. It is a claim.
Trajectory-based evaluation catches some of this. Researchers record every tool call an agent makes and check whether the sequence looks reasonable. But a reasonable-looking sequence of actions can still leave the underlying data in the wrong state. The agent in the Nashville ticket scenario made nine tool calls. Nine is not an unreasonable number of calls for a complex support resolution. The problem was not the volume of actions. It was the final field value.
How ThinkingBox Works
ThinkingBox-Bench, as reported by Microsoft and Hugging Face, constructs 507 stateful workflows across five business domains:
Retail: 98 tasks
Auto insurance: 100 tasks
Travel: 104 tasks
Neobank: 104 tasks
Consulting: 101 tasks
Every workflow is run 20 times from an identical, clean backend state. This is not 20 variations of the same task. It is the same task, reset to the same starting conditions, executed 20 separate times. The goal is to separate a model that can solve a problem from a model that will solve it reliably.
Of the 507 workflows, 477 are graded on database state alone. Researchers check whether the agent wrote the correct values to the correct fields and did not introduce unintended side effects. The remaining 30 tasks add response rubrics to the state check, because some workflows require the agent to both change data correctly and communicate something specific to the user.
The result is 121,680 valid trials across 12 models. The benchmark is available through Hugging Face and OpenEnv. Code is released under an MIT license. Data is released under CDLA-Permissive-2.0. The synthetic tasks are modeled on real enterprise patterns; the customers are not real.
Numbers in Context
Across all 121,680 trials, 79,853 failed their executable checks. That is roughly 65.6% of all runs ending in a state the benchmark treats as incorrect.
The failure breakdown is where the numbers get uncomfortable. Of those failed runs:
67.24% terminated cleanly, with no tool error of any kind
77.61% involved wrong field values
43.30% produced unintended extra effects
25.36% were missing required effects entirely
The 67.24% figure deserves a pause. Nearly two-thirds of failures produced no visible signal that anything had gone wrong. The agent finished, returned a response, and the system recorded no error. The only way to know the task failed was to inspect the database directly.
This is the core problem ThinkingBox is designed to surface. In a production environment without state verification, the majority of these failures would have looked, from the outside, like successes.
Breadth vs. Consistency
The benchmark distinguishes three metrics, and the gaps between them are instructive.
Pass@1 measures whether a single attempt succeeds. This is the number most benchmarks report and the one most deployment conversations rely on.
Pass@20 measures whether a model solved a task at least once across 20 runs. A model that can do something occasionally will show up here.
Observed 20/20 measures whether a model solved the same task every single time across all 20 runs. This is the consistency floor, and it is where the ranking changes significantly.
Claude Opus 5.5 leads on pass@1 at 67.16%. Claude Opus 5 follows at 66.50%. Both models solve the same 241 of 507 tasks successfully across all 20 runs.
Kimi-K3 presents a different profile. It solves 93.89% of tasks at least once, reaching 476 of the 507 workflows on pass@20. That is the broadest coverage in the benchmark. But it solves only 13.41% of tasks on every attempt, 68 out of 507. Kimi-K3 is also the strongest open-weight model on pass@1 at 57.37%.
The practical interpretation is that pass@20 captures capability and pass@1 captures what happens in a single deployment. A model with high pass@20 and low observed 20/20 is inconsistent. It can reach the right answer, but not reliably. Whether that matters depends entirely on the deployment. A model helping draft an email can be inconsistent without much consequence. A model updating insurance policy records probably cannot.
"A trajectory is a claim. Database state is the evidence. Repetition is the trust test."
The Cost of Dependability
The authors provide cost estimates at list rates, and the numbers vary considerably depending on what you are optimizing for.
Cost per dependable task (observed 20/20 performance):
GPT-5.4: $6.80 per dependable task, 128 tasks
GPT-6 Astra: $7.45, 231 tasks
Claude Opus 5.5: $7.80, 241 tasks
Claude Opus 5: $13.30, 241 tasks
GPT-5.4 is cheapest per dependable task but covers fewer of them. Opus 5.5 and GPT-6 Astra are close in per-task cost and both cover more ground. Opus 5 reaches the same task count as Opus 5.5 at significantly higher cost.
The cheapest model on a per-single-success basis is GPT-5.6 Sol at $0.127 per success. That figure reflects cost per individual passing run, not per reliably solved task, a distinction that matters for how an operator thinks about deployment economics.
None of these numbers represent the full cost of running an agent in production. They are list-rate estimates from the authors for what the benchmark trials cost, offered as a point of comparison rather than a procurement guide.
Failure Signatures
Across all 12 models, the benchmark categorizes failures by type:
Tool usage failures: 79.9%
Wrong state updates: 10.3%
Incomplete user resolutions: 7.0%
No state-changing action taken: 2.9%
Tool usage failures dominate. This covers cases where the agent called tools incorrectly, in the wrong order, or in ways that produced field values the benchmark treats as wrong. Wrong state updates and incomplete resolutions are smaller shares but carry significant operational weight. An agent that takes no state-changing action at all, 2.9% of failures, is at least easy to detect. An agent that updates the wrong fields confidently is not.
Domain performance varies substantially. Retail tasks average 59.52% pass@1. Auto insurance averages 33.83%. The gap suggests that some domains have structural properties that make them reliably harder for current models, whether that is greater constraint complexity, longer dependency chains, or conditional logic that requires more precise state tracking. The benchmark does not attribute the gap to specific causes.
What the Study Does Not Prove
The authors are straightforward about the limits of their work, and it is worth taking those seriously.
The tasks are synthetic. They are modeled on real enterprise workflows, but the customers are not real, and the edge cases in a production system are not always the edge cases a benchmark designer anticipates. A model that performs well here may still encounter novel failure modes in deployment that the benchmark did not include.
The authors have not measured whether the fixes they suggest actually improve performance. They propose several mitigations, but the paper does not contain ablation results showing lift from those interventions. Operators who want to know whether a specific fix helps will need to run their own evaluations.
The cost figures are estimates at list rates. Real deployment costs depend on retry logic, orchestration overhead, prompt engineering, and the specific workflows being run. A number from a benchmark trial is not a billing forecast.
The 12 models evaluated here represent a snapshot of the field as of the benchmark's publication. The methods ThinkingBox uses to grade state should remain useful as models improve, but the specific rankings will shift.
What This Means for Deployment
For anyone building or evaluating AI agents that write to real records, the benchmark offers a concrete reframing.
The question is not whether the agent can reason about the task. Most current models can construct a plausible action sequence for a support ticket, a policy update, or a transaction workflow. The question is whether the agent lands the data in the right state, every time, without requiring a human to verify the output after the fact.
ThinkingBox isolates three things worth tracking before a stateful agent goes into production: what the agent does on a single attempt, what it is capable of doing if given multiple tries, and how often it succeeds on the first try with no variance. These are different numbers and they answer different questions.
The Nashville ticket is a useful frame. The agent produced a plausible-looking resolution. It made nine tool calls. It read the policy. The system recorded no error. And the ticket ended up in the wrong state, with the customer's question unanswered. In a queue of thousands of similar tickets, that failure would have been invisible without a check on the terminal state.
Building that check in, and using a benchmark that demands it, is the practical upshot of this work.

