TL;DR
Get tech for your team delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Microsoft and Hugging Face have made ThinkingBox available through Hugging Face, a benchmark that grades AI agents on backend records and side effects rather than their answers alone. In 507 workflows tested 20 times each, the authors report frequent state-check failures, including among runs that ended without tool errors.
Microsoft and Hugging Face have released ThinkingBox, a benchmark that checks whether AI agents leave business systems in the required state, rather than judging them mainly by their tool calls or final responses. Its authors evaluated agents on 507 business workflows, with 20 runs per task, to measure both task completion and consistency.
ThinkingBox runs an agent through isolated sessions using Model Context Protocol (MCP) tools, then checks the backend records and side effects against an executable task requirement. The authors say the benchmark covers retail, auto insurance, travel, neobanking and consulting workflows. It is available through Hugging Face, and the blog describes running it through OpenEnv.
The report illustrates the test with a retail support case involving a $745 appliance delayed by a carrier exception. The agent checked the order and policy, opened and documented a support ticket, then marked it resolved and sent a generic response. But the task required the ticket to remain on hold while the carrier issue was unresolved. The example’s executable check fails on that status field, despite the agent’s orderly-looking tool sequence.
In a separate common-set analysis, the authors report 121,680 valid trials across 12 models, of which 79,853 failed executable checks. Among those failed trials, 67.24% ended without a reported final tool error after a state-changing tool was used. The checks identified wrong field values in 77.61%, unintended extra effects in 43.30%, and missing required effects in 25.36% of failures; these categories overlap. These figures are the authors’ benchmark results, not an independent audit.
Why Backend Checks Change Agent Scores
ThinkingBox highlights a gap between an agent appearing to complete a task and actually leaving a system in the requested condition. A plausible answer, a clean tool response or a completed workflow can all coexist with an incorrect record. In customer support, finance or insurance, that mismatch could affect what staff see and what action is taken next.
The repeat-run design also addresses a different question from a typical single-attempt score: does the agent succeed consistently? The authors distinguish pass@1, the share of individual attempts that pass; pass@20, whether a task passes at least once in 20 tries; and observed 20/20, whether every recorded run passes. A model that succeeds once but fails on other attempts may look capable under one metric without being dependable in routine use.
Those distinctions matter to organizations deciding whether an agent can handle workflows without close review. The benchmark offers a way to test failures such as incorrect field values or unintended changes. It does not, on its own, establish how a system will perform in a particular company’s production environment.
As an affiliate, we earn on qualifying purchases.
From Tool Calls to Task Outcomes
Many evaluations of AI agents focus on whether they select appropriate tools, follow instructions or produce a convincing final answer. ThinkingBox’s authors argue that these are proxies for the outcome: only inspection of the resulting records can establish whether a workflow’s required state was reached.
The report repeats each task 20 times, starting from an identical clean backend, and reports several measures rather than treating one successful attempt as proof of reliability. Its overall pass@1 table lists Claude Opus 5.5 at 67.16%, the highest score shown, and Kimi-K3 at 57.37% as the leading open-weight model in that table. These are the authors’ reported results for their selected tasks and models; they are not guarantees of performance on other tasks or deployments.
The publication is a joint Microsoft and Hugging Face blog post describing the benchmark and its paper. The authors say ThinkingBox can be run through Hugging Face and OpenEnv. The retail case is adapted from one benchmark task, rather than presented as a report of an actual customer incident.
“A tool call is not an outcome.”
— Microsoft and Hugging Face, in the ThinkingBox blog post
Limits of the Reported Results
The blog post reports benchmark results, but the supplied material does not establish how closely the test workflows match real deployments or how scores would change with different tools, policies or backend systems. Performance on ThinkingBox should not be read as a direct forecast of an agent’s reliability in a specific organization.
The report also does not show that every failed trial caused real-world harm; benchmark checks record task-level mismatches, including extra or missing effects. The cited error categories overlap, so they cannot be added to estimate a total number of distinct failures. The available material does not provide independent replication or a comparison against production incident rates.
Running and Validating the Benchmark
Readers can access ThinkingBox through Hugging Face and use the OpenEnv route described by the authors to run benchmark tasks. The next useful steps are to inspect the benchmark paper and task definitions, then test whether its checks fit the workflows and safeguards an organization actually uses.
Further evidence would come from independent evaluations, tests on additional models and repeated comparisons with real operational outcomes. Until then, ThinkingBox’s published scores describe performance on the benchmark’s tasks and conditions, while broader claims about production reliability remain unverified.
Key Questions
What does ThinkingBox measure?
It evaluates whether an AI agent leaves backend records and side effects in the required state for a task, including whether it does so across repeated runs.
How many workflows were tested?
The authors report 507 business workflows, with each task run 20 times. A separate common-set analysis covered 121,680 valid trials across 12 models.
Why did the delayed-appliance example fail?
The agent marked the support ticket as resolved, but the task required it to remain on hold while the carrier exception was unresolved. The check failed on the ticket’s status.
Does a ThinkingBox score predict production reliability?
Not by itself. The results apply to the benchmark’s tasks and test conditions; the supplied report does not establish how scores translate to performance in a particular production system.
Source: rss
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
