The Agent Said It Was Done. The Database Disagreed.
AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get tech for your team delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Microsoft and Hugging Face have made ThinkingBox available through Hugging Face, a benchmark that grades AI agents on backend records and side effects rather than their answers alone. In 507 workflows tested 20 times each, the authors report frequent state-check failures, including among runs that ended without tool errors.

Microsoft and Hugging Face have released ThinkingBox, a benchmark that checks whether AI agents leave business systems in the required state, rather than judging them mainly by their tool calls or final responses. Its authors evaluated agents on 507 business workflows, with 20 runs per task, to measure both task completion and consistency.

ThinkingBox runs an agent through isolated sessions using Model Context Protocol (MCP) tools, then checks the backend records and side effects against an executable task requirement. The authors say the benchmark covers retail, auto insurance, travel, neobanking and consulting workflows. It is available through Hugging Face, and the blog describes running it through OpenEnv.

The report illustrates the test with a retail support case involving a $745 appliance delayed by a carrier exception. The agent checked the order and policy, opened and documented a support ticket, then marked it resolved and sent a generic response. But the task required the ticket to remain on hold while the carrier issue was unresolved. The example’s executable check fails on that status field, despite the agent’s orderly-looking tool sequence.

In a separate common-set analysis, the authors report 121,680 valid trials across 12 models, of which 79,853 failed executable checks. Among those failed trials, 67.24% ended without a reported final tool error after a state-changing tool was used. The checks identified wrong field values in 77.61%, unintended extra effects in 43.30%, and missing required effects in 25.36% of failures; these categories overlap. These figures are the authors’ benchmark results, not an independent audit.

At a glance
announcementWhen: Announced in the Microsoft and Hugging…
The developmentMicrosoft and Hugging Face published and released ThinkingBox, a benchmark for testing whether AI agents reliably produce the required backend state across repeated business workflows.

Why Backend Checks Change Agent Scores

ThinkingBox highlights a gap between an agent appearing to complete a task and actually leaving a system in the requested condition. A plausible answer, a clean tool response or a completed workflow can all coexist with an incorrect record. In customer support, finance or insurance, that mismatch could affect what staff see and what action is taken next.

The repeat-run design also addresses a different question from a typical single-attempt score: does the agent succeed consistently? The authors distinguish pass@1, the share of individual attempts that pass; pass@20, whether a task passes at least once in 20 tries; and observed 20/20, whether every recorded run passes. A model that succeeds once but fails on other attempts may look capable under one metric without being dependable in routine use.

Those distinctions matter to organizations deciding whether an agent can handle workflows without close review. The benchmark offers a way to test failures such as incorrect field values or unintended changes. It does not, on its own, establish how a system will perform in a particular company’s production environment.

Amazon

AI workflow testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From Tool Calls to Task Outcomes

Many evaluations of AI agents focus on whether they select appropriate tools, follow instructions or produce a convincing final answer. ThinkingBox’s authors argue that these are proxies for the outcome: only inspection of the resulting records can establish whether a workflow’s required state was reached.

The report repeats each task 20 times, starting from an identical clean backend, and reports several measures rather than treating one successful attempt as proof of reliability. Its overall pass@1 table lists Claude Opus 5.5 at 67.16%, the highest score shown, and Kimi-K3 at 57.37% as the leading open-weight model in that table. These are the authors’ reported results for their selected tasks and models; they are not guarantees of performance on other tasks or deployments.

The publication is a joint Microsoft and Hugging Face blog post describing the benchmark and its paper. The authors say ThinkingBox can be run through Hugging Face and OpenEnv. The retail case is adapted from one benchmark task, rather than presented as a report of an actual customer incident.

“A tool call is not an outcome.”

— Microsoft and Hugging Face, in the ThinkingBox blog post

Limits of the Reported Results

The blog post reports benchmark results, but the supplied material does not establish how closely the test workflows match real deployments or how scores would change with different tools, policies or backend systems. Performance on ThinkingBox should not be read as a direct forecast of an agent’s reliability in a specific organization.

The report also does not show that every failed trial caused real-world harm; benchmark checks record task-level mismatches, including extra or missing effects. The cited error categories overlap, so they cannot be added to estimate a total number of distinct failures. The available material does not provide independent replication or a comparison against production incident rates.

Running and Validating the Benchmark

Readers can access ThinkingBox through Hugging Face and use the OpenEnv route described by the authors to run benchmark tasks. The next useful steps are to inspect the benchmark paper and task definitions, then test whether its checks fit the workflows and safeguards an organization actually uses.

Further evidence would come from independent evaluations, tests on additional models and repeated comparisons with real operational outcomes. Until then, ThinkingBox’s published scores describe performance on the benchmark’s tasks and conditions, while broader claims about production reliability remain unverified.

Key Questions

What does ThinkingBox measure?

It evaluates whether an AI agent leaves backend records and side effects in the required state for a task, including whether it does so across repeated runs.

How many workflows were tested?

The authors report 507 business workflows, with each task run 20 times. A separate common-set analysis covered 121,680 valid trials across 12 models.

Why did the delayed-appliance example fail?

The agent marked the support ticket as resolved, but the task required it to remain on hold while the carrier exception was unresolved. The check failed on the ticket’s status.

Does a ThinkingBox score predict production reliability?

Not by itself. The results apply to the benchmark’s tasks and test conditions; the supplied report does not establish how scores translate to performance in a particular production system.

Source: rss

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Essential Guide To AI Capabilities And Safeguards For Astra

OpenAI has announced a ‘Path to Astra’ emphasizing critical AI capabilities and frontier safety safeguards, but details remain unclear about Astra’s nature and status.

DeepSeek Takes On Anthropic’s Claude Code: A New AI Challenge Unveiled

DeepSeek has announced efforts to compete with Anthropic’s Claude Code, but details on product, performance, and release are still unclear.

The AI Architecture Behind Inside Room 107 Of 175 For Operation Sandstorm

Detailed analysis of the AI-driven design behind Room 107 of Operation Sandstorm, highlighting its immersive weather simulation technology.

Open ASR Leaderboard Expands With First Language From The Global South

The Hugging Face Open ASR Leaderboard now includes Hindi and Indian English, marking the first Indic and Global South languages on the platform, with new diverse datasets.