News

ThinkingBox: agents can look done while the database disagrees

Microsoft and Hugging Face grade agents on database state, not just tool calls. ThinkingBox-Bench: 507 workflows, Opus 5.5 leads pass@1.

ThinkingBox: agents pass tools, fail the database

Microsoft and Hugging Face published ThinkingBox on 3 October 2026. The benchmark grades AI agents on terminal backend and database state, not only on the final reply or whether a tool call looks valid. Agents can look finished while the data they were supposed to change still disagrees.

What happened

ThinkingBox-Bench covers 507 stateful business workflows across five domains: retail, auto insurance, travel, neobank, and consulting. Each task is run 20 times. The reports include pass@1, pass@20, and observed 20/20 consistency, so a single lucky run does not hide fragile behavior.

On a common-set ablation, the team logged 121,680 valid trials across 12 LLM models. Of those, 79,853 attempts failed executable checks. About two thirds of the full trial set therefore missed the intended end state. Of those failures, 67.24 percent still terminated cleanly, invoked a state-changing tool, and reported no final tool error. In plain terms, many agents look successful in the chat while the backend never ends up correct.

Claude Opus 5.5 leads overall pass@1 at 67.16 percent. Among the open-weights models listed, Kimi-K3 is strongest overall and sits within a point of GPT-6 Astra on pass@1. It has broad coverage, but lower consistency than the top closed models. The work is available through Hugging Face and OpenEnv, with the paper at arXiv:2608.19741. Microsoft also posted a companion note on the Command Line blog.

Why it matters

Tool call success is a weak proxy for business success. If your agent books a refund, updates a policy, or moves money, the database row is the ground truth. Benchmarks that only score text and tool schemas can crown systems that fail quietly in production. ThinkingBox pushes evaluation toward side effects and multi-step state, which is closer to how agents are sold for retail, insurance, travel, and banking work.

It also forces a harder honesty check on demos. A clean termination with a state-changing tool and no error message is exactly the failure mode that looks good in a screen recording and bad in an audit log.

Dany's take

If you ship agent workflows against real systems, treat "tool succeeded" as incomplete. Track pass@1 and consistency under repeated runs, and verify the final database state yourself. I care less about a single high score and more about whether the same workflow holds 20 times in a row. ThinkingBox is a useful pressure test for that bar, and I will keep watching how new models move on it.

Source: Hugging Face / Microsoft: ThinkingBox. Also: Microsoft Command Line: ThinkingBox-Bench.

Source: huggingface.co

Newsletter

The AI news that matters, in your inbox.