posts

Valid tool calls hide broken database records

When evaluation harnesses only verify tool syntax and exit codes, agents can easily pass execution checks while quietly corrupting data. In ThinkingBox, a benchmark from Microsoft and Hugging Face evaluating 507 stateful workflows across twelve frontier and open models, two-thirds of failed runs finished without a single runtime error or invalid tool call. Across nearly 80,000 recorded failures in that specific test suite, the breakdown only surfaced because an executable check inspected the database: three-quarters of those clean runs wrote incorrect values into fields, and over 40 percent triggered unintended side effects.

The benchmark showed that high task coverage can mask acute instability: Kimi-K3 solved nearly 94 percent of tasks at least once across twenty attempts, but completed all twenty runs correctly on only 13 percent. Claude Opus 5 solved fewer distinct tasks overall, yet repeated the correct end state across all twenty runs on nearly half the suite. A benchmark reporting pass@1 or pass@20 measures whether a model can hit the right trajectory once under ideal conditions. In production, an agent that resolves a ticket correctly once and corrupts the record four times out of twenty is unusable.

For teams building agentic workflows against enterprise backends, testing against conversational traces or mock tool responses is effectively grading an agent on its confidence. The only reliable evaluation boundary is the database state after execution: whether records match expected values and side effects remain bounded. If the verification layer cannot assert terminal state across repeated runs, the benchmark evaluates conversational fluency while leaving backend integrity unchecked.

Source: huggingface.co/blog/microsoft/thinkingbox

← all posts