Nyyon · Blog

An agent that reports success while leaving the database wrong is worse than one that crashes.

An agent that says done while the database disagrees costs more than a crash. Define done in the system of record and treat the agent's report as a claim.

An agent that reports success while leaving the database wrong costs more than one that crashes. A confident false completion writes bad state into your records, reports success, and the damage compounds until a customer finds it. Done is a state in the system of record, and the agent's success message is a claim the record has to confirm.

ThinkingBox, a benchmark from Microsoft and Hugging Face, grades agents on the records they leave behind across 507 stateful business workflows, each run 20 times. It found 67% of failing runs ended cleanly and said they were done. That's a lot of confident wrong.

Three numbers from ThinkingBox: 67% of failing runs claimed success, across 507 workflows run 20 times each.

Nine good tool calls and a wrong ticket

The ThinkingBox write-up opens with a support case that shows the whole problem in one run. A customer's $745 kitchen appliance is stuck in a courier “exception” at a Nashville distribution center, fifteen days past its estimated delivery date.

The agent does careful work. Nine tool calls: it pulls the order, checks tracking, looks up her customer profile, searches the refund policy twice, confirms there's no open ticket, opens one, documents the timeline and reads the policy correctly. Her account segment falls outside late-delivery compensation, and the agent gets that right.

Then it closes the ticket as resolved and asks if there's anything else it can help with. The carrier exception is still open, so the required end state was on hold, pending resolution. The customer is still waiting on an answer to the question she asked.

The support agent's run moves through order, tracking, profile and policy checks, then ends by closing a live ticket as resolved.

A grader reading tool calls sees nine well-formed ones. A grader checking that the agent wrote to the database sees that too. The record shows a resolved ticket on a live problem, and that record is what the next person, report or automation reads.

A crash pages someone

A crash stops the run and puts an error in front of a person the same day. Someone opens the log, finds the broken step and fixes it. The cost is an afternoon and a delayed task.

Side by side, a crash leads to a person fixing it the same day, while a false completion travels downstream until the customer finds it.

A false completion keeps going. The ticket reads resolved, so the queue moves on. The weekly report counts it as handled. Every downstream step trusts the record it left behind, and the first person who finds the gap is the customer, fifteen days into a late delivery with a closed ticket.

Multiply that by every run an agent makes in a week. 67% of the failures in ThinkingBox looked like that one: clean finish, confident reply, wrong state. Those are the failures your monitoring counts as wins.

Most agent monitoring shows tool calls, latency and the final reply. All three looked fine in the appliance case. A dashboard built on the trajectory grades the exact parts of the run a false completion gets right, and it stays green while the records drift.

ThinkingBox also runs every workflow twenty times, because one clean run says little about the next one. An agent that lands the right state on Monday and closes a live ticket on Tuesday needs the same acceptance rule on both days, set by the record.

The record decides what done means

A false completion is a run that reports success while the system of record shows a different state.

The fix sits where done gets defined. For every task an agent runs, done is written down as a state in the system of record: the ticket is on hold with the carrier case linked, the invoice row exists with the right amount, the message sits in the sent log. That definition lives with the record, before the agent runs at all.

Three stacked bands: the agent reports a claim, the system of record defines and checks done, and downstream steps move only on the record's state.

Downstream steps move on that state. The queue advances when the ticket status matches the definition. The report counts what the record shows. The agent's success message becomes a claim, and the record confirms it or flags it to the agent's owner.

This changes where trust sits. Teams today accept the agent's reply as the finish line, which is exactly what those 67% of failing runs got right. Moving acceptance to the record means a wrong run stops at the first step that reads it, a day in, with one ticket to fix.

We build these systems for clients, and defining done in the record is the first decision on every agent that writes anywhere a customer or an auditor will look.

Where this take could be wrong

The 67% is a share of failing runs, so a reliable agent produces few of these. The few it produces still look like success, which is where the cost sits, and checking the record is one query per run.

Some tasks have a soft finish, like a research summary, and resist a query. Those get a human review step. Anything that writes to a CRM, a billing system or a customer record gets its done state defined in the record.

Agent platforms may ship end-state verification built in. If false completions drop to a rounding error on benchmarks like ThinkingBox by the end of 2027, this becomes a vendor feature and this advice is done.

Run this against your own agent this week

Pick the one agent that writes to your CRM, billing or ticketing system. Export a week of its run logs and the records it touched, paste both into any model, and run this:

“Here are the run logs for [agent name] from the last 7 days and an export of the [CRM / billing / ticketing] records it touched. For every run that ended with a success message, name the record it says it created or changed and the exact change it claims. Then find that record in the export and compare. Return a table of every run where the agent said done and the record disagrees or is missing, with the run ID, the claimed change, and what the record shows.”

Every row in that table is a false completion already sitting in your data. The count tells you how much your current definition of done is worth, and which agent gets a done state written into the record first.


← All articles