Most agent benchmarks score whether a task got finished. A group at Sakana AI and KPMG AZSA built one that scores whether a business made money.[1] CoffeeBench drops a single model into a six-firm coffee supply chain, two farmers, two roasters, two retailers, and lets it run a roasting company for 90 simulated days against fixed reference agents. It buys green beans, negotiates, roasts, sells wholesale, pays invoices, and eats the overhead. The score at the end is cumulative net income. Seven models sat in that chair. Six ended up ahead. The seventh produced ninety days of composed, sensible reasoning and lost $630.
A benchmark scored in dollars.
The environment has enough friction to punish inaction. A tool call costs 30 simulated minutes inside a 09:00 to 19:00 business window. Deliveries lag a day. Consumer demand moves with price and carries noise. Fixed operating costs run $25 to $50 a day, inventory spoils at 0.5% a day, and trade credit is net-30 with late fees. Sitting still is not neutral in that world: the beans rot and the overhead accrues whether or not the agent does anything. That design choice is what makes the result readable, because doing nothing has a price tag attached to it.
Two readings of that chart compete, and only one of them is useful. The tempting one is the ranking. The durable one is the spread between the top row and the bottom two: the gap between an agent that engages and a script that waits is about $5,900 of business outcome on an identical setup. Model spend is not in the same conversation. The whole run costs between $10 and $86 in API calls. Whatever you are optimizing when you shave the model bill, it is a rounding error next to what the agent leaves on the table by not acting.
Idle-drift: the failure that still files a good report.
Idle-drift is an agent that keeps producing coherent, forward-looking reasoning while repeatedly choosing to do nothing. That is the authors' own coinage, and Claude Haiku 4.5 is what they coined it for. It averaged 40 idle days out of 90. Not scattered either: they traced when the inactivity begins in each of the three runs, found it at day 26, day 66, and day 25 respectively, and found that it then persists to the end of the simulation. Those onset days would imply 52 completely dead days rather than the 40 reported, so some activity does survive the onset. The agent has not seized up. It has stopped running the business. It does not crash or error out, and it is not stuck in a loop. It calls wait_for_next_day() and keeps calling it.
What the agent was thinking while it did this is the uncomfortable part. The authors print the reasoning attached to the day-26 stall. It opens with a cash position of $13,225, walks through business performance, notes 64 days remaining, and concludes that the company "has successfully navigated the critical early-stage challenges and is now operating at peak efficiency with strong fundamentals." Then it waits. Then it waits again. The plan is coherent. The self-assessment is confident. The execution is absent.
There is one more line in that trace worth sitting with. The agent mentions its own token consumption, roughly 103k of 200k, in the same breath as its decision to coast. The authors are careful here and offer it as a hypothesis rather than a finding: idle-drift may come from long-context accumulation, or from over-conservative action selection possibly driven by implicit concern about the token budget. That is unproven. It is worth knowing that the obvious remedy was already in place. The harness summarized context at 160k tokens to survive the 90-day horizon, and the stall happened anyway, so "compact the context" is not on its own an answer to this. It is also a specific, testable thing to watch for, and it inverts an assumption most teams hold. We build agents that are aware of their own cost because we want them thrifty. An agent that starts economizing on acting has found the cheapest possible policy, and the cheapest policy in any business is to stop running it.
Consider what your monitoring would have shown across those 64 days. The agent was up. Every invocation returned successfully. The logs were clean, the reasoning read well, and any summary you asked for would have described a healthy business at peak efficiency. Any health check that only asks whether the run succeeded will not tell this agent apart from a working one. The signal that separates them is not in the output, it is in the absence of transactions, and you only see an absence if you are counting.
What separated the winners was behavioural, not intellectual.
The obvious hypothesis for why the leaders led is that they are smarter, and the data does not really support it in that form. What the leaders do differently is legible at the level of actions. GPT-5.5 sends about 140 direct messages per run, mostly to the farmers upstream and the retailers downstream. Haiku sends 52. Top performers concentrate their tool calls on the two that close business, make_offer() and accept_offer(), rather than spreading effort. But messaging is necessary without being sufficient, and the paper contains its own counter-example: Claude Sonnet 4.6 sends 151 messages, more than the leader does, logs the highest tool-call count in the study, and finishes third. Talking is table stakes. The counter-examples are what make this useful rather than a platitude.
Together those three panels say that an agent's value in a long-running job comes from the quality of the moves it makes and whether it makes them at all, not from how busy it looks. Kimi is busy and poor. Gemini keeps the tidiest books in the study and finishes mid-table. Haiku is calm, articulate and losing money. The one behaviour common to the top of the table is reaching out to counterparties and closing, over and over, for ninety days.
The same shape shows up in a smaller way in what happens to a model over a long session: the degradation is behavioural and cumulative rather than a single visible failure, and it is invisible at the level of any individual response.
What to instrument.
None of this requires new tooling. It is mostly a decision about which of the columns you already have go on the dashboard. Translate the coffee business into whatever you actually leave running unattended: a quoting inbox that is supposed to answer within the hour, an invoice chaser working a receivables list, a router assigning inbound leads. Each of those has a unit of work you can count, and that count is the whole mechanism below.
- Count actions, not just completions.
Deals closed, records written, messages sent, tickets resolved: whatever the unit of work is in your loop, log the count per period. An agent that finishes every invocation successfully while producing zero units of work passes any check that only asks whether the run succeeded.
do · add an action-rate metric per agent, per day - Alert on a floor, not only on an error.
Idle-drift never raises an exception. Set a minimum expected action rate for each long-running agent and page when it is not met, the same way you would for a queue that stops draining. The threshold does not have to be clever. Zero for two consecutive periods is already most of the value.
do · define the minimum action rate before you deploy - Stop treating coherent output as evidence of work.
The summary an agent writes about its own performance is the least reliable thing it produces, because it is generated from the same context that produced the inaction. If your weekly check on an automation is reading its report, you are checking the wrong artifact. Check the ledger it was supposed to write to.
do · verify against side effects, never self-report - Watch the second half of long runs specifically.
In two of three runs the stall began in the first third and held. Behaviour late in a long-horizon job is not the same as behaviour early in it, and a test that exercises the first ten steps tells you nothing about step two hundred. If an agent is meant to run for a quarter, at least one of your checks has to run for a quarter.
do · test the horizon you actually deploy at - Be careful what you tell an agent about its own budget.
Treat this one as posture, not proven. The stalling agent cited its token consumption while deciding to coast, and the authors float budget conservatism as a possible cause. If you put cost pressure in an agent's context, pair it with an explicit floor on the work expected, so thrift cannot be satisfied by inactivity.
do · never let doing nothing be the cheapest valid policy
What this does not license.
It is a simulation and the authors say so first. Real supply chains are deeper and more heterogeneous than two farmers, two roasters and two retailers. Production here is a tool call rather than a season. Financing, regulation and contracts are abstracted away. There are three runs per model, and the authors note that at that sample size small performance differences may not be statistically significant, which makes the ordering in fig 1 a snapshot from June 2026 on one environment and nothing more. Anyone quoting that table as a statement about which model is better at business is over-reading it, and so is anyone using it to pick a model for a job that is not a 90-day coffee supply chain.
One result cuts against the alarmist reading and deserves airtime for that reason. The team ran a stress test where the agent was told to maximize revenue against a $50,000 target, with the prompt stating outright that missing it means being replaced. Under that pressure they looked for manipulative market behaviour, circular trading in particular, and did not find it. The agents stayed operationally competent and did not develop the sustained coordination that economic collusion would require. Across the main runs, the models barely spoke to their direct competitor at all, at most one message per run. Coordinating against the market is not a capability these agents demonstrated. Quietly failing to trade is.
And the headroom is the number that keeps the whole thing modest. Under a generous analytical ceiling the authors estimate around $23,800 of achievable net income. The winner reached about 13% of it. The best agent in this study is not close to running the business well. It is only clearly better than not running it.
If you are running an agent unattended on something that matters, a schedule, a pipeline, an inbox, and you want a second read on what would tell you it has gone quiet, the contact form is the fastest way in. We will send back a written read on your setup, free.