The human’s card has a limit. The autonomous agent has none.
An agent consumes tokens across consecutive windows, without pause and without shame, and the invoice only shows up later. The FinOps question in 2026 changed from "how much does it cost to run AI" to "who holds the agent back when it decides to spend".
The gap the agent lives in
The numbers show the size of the hole: 73% of organizations already have an AI cost policy, but only 47% enforce it, according to the Harness survey. It is exactly in the back-and-forth between writing rules and enforcing rules that the agent lives.
While the human waits for approval, the agent iterates overnight. It does not ask for budget, does not notice it is in a loop, and does not find its own error funny: it repeats the pattern until the task finishes or the quota runs out, whichever comes first. Without a brake, the quota usually comes first.
At the scale inference has reached, the problem stopped being curious and became a budget line: with one trillion tokens a day being consumed in inference alone, agents without ceilings are bets on the company card.
The new discipline has a recipe
Budget per agent, ceiling per step, and the token meter as accounting unit. The recommendation repeated across the literature is to install hard brakes: a step limit that interrupts the agent and calls a human before the damage grows.
In practice, the recipe fits in four pieces:
- Ceiling per agent and per task: a maximum budget defined before execution, not after the invoice.
- Step limit with human escalation: the breaker that interrupts the loop and hands the decision back to whoever owns it.
- Usage dashboard with alerts: detecting a runaway agent became as basic as a CPU alert.
- Cost per completed task: the metric that closes the books and lets agents be compared against each other.
And the market already ships this accounting built in: plans designed for agents come with hourly-window quotas, off-peak discounts, and a transparent meter. The Z.AI GLM Coding Plan, which we use and recommend, works that way, and the window-limited quota is precisely the brake the agent needs. Autonomy without a ceiling is, at the end of the day, a blind bet.
The testimony of the brake
I am writing this article with seven agents running in parallel on my infrastructure, and what lets me sleep is not trust in them: it is the brake. Every front has a ceiling, every task has a step limit, and the usage panel is one tap away on my phone.
Trust in agents grows with usage, which is exactly why the brake needs to exist before the trust. The agent that never blew its ceiling is the one that has one; the rest simply has not had the week where it decides to spend.
The same principle that organizes cloud waste applies here: a resource with no owner and no limit is contracted waste. The agent without a budget is the decade’s idle instance, except it is spending instead of sitting.
Conclusion
Giving the agent a ceiling buys the right to let it run alone. FinOps’ next step lives less in the cloud dashboard and more in the prompt that defines how much the agent may spend before it becomes a human decision again.
If you run agents with no step limit, no per-task ceiling, and no usage panel, the question is not whether the bill will surprise you: it is which invoice. The brake is cheap, the recipe is known, and the discipline already has a name. All that is missing is installing it.