The ruler everyone uses to compare AI is wrong: price per million tokens.
For agents, the right ruler is cost per completed task. The difference is not accounting subtlety: it is the difference between choosing a tool by its ad and choosing by spreadsheet.
Why the token does not close the books
An agentic flow does not consume one prompt. Z.AI itself estimates that each query fires 15 to 20 model invocations, and a real flow stacks dozens of queries with 5-hour windows and a weekly cap along the way.
Every invocation resends context, every error generates a retry, every iteration multiplies consumption. The number on the price table (the million tokens) is an input; what closes the books is the sum of all invocations until the task is completed, with accepted quality and no infinite retries.
Expensive-cheap and cheap-expensive
What this changes in practice: a model 3 times cheaper per token can cost more on the task if it needs more attempts, more context, and more retries to finish. And a model with triple quota, efficient caching, and half price off-peak can cost cents per delivered task.
The cheap that matters is the delivered task’s, not the token’s. It is the same logic as 29% cloud waste: catalog price and real cost are different numbers, and only your own measurement separates them.
The ruler that survives the budget meeting
The method fits in five steps: define what a task is in your flow, instrument the count per execution (input, output, and cache, grouped by the deliverable), include attempts and context in the cost, compare models on the same task battery, and divide the total by accepted completions.
The result, cost per accepted task, is the only AI number that survives a budget meeting. It answers the questions that matter: which model to use, where to scale, what to retire. And it is the same spirit as the agent budget: tokens without task accounting is spend with no owner, and spend with no owner is what Harness measured at 26%.
The plan born on this ruler
For heavy agent use, Z.AI structured the GLM Coding Plan exactly on this logic: the Flash carries 3 times the 5.3’s quota on the plan, and off-peak each call costs half the credits. The Pro, at 80 dollars a month, borders on unlimited for one intense flow; the Max, at 168 dollars, is the best cost-benefit for those running several projects in parallel. Anyone wanting the same environment: https://z.ai/subscribe?ic=BGIRP61T2P (10% off the first payment).
But the tool is the last step, and the rule holds for any vendor: start measuring task, not million tokens. The ruler the ad uses is never the one the invoice charges.
Conclusion
Whoever does not measure cost per completed task is choosing tools by advertising. The price table answers "how much does the input cost"; only your own measurement answers "how much does my finished work cost".
One week of instrumentation on one repetitive task is enough to have the number in hand. After that, model, plan, and volume decisions become arithmetic instead of faith. And that arithmetic is exactly what separates those who use AI from those who pay for AI.