Skip to main content
Close
FinOps

Cost per task: the AI metric the token price hides

Gabriel Ferraresi· CEO | Tech86October 9, 20263 min
finopsaiagentstokenscost

The ruler everyone uses to compare AI is wrong: price per million tokens.

For agents, the right ruler is cost per completed task. The difference is not accounting subtlety: it is the difference between choosing a tool by its ad and choosing by spreadsheet.

Why the token does not close the books

An agentic flow does not consume one prompt. Z.AI itself estimates that each query fires 15 to 20 model invocations, and a real flow stacks dozens of queries with 5-hour windows and a weekly cap along the way.

Every invocation resends context, every error generates a retry, every iteration multiplies consumption. The number on the price table (the million tokens) is an input; what closes the books is the sum of all invocations until the task is completed, with accepted quality and no infinite retries.

Expensive-cheap and cheap-expensive

What this changes in practice: a model 3 times cheaper per token can cost more on the task if it needs more attempts, more context, and more retries to finish. And a model with triple quota, efficient caching, and half price off-peak can cost cents per delivered task.

The cheap that matters is the delivered task’s, not the token’s. It is the same logic as 29% cloud waste: catalog price and real cost are different numbers, and only your own measurement separates them.

The ruler that survives the budget meeting

The method fits in five steps: define what a task is in your flow, instrument the count per execution (input, output, and cache, grouped by the deliverable), include attempts and context in the cost, compare models on the same task battery, and divide the total by accepted completions.

The result, cost per accepted task, is the only AI number that survives a budget meeting. It answers the questions that matter: which model to use, where to scale, what to retire. And it is the same spirit as the agent budget: tokens without task accounting is spend with no owner, and spend with no owner is what Harness measured at 26%.

The plan born on this ruler

For heavy agent use, Z.AI structured the GLM Coding Plan exactly on this logic: the Flash carries 3 times the 5.3’s quota on the plan, and off-peak each call costs half the credits. The Pro, at 80 dollars a month, borders on unlimited for one intense flow; the Max, at 168 dollars, is the best cost-benefit for those running several projects in parallel. Anyone wanting the same environment: https://z.ai/subscribe?ic=BGIRP61T2P (10% off the first payment).

But the tool is the last step, and the rule holds for any vendor: start measuring task, not million tokens. The ruler the ad uses is never the one the invoice charges.

Conclusion

Whoever does not measure cost per completed task is choosing tools by advertising. The price table answers "how much does the input cost"; only your own measurement answers "how much does my finished work cost".

One week of instrumentation on one repetitive task is enough to have the number in hand. After that, model, plan, and volume decisions become arithmetic instead of faith. And that arithmetic is exactly what separates those who use AI from those who pay for AI.

Need expert guidance?

Schedule a consultation with our specialists.

AI FinOps for your company

Frequently Asked Questions

Because an agentic flow does not consume one prompt: Z.AI itself estimates that each query fires 15 to 20 model invocations, and a real flow stacks dozens of queries with 5-hour windows and a weekly cap along the way. The real task cost is the sum of all those invocations, including attempts and resent context, and it never appears on the per-million price table.

Yes. A model 3 times cheaper per token can cost more per task if it needs more attempts, more context, and more retries to finish. And a model with a bigger quota, efficient caching, and half price off-peak can cost cents per delivered task. The cheap that matters is the delivered task’s, not the token’s.

Plans designed for agents change the math without changing the advertised price: Z.AI’s GLM Coding Plan, for example, gives triple quota on the Flash model compared to the 5.3 and charges half the credits off-peak. Two levers that only appear when the ruler is cost per completed task. Want the plan: https://z.ai/subscribe?ic=BGIRP61T2P (link with 10% off the first payment).

On the team’s most repetitive task: code review, report generation, or ticket triage. Define the task, instrument one week, compute total cost divided by accepted completions, and repeat with a second model. Two rows on a spreadsheet and the budget conversation changes level.

For agents it is the most glaring (they iterate nonstop and [need their own ceiling](https://www.tech86.com.br/en/blog/orcamento-para-agentes-ia-finops-kill-switch)), but the ruler applies to any use: chatbot, content generation, and analysis all consume multiple invocations per delivery. Wherever there is a deliverable, there is a cost per deliverable, and that is what closes the books.

Blog, Get in Touch

Have a question about our articles or services? Our team is ready to help.

Schedule a Meeting

Book a time slot.

Schedule Now

Email

Send us a message.

[email protected]

WhatsApp

Quick conversation.

Address

Avenida Paulista, 1636 - São Paulo - SP - 01310-200

Tech86 Specialist

Online now

Hello! How can we help scale your business today?

Tech86 Engineering

We Value Your Privacy

We use cookies and similar technologies to optimize your experience, analyze site traffic, and personalize content. By clicking "Accept All", you agree to the use of all cookies. Read our Privacy Policy.