Pular para o conteúdo principal
Close
FinOps

Kubernetes at 8%: The Utilization Crisis That Costs 27K Reais per Month in Idle GPU

Gabriel Ferraresi· CEO | Tech86July 18, 20264 min
finopskubernetesgpucloudcostoverprovisioningkarpenterchargeback

23,000 Kubernetes clusters. Average CPU: 8%. The Cast AI 2026 report analyzed clusters in production and the number dropped from 10% to 8% in one year. You provision 12x more compute than you use. GPU is worse: average utilization of 5%, which means 20x more capacity allocated than consumed. We analyzed the data and the signal is clear — the problem is not Kubernetes. It is provisioning for peak and paying for idle.

The scale: 23,000 clusters, CPU at 8%, GPU at 5%

According to the Cast AI 2026 report, which analyzed 23,000 clusters in production, average CPU utilization dropped from 10% to 8% in one year. You provision 12x more compute than you use. GPU is worse: average utilization of 5%, which equals allocating 20x more GPU capacity than is consumed.

CPU overprovisioning jumped from 40% to 69%. Memory overprovisioning: 79%. Engineers request resources for peak. The peak never arrives. The Cluster Autoscaler provisions nodes to match inflated requests, not real usage. Every node provisioned for a peak that never occurs is idle money.

The cost of idle: 83% of the container bill

According to Datadog, 83% of container cost is idle. 54% of the cluster idle. 29% of the workload idle. The cost is not in the workload that runs — it is in the workload that does not run, in the node that waits, in the GPU that nobody consumes.

Only 14% of companies do Kubernetes chargeback, according to the Cast AI 2026 report. 24% do not monitor K8s cost at all. You cannot cut what you do not measure. And what is not measured grows — because every team asks for more resources and nobody gives any back.

And the price went up. AWS raised the price of GPU Capacity Blocks by 15% in January 2026. The first time in 20 years that AWS raised GPU prices instead of lowering them. Idle GPU got more expensive at the exact moment when more GPU is idle.

Brazil: R$27,742 per month of idle GPU

In Brazil, each dollar of cloud costs 5.60 reais. An idle H100 cluster costs $4,954 per month. In reais: R$27,742 per month of idle GPU. It is not an isolated case — it is the average for anyone who provisions H100 for inference and never comes close to saturating it.

According to Brazilian market data, 65% of Brazilian companies have no visibility into cloud spend. Only 19.5% adopted FinOps. The result is predictable: without visibility, without rightsizing, without chargeback, the bill grows until someone asks where the number comes from. By the time the question arrives, the waste has already become habit.

What Tech86 implements

We do not sell generic monitoring. We implement five levers with measurable effect on the bill:

  1. Automated rightsizing. Cuts provisioned CPU by 50% without OOM kills. We analyze real usage and recommend requests at the 95th percentile.
  2. Spot for GPU. Less than 2% of GPUs run on Spot today. Spot cuts 60 to 90% of the cost of interruption-tolerant workloads.
  3. Karpenter. Provisions nodes in 45 seconds. The Cluster Autoscaler takes 3 to 4 minutes. According to Salesforce, the replacement across 1,000 EKS clusters cut 70% of cost.
  4. Turn off dev outside business hours. 20 to 30% of the bill comes from environments nobody uses at night.
  5. FinOps with chargeback by namespace. When the team sees the cost, it reduces 15 to 20%.

Each lever has a number. It is not a promise — it is what the Cast AI 2026 report, Datadog, and Salesforce have already measured in production.

The ELO case: chargeback in reais

ELO, the Brazilian card brand, launched Kubernetes chargeback in January 2026. It standardized reports in reais. It recovered 5 analyst days per month that were previously spent reconciling dollar invoices with real costs. The gain was not only financial — it was operational. The FinOps team stopped fighting exchange rates and started fighting rightsizing.

The ELO case confirms the pattern: when cost becomes a namespace responsibility, the team acts. When cost stays in a global spreadsheet, nobody acts.

Conclusion: provisioning for peak is the problem

Provisioning for peak and paying for idle is the problem. Kubernetes is the tool. The tool is not the cause of the waste — the provisioning model is. As long as requests are inflated for a peak that never arrives, the Cluster Autoscaler will provision nodes for an event that never occurs and the bill will grow alongside the fear of OOM kills.

At Tech86, we implement automated rightsizing, Spot for GPU, Karpenter, dev shutdown outside business hours, and chargeback by namespace. The result is less idle, fewer nodes, and fewer reais sitting still. R$27,742 per month of idle GPU is not a destination — it is a symptom of a provisioning model that needs to change.

blog.cta_consulting_title

blog.cta_consulting_subtitle

FinOps for Kubernetes and GPU

Frequently Asked Questions

According to the Cast AI 2026 report, which analyzed 23,000 clusters in production, average CPU utilization dropped from 10% to 8% in one year. That means companies provision 12x more compute than they use. Average GPU utilization is even worse: 5%, which equals allocating 20x more GPU capacity than is consumed.

In Brazil, each dollar of cloud costs 5.60 reais. An idle H100 cluster costs $4,954 per month, which equals R$27,742 per month of idle GPU. According to Datadog, 83% of container cost is idle — 54% of the cluster idle and 29% of the workload idle. AWS also raised the price of GPU Capacity Blocks by 15% in January 2026, the first increase in 20 years.

According to the Cast AI 2026 report, CPU overprovisioning jumped from 40% to 69% and memory overprovisioning stands at 79%. Engineers request resources for peak, but the peak never arrives. The Cluster Autoscaler provisions nodes to match inflated requests, not real usage. The result is that you pay for capacity reserved for an event that never occurs — and the bill grows alongside the fear of OOM kills.

Karpenter is a Kubernetes node autoscaler that provisions instances in 45 seconds, versus 3 to 4 minutes for the Cluster Autoscaler. Instead of waiting for the Cluster Autoscaler to iterate over node groups, Karpenter directly picks the cheapest instance type that satisfies pending pods, consolidates workloads onto fewer nodes, and terminates underutilized nodes. According to Salesforce, the replacement across 1,000 EKS clusters cut 70% of infrastructure cost.

When the team sees the cost of its own namespace, it reduces 15 to 20%. According to the Cast AI 2026 report, only 14% of companies do Kubernetes chargeback and 24% do not monitor K8s cost at all. Chargeback by namespace assigns each workload's bill to the responsible team, turns invisible cost into visible responsibility, and creates an incentive to rightsize. ELO, the Brazilian card brand, launched Kubernetes chargeback in January 2026 and recovered 5 analyst days per month.

Blog — Get in Touch

Have a question about our articles or services? Our team is ready to help.

Schedule a Meeting

Book a time slot.

Schedule Now

Email

Send us a message.

[email protected]

WhatsApp

Quick conversation.

Address

Avenida Paulista, 1636 - São Paulo - SP - 01310-200

Tech86 Specialist

Online now

Hello! How can we help scale your business today?

Tech86 Engineering

We Value Your Privacy

We use cookies and similar technologies to optimize your experience, analyze site traffic, and personalize content. By clicking "Accept All", you agree to the use of all cookies. Read our Privacy Policy.