23,000 Kubernetes clusters. Average CPU: 8%. The Cast AI 2026 report analyzed clusters in production and the number dropped from 10% to 8% in one year. You provision 12x more compute than you use. GPU is worse: average utilization of 5%, which means 20x more capacity allocated than consumed. We analyzed the data and the signal is clear — the problem is not Kubernetes. It is provisioning for peak and paying for idle.
The scale: 23,000 clusters, CPU at 8%, GPU at 5%
According to the Cast AI 2026 report, which analyzed 23,000 clusters in production, average CPU utilization dropped from 10% to 8% in one year. You provision 12x more compute than you use. GPU is worse: average utilization of 5%, which equals allocating 20x more GPU capacity than is consumed.
CPU overprovisioning jumped from 40% to 69%. Memory overprovisioning: 79%. Engineers request resources for peak. The peak never arrives. The Cluster Autoscaler provisions nodes to match inflated requests, not real usage. Every node provisioned for a peak that never occurs is idle money.
The cost of idle: 83% of the container bill
According to Datadog, 83% of container cost is idle. 54% of the cluster idle. 29% of the workload idle. The cost is not in the workload that runs — it is in the workload that does not run, in the node that waits, in the GPU that nobody consumes.
Only 14% of companies do Kubernetes chargeback, according to the Cast AI 2026 report. 24% do not monitor K8s cost at all. You cannot cut what you do not measure. And what is not measured grows — because every team asks for more resources and nobody gives any back.
And the price went up. AWS raised the price of GPU Capacity Blocks by 15% in January 2026. The first time in 20 years that AWS raised GPU prices instead of lowering them. Idle GPU got more expensive at the exact moment when more GPU is idle.
Brazil: R$27,742 per month of idle GPU
In Brazil, each dollar of cloud costs 5.60 reais. An idle H100 cluster costs $4,954 per month. In reais: R$27,742 per month of idle GPU. It is not an isolated case — it is the average for anyone who provisions H100 for inference and never comes close to saturating it.
According to Brazilian market data, 65% of Brazilian companies have no visibility into cloud spend. Only 19.5% adopted FinOps. The result is predictable: without visibility, without rightsizing, without chargeback, the bill grows until someone asks where the number comes from. By the time the question arrives, the waste has already become habit.
What Tech86 implements
We do not sell generic monitoring. We implement five levers with measurable effect on the bill:
- Automated rightsizing. Cuts provisioned CPU by 50% without OOM kills. We analyze real usage and recommend requests at the 95th percentile.
- Spot for GPU. Less than 2% of GPUs run on Spot today. Spot cuts 60 to 90% of the cost of interruption-tolerant workloads.
- Karpenter. Provisions nodes in 45 seconds. The Cluster Autoscaler takes 3 to 4 minutes. According to Salesforce, the replacement across 1,000 EKS clusters cut 70% of cost.
- Turn off dev outside business hours. 20 to 30% of the bill comes from environments nobody uses at night.
- FinOps with chargeback by namespace. When the team sees the cost, it reduces 15 to 20%.
Each lever has a number. It is not a promise — it is what the Cast AI 2026 report, Datadog, and Salesforce have already measured in production.
The ELO case: chargeback in reais
ELO, the Brazilian card brand, launched Kubernetes chargeback in January 2026. It standardized reports in reais. It recovered 5 analyst days per month that were previously spent reconciling dollar invoices with real costs. The gain was not only financial — it was operational. The FinOps team stopped fighting exchange rates and started fighting rightsizing.
The ELO case confirms the pattern: when cost becomes a namespace responsibility, the team acts. When cost stays in a global spreadsheet, nobody acts.
Conclusion: provisioning for peak is the problem
Provisioning for peak and paying for idle is the problem. Kubernetes is the tool. The tool is not the cause of the waste — the provisioning model is. As long as requests are inflated for a peak that never arrives, the Cluster Autoscaler will provision nodes for an event that never occurs and the bill will grow alongside the fear of OOM kills.
At Tech86, we implement automated rightsizing, Spot for GPU, Karpenter, dev shutdown outside business hours, and chargeback by namespace. The result is less idle, fewer nodes, and fewer reais sitting still. R$27,742 per month of idle GPU is not a destination — it is a symptom of a provisioning model that needs to change.