Cloud migration was supposed to be a one-way door. The promise lasted a decade: everything goes up, nothing comes down. Infinite elasticity, CapEx becoming OpEx, end of the on-prem datacenter. In March 2026, the door opened the other way. We have tracked this movement across corporate clients and the signal is structural: the cloud economics equation changed, and whoever treats cloud as religion loses the cost window.
93% of companies evaluating AI workload repatriation
The most striking data point comes from Cloudian. According to the Cloudian Enterprise AI Infrastructure Survey 2026, 93% of companies have already repatriated, are repatriating, or are evaluating repatriation of AI workloads from public cloud. The survey polled 203 enterprise IT decision-makers via Centiment in February 2026. The cut is specifically about AI workloads: 79% have already moved some — 26% significantly, 53% partially or in process.
Independent data corroborates. According to Broadcom's Private Cloud Outlook 2026, published in June, 50% of companies have already repatriated some workload — 15 percentage points more than in 2025. 83% are considering repatriation. The sample is larger: 1,800 IT decision-makers, companies with more than 1,000 employees, 8 countries. Two independent surveys, same direction.
What changed the equation was inference
Training in the cloud makes sense for bursts. Production inference is continuous, predictable, expensive on On-Demand. That distinction is what explains the turn. According to Broadcom's Private Cloud Outlook 2026, public cloud use for inference fell 15 percentage points in one year: from 56% to 41%. 56% of companies run or plan to run production inference on private cloud. 43% are repatriating AI workloads specifically — a category that did not exist in the 2025 study.
Inference changed the equation because it is the opposite of the use case that justified the cloud. Elasticity serves unpredictability. Production inference is predictable by definition — you know the volume, you know the window, you know the latency. Paying On-Demand for continuous load is paying premium for something that needs no premium.
The driver is cost — and the math stopped closing
According to the Broadcom report, cost surpassed security as the top concern with public cloud: 31% vs 26% in 2025. 97% of IT leaders believe part of cloud spend is wasted. 52% say more than 25% of the budget goes down the drain.
In the Cloudian Enterprise AI Infrastructure Survey 2026, 84% are over budget on cloud storage. 19% more than 30% over. 40% had cloud AI spending above projected. The cloud stopped being the cheap option — it became the option whose cost nobody can predict.
The math closes on the outside. According to Lenovo's whitepaper from February 2026, the break-even of on-prem infra vs Azure On-Demand reaches ~5.2 months for high-utilization inference. A 6x advantage in cost per million tokens. This number is specific to high-utilization inference — it is not universal. Workloads with low utilization or sporadic spikes may still make more sense in the cloud. The math is done in real TCO, workload by workload.
Repatriation is an architectural decision, not a religion
Repatriation is not "cloud is bad." It is separating what belongs in the cloud from what belongs on-prem. LLM training with GPU H100 bursts still makes sense in public cloud. Production inference at continuous volume makes sense on-prem. Hot storage with high egress makes sense on-prem. Cold storage with sporadic access makes sense in the cloud.
This is where Tech86 has been working with corporate clients: measure before moving. Without real TCO per workload, repatriation becomes migration in the dark — the same mistake as the first wave of cloud migration, only in the opposite direction. A pilot with one inference workload validates the model before scaling. Without a pilot, you repatriate waste along with useful load.
The door opened both ways
The door opened both ways. Whoever treats cloud as religion loses the cost window. Public cloud is neither right nor wrong — it is a tool. The same goes for on-prem. The decision is architectural, workload by workload, with real TCO and measurement before moving.
At Tech86, we help companies do exactly that: separate what belongs in the cloud from what belongs on-prem, calculate real TCO per workload, pilot repatriation with inference before scaling, and implement FinOps with chargeback so the waste that drove repatriation does not reinstall itself on the other side. Cloud is not religion. It is math.