Individuals who assume present AI use is unsustainable usually depend on the declare that inference GPUs solely final “three years on the most” below load. The concept right here is that after the AI bubble cash drains away, present infrastructure will quickly grow to be out of date, and there received’t be sufficient cash floating round to purchase a complete slate of brand-new GPUs. Inference prices would thus quickly grow to be manner too costly for present AI merchandise to make any monetary sense.
The place does this “three years on the most” declare come from? Is it believable?
Sourcing the quote
The unique Tom’s {Hardware} article quotes this tweet from Tech Fund, an nameless former PM and tech investor, who quotes an nameless “GenAI principal architect” at Google as saying “when you’ve got a excessive utilization price, then fixed excessive utilization price for a yr or two, I feel the lifespan will likely be three years at most”.

This screenshot seems to be prefer it was from an interview. What interview? I scrolled again to October 2024 on Tech Fund’s Twitter feed and noticed a bunch of similarly-formatted screenshots, a few of which had been cited as coming from Tegus. Tegus is outwardly an organization with a enterprise mannequin of reaching out to insiders (on this case, AI firm workers) and paying them tons of of {dollars} an hour to be able to reply particular technical questions. It’s primarily gig work for almost-but-not-quite insider buying and selling: the extra knowledgeable and assured you sound, the extra probably Tegus analysts will choose you for future interviews.
I’m certain the supply for this tweet is actually a GenAI principal architect, since Tegus would have presumably requested for some proof of that earlier than they paid them out. But it surely’s fairly clear that the incentives listed below are to sound assured and authoritative, even on questions that you just’re unsure about. With that in thoughts, the quote itself additionally reads a bit suspiciously. I’ve labored with sufficient principal engineers and designers to take their informal back-of-envelope estimates with a grain of salt. In the event that they knew the precise price at which GPUs fail and get retired in Google datacenters, wouldn’t they’ve simply mentioned that?
Proof for an extended lifespan
We’ve some anecdotal proof that factors the opposite manner. Google has publicly claimed to have eight yr outdated TPUs (their model of GPUs) operating in manufacturing at “100% utilization”. Nvidia solely made A100 GPUs from 2020-2024, however in February 2026 the AWS CEO claimed that AWS had by no means retired an A100 server (and you may nonetheless simply hire A100s for AI work). AI GPU utilization isn’t precisely like crypto mining GPU utilization, but it surely definitely looks as if years-old ex-crypto GPUs are useful. There’s additionally this remark from Hacker Information I seen the place somebody claims that their GPU cluster in academia has lasted six years with lower than 20% failure price.
What about laborious knowledge? It’s laborious to get concrete knowledge on the lifespan of AI GPUs, as a result of fashionable AI datacenters have solely existed for a handful of years. However an fascinating case research could be latest supercomputer clusters like Oak Ridge’s Summit, which had over 27 thousand Nvidia V100s operating from 2018 to 2024, or its predecessor, the Cray Titan supercomputer that ran from 2012 to 2019. I couldn’t discover any proof that Summit had to purchase an extra 27,000 GPUs to exchange their outdated ones, and GPU failures in Titan have been fastidiously studied:

These cages of GPUs are stacked vertically, and chilly air is pumped in from the underside, which explains why cage 0 (on the backside) has higher survival charges than cage 2 (on the high). Let’s think about cage 0, so we’re simply wanting on the GPU lifespan as an alternative of on the lifespan of improperly-cooled GPUs. At three years, over 95% of GPUs survived. At six years, nodes 2 and three (the GPUs closest to the underside of the cage) had been nonetheless at above 90% survival price, and the best nodes had been over 60%.
It’s potential that newer Nvidia GPUs are much less dependable than older ones (they definitely draw extra energy), or that AI datacenters are under-cooled, or that one thing about LLM utilization is extra anxious than the workloads that ran on conventional GPU datacenters. However that is not less than circumstantial proof that GPUs can survive below load for a lot longer than three years.
Financial lifespans
This dialogue is sophisticated by the truth that GPUs could have a brief financial lifespan. Supposedly a B100 GPU attracts twice as a lot energy as an A100, however can do 5 occasions as a lot work. For some AI suppliers, which may imply that A100s are solely value operating till they are often changed with B100s (when you’re bottlenecked on electrical energy, it’s best to spend all of it on B100s and throw out your out of date A100s). Because of this the Titan supercomputer was decommissioned in favor of Summit: it might have continued to function, but it surely was extra worthwhile to spend the cash and upkeep effort on newer {hardware}.
It must be apparent that this doesn’t assist the “inference will grow to be dearer when the bubble pops” argument. As long as A100s are worthwhile proper now, cash-poor AI suppliers can proceed profitably serving inference from them, even when there are extra environment friendly choices out there for these with the capital to improve.
On high of that, GPUs solely symbolize one a part of AI datacenter infrastructure spending. In case your GPUs put on out, you don’t should go and construct a wholly new datacenter. About 30-50% of datacenter spend goes to land, energy, cooling, and so forth. The remaining 50-70% is the price of all the server rack, which features a bunch of issues that aren’t GPUs.
Conclusion
Like the concept AI inference requires utilizing large quantities of water, the concept AI GPUs solely dwell a yr or two is common as a result of it’s a helpful thought for AI skeptics, not as a result of it’s true. It comes from a pseudonymous tweet quoting an nameless supply who’s being paid tons of of {dollars} to sound like a reputable professional on AI. Different public communications from AI inference suppliers cite a lot greater lifespan numbers, and the statistics from supercomputers (the normal examples of enormous GPU clusters) don’t bear out the declare that the utmost lifespan is three years.
It is likely to be true that the financial lifespan is three years, in a world the place new GPUs come out each eighteen months and GPU suppliers are flush with money to improve, however that doesn’t inform us a lot in regards to the economics of inference in an AI winter. If cash turns into much more scarce, it’s probably that AI datacenters will proceed profitably operating their B300s (or their H100s and even A100s) for six years or longer.
Here is a preview of a associated submit that shares tags with this one.

![[2508.00126] Environment friendly and easy Gibbs state preparation of the 2D toric code through duality to classical Ising chains [2508.00126] Environment friendly and easy Gibbs state preparation of the 2D toric code through duality to classical Ising chains](http://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png)