Many individuals declare that AI inference is unprofitable to serve, and thus have to be backed by an ocean of dumb cash from buyers who imagine that some future AI mannequin will come to dominate the world economic system. When that dumb cash goes away, so will AI merchandise. In response to this view, LLMs are simply inherently too costly (by way of cash, energy, and water) for use in shopper merchandise. The truth is, they’ll solely be used right this moment by externalizing the prices: cash onto VC funds and now retail ETF buyers, energy onto electrical utility shoppers, and water onto the communities the place datacenters are constructed.
There are good causes to dislike AI, however this actually isn’t certainly one of them. The truth is, AI inference is clearly worthwhile.
Doing the maths demonstrates that inference is worthwhile
Frontier AI suppliers are reporting 70%-80% gross margins on inference, however perhaps we will’t belief them. Let’s do some very tough estimates on the precise price.
A Nvidia A100 consumes 400W of energy beneath full load. In apply, even a carefully-tuned inference server is not going to be at full load on a regular basis, but it surely’s at the least an higher certain. Suppose you’re working a dense 70B mannequin, which is able to match comfortably (unquantized) on 4 A100s at round 2M tokens per hour. At industrial energy costs, that’s about 13c/hr within the USA. Suppose (pessimistically) cooling is identical price. That’s about 13 cents per million output tokens.
Let’s amortize the price of the GPUs, since that’s going to be the costliest half. An A100 prices about $20k. If every A100 lasts round 5 years, you’ll should make 16k/yr in revenue to recoup your capital funding (or $1.80 per hour). At decrease utilization, it’ll take longer to recoup, however your GPUs may also last more. Both method, your total inference prices are at about one greenback per million tokens.
GPT-5.4-mini costs $4.50 per million tokens, and stronger OpenAI or Anthropic fashions are three to 6 occasions as costly. It’s exhausting to make a direct comparability as a result of we don’t know the dimensions of OpenAI or Anthropic fashions, however the claimed 70% or 80% revenue margin is extraordinarily believable.
Open LLMs exhibit that inference is worthwhile
What in case you don’t belief my estimates both? Let’s have a look at the pricing of open-weights Chinese language LLMs. DeepSeek have claimed a bit over 80% revenue margin on inference for DeepSeek-R1. Since their API pricing for R1 is lower than half that of OpenAI or Anthropic, that implies that my estimates above for inference price is likely to be too costly. Cooling at scale might be cheaper than energy, R1 solely has half the lively parameters of a dense 70B mannequin, fashionable GPUs are extra environment friendly than the A100, and there are vital economies of scale in inference.
Since DeepSeek’s fashions can be found for anybody to obtain, they’ll’t get away with extracting a big revenue margin. One of many different inference suppliers would undercut them with the identical mannequin. Inference prices for DeepSeek-V4-Professional in the marketplace are round 87 cents per million output tokens, which might be fairly near the precise price of serving the mannequin.
For AI labs, inference should subsidize coaching
All of this doesn’t imply that OpenAI or Anthropic are worthwhile. These firms are making big capital investments that will or might not pan out, and are spending huge quantities of cash on expertise and compute to coach brand-new fashions and retain customers.
They’re doing loopy issues like providing per-month subscription fashions for practically limitless inference, which is sort of actually not worthwhile. For those who used an API token as an alternative of your Anthropic subscription in Claude Code, you’d pay ten occasions the price. However that doesn’t imply API-based Claude Code couldn’t be a great deal. Some persons are already utilizing DeepSeek’s inference API for agentic coding, as a result of as soon as you are taking away the large revenue margin it’s cheaper than the relative per-month subscription.
Why gained’t OpenAI or Anthropic decrease their costs? Supposedly OpenAI has considered it, however for an AI lab, inference has to subsidize coaching prices. An organization like OpenAI has to fund the manufacturing of latest fashions from the inference margins on present fashions (at the least partially). That’s why the margins on inference are so excessive: the AI labs try to squeeze out each greenback to allow them to keep alive within the coaching arms race.
Nonetheless, inference solely has to subsidize coaching prices for an AI lab. For those who’re merely an inference supplier, you don’t should do any coaching in any respect. Subsequently, even when OpenAI and Anthropic exit of enterprise, whoever snaps up the rights to their frontier fashions will be capable to proceed promoting Opus and GPT inference at a revenue. The AI bubble popping is not going to imply the top of the inference enterprise, as a result of AI inference is clearly worthwhile.
Here is a preview of a associated submit that shares tags with this one.

