Wednesday, September 16, 2026
No Result
View All Result
Future News 24
Advertisement
  • Home
  • AI Research
  • Platforms
  • Ethics
  • Developer AI
  • Industry
  • Data Science
  • Emerging Tech
  • Quantum
  • BioTech
  • Decentralized
  • Home
  • AI Research
  • Platforms
  • Ethics
  • Developer AI
  • Industry
  • Data Science
  • Emerging Tech
  • Quantum
  • BioTech
  • Decentralized
No Result
View All Result
Future News 24
No Result
View All Result
Home Developer AI & Open-Source Ecosystem

AI inference is clearly worthwhile

Future News 24 by Future News 24
July 1, 2026
in Developer AI & Open-Source Ecosystem
0 0
0
AI inference is clearly worthwhile
0
SHARES
0
VIEWS
Share on FacebookShare on Twitter


Many individuals declare that AI inference is unprofitable to serve, and thus have to be backed by an ocean of dumb cash from buyers who imagine that some future AI mannequin will come to dominate the world economic system. When that dumb cash goes away, so will AI merchandise. In response to this view, LLMs are simply inherently too costly (by way of cash, energy, and water) for use in shopper merchandise. The truth is, they’ll solely be used right this moment by externalizing the prices: cash onto VC funds and now retail ETF buyers, energy onto electrical utility shoppers, and water onto the communities the place datacenters are constructed.

There are good causes to dislike AI, however this actually isn’t certainly one of them. The truth is, AI inference is clearly worthwhile.

Doing the maths demonstrates that inference is worthwhile

Frontier AI suppliers are reporting 70%-80% gross margins on inference, however perhaps we will’t belief them. Let’s do some very tough estimates on the precise price.

A Nvidia A100 consumes 400W of energy beneath full load. In apply, even a carefully-tuned inference server is not going to be at full load on a regular basis, but it surely’s at the least an higher certain. Suppose you’re working a dense 70B mannequin, which is able to match comfortably (unquantized) on 4 A100s at round 2M tokens per hour. At industrial energy costs, that’s about 13c/hr within the USA. Suppose (pessimistically) cooling is identical price. That’s about 13 cents per million output tokens.

Let’s amortize the price of the GPUs, since that’s going to be the costliest half. An A100 prices about $20k. If every A100 lasts round 5 years, you’ll should make 16k/yr in revenue to recoup your capital funding (or $1.80 per hour). At decrease utilization, it’ll take longer to recoup, however your GPUs may also last more. Both method, your total inference prices are at about one greenback per million tokens.

GPT-5.4-mini costs $4.50 per million tokens, and stronger OpenAI or Anthropic fashions are three to 6 occasions as costly. It’s exhausting to make a direct comparability as a result of we don’t know the dimensions of OpenAI or Anthropic fashions, however the claimed 70% or 80% revenue margin is extraordinarily believable.

Open LLMs exhibit that inference is worthwhile

What in case you don’t belief my estimates both? Let’s have a look at the pricing of open-weights Chinese language LLMs. DeepSeek have claimed a bit over 80% revenue margin on inference for DeepSeek-R1. Since their API pricing for R1 is lower than half that of OpenAI or Anthropic, that implies that my estimates above for inference price is likely to be too costly. Cooling at scale might be cheaper than energy, R1 solely has half the lively parameters of a dense 70B mannequin, fashionable GPUs are extra environment friendly than the A100, and there are vital economies of scale in inference.

Since DeepSeek’s fashions can be found for anybody to obtain, they’ll’t get away with extracting a big revenue margin. One of many different inference suppliers would undercut them with the identical mannequin. Inference prices for DeepSeek-V4-Professional in the marketplace are round 87 cents per million output tokens, which might be fairly near the precise price of serving the mannequin.

For AI labs, inference should subsidize coaching

All of this doesn’t imply that OpenAI or Anthropic are worthwhile. These firms are making big capital investments that will or might not pan out, and are spending huge quantities of cash on expertise and compute to coach brand-new fashions and retain customers.

They’re doing loopy issues like providing per-month subscription fashions for practically limitless inference, which is sort of actually not worthwhile. For those who used an API token as an alternative of your Anthropic subscription in Claude Code, you’d pay ten occasions the price. However that doesn’t imply API-based Claude Code couldn’t be a great deal. Some persons are already utilizing DeepSeek’s inference API for agentic coding, as a result of as soon as you are taking away the large revenue margin it’s cheaper than the relative per-month subscription.

Why gained’t OpenAI or Anthropic decrease their costs? Supposedly OpenAI has considered it, however for an AI lab, inference has to subsidize coaching prices. An organization like OpenAI has to fund the manufacturing of latest fashions from the inference margins on present fashions (at the least partially). That’s why the margins on inference are so excessive: the AI labs try to squeeze out each greenback to allow them to keep alive within the coaching arms race.

Nonetheless, inference solely has to subsidize coaching prices for an AI lab. For those who’re merely an inference supplier, you don’t should do any coaching in any respect. Subsequently, even when OpenAI and Anthropic exit of enterprise, whoever snaps up the rights to their frontier fashions will be capable to proceed promoting Opus and GPT inference at a revenue. The AI bubble popping is not going to imply the top of the inference enterprise, as a result of AI inference is clearly worthwhile.

Here is a preview of a associated submit that shares tags with this one.



Source link

Tags: inferenceprofitable
Previous Post

Run a vLLM Server on HF Jobs in One Command

Next Post

Shortly apply LUTs (coloration grading) with ffmpeg

Next Post
Shortly apply LUTs (coloration grading) with ffmpeg

Shortly apply LUTs (coloration grading) with ffmpeg

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Fetching latest news…
FUTURENEWS24
Live Feed
All
AI
Dev
Industry
Frontier
Updates in 60s
FN24 AI & Tech
View All →
Future News 24

The world's leading source for AI research, emerging technology, and the people building the future. Independent, rigorous, and always ahead.

CATEGORIES

  • AI Platforms & Apps
  • AI Research & Breakthroughs
  • BioTechnology
  • Data Science & MLOps
  • Decentralized Technology
  • Developer AI & Open-Source Ecosystem
  • Emerging Technologies & Innovations
  • Ethics & Policy
  • Industry & Business
  • Quantum Computing
  • Uncategorized

LATEST

  • [2602.13312] PeroMAS: A Multi-agent System of Perovskite Materials Discovery
  • GPT-6 Astra overview: code overview good points, privateness, and value
  • GPT-6 Astra: Options, Benchmarks, Pricing, and What’s New
  • About Us
  • Advertise with Us
  • Disclaimer
  • Privacy Policy
  • DMCA 
  • Cookie Policy
  • Terms and Conditions
  • Contact us

© 2026 Future News 24. All rights reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • AI Research
  • Platforms
  • Ethics
  • Developer AI
  • Industry
  • Data Science
  • Emerging Tech
  • Quantum
  • BioTech
  • Decentralized

© 2026 Future News 24. All rights reserved.

Website security powered by MilesWeb