Energy can account for 40% of the working bills (OpEx) to run an AI manufacturing facility. Every watt could be spent on overhead, knowledge ingestion, coaching, or producing tokens for patrons. And most websites are capped at a set energy stage offered by a regional supplier. Below these situations, efficiency per watt turns into a key effectivity metric that straight interprets to token prices.
NVIDIA delivers the bottom price per token for AI inference workloads and the bottom price to coach massive fashions. That is potential via excessive co-design with energy, cooling, and system infrastructure and deep collaboration with the OEM, ODM, CSP, NCP, methods integrator, ISV, and mannequin ecosystems companions.
This publish explores the levers that an operator can use to maximise efficiency per watt and decrease token price in an AI manufacturing facility.
Why is inference optimization essential for AI factories?
Inference drives income, so it’s the key workload to optimize. When operators improve inference throughput per watt, they straight improve the variety of tokens they’ll promote or insights they’ll create. This additionally interprets to further income per unit of time.
On the hundred megawatt to gigawatt scale, even a couple of proportion factors of throughput enchancment per megawatt can translate into significant good points in revenue.
Mannequin structure can be essential. Combination-of-experts (MoE) fashions are sometimes extra vitality environment friendly per unit of intelligence in comparison with dense fashions with related whole parameters as a result of solely a subset of specialists is lively per token. For instance, DeepSeek-R1 has a big parameter rely, a fraction of which is activated for every token. It achieves increased activity efficiency at the same or decrease per‑token compute price than dense predecessors. In different phrases, the MoE design delivers extra intelligence for a similar or much less vitality spent producing every token.
The best way to optimize for system-level vitality use and efficiency per watt
NVIDIA architectures and platforms are engineered to extend the quantity of intelligence produced per watt with every era. Throughout six structure generations, NVIDIA has improved inference throughput per megawatt by 1,000,000x.
The NVIDIA GB200 NVL72 rack-scale system will increase vitality effectivity via excessive co-design, with dense, direct-to-chip liquid-cooled structure that delivers extra throughput per watt. It makes use of in-rack energy smoothing to flatten peak present spikes, enabling operators to securely deploy extra GPUs inside the similar energy and infrastructure funds.
As well as, NVIDIA DSX is an open, AI factory-scale platform that drives dynamic energy allocation, real-time telemetry, and making use of superior rack-level controls that get well stranded energy and improve tokens per watt.
Floating level precision provides one other layer: increased‑precision calculations are usually slower and devour extra vitality, whereas narrow-precision codecs like NVFP4 are extra vitality‑environment friendly and may ship increased throughput, at equal accuracy to FP8.Equally essential, NVIDIA Dynamo and NVIDIA TensorRT-LLM assist translate these good points into real-world inference efficiency by boosting throughput, reducing prices, and scaling reasoning fashions extra effectively throughout GPU infrastructure.


General vitality use is ruled by the quantity of computation, {hardware} effectivity, GPU utilization, and the place the system operates on the velocity/vitality tradeoff frontier. In consequence, system design, eradicating non‑GPU bottlenecks, and tuning batch measurement to be used case, reminiscence, and parallelism are key levers for optimizing vitality use and throughput per watt.
Optimizing vitality effectivity in LLM coaching
Giant mannequin coaching requires the distribution of labor throughout a number of GPUs utilizing a mix of a number of parallelization strategies. Throughout coaching, pushing for max iteration velocity comes at the price of very massive vitality consumption.
Additional, particular person GPU workload allocation is just not completely balanced, resulting in a number of GPUs in idle state whereas few GPUs end computations. Power is wasted if all GPUs dash to the end to finish a activity solely to take a seat idle ready for others to complete theirs and sync.
Researchers from the ML.ENERGY Initiative on the College of Michigan have proven that tuning the processing velocity for particular person GPUs can scale back vitality bloat in massive mannequin coaching. These with extra work are on the essential path (the slowest chain of duties within the pipeline) and run at most velocity, whereas these with much less work are deliberately slowed down.
This achieves the next:
Idle time from GPUs ending early is minimized
GPUs working at decrease velocity use much less vitality
Finish-to-end coaching time stays unchanged


Megatron-LM is the NVIDIA open supply reference implementation for coaching large-scale language fashions. In collaboration with the ML.ENERGY workforce, NVIDIA continues to advance Megatron-LM coaching vitality effectivity by profiling energy and efficiency habits on the kernel, scheduling, and parallelism ranges, after which utilizing these measurements to information focused, vitality‑conscious optimizations.
This work contains:
Implementing positive‑grained kernel and part‑stage vitality profiling to determine compute, reminiscence, communication, and energy‑restricted areas
Analyzing how parallelism configurations, pipeline imbalance, and communication overlap impression efficiency‑per‑watt
These insights are used to design vitality‑conscious scheduling and GPU frequency/energy‑cap tuning aligned with the true essential path (the slowest chain of duties within the pipeline) of coaching iterations. The subsequent step is to stipulate how these methods can be utilized to bigger scale Megatron-LM coaching.
This work goals to extend vitality effectivity in order that mannequin coaching could be accomplished sooner inside the similar energy envelope or obtain the identical coaching throughput with much less vitality. In consequence, energy could be redirected to further coaching runs or from coaching to inference on the identical optimized infrastructure—rising token era with out elevating whole web site energy. To be taught extra, see Kareus: Joint Discount of Dynamic and Static Power in Giant Mannequin Coaching.


How does NVIDIA DSX optimize AI manufacturing facility efficiency?
The ML.ENERGY Initiative has developed a leaderboard and benchmark for sharing observations from their measurements and a reasoning framework that explains why they observe sure vitality behaviors.
These benchmarks could be tied into vitality conscious operations- telemetry-driven methods that present tips on how to run an AI manufacturing facility below actual deployment constraints, together with energy price, carbon depth, thermals, cooling capability, and grid limits.
NVIDIA DSX supplies these energy-aware operations. The platform delivers a coordinated view throughout compute, racks, cooling, facility energy, and workload scheduling. It supplies a standard operational structure that may join design-time simulation with runtime telemetry, serving to operators perceive the place energy is getting used, the place it’s stranded, and the way a lot further helpful compute can match inside a set web site envelope.
DSX defines how AI factories are designed, constructed, and optimized throughout the total stack, from chips and methods to infrastructure software program, amenities, digital twins, and accomplice applied sciences. It combines open software program libraries, workflow guides, and reference designs with NVIDIA compute platforms and co-designed OEM infrastructure to allow a broad ecosystem of software program and {hardware} options.
By aligning each layer via a standard structure, DSX improves tokens per watt, accelerates deployment, and strengthens operational reliability and resiliency.
DSX manages energy effectivity and behaviors inside the rack, on the AI manufacturing facility stage, and between the AI manufacturing facility and the grid. DSX MaxLPS operates contained in the AI manufacturing facility, whereas DSX Flex operates between the grid and the manufacturing facility.
DSX MaxLPS is a set of applied sciences for maximizing AI manufacturing facility throughput, together with:
45°C liquid cooling: By leveraging built-in chip, thermal, and system-level improvements, operators can make the most of increased 45°C inlet temperatures to enhance energy utilization effectiveness (PUE), guaranteeing {that a} bigger portion of AI manufacturing facility energy is redirected towards revenue-generating compute.
Dynamic energy allocation: Software program repeatedly displays GPU and rack-level energy consumption, reallocating it the place wanted to unlock stranded capability and optimize total utilization. It operates inside outlined energy budgets, adapts to funds adjustments in actual time, and ensures protected, compliant execution.
Superior methods: Built-in straight into NVIDIA GPUs, superior methodologies enhance efficiency per watt at iso-performance. These embody energy steering, optimized workload profiles for speedy GPU configuration, and software program equivalent to NVIDIA Dynamo for orchestrating inter-rack energy and efficiency optimization.
DSX Flex is the grid-aware energy orchestration layer that connects the AI manufacturing facility to grid indicators and exterior vitality sources.
With energy, cooling, and grid integration optimized finish to finish, consideration can shift to extracting most effectivity from the workloads themselves.
The important thing alternative is to make use of benchmarks to information mannequin, batching, and precision decisions on prime of the optimized AI manufacturing facility. By aligning workload placement, scheduling, and energy allocation with probably the most environment friendly compute and cooling zones, operators can stack workload-level optimizations on prime of infrastructure-level good points.
This contains rebalancing workloads below a set energy funds, figuring out workloads the place energy could be decreased via extra environment friendly configurations or mannequin households, and prioritizing workloads that justify increased energy budgets as a result of they generate extra income per token. In doing so, we repeatedly steer the AI manufacturing facility towards most tokens per watt, driving down price per token over time.
Wanting forward, AI tokenomics metrics needs to be thought to be first‑class design objectives. Groups ought to discover combining digital‑twin‑pushed infrastructure optimization with benchmark‑pushed workload tuning.
This method turns constrained energy right into a function‑constructed aggressive benefit in each token capability and income.


Study extra
AI factories are basically restricted by energy, making efficiency per watt a key driver of token price and profitability. Optimizing inference is essential as a result of it straight will increase income via increased token output, whereas full-stack enhancements throughout {hardware}, software program, and mannequin design enhance effectivity.
Coaching can be made extra energy-efficient with out compromising velocity by decreasing idle GPU time. NVIDIA DSX permits real-time, energy-aware optimization throughout infrastructure, maximizing tokens per watt and income per megawatt.
To be taught extra about power-constrained AI manufacturing facility design, simulation, operations, and NVIDIA DSX, go to the NVIDIA sales space at ISC 2026.
Acknowledgments
We’d prefer to thank Mosharaf Chowdhury, Jae-Gained Chung, and Ruofan Wu from the ML Power initiative on the College of Michigan for his or her contributions.

