AI brokers have essentially modified the complexity of inference workloads. Till now, the business has struggled to outline a regular for measuring how inference methods carry out underneath these circumstances. Synthetic Evaluation AgentPerf (AA-AgentPerf) gives the business’s first multi-vendor open benchmarks profiling trajectories which might be consultant of real-world AI agent coding duties.
This put up explains how AA-AgentPerf units a brand new normal for measuring agentic workload efficiency, and the way NVIDIA excessive co-design helps ship as much as 20x higher agentic coding efficiency than earlier generations.
What’s AA-AgentPerf?
AA-AgentPerf is a {hardware} benchmark created by Synthetic Evaluation that measures the variety of concurrent AI brokers an inference system can assist whereas assembly predefined, model-specific efficiency service degree goal (SLO) tiers. An SLO is outlined as a selected threshold of output token velocity and time-to-first-token (TTFT). The benchmark outcomes are normalized per accelerator and per megawatt to allow comparability throughout {hardware} configurations.


Measuring consultant agentic coding efficiency
Agentic workloads are distinctive as a result of LLM-driven selections typically produce non-deterministic sequences of requests and gear calls. Probably the most tough a part of measuring agent efficiency is to precisely seize this non-determinism in a consultant agent trajectory—the entire sequence of actions, selections, and observations made by an agent because it traverses via a job from starting to finish (Determine 2).


AA-AgentPerf captures this by measuring GPU efficiency throughout prerecorded agentic coding trajectories with interleaved reasoning and gear use, whereas simulating interturn latency with a consultant baseline for CPU tool-call efficiency. These trajectories are constructed round fixing points in public code repositories throughout a number of use-cases,12+ programming languages, and response from frontier fashions. Along with rigorous definition of the trajectories, the Synthetic Evaluation group additionally:
Leveraged consultant cached, enter, and output sequence lengths for requests, starting from 5K to 131K with a imply of roughly 27K.
Mapped instrument calls to consultant CPU-side duties in agentic coding workflows and simulated instrument calls throughout a distribution with a one-second median delay time. The identical CPU tool-call baseline was then utilized throughout all methods examined.
Retains the test-set personal to forestall benchmark-targeted optimization.
AA-AgentPerf testing and measurement methodology
The AA-AgentPerf harness measures the variety of concurrent brokers an inference system can assist whereas assembly SLO necessities (Determine 3). At launch, this benchmark focuses on testing DeepSeek-V4-Professional throughout a number of SLO tiers derived from Synthetic Evaluation serverless API benchmarking information. This ensures that the benchmarks replicate quality-of-service ranges noticed in manufacturing suppliers at this time.


Throughout a benchmarking run, AA-AgentPerf sends GPUs 1000’s of concurrent requests drawn from its prerecorded agent trajectory dataset. To make sure unbiased outcomes for every run, dynamic prefixes are added at first of each trajectory section. Strict SLO thresholds are enforced all through the trajectory, and the best concurrency degree that satisfies these necessities is recorded because the official benchmark outcome for a given SLO (Determine 3). This course of is then repeated throughout a number of SLO tiers to seize totally different person expertise targets (Desk 1).
Find out how to interpret AA-AgentPerf outcomes
The core AA-AgentPerf metric is runtime energy per megawatt—a sensible normalization for representing information middle scale efficiency. Desk 2 outlines learn how to leverage the reported efficiency to estimate what number of agentic periods could possibly be supported for a given energy funds.
On launch day, NVIDIA GB300 NVL72 delivers as much as 20x extra concurrent brokers per megawatt than the earlier era, NVIDIA H200 (Determine 4).


This efficiency highlights how GB300 NVL72 is ready to ship throughout large-scale agentic coding workloads, from routing long-lived periods effectively to preserving combination of consultants (MoEs) and GPUs totally utilized throughout many concurrent agent periods..
SGLang, TensorRT LLM, or vLLM: Agent runtimes apply optimizations akin to WideEP and DeepEP to unfold MoE skilled execution throughout the complete NVL72 area, maximizing efficient batch sizes and scaling successfully to 1000’s of brokers.
DeepGEMM and Mega MoE optimizations: MXFP4/MXFP8 kernels and fused MoE overlap NVLink communication with tensor core compute to spice up token throughput for reasoning and code era.
NVIDIA NVLink scale-up area: GB300 NVL72 hyperlinks 72 GPUs right into a single high-bandwidth NVLink material, so each GPU can quickly share parameters, KV cache, and intermediate outcomes—essential for quick, coordinated execution of agentic coding methods.
Wanting ahead: NVIDIA Vera Rubin platform
AA-AgentPerf establishes the usual for evaluating agentic inference, and the outcomes spotlight how tightly built-in {hardware} and software program can unlock step-function features in concurrency and effectivity. NVIDIA GB300 NVL72 demonstrates as much as 20x larger agentic coding efficiency.
The NVIDIA Vera Rubin platform is anticipated to increase these features by leveraging 50 PFLOPs of NVFP4 compute and leveraging the Vera CPU to speed up LLM instrument calls and enhance end-to-end efficiency, economics, and effectivity for agentic workflows.
To be taught extra about why agentic workloads place distinctive calls for on inference infrastructure and the way the NVIDIA Vera Rubin platform optimizes efficiency, see Constructing for the Rising Complexity of Agentic Methods with Excessive Co-Design.
Acknowledgments
This work was made doable via the experience and engineering contributions of Jatin Gangani, Iman Tabrizian, Xiaoming Chen, Peiheng Hu, Taizhong Wu, Shichen Li, Manu Maheswari, and plenty of different proficient NVIDIA engineers.

