InferenceMAX Benchmarks and NVIDIA Blackwell's Dominance

Last update: 10/10/2025
Author Isaac
  • InferenceMAX v1 measures real-world performance and economy with reproducible overnight testing.
  • NVIDIA Blackwell leads in tokens/s, cost per million tokens, and tokens per MW.
  • Continuous software (TensorRT-LLM, Dynamo, SGLang, vLLM) drives 5x-15x improvements.
  • GB200 NVL72 achieves 15x ROI and minimum TCO in dense loads and MoE models.

InferenceMAX AI Benchmarks

The conversation about AI inference performance has accelerated, and rightly so: InferenceMAX v1 has brought order to the discussion with verifiable and up-to-date data that looks beyond raw speed to assess real economics. In this context, NVIDIA's Blackwell platform has not only set the pace, it has blown it away with unprecedented efficiency and cost-per-token results.

In short, we're talking about a paradigm shift: from "how fast it runs" to "how much it delivers per euro and per watt in production ." The combination of Blackwell hardware (B200 and GB200 NVL72), fifth-generation NVLink interconnect, NVFP4 low-precision, and continuous software optimizations (TensorRT-LLM, Dynamo, SGLang, vLLM) raises the bar in tokens per second, cost per million tokens, and effective ROI in real-world scenarios.

What is InferenceMAX v1 and why it matters

The biggest complaint from the industry was that traditional benchmarks become outdated quickly and often favor unrealistic configurations . InferenceMAX v1 breaks with that: it's an open-source, automated benchmark with nightly runs under the Apache 2.0 license that re-evaluates popular frameworks and models daily to capture real-world software progress.

For each model and hardware combination, the system performs sweeps of tensor parallelism and concurrency sizes , and presents performance curves that balance throughput and latency. Furthermore, CI results are published daily , and multiple frameworks (SGLang, TensorRT-LLM, and vLLM) are tested, allowing users to see how recent optimizations are pushing the Pareto frontier in near real time.

Methodologically, the tests cover single-node and multi-node environments with Expert Parallelism (EP) , and include variable input/output sequence lengths (80%-100% of ISL/OSL combinations) to simulate real-world workloads for reasoning, document processing, summarizing, and chat . The result is a continuous snapshot of latency, throughput, batch sizes, and input/output ratios that reflects actual operational economics, not just theoretical performance.

Blackwell leads: performance, efficiency, and economies of scale

The published data leaves little room for doubt: NVIDIA Blackwell dominates InferenceMAX v1 in inference performance and efficiency across the entire workload range. Compared to the Hopper generation (HGX H200), the jump to B200 and GB200 NVL72 represents order-of-magnitude improvements in compute-per-watt and memory bandwidth , in addition to a dramatic drop in cost per million tokens.

Specifically, the GB200 NVL72 system achieves a 15x ROI : a $5 million investment can generate $75 million in token revenue . This isn't an accounting trick; it's a result of the combination of NVFP4 for native low precision, NVLink and fifth-generation NVLink Switch , and the maturity of TensorRT-LLM and NVIDIA Dynamo in the software stack.

The same pattern is playing out with the cost per token. In gpt-oss, B200's optimizations have reduced the cost to two cents per million tokens , a fivefold decrease in just two months. This trend, supported by ongoing software improvements, completely changes the economic viability of new use cases.

Methodology that captures the reality of production

InferenceMAX v1 doesn't just measure tokens per second. It maps throughput to latency on a Pareto frontier , helping determine where trading is most profitable based on interactivity SLAs and TCO targets. The key takeaway is how Blackwell maintains its edge across the entire range , not just in a single optimal corner.

To ensure representativeness, the tests include concurrency levels from 4 to 64 (and scenarios beyond these limits in supplementary analyses), various EP and DEP configurations , and community reference models , ranging from gpt-oss 120B to Llama 3.3 70B or DeepSeek-R1. All of this is available in an open repository with reproducible recipes so anyone can validate results.

Pure performance: tokens/s per GPU and interactivity

The Blackwell B200 sets the pace with figures that seemed like science fiction a year ago. With the latest NVIDIA TensorRT-LLM stack , it reports 60.000 tokens per second per GPU and up to 1.000 tokens per second per user in gpt-oss, maintaining interactivity without sacrificing the user experience.

  How to perform stability testing with OCCT on CPU, GPU, RAM, and PSU

In dense models like Llama 3.3 70B , which activate all inference parameters, Blackwell achieves 10.000 tokens/s per GPU at 50 TPS/user in InferenceMAX v1 , more than 4x faster than H200. This improvement is supported by NVFP4, fifth-generation Tensor Cores , and bidirectional NVLink bandwidth of 1.800 GB/s , preventing bottlenecks between GPUs.

Efficiency is also measured in tokens per watt and cost per million tokens . For AI factories with power constraints, Blackwell delivers 10x more throughput per megawatt compared to the previous generation. Furthermore, it has reduced the cost per million tokens by 15x , opening the door to much more cost-effective mass deployments.

Software that improves every week: from 6K to 30K tokens/s per GPU

Beyond the hardware, speed is the defensive moat . After the release of gpt-oss-120b on August 5th, the B200 in InferenceMAX v1 was already performing well with TensorRT-LLM, but subsequent optimizations have doubled and then multiplied the initial numbers. At around 100 TPS/user, the throughput per GPU nearly doubled in a short time compared to launch day.

With the October 9th release of TensorRT-LLM, EP and DEP parallelism allocations arrived , and performance at 100 TPS/user increased by up to 5x compared to the initial version , going from ~6K to ~30K tokens/s per GPU. Part of this leap is achieved with higher concurrency levels than those tested by InferenceMAX out of the box (4-64), demonstrating how much more potential there is to be unlocked in advanced configurations.

The masterstroke has been enabling speculative decoding for gpt-oss-120b with the gpt-oss-120b-Eagle3-v2 model . With EAGLE, GPU throughput at 100 TPS/user triples compared to published results, jumping from 10 to 30 tokens/s . Even better, the cost per million tokens at 100 TPS/user has dropped from $0,11 to $0,02 in just two months. Even at 400 TPS/user, it remains around $0,12 , making multi-agent scenarios and complex reasoning viable.

Real economy: 15x ROI and minimum TCO with GB200 NVL72

In the DeepSeek-R1 reasoning model , the InferenceMAX v1 curves show that GB200 NVL72 significantly reduces the cost per million tokens compared to H200 at all interactivity levels. At approximately 75 TPS/user, H200 costs $1,56 , while GB200 NVL72 drops to just over $0,10 , a 15x reduction . Furthermore, GB200's cost curve remains flat for longer periods , allowing it to serve over 100 TPS/user without significantly impacting costs.

For large-scale deployments, this translates to AI factories being able to serve more users with improved SLAs without triggering OPEX or sacrificing throughput. Add to that the fact that a $5 million investment can generate $75 million in token revenue , and the message is clear: inference is where AI delivers value every day , and Blackwell's full-stack approach gives it a competitive edge.

Architecture that enables the jump: NVFP4, NVLink 5 and NVLink Switch

Blackwell's dominance didn't come from nowhere. The stack is based on extreme hardware-software code-design : NVFP4 precision for efficiency without sacrificing accuracy, fifth-generation NVIDIA NVLink , and an NVLink Switch that allows 72 GPUs to be treated as a single macro-GPU , enabling extremely high concurrency with tensor, expert, and data parallelism.

This approach complements an annual hardware cadence and continuous software improvements that, on their own, have more than doubled Blackwell's performance since launch . Integration with TensorRT-LLM, NVIDIA Dynamo, SGLang, and vLLM completes the picture, supported by a massive ecosystem of millions of GPUs, CUDA developers, and hundreds of open-source projects.

MoE at full power: disaggregated serving with GB200, Dynamo, and TensorRT-LLM

Verified tests demonstrate that the combination of GB200 NVL72, Dynamo, and TensorRT-LLM significantly boosts the throughput of MoE models like DeepSeek-R1 under very different SLAs, outperforming Hopper-based systems. The NVL72's scale-up design interconnects 72 GPUs with NVLink in a single domain, providing up to 130 TB/s of inter-GPU bandwidth , crucial for routing expert tokens without traditional interconnection bottlenecks .

  Comet, Perplexity's browser: Advanced AI, new features, and privacy controversy as it arrives on Windows

Dynamo's disaggregated serving separates prefill and decode onto different nodes, optimizing each phase with different GPU and EP allocations . This allows the memory-constrained decode phase to leverage wide EP for expert users without hindering the more compute-intensive prefill phase.

To prevent idle GPUs in large EP deployments, TensorRT-LLM monitors the workload of experts , distributes the most frequently used ones, and can replicate them for load balancing. The result: high and stable utilization , with net gains in effective throughput.

Open collaboration: SGLang, vLLM and FlashInfer

Beyond Dynamo and TensorRT-LLM, NVIDIA has co-developed kernels and optimizations for Blackwell with SGLang and vLLM, delivered via FlashInfer . These include kernel improvements for Attention Prefill and Decode, Communication, GEMM, MNNVL, MLA, and MoE , as well as runtime optimizations.

SGLang now includes Multi-Token Prediction (MTP) and disaggregation capabilities for DeepSeek-R1. vLLM now features overlapping asynchronous schedulers to reduce host overhead, automatic graph merging , and performance and functionality improvements for gpt-oss, Llama 3.3, and general architectures . All of these enhancements help Blackwell maximize its efficiency across the most widely used open-source frameworks.

Comparisons and additional technical details of the ecosystem

In technical analysis, the Blackwell architecture stands out as a significant advancement for low-latency, high-throughput inference . Key features include mixed FP8/FP4 execution on fifth-generation Tensor Cores, along with NVLink 5 offering up to 1,8 TB/s per GPU for throttling-free communication between multiple drives.

On DGX B200 nodes with NVSwitch, configurations of up to eight GPUs with unified HBM3e memory approaching 1,44 TB in aggregate are cited , along with inference pipelines that reflect real-world usage : initial prefill and subsequent auto-reverse decoding. The suite measures tokens/s, latency per request, and efficiency in FLOPS , with kernel-level optimizations and specialized TensorRT-LLM engines.

Compared to the H100 (Hopper), Blackwell achieves 4x the throughput on Llama 2/3 70B in a similar node, attributable to more tensor cores and memory bandwidth improvements (up to 5 TB/s per GPU in some tests) . Linear scalability in clusters of hundreds of GPUs is also mentioned , maintaining high efficiency in HBM3e usage and avoiding costly paging to host memory.

In terms of energy efficiency, improvements of up to 2,5x compared to the H100 are reported , with power consumption in high-load scenarios ranging from 700W to 1.000W per GPU depending on the configuration, and peak FP4 performance that clearly surpasses the previous generation in FLOPS per watt . Tools such as DCGM and telemetry with Prometheus/Grafana facilitate top-tier observability.

Operating economics, sustainability and compliance

InferenceMAX v1's focus on metrics like tokens per megawatt and cost per million tokens isn't just for show: it influences capex and opex decisions . Blackwell achieves 10x more throughput per MW than the previous generation and has reduced the cost per million tokens by 15x , with direct implications for service expansion and sustainability.

Practices geared towards renewable energy in DGX systems are described , along with regulatory references such as the EU AI Act, GDPR, and NIST SP 800-53 . Furthermore, Blackwell incorporates Confidential Computing with secure enclaves and memory encryption to protect data in highly regulated sectors such as banking and healthcare.

Use cases: security, IT and even blockchain

The combination of high performance and interactivity enables a shift from pilot programs to real-time security systems , from log analysis to anomaly detection in petabyte-scale networks with sub-second latencies . In IT, hyperscalers are integrating Blackwell into offerings for hybrid workloads with distributed storage and 5G networks , leveraging RoCE for minimal latency at the edge, and companies like ByteDance are reinforcing their commitment to NVIDIA chips.

Even on blockchain, decentralized AI oracles and accelerated ZK proofs are being considered on networks like Ethereum or Solana thanks to tensor parallelism. Operationally, reductions of up to 40% in inference TCO are reported due to higher rack density and advanced liquid cooling, maintaining temperatures below 85°C under sustained load.

  ChatGPT finally integrates voice mode into the chat itself.

Good practices and migration challenges

It's not all smooth sailing: migrating from Hopper requires recompiling CUDA kernels and can uncover bugs in legacy pipelines. NVIDIA's best practices guides for LLM inference recommend profiling with Nsight Systems , detecting attention and decoding bottlenecks , and applying sharding techniques with Megatron-LM to balance the workload across GPUs.

For security, it's advisable to enable secure boot and runtime protections in TensorRT to prevent code injection . In decentralized deployments, latency is managed with sidechains and offloading compute to dedicated GPUs, preserving integrity with cryptographic proofs.

Community, resources and transparency

InferenceMAX v1 is a community effort. Thanks are due to AMD (MI355X and CDNA3) for hardware for the project and to NVIDIA for access to GB200 NVL72 (via OCI) and B200 . Thanks are also due to the inference teams and Dynamo , as well as computing providers such as Crusoe, CoreWeave, Nebius, TensorWave, Oracle, and TogetherAI for supporting open source with real resources.

The platform publishes a live dashboard at inferencemax.ai with updated results and provides containers and configurations for running benchmarks. Given the speed at which AI software evolves, nightly tests are the honest way to show where performance stands today, not months ago.

Industry voices and career opportunities

Infrastructure managers and scientists acknowledge that the gap between theoretical peak throughput and actual throughput is determined by system software , distributed strategies, and low-level kernels . Therefore, they value open and reproducible benchmarks that demonstrate how optimizations perform on different hardware and that transparently reveal tokens per second, cost per dollar, and tokens per megawatt .

In addition, the project is seeking talent for a special projects team . Key responsibilities include:

  • Design and execute large-scale benchmarks across multiple vendors (AMD, NVIDIA, TPU, Trainium, etc.).
  • Build reproducible CI/CD pipelines to automate executions.
  • Ensure reliability and scalability of systems shared with industry partners.

Collaborations with open models and ecosystems

NVIDIA maintains open collaborations with the community and with teams such as OpenAI (gpt-oss 120B), Meta (Llama 3 70B), and DeepSeek AI ( DeepSeek R1 ) , in addition to contributions with FlashInfer, SGLang, and vLLM . This ensures that the latest models are optimized for the world's largest inference infrastructure and that kernel and runtime improvements are integrated at scale.

For businesses, NVIDIA's Think SMART framework helps navigate the transition from AI pilots to full-scale AI operations , refining platform decisions, cost per token , latency SLAs, and utilization based on changing workloads. In a world moving from one-shot solutions to multi-stage reasoning and tool usage , this guide becomes strategic.

Practical note: Some content shared on social networks like X may require JavaScript to be enabled to view; otherwise, the site's help and policies will be displayed . It's a minor detail, but useful if you want to keep track of ads in real time.

Anyone wondering whether it's worth taking a closer look at the recipes in InferenceMAX v1 should know that they are open for anyone to replicate Blackwell's leadership in very different inference scenarios. This is precisely the kind of transparency that accelerates progress across the entire community.

After reviewing the data, software improvements, and open collaborations, one key takeaway emerges: inference is where AI translates performance into business on a daily basis . With flat cost curves at high levels of interactivity, elegantly scaling tokens per second per GPU, and an ecosystem that continuously optimizes kernels and runtimes, Blackwell is solidifying its position as the benchmark platform for those looking to build efficient, fast, and cost-effective AI factories.

what is nvidia project digits-1
Related articles:
NVIDIA Project DIGITS: The AI ​​revolution from your desktop