country_code

New NVIDIA Neural Graphics SDKs Make Metaverse Content Creation Available to All

A dozen tools and programs — including new releases NeuralVDB and Kaolin Wisp — enable easy, fast 3D content creation for millions of designers and creators.
by

The creation of 3D objects for building scenes for games, virtual worlds including the metaverse, product design or visual effects is traditionally a meticulous process, where skilled artists balance detail and photorealism against deadlines and budget pressures.

It takes a long time to make something that looks and acts as it would in the physical world. And the problem gets harder when multiple objects and characters need to interact in a virtual world. Simulating physics becomes just as important as simulating light. A robot in a virtual factory, for example, needs to have not only the same look, but also the same weight capacity and braking capability as its physical counterpart.

It’s hard. But the opportunities are huge, affecting trillion-dollar industries as varied as transportation, healthcare, telecommunications and entertainment, in addition to product design. Ultimately, more content will be created in the virtual world than in the physical one.

To simplify and shorten this process, NVIDIA today released new research and a broad suite of tools that apply the power of neural graphics to the creation and animation of 3D objects and worlds.

These SDKs — including NeuralVDB, a ground-breaking update to industry standard OpenVDB, and Kaolin Wisp, a PyTorch library establishing a framework for neural fields research — ease the creative process for designers while making it easy for millions of users who aren’t design professionals to create 3D content.

Neural graphics is a new field intertwining AI and graphics to create an accelerated graphics pipeline that learns from data. Integrating AI enhances results, helps automate design choices and provides new, yet to be imagined opportunities for artists and creators. Neural graphics will redefine how virtual worlds are created, simulated and experienced by users.

These SDKs and research contribute to each stage of the content creation pipeline, including:

3D Content Creation

  • Kaolin Wisp – an addition to Kaolin, a PyTorch library enabling faster 3D deep learning research by reducing the time needed to test and implement new techniques from weeks to days. Kaolin Wisp is a research-oriented library for neural fields, establishing a common suite of tools and a framework to accelerate new research in neural fields.
  • Instant Neural Graphics Primitives – a new approach to capturing the shape of real-world objects, and the inspiration behind NVIDIA Instant NeRF, an inverse rendering model that turns a collection of still images into a digital 3D scene. This technique and associated GitHub code accelerate the process by up to 1,000x.
  • 3D MoMa – a new inverse rendering pipeline that allows users to quickly import a 2D object into a graphics engine to create a 3D object that can be modified with realistic materials, lighting and physics.
  • GauGAN360 – the next evolution of NVIDIA GauGAN, an AI model that turns rough doodles into photorealistic masterpieces. GauGAN360 generates 8K, 360-degree panoramas that can be ported into Omniverse scenes.
  • Omniverse Avatar Cloud Engine (ACE) – a new collection of cloud APIs, microservices and tools to create, customize and deploy digital human applications. ACE is built on NVIDIA’s Unified Compute Framework, allowing developers to seamlessly integrate core NVIDIA AI technologies into their avatar applications.

Physics and Animation

  • NeuralVDB – a groundbreaking improvement on OpenVDB, the current industry standard for volumetric data storage. Using machine learning, NeuralVDB introduces compact neural representations, dramatically reducing memory footprint to allow for higher-resolution 3D data.
  • Omniverse Audio2Face – an AI technology that generates expressive facial animation from a single audio source. It’s useful for interactive real-time applications and as a traditional facial animation authoring tool.
  • ASE: Animation Skills Embedding – an approach enabling physically simulated characters to act in a more responsive and life-like manner in unfamiliar situations. It uses deep learning to teach characters how to respond to new tasks and actions.
  • TAO Toolkit – a framework to enable users to create an accurate, high-performance pose estimation model, which can evaluate what a person might be doing in a scene using computer vision much more quickly than current methods.

Experience

  • Image Features Eye Tracking – a research model linking the quality of pixel rendering to a user’s reaction time. By predicting the best combination of rendering quality, display properties and viewing conditions for the least latency, It will allow for better performance in fast-paced, interactive computer graphics applications such as competitive gaming.
  • Holographic Glasses for Virtual Reality – a collaboration with Stanford University on a new VR glasses design that delivers full-color 3D holographic images in a groundbreaking 2.5-mm-thick optical stack.

Join NVIDIA at SIGGRAPH to see more of the latest research and technology breakthroughs in graphics, AI and virtual worlds. Check out the latest innovations from NVIDIA Research, and access the full suite of NVIDIA’s SDKs, tools and libraries.

How XPUs Meet a World-Class AI Factory

Deploying custom silicon with leading AI infrastructure enables hyperscalers and AI-native companies to build flexible AI factories that combine specialization with scale.
by

To generate intelligence at scale, AI factories run continuously, and their economics are defined by delivered output: tokens per second, tokens per watt, cost per token, utilization and uptime. 

That requires AI infrastructure designed and built as a full factory, not a collection of individual accelerators.

Hyperscalers and AI-native companies building custom XPUs must consider not just XPU design, but the design and development of the entire AI platform, including scale-up and scale-out networking, rack-scale architecture, production factory software and a robust supplier ecosystem. 

At AI factory scale, this path is complex and costly, and represents a fundamental obstacle to getting XPUs to market quickly. 

Breaking the constraint means combining custom XPUs with proven, mature infrastructure — allowing builders to focus innovation where it matters most while harnessing established technology for the rest. 

NVLink Fusion delivers on that need, connecting XPUs to NVIDIA’s world-leading AI infrastructure to increase performance, accelerate time to market and mitigate risk for semi-custom AI factories.

Unlock XPU Performance With Fast Scale-Up

For modern workloads such as running trillion-parameter models, mixture-of-experts architectures and agentic AI, if the scale-up fabric cannot keep up, utilization drops and cost per token rises.

A scale-up networking solution must excel on three dimensions: 

  • Delivered performance: End-to-end network performance, in-network compute and  mature software integration.
  • Factory resiliency: Uptime, continuous health monitoring and telemetry, and component-level serviceability while the factory keeps running.
  • Platform maturity: Reduced operational risk by using a mature technology stack with a demonstrated track record of large-scale deployments and realized return on investment. 

As an example, NVLink Fusion brings XPUs into the NVIDIA NVLink scale-up domain. Sixth-generation NVLink provides leading high-bandwidth, low-latency networking across a 72-XPU domain. The end-to-end latency for XPU-to-XPU transfers is 3x lower than alternative solutions based on off-the-shelf Ethernet, and the packet rate is 10x higher.

For end-to-end performance, NVIDIA GB300 NVL72 systems help deliver significantly higher throughput and better interactivity compared with configurations that don’t use NVL72, and future NVLink roadmap configurations include domains of up to 1,152 accelerators and co-packaged optics.

A Pareto chart comparing GB300 NVL72 with B300 inference throughput performance in tokens per second per GPU on DeepSeek-V4-Pro at ISL=1K and OSL=1K sampled at various interactivity points in tokens per second per user. GB300 NVL72 is more than 10x the throughput of B300 in the middle of the Pareto between 70 and 100 tokens per second per user.
The 72-GPU NVLink scale-up domain enables GB300 NVL72 to deliver higher per-GPU throughput and interactivity compared with NVIDIA B300. Results from NVIDIA’s AI Inference Performance Benchmarks page.

NVLink Fusion also includes NVIDIA NVLink-C2C for connecting XPUs to NVIDIA Vera CPUs or other ecosystem CPUs, delivering up to 6x the energy efficiency of a PCIe interface — helping remove barriers between control and compute for agentic systems.

A Proven Stack and Ecosystem for Development and Deployment

Teams developing custom XPUs often underestimate the effort and complexity of turning XPU innovation into data center deployment. This includes:

  • Integrating high-speed CPU and scale-up interfaces
  • Sourcing and validating a scale-up network solution
  • Designing compute and switch trays
  • Designing and validating a rack architecture, including cooling and power
  • Integrating security and storage
  • Managing a complex supplier ecosystem

The ideal platform provides all of this, allowing teams to focus on targeted innovation while using proven solutions for the rest.

NVLink Fusion is supported by an ecosystem designed for rapid development, integration and deployment, spanning ASIC design, CPU, and IP and optical interconnect partners.

“NVLink Fusion gives customers the ability to choose the CPU architecture, the performance level, the software capabilities that best meet their needs for the workloads that they care about,” said Tim Wilson, vice president and general manager of data center silicon engineering at Intel.

NVLink Fusion adopters can also use the NVIDIA MGX rack-scale architecture and the same supply chain used for MGX-based systems such as NVIDIA Vera Rubin NVL72. Manufacturing partners manage design and integration, while MGX suppliers provide the building blocks for rack, cooling, power and emerging 800 VDC designs.

“With Vera Rubin [NVL72], we are looking at almost 100% automation of system builds in the manufacturing line,” said Jack Luoh, head of product and solution at QCT and Quanta Computer. “Most of those investments can be leveraged if the XPU leverages NVLink Fusion.”

The NVIDIA AI infrastructure platform is vertically integrated and horizontally open. NVLink Fusion adopters can optionally incorporate NVIDIA Rubin GPUs, Vera CPUs, co-packaged optics switches, ConnectX SuperNICs, BlueField DPUs, Mission Control software and full-rack solutions including NVIDIA Vera Rubin NVL72, Vera CPU Rack, LPX, STX and SPX.

Managing Risk With Infrastructure Standardization

AI factory planning doesn’t wait for silicon. Power procurement, facility design, cooling, rack layout and network architecture begin long before the final accelerator mix is available. A data center locked to one chip can become a schedule risk.

Different workloads may favor different accelerators, including XPUs, GPUs, CPUs and LPUs. GPU systems may work alongside semi-custom systems for training, post-training, reasoning, retrieval and serving.

“The value of the NVLink Fusion program is … [customers] can deploy their rack-level solution with the NVIDIA GPU, and then they can decouple the development of their XPU and put it at a different pace,” said Vince Hu, corporate senior vice president and general manager of the data center and computing business group at MediaTek.

NVLink Fusion addresses these challenges  through a unified architecture. XPU- and GPU-based systems such as Vera Rubin NVL72 can share rack footprints, networking, cooling, power delivery and management systems. Operators can move forward with buildout while deferring the precise silicon mix, then reprovision capacity as workload demand, silicon supply and business priorities change. 

“NVLink Fusion allows the hyperscalers or the custom ASIC designers to integrate their own custom CPU or XPU and bridges the NVIDIA technology with a third-party process to create a unified rack-scale architecture,” said Lie-Szu Juang, chair and chief strategy officer at GUC.

Designed, Validated and Operated as a Factory

Factory buildout is expensive, and mistakes can require costly rework. Infrastructure must be validated before construction begins. NVLink Fusion aligns with the NVIDIA DSX reference architecture for AI factories: codesigning buildings, power, cooling, compute and networking. The NVIDIA Omniverse DSX AI Factory Blueprint provides a digital twin and open reference design for gigawatt-scale AI factories, enabling partners to model facilities and technology together before deployment.

At the rack level, serviceability is part of performance. Reference compute trays feature 100% liquid cooling with no fans, cables or hoses, and allow trays to be removed while the rest of the rack remains operational. NVLink Switch trays are also liquid cooled and support continued operation during service.

“With NVLink Fusion we can use proven NVL72 rack design to have time-to-market, and we can have access to multiple suppliers to help us to deliver more into the hands of our customers,” said CC Lee, senior hardware development manager at Annapurna Labs, an Amazon company.

Software completes the factory. NVIDIA NCCL for distributed workloads, NVIDIA Dynamo and NIXL for disaggregation and NVIDIA Mission Control for cluster management, telemetry and debugging help operators run mixed AI infrastructure as a coordinated system.

With NVLink Fusion, XPUs can now meet a world-class AI platform, enabling hyperscalers and AI-native companies to build unified, semi-custom AI factories that combine the strengths of many builders into infrastructure no one company could build alone.

Learn more about NVLink Fusion.

With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents

by

The next era of AI inference won’t be defined by a single breakthrough chip, network or system. It’ll be defined by how every layer of the AI factory works together. That’s why NVIDIA is extending Vera Rubin NVL72 with fast token generation for agentic systems.

Announced today, the NVIDIA Vera Rubin rack-scale system NVIDIA Groq 3 LPX is in full production. In an Artificial Analysis benchmark running Gemma 4 31B, an open source agentic model, it delivered 3,400 output tokens per second for 100,000-token long-context use cases critical to agentic systems, 4x faster than the nearest alternative platform. 

Industry partners worldwide are adopting Vera Rubin platform solutions. SpaceXAI announced that NVIDIA Vera CPUs will power its next generation of agentic AI. CoreWeave has deployed into production Spectrum-X Multiplane, which connects NVIDIA Vera Rubin racks using multiple parallel switches to provide high-bandwidth, flat and lossless AI networks. Nebius is the first AI cloud to adopt NVIDIA Groq 3 LPX.

As AI shifts from training to reasoning and agentic, inference has become the new frontier. Agentic AI systems are generating more tokens, processing dramatically larger context windows and increasingly collaborating with other AI systems to solve complex problems. 

These workloads demand a new class of infrastructure optimized not just for performance but for throughput, responsiveness and economics at unprecedented scale. 

At the Hot Chips conference this week in Palo Alto, California, NVIDIA is showcasing how extreme codesign is reshaping the AI factory from end to end. By architecting compute, networking and inference acceleration as a unified system, NVIDIA is helping customers build infrastructure purpose-built for the emerging demands of long-context inference and multi-agent systems.

Extreme Codesign Optimizes for Performance

Extreme codesign is the guiding principle behind NVIDIA platforms. Vera Rubin is engineered to accelerate inference as agents reason over increasingly long sequences. 

NVIDIA Spectrum-X Ethernet moves those massive data flows efficiently across AI factories, and NVIDIA Groq 3 LPX is built to generate tokens at ultrafast speeds. Together, they show how NVIDIA is optimizing every stage of the AI pipeline, from context and communication to generation, as part of a single, integrated AI factory architecture.

NVIDIA Groq 3 LPX brings a new low-latency inference architecture designed to work alongside Vera Rubin NLV72, the most versatile AI factory platform, helping enterprises and cloud providers deliver the low latency, extreme throughput and scalable economics required for agentic applications.

Breakthrough performance comes not from optimizing individual components in isolation, but from codesigning every layer of the stack. From networking and context processing to large-scale inference, NVIDIA’s full-stack platform turns AI factories into integrated engines for intelligence, built to turn ever-growing volumes of tokens into revenue.


Tuesday, Aug. 24, 8:00 a.m. PT 🔗

NVIDIA Partners Adopt Vera Rubin for Lowest Token Costs

Nebius, a leading AI cloud, is first to adopt NVIDIA Groq 3 LPX, giving developers access to leading token generation speeds for highly responsive agentic AI applications. 

Adding NVIDIA Groq 3 LPX to NVIDIA Vera Rubin NVL72 in  Nebius Token Factory will boost inference performance so developers can build highly interactive agents, coding systems and other real-time AI experiences at scale.

Connecting NVIDIA Vera Rubin racks, CoreWeave is deploying Spectrum-X Multiplane in production, unlocking advances for its AI cloud infrastructure.


Tuesday, Aug. 24, 8:00 a.m. PT 🔗

SpaceXAI Adopts NVIDIA Vera CPUs for Agentic AI

SpaceXAI plans to build and scale its future AI architecture around NVIDIA Vera Rubin, from data centers on Earth to orbital satellites. The company plans to deploy NVIDIA Vera CPUs to accelerate the CPU-intensive work behind agentic AI, including orchestration, tool use, code execution, data processing and simulation. 

The SpaceXAI partnership extends NVIDIA’s full-stack AI platform to SpaceXAI, bringing together Vera CPUs, NVIDIA accelerated computing, networking and software to advance AI at unprecedented scale.

Designed for the agentic era, Vera Rubin provides leading per-core performance, exceptional memory bandwidth and predictable performance under load, helping agents complete tasks faster and keeping valuable GPU infrastructure fully utilized.


Tuesday, Aug. 24, 8:00 a.m. PT 🔗

NVIDIA Groq 3 LPX: The Interactive AI Inference Accelerator

Codesigned with the Vera Rubin NVL72 platform, NVIDIA Groq 3 LPX is helping AI factories deliver tokens at the lowest latency for agentic workloads.

Agentic AI is creating a new performance challenge: decode latency. As AI agents reason, use tools and interact with other systems, they generate responses one token at a time, causing even tiny delays to multiply across complex chains of work. To keep agents operating at the pace users expect, NVIDIA Groq 3 LPX extends the Vera Rubin NVL72 platform with specialized acceleration for token generation. 

NVIDIA Rubin GPUs handle large-scale context processing while LPX accelerates latency-sensitive decode workloads. The result is faster, more predictable token generation that helps AI factories deliver responsive reasoning, smoother agent interactions and greater infrastructure efficiency. 

Together, Rubin GPUs and LPUs are designed to eliminate the traditional tradeoff between speed and throughput, helping AI providers deliver responsive, large-scale inference for the next generation of agentic AI applications.

Building the Token Factory

As the industry shifts from model training to serving intelligence at scale, infrastructure must evolve into what NVIDIA describes as a “token factory” capable of delivering performance, throughput, intelligence integrity and economic efficiency simultaneously. Agentic AI systems increasingly communicate with other AI systems, access multiple data sources and maintain large amounts of context, creating unprecedented demand for fast inference.

NVIDIA Groq 3 LPX was designed for exactly these workloads. As an extension of the Vera Rubin NVL72, it enables ultrafast responsiveness even across massive context windows while helping service providers maximize throughput and infrastructure utilization. 

Extreme Codesign for Inference

Unlike standalone accelerators, NVIDIA Groq 3 LPX combines the strengths of GPUs and LPUs through extreme codesign. Rubin GPUs and LPUs jointly compute every layer of an AI model, enabling new levels of inference performance for agentic workloads. 

At scale, fleets of LPUs operate as a giant processor optimized for deterministic inference. A rack-scale NVIDIA Groq 3 LPX deployment can include 256 LP30 accelerators connected through direct chip-to-chip links, creating a highly efficient inference engine built for modern AI factories.

Designed for the Agentic AI Era

As reasoning models grow and agentic workflows generate ever more tokens, the infrastructure required to serve them must evolve. NVIDIA Groq 3 LPX extends the Vera Rubin NVL72 platform with a purpose-built inference architecture designed to maximize responsiveness, throughput and efficiency, helping power the next generation of AI factories.

And this is only the beginning, more optimizations, more models, more performance when paired with Vera Rubin NVL72 — new levels of throughput and interactivity are coming. Stay tuned. 


Tuesday, Aug. 24, 8:00 a.m. PT 🔗

NVIDIA Spectrum-X Multiplane Enables Massive AI Factory Scale on a Flatter, More Resilient Network​

As AI factories grow massive, the network has become a critical engine of performance. At Hot Chips, NVIDIA is spotlighting Spectrum-X Multiplane — the latest in the hardware-accelerated Spectrum-X Ethernet architecture that lets Ethernet scale to unprecedented size while avoiding the latency, jitter and cost of adding another network tier.

NVIDIA Spectrum-X Ethernet is designed as an end-to-end, AI-optimized Ethernet platform, combining NVIDIA Spectrum-X Ethernet switches, SuperNICs and software to improve the performance and efficiency of Ethernet-based AI infrastructure for AI factories and clouds. The platform is designed to deliver 1.6x better AI networking performance compared with off-the-shelf Ethernet, while providing consistent, predictable performance in multi-tenant environments.

Multiplane Unlocks Scale Without the Tradeoffs of a New Tier

Scaling an AI factory beyond today’s largest clusters traditionally means adding a third network tier, which adds latency, slows things down unpredictably and drives up the cost of cabling, optics and power. Spectrum-X Multiplane takes a simpler approach: It splits each server’s network connection into several independent paths, or “planes,” each running its own lightweight two-tier network. The result is a flat, simple network that scales to 512,000 GPUs, without the added cost and complexity of a third tier.

This all happens automatically. A dedicated hardware engine inside the NVIDIA ConnectX SuperNIC manages traffic across the planes and instantly reroutes around any failure, so applications and software simply see one fast, reliable connection. In an eight-plane topology, if one plane fails, the network still maintains about 90% of its total bandwidth, with hardware recovery that’s 11x faster than software-based multiplane load balancing. This translates to 1.6x higher AI factory output.

Built Through Extreme Codesign

That reliability comes from extreme codesign of Vera Rubin NVL72, spanning switch silicon, SuperNICs and software. Spectrum-X SN6000 series switches, based on the 102.4Tb/s Spectrum-6 Ethernet ASIC and ConnectX-9 SuperNICs, supporting up to 1,600Gb/s per GPU, are purpose-built for Vera Rubin NVL72 AI factories. Spectrum-XGS Ethernet extends that same codesign across data centers, letting multiple facilities function as a single AI super-factory and accelerating multi-site NCCL collectives by 1.9x.


Tuesday, Aug. 24, 8:00 a.m. PT 🔗

NVIDIA Introduces Scale-In Infrastructure for Agentic AI Factories, Powered by BlueField-4, DOCA

NVIDIA is introducing NVIDIA Scale-In, the fifth pillar of NVIDIA AI networking and a new class of accelerated network infrastructure for agentic AI factories. Scale-In extends purpose-built acceleration to the infrastructure services that secure, manage and operate the AI factory.

Powered by the NVIDIA BlueField-4 processor and NVIDIA DOCA software platform and connected over NVIDIA Spectrum-X Ethernet, NVIDIA Scale-In transforms the traditional north-south access network into a unified, accelerated infrastructure domain.

Cloud computing brought software-defined networking, composability and elasticity to the data center, enabling users, applications, data and services to scale dynamically. 

Agentic AI represents the next platform shift. AI factories bring together massive accelerated compute with growing numbers of users, applications and autonomous agents, all continuously interacting with data, storage and services. This transforms the demands on the infrastructure that brings AI to life. Networking, storage, cybersecurity and operations must now be accelerated alongside AI compute, combining software-defined flexibility with purpose-built hardware acceleration and full-stack codesign. 

NVIDIA Scale-In delivers multi-tenant networking, high-performance storage access, in-silicon security, elastic provisioning and real-time observability, while keeping infrastructure processing independent of host compute resources. By accelerating and codesigning these services as part of the AI factory, Scale-In helps security, data access and operations scale alongside AI compute. The result is secure, efficient and manageable shared infrastructure for deploying and operating agentic AI at massive scale.


Tuesday, Aug. 24, 8:00 a.m. PT 🔗

NVIDIA NVLink Fusion Connects XPUs to NVIDIA’s Leading AI Platform

NVIDIA NVLink Fusion brings custom silicon into NVIDIA’s world-leading AI infrastructure platform, enabling hyperscalers and AI-native companies to build semi-custom AI factories with greater performance, flexibility and speed.

As AI models grow in size and complexity, raw compute alone is not enough. AI factories require high-bandwidth, low-latency scale-up networking, proven rack-scale architectures and a full ecosystem spanning power, cooling, management software and supply chain. NVLink Fusion addresses these challenges by connecting custom XPUs and CPUs to NVIDIA’s scale-up and scale-out technology stack.

The platform includes sixth-generation NVIDIA NVLink and NVLink Switch purpose-built scale-up networking, as well as NVLink-C2C for energy-efficient connectivity between XPUs and CPUs. Through the NVIDIA MGX ecosystem, adopters can also use production-proven rack designs, components, manufacturing partner solutions and open, extensible software for distributed computing, disaggregated workloads and cluster management.

By standardizing GPU- and XPU-based systems on a unified architecture, NVLink Fusion helps decouple data center buildout from silicon readiness. Operators can share rack footprints, networking, cooling, power delivery and management systems, then adjust the mix of GPUs and XPUs as supply and workload requirements evolve.

NVLink Fusion extends the NVIDIA AI platform’s vertically integrated, horizontally open approach to custom silicon. It gives partners the freedom to innovate where they differentiate while drawing on NVIDIA technologies across compute, networking, infrastructure and software — creating a single, flexible AI factory that no one company could build alone.

Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents

New on-silicon performance data measured by NVIDIA using real-world agentic coding trajectories shows Vera Rubin NVL72 systems deliver 30x higher throughput per megawatt and 35x lower token costs than NVIDIA GB300 NVL72.
by

According to OpenRouter data, agentic AI workloads consume 15x more tokens than a simple chat request. Why? 

Consider what happens when an AI agent researches a company for an investment decision. The agent queries financial databases, searches news and filings, invokes a sub-agent to run peer comparisons and model valuations, then synthesizes everything into a recommendation. Agents and sub-agents keep reasoning until the task is done, driving increased token demand. With every step, the accumulated tokens become the input to the next, making long-context handling central to agentic AI performance.

The same pattern plays out across every agentic use case, from software development to customer service to deep research. 

As agentic AI moves into production across industries, the infrastructure running it needs to meet that token demand efficiently. 

New measured performance data shows NVIDIA Vera Rubin NVL72 systems deliver up to 30x higher throughput per megawatt than NVIDIA GB300 NVL72 on agentic workloads. NVIDIA measured this inference throughput data using the SemiAnalysis AgentX workload, consisting of recorded real-world agentic coding sessions, with actual context growth, tool calls and sub-agent spawning preserved. For power-constrained AI factories, that translates directly into 30x more agentic work for the same energy footprint.

These early results for Vera Rubin NVL72 demonstrate NVIDIA’s accelerated pace of innovation. With continuous software optimizations, performance across both Vera Rubin NVL72 and GB300 NVL72 will continue to improve.

Vera Rubin NVL72: 30x Higher Throughput per Megawatt and 35x Lower Token Cost

Agentic workloads look fundamentally different from chat or document summarization, where input and output sequences typically range from 1K to 8K tokens. In agentic sessions, context accumulates across steps and can reach hundreds of thousands of input tokens, with wide variability in both input and output lengths across requests. Performance measurement must evolve to capture the full agent workflow rather than a single inference request.

The results below reflect performance measured on real-world agentic coding trajectories.

In SemiAnalysis AgentX, the NVIDIA Blackwell platform delivers leading performance across multiple agentic models including Kimi K3, MiniMax M3, GLM5.3, Qwen3.5 and DeepSeek V4 Pro. 

For example, GB300 NVL72 delivers up to 15x better throughput per megawatt than the NVIDIA Hopper architecture on the DeepSeek V4 Pro model, giving customers a high-performance foundation to run agentic workloads. This leap reflects the advantage of a larger scale-up GPU domain and codesigned software in delivering significantly better inference efficiency.

Vera Rubin extends that advantage, lifting the performance across the entire Pareto curve, to deliver as much as 30x higher throughput per megawatt than GB300 NVL72 on the DeepSeek V4 Pro model. These early results, measured using the SemiAnalysis AgentX workload and currently pending SemiAnalysis review, don’t yet reflect Vera CPU performance for tool calling. 

NVIDIA DSX MaxLPS technologies manages power across the GPU, rack and workload levels to provision up to 40% more GPUs within the same megawatt budget, pushing throughput per megawatt further at AI factory scale. 

Throughput per megawatt also directly impacts the cost of every token produced. At up to 35x lower cost per million tokens than GB300 NVL72, Vera Rubin NVL72 can run agents continuously, at scale, across the full breadth of customers’ workloads. 

For power-constrained AI factories, throughput per megawatt determines AI factory revenue and cost per million tokens determines the profit margin on that revenue.

Extreme Codesign for Agentic Scale

Modern inference optimization spans a range of techniques that are especially critical for agentic AI. NVIDIA Vera Rubin NVL72 enables all of these and more through extreme codesign across every layer of the platform to deliver multifold performance gains.

  • Disaggregated serving separates context processing (prefill) from response generation (decode) so each scales independently.
  • Rate matching synchronizes the speeds at which prefill GPUs and decode GPUs produce tokens to maximize efficiency.
  • Large-scale expert parallelism distributes expert sub-networks in mixture-of-experts models across the scale-up GPU domain. 
  • Distributed KV-caching extends memory across the scale-up GPU domain, while KV-cache offloading tiers less-active context to host and storage, keeping previously processed context accessible without recomputation.
  • KV-aware routing directs incoming requests to the GPUs that already hold the relevant cached context, reducing redundant computation across long sessions. 
  • Fused CUDA kernels like MegaMoE combine many computation and inter-GPU communication operations into a single execution pass, keeping GPUs active rather than waiting for data.

NVIDIA Rubin GPUs’ enhanced fifth-generation Tensor Cores and the third-generation Transformer Engine accelerate both prefill and decode stages of inference. NVFP4 quantization compresses model weights to 4-bit precision, reducing memory footprint and increasing throughput without sacrificing output quality. 

The NVL72 scale-up domain, a defining architecture across Vera Rubin and Grace Blackwell, enables the high-bandwidth and low-latency inter-GPU communication essential for techniques such as large-scale expert parallelism and distributed KV-caching. Purpose-built to power this scale-up domain, NVIDIA NVLink interconnect technology and NVLink Switches, now in their sixth generation, deliver 10x higher packet rates and 3x lower latency than off-the-shelf Ethernet alternatives.

Spanning optimized CUDA kernels, inference runtimes like NVIDIA TensorRT LLM and serving frameworks like NVIDIA Dynamo, NVIDIA’s software stack is codesigned with the hardware to enable inference optimizations.

While the results above reflect current Vera Rubin NVL72 performance, the full platform is a seven-chip architecture that also includes the NVIDIA Vera CPU, Groq 3 LPU, NVLink 6 Switch, BlueField-4 DPU, Spectrum-6 SPX and ConnectX-9 SuperNIC, all purpose-built for AI factories deploying agents at scale.

Extreme codesign also extends to NVIDIA’s co-engineering with its partners. Vera Rubin is in full production and is scaling across the ecosystem. 

Learn more about the NVIDIA Vera Rubin platform.