country_code

Top Israel Medical Center Partners with AI Startups to Help Detect Brain Bleeds, Other Critical Cases

Assuta Medical Centers, Aidoc and Rhino Health are integrating tools built on NVIDIA AI technology with hospital research and clinical workflows to improve patient outcomes.
by
Aidoc's AI system

Israel’s largest private medical center is working with startups and researchers to bring potentially life-saving AI solutions to real-world healthcare workflows.

With more than 1.5 million patients across eight medical centers, Assuta Medical Centers conduct over 100,000 surgeries, 800,000 imaging tests and hundreds of thousands of other health diagnostics and treatments each year. These create huge amounts of de-identified data that Assuta is securely sharing with more than 20 startups through its innovation arm, RISE, launched last year working in collaboration with NVIDIA.

One of the startups, Aidoc, is helping Assuta alert imaging technicians with AI-based insights of possible bleeding in the brain and other critical conditions in a patient’s scan within minutes. Another, Rhino Health, is using federated learning powered by NVIDIA FLARE to make AI development on diverse medical datasets from hospitals across the globe more accessible to Assuta’s collaborators.

Both companies are members of NVIDIA Inception, a global program designed to support cutting-edge startups with go-to-market support, expertise and technology.

“We’re building a hub to serve innovators with the infrastructure they need to develop, test and deploy new AI technology for image analysis and other data-heavy computations in radiology, pathology, genomics and more,” said Daniel Rabina, director of innovation at RISE. “We want to make collaboration with companies, research institutes, hospitals and universities possible while maintaining patient data privacy.”

To support AI development, testing and deployment, Assuta has installed NVIDIA DGX A100 systems on premises and adopted the NVIDIA Clara Holoscan platform, plus software libraries including MONAI for healthcare imaging and NVIDIA FLARE for federated learning.

NVIDIA and RISE are collaborating on RISE with US, a program built to introduce selected Israeli entrepreneurs and early-stage startups working on digital and computational health solutions to the U.S. market. Applications to join the program are open until August 28.

Aidoc Flags Urgent Cases for Radiologist Review

Aidoc, which is New York-based with a research branch in Israel, has developed FDA-cleared AI solutions to flag acute conditions including brain hemorrhages, pulmonary embolisms and strokes from imaging scans.

Aidoc desktop and mobile interfaceFounded in 2016 by a group of veterans from the Israel Defense Forces, the startup has deployed its AI to analyze millions of cases across more than 1,000 medical facilities, primarily in the U.S., Europe and Israel.

Its algorithms integrate seamlessly with the PACS imaging workflow used by radiologists worldwide, working behind the scenes to analyze each imaging study and flag urgent findings — bringing potentially critical cases to the radiologist’s attention for review.

Aidoc’s tools can help address the growing shortage of radiologists globally by reducing the time a radiologist needs to spend on each case, enabling care for more patients. And by pushing potentially critical cases to the top of a radiologist’s pile, the AI can help clinicians catch important findings sooner, improving patient outcomes.

The startup uses NVIDIA Tensor Core GPUs in the cloud through AWS for AI training and inference. Adopting NVIDIA GPUs helped reduce model training time from days to a couple hours.

Immediate Impact at Assuta Medical Centers 

Assuta is a private chain of hospitals that provides elective care — typically dealing with routine screenings rather than emergency room patients — but it adopted Aidoc’s solution to help imaging technicians spot critical cases that need urgent attention among its roughly 200,000 CT tests conducted annually.

When a radiology scan isn’t urgent, it may take a couple days for a doctor to review the case. Aidoc can shrink this time to minutes by identifying concerning cases as soon as the scans are captured by radiology staff. Assuta facilities

At Assuta, urgent findings are typically found among cancer patients, or people who have recently undergone surgery and need follow-up scans. The healthcare organization is using Aidoc’s AI tools to detect intracranial hemorrhages and two kinds of pulmonary embolism.

“We saw the impact right away,” said Dr. Michal Guindy, head of medical imaging and head of RISE at Assuta. “Just a couple days after installing Aidoc at Assuta, a patient came in for a follow-up scan after a brain procedure and had an intracranial hemorrhage. Because Aidoc alerted the imaging technician to flag it for further review, our doctors were able to call the patient while they were on their way home and immediately redirect them to the hospital for treatment.”

Rhino Health Fosters Collaboration With Federated Learning

In addition to deploying AI models in full-scale, real-world settings, Assuta is supporting innovators who are developing, testing or validating new medical AI solutions by sharing the healthcare organization’s data, while also using federated learning through Rhino Health.

Assuta has millions of radiology cases digitized — a desirable resource for researchers and startups looking for robust, diverse datasets to train or validate their AI models. But because of data privacy protection, it’s important that patient information stays safely within the firewall of medical centers like Assuta.

“Data diversity is necessary to develop AI models meant for the use of medical teams. Without optimal computing resources, it would be extremely difficult to use our data and make the magic happen,” said Rabina. “That’s why we need federated learning enabled by both NVIDIA and Rhino Health.”

Federated learning allows companies, healthcare institutions and universities to work together by training and validating AI models across multiple organizations’ datasets while maintaining each organization’s data privacy. Rhino Health provides a neutral platform — available through the NVIDIA AI Enterprise software suite — that enables secure collaboration, powered by NVIDIA A100 GPUs in the cloud and the NVIDIA FLARE federated learning framework.

With Rhino Health, Assuta aims to help its collaborators develop AI models across hospitals internationally, resulting in more generalizable algorithms that perform more accurately across different patient populations.

Register for NVIDIA GTC, running online Sept. 19-22, to hear more from leaders in healthcare AI.

Subscribe to NVIDIA healthcare news and watch on demand as Assuta, Aidoc and Rhino Health speak at an GTC panel.

How XPUs Meet a World-Class AI Factory

Deploying custom silicon with leading AI infrastructure enables hyperscalers and AI-native companies to build flexible AI factories that combine specialization with scale.
by

To generate intelligence at scale, AI factories run continuously, and their economics are defined by delivered output: tokens per second, tokens per watt, cost per token, utilization and uptime. 

That requires AI infrastructure designed and built as a full factory, not a collection of individual accelerators.

Hyperscalers and AI-native companies building custom XPUs must consider not just XPU design, but the design and development of the entire AI platform, including scale-up and scale-out networking, rack-scale architecture, production factory software and a robust supplier ecosystem. 

At AI factory scale, this path is complex and costly, and represents a fundamental obstacle to getting XPUs to market quickly. 

Breaking the constraint means combining custom XPUs with proven, mature infrastructure — allowing builders to focus innovation where it matters most while harnessing established technology for the rest. 

NVLink Fusion delivers on that need, connecting XPUs to NVIDIA’s world-leading AI infrastructure to increase performance, accelerate time to market and mitigate risk for semi-custom AI factories.

Unlock XPU Performance With Fast Scale-Up

For modern workloads such as running trillion-parameter models, mixture-of-experts architectures and agentic AI, if the scale-up fabric cannot keep up, utilization drops and cost per token rises.

A scale-up networking solution must excel on three dimensions: 

  • Delivered performance: End-to-end network performance, in-network compute and  mature software integration.
  • Factory resiliency: Uptime, continuous health monitoring and telemetry, and component-level serviceability while the factory keeps running.
  • Platform maturity: Reduced operational risk by using a mature technology stack with a demonstrated track record of large-scale deployments and realized return on investment. 

As an example, NVLink Fusion brings XPUs into the NVIDIA NVLink scale-up domain. Sixth-generation NVLink provides leading high-bandwidth, low-latency networking across a 72-XPU domain. The end-to-end latency for XPU-to-XPU transfers is 3x lower than alternative solutions based on off-the-shelf Ethernet, and the packet rate is 10x higher.

For end-to-end performance, NVIDIA GB300 NVL72 systems help deliver significantly higher throughput and better interactivity compared with configurations that don’t use NVL72, and future NVLink roadmap configurations include domains of up to 1,152 accelerators and co-packaged optics.

A Pareto chart comparing GB300 NVL72 with B300 inference throughput performance in tokens per second per GPU on DeepSeek-V4-Pro at ISL=1K and OSL=1K sampled at various interactivity points in tokens per second per user. GB300 NVL72 is more than 10x the throughput of B300 in the middle of the Pareto between 70 and 100 tokens per second per user.
The 72-GPU NVLink scale-up domain enables GB300 NVL72 to deliver higher per-GPU throughput and interactivity compared with NVIDIA B300. Results from NVIDIA’s AI Inference Performance Benchmarks page.

NVLink Fusion also includes NVIDIA NVLink-C2C for connecting XPUs to NVIDIA Vera CPUs or other ecosystem CPUs, delivering up to 6x the energy efficiency of a PCIe interface — helping remove barriers between control and compute for agentic systems.

A Proven Stack and Ecosystem for Development and Deployment

Teams developing custom XPUs often underestimate the effort and complexity of turning XPU innovation into data center deployment. This includes:

  • Integrating high-speed CPU and scale-up interfaces
  • Sourcing and validating a scale-up network solution
  • Designing compute and switch trays
  • Designing and validating a rack architecture, including cooling and power
  • Integrating security and storage
  • Managing a complex supplier ecosystem

The ideal platform provides all of this, allowing teams to focus on targeted innovation while using proven solutions for the rest.

NVLink Fusion is supported by an ecosystem designed for rapid development, integration and deployment, spanning ASIC design, CPU, and IP and optical interconnect partners.

“NVLink Fusion gives customers the ability to choose the CPU architecture, the performance level, the software capabilities that best meet their needs for the workloads that they care about,” said Tim Wilson, vice president and general manager of data center silicon engineering at Intel.

NVLink Fusion adopters can also use the NVIDIA MGX rack-scale architecture and the same supply chain used for MGX-based systems such as NVIDIA Vera Rubin NVL72. Manufacturing partners manage design and integration, while MGX suppliers provide the building blocks for rack, cooling, power and emerging 800 VDC designs.

“With Vera Rubin [NVL72], we are looking at almost 100% automation of system builds in the manufacturing line,” said Jack Luoh, head of product and solution at QCT and Quanta Computer. “Most of those investments can be leveraged if the XPU leverages NVLink Fusion.”

The NVIDIA AI infrastructure platform is vertically integrated and horizontally open. NVLink Fusion adopters can optionally incorporate NVIDIA Rubin GPUs, Vera CPUs, co-packaged optics switches, ConnectX SuperNICs, BlueField DPUs, Mission Control software and full-rack solutions including NVIDIA Vera Rubin NVL72, Vera CPU Rack, LPX, STX and SPX.

Managing Risk With Infrastructure Standardization

AI factory planning doesn’t wait for silicon. Power procurement, facility design, cooling, rack layout and network architecture begin long before the final accelerator mix is available. A data center locked to one chip can become a schedule risk.

Different workloads may favor different accelerators, including XPUs, GPUs, CPUs and LPUs. GPU systems may work alongside semi-custom systems for training, post-training, reasoning, retrieval and serving.

“The value of the NVLink Fusion program is … [customers] can deploy their rack-level solution with the NVIDIA GPU, and then they can decouple the development of their XPU and put it at a different pace,” said Vince Hu, corporate senior vice president and general manager of the data center and computing business group at MediaTek.

NVLink Fusion addresses these challenges  through a unified architecture. XPU- and GPU-based systems such as Vera Rubin NVL72 can share rack footprints, networking, cooling, power delivery and management systems. Operators can move forward with buildout while deferring the precise silicon mix, then reprovision capacity as workload demand, silicon supply and business priorities change. 

“NVLink Fusion allows the hyperscalers or the custom ASIC designers to integrate their own custom CPU or XPU and bridges the NVIDIA technology with a third-party process to create a unified rack-scale architecture,” said Lie-Szu Juang, chair and chief strategy officer at GUC.

Designed, Validated and Operated as a Factory

Factory buildout is expensive, and mistakes can require costly rework. Infrastructure must be validated before construction begins. NVLink Fusion aligns with the NVIDIA DSX reference architecture for AI factories: codesigning buildings, power, cooling, compute and networking. The NVIDIA Omniverse DSX AI Factory Blueprint provides a digital twin and open reference design for gigawatt-scale AI factories, enabling partners to model facilities and technology together before deployment.

At the rack level, serviceability is part of performance. Reference compute trays feature 100% liquid cooling with no fans, cables or hoses, and allow trays to be removed while the rest of the rack remains operational. NVLink Switch trays are also liquid cooled and support continued operation during service.

“With NVLink Fusion we can use proven NVL72 rack design to have time-to-market, and we can have access to multiple suppliers to help us to deliver more into the hands of our customers,” said CC Lee, senior hardware development manager at Annapurna Labs, an Amazon company.

Software completes the factory. NVIDIA NCCL for distributed workloads, NVIDIA Dynamo and NIXL for disaggregation and NVIDIA Mission Control for cluster management, telemetry and debugging help operators run mixed AI infrastructure as a coordinated system.

With NVLink Fusion, XPUs can now meet a world-class AI platform, enabling hyperscalers and AI-native companies to build unified, semi-custom AI factories that combine the strengths of many builders into infrastructure no one company could build alone.

Learn more about NVLink Fusion.

With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents

by

The next era of AI inference won’t be defined by a single breakthrough chip, network or system. It’ll be defined by how every layer of the AI factory works together. That’s why NVIDIA is extending Vera Rubin NVL72 with fast token generation for agentic systems.

Announced today, the NVIDIA Vera Rubin rack-scale system NVIDIA Groq 3 LPX is in full production. In an Artificial Analysis benchmark running Gemma 4 31B, an open source agentic model, it delivered 3,400 output tokens per second for 100,000-token long-context use cases critical to agentic systems, 4x faster than the nearest alternative platform. 

Industry partners worldwide are adopting Vera Rubin platform solutions. SpaceXAI announced that NVIDIA Vera CPUs will power its next generation of agentic AI. CoreWeave has deployed into production Spectrum-X Multiplane, which connects NVIDIA Vera Rubin racks using multiple parallel switches to provide high-bandwidth, flat and lossless AI networks. Nebius is the first AI cloud to adopt NVIDIA Groq 3 LPX.

As AI shifts from training to reasoning and agentic, inference has become the new frontier. Agentic AI systems are generating more tokens, processing dramatically larger context windows and increasingly collaborating with other AI systems to solve complex problems. 

These workloads demand a new class of infrastructure optimized not just for performance but for throughput, responsiveness and economics at unprecedented scale. 

At the Hot Chips conference this week in Palo Alto, California, NVIDIA is showcasing how extreme codesign is reshaping the AI factory from end to end. By architecting compute, networking and inference acceleration as a unified system, NVIDIA is helping customers build infrastructure purpose-built for the emerging demands of long-context inference and multi-agent systems.

Extreme Codesign Optimizes for Performance

Extreme codesign is the guiding principle behind NVIDIA platforms. Vera Rubin is engineered to accelerate inference as agents reason over increasingly long sequences. 

NVIDIA Spectrum-X Ethernet moves those massive data flows efficiently across AI factories, and NVIDIA Groq 3 LPX is built to generate tokens at ultrafast speeds. Together, they show how NVIDIA is optimizing every stage of the AI pipeline, from context and communication to generation, as part of a single, integrated AI factory architecture.

NVIDIA Groq 3 LPX brings a new low-latency inference architecture designed to work alongside Vera Rubin NLV72, the most versatile AI factory platform, helping enterprises and cloud providers deliver the low latency, extreme throughput and scalable economics required for agentic applications.

Breakthrough performance comes not from optimizing individual components in isolation, but from codesigning every layer of the stack. From networking and context processing to large-scale inference, NVIDIA’s full-stack platform turns AI factories into integrated engines for intelligence, built to turn ever-growing volumes of tokens into revenue.


Tuesday, Aug. 24, 8:00 a.m. PT 🔗

NVIDIA Partners Adopt Vera Rubin for Lowest Token Costs

Nebius, a leading AI cloud, is first to adopt NVIDIA Groq 3 LPX, giving developers access to leading token generation speeds for highly responsive agentic AI applications. 

Adding NVIDIA Groq 3 LPX to NVIDIA Vera Rubin NVL72 in  Nebius Token Factory will boost inference performance so developers can build highly interactive agents, coding systems and other real-time AI experiences at scale.

Connecting NVIDIA Vera Rubin racks, CoreWeave is deploying Spectrum-X Multiplane in production, unlocking advances for its AI cloud infrastructure.


Tuesday, Aug. 24, 8:00 a.m. PT 🔗

SpaceXAI Adopts NVIDIA Vera CPUs for Agentic AI

SpaceXAI plans to build and scale its future AI architecture around NVIDIA Vera Rubin, from data centers on Earth to orbital satellites. The company plans to deploy NVIDIA Vera CPUs to accelerate the CPU-intensive work behind agentic AI, including orchestration, tool use, code execution, data processing and simulation. 

The SpaceXAI partnership extends NVIDIA’s full-stack AI platform to SpaceXAI, bringing together Vera CPUs, NVIDIA accelerated computing, networking and software to advance AI at unprecedented scale.

Designed for the agentic era, Vera Rubin provides leading per-core performance, exceptional memory bandwidth and predictable performance under load, helping agents complete tasks faster and keeping valuable GPU infrastructure fully utilized.


Tuesday, Aug. 24, 8:00 a.m. PT 🔗

NVIDIA Groq 3 LPX: The Interactive AI Inference Accelerator

Codesigned with the Vera Rubin NVL72 platform, NVIDIA Groq 3 LPX is helping AI factories deliver tokens at the lowest latency for agentic workloads.

Agentic AI is creating a new performance challenge: decode latency. As AI agents reason, use tools and interact with other systems, they generate responses one token at a time, causing even tiny delays to multiply across complex chains of work. To keep agents operating at the pace users expect, NVIDIA Groq 3 LPX extends the Vera Rubin NVL72 platform with specialized acceleration for token generation. 

NVIDIA Rubin GPUs handle large-scale context processing while LPX accelerates latency-sensitive decode workloads. The result is faster, more predictable token generation that helps AI factories deliver responsive reasoning, smoother agent interactions and greater infrastructure efficiency. 

Together, Rubin GPUs and LPUs are designed to eliminate the traditional tradeoff between speed and throughput, helping AI providers deliver responsive, large-scale inference for the next generation of agentic AI applications.

Building the Token Factory

As the industry shifts from model training to serving intelligence at scale, infrastructure must evolve into what NVIDIA describes as a “token factory” capable of delivering performance, throughput, intelligence integrity and economic efficiency simultaneously. Agentic AI systems increasingly communicate with other AI systems, access multiple data sources and maintain large amounts of context, creating unprecedented demand for fast inference.

NVIDIA Groq 3 LPX was designed for exactly these workloads. As an extension of the Vera Rubin NVL72, it enables ultrafast responsiveness even across massive context windows while helping service providers maximize throughput and infrastructure utilization. 

Extreme Codesign for Inference

Unlike standalone accelerators, NVIDIA Groq 3 LPX combines the strengths of GPUs and LPUs through extreme codesign. Rubin GPUs and LPUs jointly compute every layer of an AI model, enabling new levels of inference performance for agentic workloads. 

At scale, fleets of LPUs operate as a giant processor optimized for deterministic inference. A rack-scale NVIDIA Groq 3 LPX deployment can include 256 LP30 accelerators connected through direct chip-to-chip links, creating a highly efficient inference engine built for modern AI factories.

Designed for the Agentic AI Era

As reasoning models grow and agentic workflows generate ever more tokens, the infrastructure required to serve them must evolve. NVIDIA Groq 3 LPX extends the Vera Rubin NVL72 platform with a purpose-built inference architecture designed to maximize responsiveness, throughput and efficiency, helping power the next generation of AI factories.

And this is only the beginning, more optimizations, more models, more performance when paired with Vera Rubin NVL72 — new levels of throughput and interactivity are coming. Stay tuned. 


Tuesday, Aug. 24, 8:00 a.m. PT 🔗

NVIDIA Spectrum-X Multiplane Enables Massive AI Factory Scale on a Flatter, More Resilient Network​

As AI factories grow massive, the network has become a critical engine of performance. At Hot Chips, NVIDIA is spotlighting Spectrum-X Multiplane — the latest in the hardware-accelerated Spectrum-X Ethernet architecture that lets Ethernet scale to unprecedented size while avoiding the latency, jitter and cost of adding another network tier.

NVIDIA Spectrum-X Ethernet is designed as an end-to-end, AI-optimized Ethernet platform, combining NVIDIA Spectrum-X Ethernet switches, SuperNICs and software to improve the performance and efficiency of Ethernet-based AI infrastructure for AI factories and clouds. The platform is designed to deliver 1.6x better AI networking performance compared with off-the-shelf Ethernet, while providing consistent, predictable performance in multi-tenant environments.

Multiplane Unlocks Scale Without the Tradeoffs of a New Tier

Scaling an AI factory beyond today’s largest clusters traditionally means adding a third network tier, which adds latency, slows things down unpredictably and drives up the cost of cabling, optics and power. Spectrum-X Multiplane takes a simpler approach: It splits each server’s network connection into several independent paths, or “planes,” each running its own lightweight two-tier network. The result is a flat, simple network that scales to 512,000 GPUs, without the added cost and complexity of a third tier.

This all happens automatically. A dedicated hardware engine inside the NVIDIA ConnectX SuperNIC manages traffic across the planes and instantly reroutes around any failure, so applications and software simply see one fast, reliable connection. In an eight-plane topology, if one plane fails, the network still maintains about 90% of its total bandwidth, with hardware recovery that’s 11x faster than software-based multiplane load balancing. This translates to 1.6x higher AI factory output.

Built Through Extreme Codesign

That reliability comes from extreme codesign of Vera Rubin NVL72, spanning switch silicon, SuperNICs and software. Spectrum-X SN6000 series switches, based on the 102.4Tb/s Spectrum-6 Ethernet ASIC and ConnectX-9 SuperNICs, supporting up to 1,600Gb/s per GPU, are purpose-built for Vera Rubin NVL72 AI factories. Spectrum-XGS Ethernet extends that same codesign across data centers, letting multiple facilities function as a single AI super-factory and accelerating multi-site NCCL collectives by 1.9x.


Tuesday, Aug. 24, 8:00 a.m. PT 🔗

NVIDIA Introduces Scale-In Infrastructure for Agentic AI Factories, Powered by BlueField-4, DOCA

NVIDIA is introducing NVIDIA Scale-In, the fifth pillar of NVIDIA AI networking and a new class of accelerated network infrastructure for agentic AI factories. Scale-In extends purpose-built acceleration to the infrastructure services that secure, manage and operate the AI factory.

Powered by the NVIDIA BlueField-4 processor and NVIDIA DOCA software platform and connected over NVIDIA Spectrum-X Ethernet, NVIDIA Scale-In transforms the traditional north-south access network into a unified, accelerated infrastructure domain.

Cloud computing brought software-defined networking, composability and elasticity to the data center, enabling users, applications, data and services to scale dynamically. 

Agentic AI represents the next platform shift. AI factories bring together massive accelerated compute with growing numbers of users, applications and autonomous agents, all continuously interacting with data, storage and services. This transforms the demands on the infrastructure that brings AI to life. Networking, storage, cybersecurity and operations must now be accelerated alongside AI compute, combining software-defined flexibility with purpose-built hardware acceleration and full-stack codesign. 

NVIDIA Scale-In delivers multi-tenant networking, high-performance storage access, in-silicon security, elastic provisioning and real-time observability, while keeping infrastructure processing independent of host compute resources. By accelerating and codesigning these services as part of the AI factory, Scale-In helps security, data access and operations scale alongside AI compute. The result is secure, efficient and manageable shared infrastructure for deploying and operating agentic AI at massive scale.


Tuesday, Aug. 24, 8:00 a.m. PT 🔗

NVIDIA NVLink Fusion Connects XPUs to NVIDIA’s Leading AI Platform

NVIDIA NVLink Fusion brings custom silicon into NVIDIA’s world-leading AI infrastructure platform, enabling hyperscalers and AI-native companies to build semi-custom AI factories with greater performance, flexibility and speed.

As AI models grow in size and complexity, raw compute alone is not enough. AI factories require high-bandwidth, low-latency scale-up networking, proven rack-scale architectures and a full ecosystem spanning power, cooling, management software and supply chain. NVLink Fusion addresses these challenges by connecting custom XPUs and CPUs to NVIDIA’s scale-up and scale-out technology stack.

The platform includes sixth-generation NVIDIA NVLink and NVLink Switch purpose-built scale-up networking, as well as NVLink-C2C for energy-efficient connectivity between XPUs and CPUs. Through the NVIDIA MGX ecosystem, adopters can also use production-proven rack designs, components, manufacturing partner solutions and open, extensible software for distributed computing, disaggregated workloads and cluster management.

By standardizing GPU- and XPU-based systems on a unified architecture, NVLink Fusion helps decouple data center buildout from silicon readiness. Operators can share rack footprints, networking, cooling, power delivery and management systems, then adjust the mix of GPUs and XPUs as supply and workload requirements evolve.

NVLink Fusion extends the NVIDIA AI platform’s vertically integrated, horizontally open approach to custom silicon. It gives partners the freedom to innovate where they differentiate while drawing on NVIDIA technologies across compute, networking, infrastructure and software — creating a single, flexible AI factory that no one company could build alone.

Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents

New on-silicon performance data measured by NVIDIA using real-world agentic coding trajectories shows Vera Rubin NVL72 systems deliver 30x higher throughput per megawatt and 35x lower token costs than NVIDIA GB300 NVL72.
by

According to OpenRouter data, agentic AI workloads consume 15x more tokens than a simple chat request. Why? 

Consider what happens when an AI agent researches a company for an investment decision. The agent queries financial databases, searches news and filings, invokes a sub-agent to run peer comparisons and model valuations, then synthesizes everything into a recommendation. Agents and sub-agents keep reasoning until the task is done, driving increased token demand. With every step, the accumulated tokens become the input to the next, making long-context handling central to agentic AI performance.

The same pattern plays out across every agentic use case, from software development to customer service to deep research. 

As agentic AI moves into production across industries, the infrastructure running it needs to meet that token demand efficiently. 

New measured performance data shows NVIDIA Vera Rubin NVL72 systems deliver up to 30x higher throughput per megawatt than NVIDIA GB300 NVL72 on agentic workloads. NVIDIA measured this inference throughput data using the SemiAnalysis AgentX workload, consisting of recorded real-world agentic coding sessions, with actual context growth, tool calls and sub-agent spawning preserved. For power-constrained AI factories, that translates directly into 30x more agentic work for the same energy footprint.

These early results for Vera Rubin NVL72 demonstrate NVIDIA’s accelerated pace of innovation. With continuous software optimizations, performance across both Vera Rubin NVL72 and GB300 NVL72 will continue to improve.

Vera Rubin NVL72: 30x Higher Throughput per Megawatt and 35x Lower Token Cost

Agentic workloads look fundamentally different from chat or document summarization, where input and output sequences typically range from 1K to 8K tokens. In agentic sessions, context accumulates across steps and can reach hundreds of thousands of input tokens, with wide variability in both input and output lengths across requests. Performance measurement must evolve to capture the full agent workflow rather than a single inference request.

The results below reflect performance measured on real-world agentic coding trajectories.

In SemiAnalysis AgentX, the NVIDIA Blackwell platform delivers leading performance across multiple agentic models including Kimi K3, MiniMax M3, GLM5.3, Qwen3.5 and DeepSeek V4 Pro. 

For example, GB300 NVL72 delivers up to 15x better throughput per megawatt than the NVIDIA Hopper architecture on the DeepSeek V4 Pro model, giving customers a high-performance foundation to run agentic workloads. This leap reflects the advantage of a larger scale-up GPU domain and codesigned software in delivering significantly better inference efficiency.

Vera Rubin extends that advantage, lifting the performance across the entire Pareto curve, to deliver as much as 30x higher throughput per megawatt than GB300 NVL72 on the DeepSeek V4 Pro model. These early results, measured using the SemiAnalysis AgentX workload and currently pending SemiAnalysis review, don’t yet reflect Vera CPU performance for tool calling. 

NVIDIA DSX MaxLPS technologies manages power across the GPU, rack and workload levels to provision up to 40% more GPUs within the same megawatt budget, pushing throughput per megawatt further at AI factory scale. 

Throughput per megawatt also directly impacts the cost of every token produced. At up to 35x lower cost per million tokens than GB300 NVL72, Vera Rubin NVL72 can run agents continuously, at scale, across the full breadth of customers’ workloads. 

For power-constrained AI factories, throughput per megawatt determines AI factory revenue and cost per million tokens determines the profit margin on that revenue.

Extreme Codesign for Agentic Scale

Modern inference optimization spans a range of techniques that are especially critical for agentic AI. NVIDIA Vera Rubin NVL72 enables all of these and more through extreme codesign across every layer of the platform to deliver multifold performance gains.

  • Disaggregated serving separates context processing (prefill) from response generation (decode) so each scales independently.
  • Rate matching synchronizes the speeds at which prefill GPUs and decode GPUs produce tokens to maximize efficiency.
  • Large-scale expert parallelism distributes expert sub-networks in mixture-of-experts models across the scale-up GPU domain. 
  • Distributed KV-caching extends memory across the scale-up GPU domain, while KV-cache offloading tiers less-active context to host and storage, keeping previously processed context accessible without recomputation.
  • KV-aware routing directs incoming requests to the GPUs that already hold the relevant cached context, reducing redundant computation across long sessions. 
  • Fused CUDA kernels like MegaMoE combine many computation and inter-GPU communication operations into a single execution pass, keeping GPUs active rather than waiting for data.

NVIDIA Rubin GPUs’ enhanced fifth-generation Tensor Cores and the third-generation Transformer Engine accelerate both prefill and decode stages of inference. NVFP4 quantization compresses model weights to 4-bit precision, reducing memory footprint and increasing throughput without sacrificing output quality. 

The NVL72 scale-up domain, a defining architecture across Vera Rubin and Grace Blackwell, enables the high-bandwidth and low-latency inter-GPU communication essential for techniques such as large-scale expert parallelism and distributed KV-caching. Purpose-built to power this scale-up domain, NVIDIA NVLink interconnect technology and NVLink Switches, now in their sixth generation, deliver 10x higher packet rates and 3x lower latency than off-the-shelf Ethernet alternatives.

Spanning optimized CUDA kernels, inference runtimes like NVIDIA TensorRT LLM and serving frameworks like NVIDIA Dynamo, NVIDIA’s software stack is codesigned with the hardware to enable inference optimizations.

While the results above reflect current Vera Rubin NVL72 performance, the full platform is a seven-chip architecture that also includes the NVIDIA Vera CPU, Groq 3 LPU, NVLink 6 Switch, BlueField-4 DPU, Spectrum-6 SPX and ConnectX-9 SuperNIC, all purpose-built for AI factories deploying agents at scale.

Extreme codesign also extends to NVIDIA’s co-engineering with its partners. Vera Rubin is in full production and is scaling across the ecosystem. 

Learn more about the NVIDIA Vera Rubin platform.