country_code

NVIDIA and Zoox Pave the Way for Autonomous Ride-Hailing

‘The world has never seen a robotics company like this before,’ NVIDIA founder and CEO Jensen Huang said in a fireside chat with Zoox CEO Aicha Evans and Zoox cofounder and CTO Jesse Levinson.
by

In celebration of Zoox’s 10th anniversary, NVIDIA founder and CEO Jensen Huang recently joined the robotaxi company’s CEO, Aicha Evans, and its cofounder and CTO, Jesse Levinson, to discuss the latest in autonomous vehicle (AV) innovation and experience a ride in the Zoox robotaxi.

In a fireside chat at Zoox’s headquarters in Foster City, Calif., the trio reflected on the two companies’ decade of collaboration. Evans and Levinson highlighted how Zoox pioneered the concept of a robotaxi purpose-built for ride-hailing and created groundbreaking innovations along the way, using NVIDIA technology.

“The world has never seen a robotics company like this before,” said Huang. “Zoox started out solely as a sustainable robotics company that delivers robots into the world as a fleet.”

Since 2014, Zoox has been on a mission to create fully autonomous, bidirectional vehicles purpose-built for ride-hailing services. This sets it apart in an industry largely focused on retrofitting existing cars with self-driving technology.

A decade later, the company is operating its robotaxi, powered by NVIDIA GPUs, on public roads.

Computing at the Core

Zoox robotaxis are, at their core, supercomputers on wheels. They’re built on multiple NVIDIA GPUs dedicated to processing the enormous amounts of data generated in real time by their sensors.

The sensor array includes cameras, lidar, radar, long-wave infrared sensors and microphones. The onboard computing system rapidly processes the raw sensor data collected and fuses it to provide a coherent understanding of the vehicle’s surroundings.

The processed data then flows through a perception engine and prediction module to planning and control systems, enabling the vehicle to navigate complex urban environments safely.

NVIDIA GPUs deliver the immense computing power required for the Zoox robotaxis’ autonomous capabilities and continuous learning from new experiences.

Using Simulation as a Virtual Proving Ground

Key to Zoox’s AV development process is its extensive use of simulation. The company uses NVIDIA GPUs and software tools to run a wide array of simulations, testing its autonomous systems in virtual environments before real-world deployment.

These simulations range from synthetic scenarios to replays of real-world scenarios created using data collected from test vehicles. Zoox uses retrofitted Toyota Highlanders equipped with the same sensor and compute packages as its robotaxis to gather driving data and validate its autonomous technology.

This data is then fed back into simulation environments, where it can be used to create countless variations and replays of scenarios and agent interactions.

Zoox also uses what it calls “adversarial simulations,” carefully crafted scenarios designed to test the limits of the autonomous systems and uncover potential edge cases.

The company’s comprehensive approach to simulation allows it to rapidly iterate and improve its autonomous driving software, bolstering AV safety and performance.

“We’ve been using NVIDIA hardware since the very start,” said Levinson. “It’s a huge part of our simulator, and we rely on NVIDIA GPUs in the vehicle to process everything around us in real time.”

A Neat Way to Seat

Zoox’s robotaxi, with its unique bidirectional design and carriage-style seating, is optimized for autonomous operation and passenger comfort, eliminating traditional concepts of a car’s “front” and “back” and providing equal comfort and safety for all occupants.

“I came to visit you when you were zero years old, and the vision was compelling,” Huang said, reflecting on Zoox’s evolution over the years. “The challenge was incredible. The technology, the talent — it is all world-class.”

Using NVIDIA GPUs and tools, Zoox is poised to redefine urban mobility, pioneering a future of safe, efficient and sustainable autonomous transportation for all.

From Testing Miles to Market Projections

As the AV industry gains momentum, recent projections highlight the potential for explosive growth in the robotaxi market. Guidehouse Insights forecasts over 5 million robotaxi deployments by 2030, with numbers expected to surge to almost 34 million by 2035.

The regulatory landscape reflects this progress, with 38 companies currently holding valid permits to test AVs with safety drivers in California. Zoox is currently one of only six companies permitted to test AVs without safety drivers in the state.

As the industry advances, Zoox has created a next-generation robotaxi by combining cutting-edge onboard computing with extensive simulation and development.

In the image at top, NVIDIA founder and CEO Jensen Huang stands with Zoox CEO Aicha Evans and Zoox cofounder and CTO Jesse Levinson in front of a Zoox robotaxi.

With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents

by

The next era of AI inference won’t be defined by a single breakthrough chip, network or system. It’ll be defined by how every layer of the AI factory works together. That’s why NVIDIA is extending Vera Rubin NVL72 with fast token generation for agentic systems.

Announced today, the NVIDIA Vera Rubin rack-scale system NVIDIA Groq 3 LPX is in full production. In an Artificial Analysis benchmark running Gemma 4 31B, an open source agentic model, it delivered 3,400 output tokens per second for 100,000-token long-context use cases critical to agentic systems, 4x faster than the nearest alternative platform. 

Industry partners worldwide are adopting Vera Rubin platform solutions. SpaceXAI announced that NVIDIA Vera CPUs will power its next generation of agentic AI. CoreWeave has deployed into production Spectrum-X Multiplane, which connects NVIDIA Vera Rubin racks using multiple parallel switches to provide high-bandwidth, flat and lossless AI networks. Nebius is the first AI cloud to adopt NVIDIA Groq 3 LPX.

As AI shifts from training to reasoning and agentic, inference has become the new frontier. Agentic AI systems are generating more tokens, processing dramatically larger context windows and increasingly collaborating with other AI systems to solve complex problems. 

These workloads demand a new class of infrastructure optimized not just for performance but for throughput, responsiveness and economics at unprecedented scale. 

At the Hot Chips conference this week in Palo Alto, California, NVIDIA is showcasing how extreme codesign is reshaping the AI factory from end to end. By architecting compute, networking and inference acceleration as a unified system, NVIDIA is helping customers build infrastructure purpose-built for the emerging demands of long-context inference and multi-agent systems.

Extreme Codesign Optimizes for Performance

Extreme codesign is the guiding principle behind NVIDIA platforms. Vera Rubin is engineered to accelerate inference as agents reason over increasingly long sequences. 

NVIDIA Spectrum-X Ethernet moves those massive data flows efficiently across AI factories, and NVIDIA Groq 3 LPX is built to generate tokens at ultrafast speeds. Together, they show how NVIDIA is optimizing every stage of the AI pipeline, from context and communication to generation, as part of a single, integrated AI factory architecture.

NVIDIA Groq 3 LPX brings a new low-latency inference architecture designed to work alongside Vera Rubin NLV72, the most versatile AI factory platform, helping enterprises and cloud providers deliver the low latency, extreme throughput and scalable economics required for agentic applications.

Breakthrough performance comes not from optimizing individual components in isolation, but from codesigning every layer of the stack. From networking and context processing to large-scale inference, NVIDIA’s full-stack platform turns AI factories into integrated engines for intelligence, built to turn ever-growing volumes of tokens into revenue.


Tuesday, Aug. 24, 8:00 a.m. PT 🔗

NVIDIA Partners Adopt Vera Rubin for Lowest Token Costs

Nebius, a leading AI cloud, is first to adopt NVIDIA Groq 3 LPX, giving developers access to leading token generation speeds for highly responsive agentic AI applications. 

Adding NVIDIA Groq 3 LPX to NVIDIA Vera Rubin NVL72 in  Nebius Token Factory will boost inference performance so developers can build highly interactive agents, coding systems and other real-time AI experiences at scale.

Connecting NVIDIA Vera Rubin racks, CoreWeave is deploying Spectrum-X Multiplane in production, unlocking advances for its AI cloud infrastructure.


Tuesday, Aug. 24, 8:00 a.m. PT 🔗

SpaceXAI Adopts NVIDIA Vera CPUs for Agentic AI

SpaceXAI plans to build and scale its future AI architecture around NVIDIA Vera Rubin, from data centers on Earth to orbital satellites. The company plans to deploy NVIDIA Vera CPUs to accelerate the CPU-intensive work behind agentic AI, including orchestration, tool use, code execution, data processing and simulation. 

The SpaceXAI partnership extends NVIDIA’s full-stack AI platform to SpaceXAI, bringing together Vera CPUs, NVIDIA accelerated computing, networking and software to advance AI at unprecedented scale.

Designed for the agentic era, Vera Rubin provides leading per-core performance, exceptional memory bandwidth and predictable performance under load, helping agents complete tasks faster and keeping valuable GPU infrastructure fully utilized.


Tuesday, Aug. 24, 8:00 a.m. PT 🔗

NVIDIA Groq 3 LPX: The Interactive AI Inference Accelerator

Codesigned with the Vera Rubin NVL72 platform, NVIDIA Groq 3 LPX is helping AI factories deliver tokens at the lowest latency for agentic workloads.

Agentic AI is creating a new performance challenge: decode latency. As AI agents reason, use tools and interact with other systems, they generate responses one token at a time, causing even tiny delays to multiply across complex chains of work. To keep agents operating at the pace users expect, NVIDIA Groq 3 LPX extends the Vera Rubin NVL72 platform with specialized acceleration for token generation. 

NVIDIA Rubin GPUs handle large-scale context processing while LPX accelerates latency-sensitive decode workloads. The result is faster, more predictable token generation that helps AI factories deliver responsive reasoning, smoother agent interactions and greater infrastructure efficiency. 

Together, Rubin GPUs and LPUs are designed to eliminate the traditional tradeoff between speed and throughput, helping AI providers deliver responsive, large-scale inference for the next generation of agentic AI applications.

Building the Token Factory

As the industry shifts from model training to serving intelligence at scale, infrastructure must evolve into what NVIDIA describes as a “token factory” capable of delivering performance, throughput, intelligence integrity and economic efficiency simultaneously. Agentic AI systems increasingly communicate with other AI systems, access multiple data sources and maintain large amounts of context, creating unprecedented demand for fast inference.

NVIDIA Groq 3 LPX was designed for exactly these workloads. As an extension of the Vera Rubin NVL72, it enables ultrafast responsiveness even across massive context windows while helping service providers maximize throughput and infrastructure utilization. 

Extreme Codesign for Inference

Unlike standalone accelerators, NVIDIA Groq 3 LPX combines the strengths of GPUs and LPUs through extreme codesign. Rubin GPUs and LPUs jointly compute every layer of an AI model, enabling new levels of inference performance for agentic workloads. 

At scale, fleets of LPUs operate as a giant processor optimized for deterministic inference. A rack-scale NVIDIA Groq 3 LPX deployment can include 256 LP30 accelerators connected through direct chip-to-chip links, creating a highly efficient inference engine built for modern AI factories.

Designed for the Agentic AI Era

As reasoning models grow and agentic workflows generate ever more tokens, the infrastructure required to serve them must evolve. NVIDIA Groq 3 LPX extends the Vera Rubin NVL72 platform with a purpose-built inference architecture designed to maximize responsiveness, throughput and efficiency, helping power the next generation of AI factories.

And this is only the beginning, more optimizations, more models, more performance when paired with Vera Rubin NVL72 — new levels of throughput and interactivity are coming. Stay tuned. 


Tuesday, Aug. 24, 8:00 a.m. PT 🔗

NVIDIA Spectrum-X Multiplane Enables Massive AI Factory Scale on a Flatter, More Resilient Network​

As AI factories grow massive, the network has become a critical engine of performance. At Hot Chips, NVIDIA is spotlighting Spectrum-X Multiplane — the latest in the hardware-accelerated Spectrum-X Ethernet architecture that lets Ethernet scale to unprecedented size while avoiding the latency, jitter and cost of adding another network tier.

NVIDIA Spectrum-X Ethernet is designed as an end-to-end, AI-optimized Ethernet platform, combining NVIDIA Spectrum-X Ethernet switches, SuperNICs and software to improve the performance and efficiency of Ethernet-based AI infrastructure for AI factories and clouds. The platform is designed to deliver 1.6x better AI networking performance compared with off-the-shelf Ethernet, while providing consistent, predictable performance in multi-tenant environments.

Multiplane Unlocks Scale Without the Tradeoffs of a New Tier

Scaling an AI factory beyond today’s largest clusters traditionally means adding a third network tier, which adds latency, slows things down unpredictably and drives up the cost of cabling, optics and power. Spectrum-X Multiplane takes a simpler approach: It splits each server’s network connection into several independent paths, or “planes,” each running its own lightweight two-tier network. The result is a flat, simple network that scales to 512,000 GPUs, without the added cost and complexity of a third tier.

This all happens automatically. A dedicated hardware engine inside the NVIDIA ConnectX SuperNIC manages traffic across the planes and instantly reroutes around any failure, so applications and software simply see one fast, reliable connection. In an eight-plane topology, if one plane fails, the network still maintains about 90% of its total bandwidth, with hardware recovery that’s 11x faster than software-based multiplane load balancing. This translates to 1.6x higher AI factory output.

Built Through Extreme Codesign

That reliability comes from extreme codesign of Vera Rubin NVL72, spanning switch silicon, SuperNICs and software. Spectrum-X SN6000 series switches, based on the 102.4Tb/s Spectrum-6 Ethernet ASIC and ConnectX-9 SuperNICs, supporting up to 1,600Gb/s per GPU, are purpose-built for Vera Rubin NVL72 AI factories. Spectrum-XGS Ethernet extends that same codesign across data centers, letting multiple facilities function as a single AI super-factory and accelerating multi-site NCCL collectives by 1.9x.


Tuesday, Aug. 24, 8:00 a.m. PT 🔗

NVIDIA Introduces Scale-In Infrastructure for Agentic AI Factories, Powered by BlueField-4, DOCA

NVIDIA is introducing NVIDIA Scale-In, the fifth pillar of NVIDIA AI networking and a new class of accelerated network infrastructure for agentic AI factories. Scale-In extends purpose-built acceleration to the infrastructure services that secure, manage and operate the AI factory.

Powered by the NVIDIA BlueField-4 processor and NVIDIA DOCA software platform and connected over NVIDIA Spectrum-X Ethernet, NVIDIA Scale-In transforms the traditional north-south access network into a unified, accelerated infrastructure domain.

Cloud computing brought software-defined networking, composability and elasticity to the data center, enabling users, applications, data and services to scale dynamically. 

Agentic AI represents the next platform shift. AI factories bring together massive accelerated compute with growing numbers of users, applications and autonomous agents, all continuously interacting with data, storage and services. This transforms the demands on the infrastructure that brings AI to life. Networking, storage, cybersecurity and operations must now be accelerated alongside AI compute, combining software-defined flexibility with purpose-built hardware acceleration and full-stack codesign. 

NVIDIA Scale-In delivers multi-tenant networking, high-performance storage access, in-silicon security, elastic provisioning and real-time observability, while keeping infrastructure processing independent of host compute resources. By accelerating and codesigning these services as part of the AI factory, Scale-In helps security, data access and operations scale alongside AI compute. The result is secure, efficient and manageable shared infrastructure for deploying and operating agentic AI at massive scale.


Tuesday, Aug. 24, 8:00 a.m. PT 🔗

NVIDIA NVLink Fusion Connects XPUs to NVIDIA’s Leading AI Platform

NVIDIA NVLink Fusion brings custom silicon into NVIDIA’s world-leading AI infrastructure platform, enabling hyperscalers and AI-native companies to build semi-custom AI factories with greater performance, flexibility and speed.

As AI models grow in size and complexity, raw compute alone is not enough. AI factories require high-bandwidth, low-latency scale-up networking, proven rack-scale architectures and a full ecosystem spanning power, cooling, management software and supply chain. NVLink Fusion addresses these challenges by connecting custom XPUs and CPUs to NVIDIA’s scale-up and scale-out technology stack.

The platform includes sixth-generation NVIDIA NVLink and NVLink Switch purpose-built scale-up networking, as well as NVLink-C2C for energy-efficient connectivity between XPUs and CPUs. Through the NVIDIA MGX ecosystem, adopters can also use production-proven rack designs, components, manufacturing partner solutions and open, extensible software for distributed computing, disaggregated workloads and cluster management.

By standardizing GPU- and XPU-based systems on a unified architecture, NVLink Fusion helps decouple data center buildout from silicon readiness. Operators can share rack footprints, networking, cooling, power delivery and management systems, then adjust the mix of GPUs and XPUs as supply and workload requirements evolve.

NVLink Fusion extends the NVIDIA AI platform’s vertically integrated, horizontally open approach to custom silicon. It gives partners the freedom to innovate where they differentiate while drawing on NVIDIA technologies across compute, networking, infrastructure and software — creating a single, flexible AI factory that no one company could build alone.

Securing the Infrastructure of Intelligence

Land, power and shell: The next critical resource for AI factories.
by

AI factories are the defining infrastructure of the AI era — where compute transforms energy and data into intelligence that powers every business, industry and country.

In the AI economy, compute is revenue.

AI factories require a full stack of critical resources: advanced chips, packaging, memory and networking — as well as land, power and shell.

Just as NVIDIA has used its scale, long-term visibility and supply-chain partnerships to secure critical semiconductor resources, we are now applying that same discipline to secure LPS capacity exclusively for NVIDIA AI factories.

Today, we are partnering with SB Energy to secure LPS capacity at the exceptional PORTS-Pike Technology Campus in Portsmouth, Ohio, to host NVIDIA compute. OpenAI will be the tenant.

LPS: The Next Strategic Resource

For the vast majority of NVIDIA customers, securing LPS has long been a part of their infrastructure strategy.

The world’s largest cloud service providers and investment-grade enterprises have balance sheets, infrastructure expertise and long-term contracts to secure LPS independently. They build and operate AI factories using NVIDIA accelerated computing, networking, systems and software.

This model will continue to represent most of NVIDIA’s business.

But frontier AI labs are different.

Frontier AI labs have extraordinary demand for training and inference compute, but many are growing faster than their balance sheets and long-term credit profiles can support. They may have strong customer demand and rapidly growing revenue yet still lack the decades-long infrastructure contracts and investment-grade financing capacity needed to secure the AI factory infrastructure independently.

Their growth is increasingly constrained not by algorithms or customer demand, but by the availability of compute.

For these companies, more compute means more intelligence, more products, more users and more revenue. NVIDIA is helping provide the infrastructure that powers this flywheel.

PORTS-Pike: A Site for Generations of NVIDIA Compute

OpenAI will build and operate a world-class AI factory at PORTS-Pike. The AI factory will use NVIDIA’s full-stack DSX AI factory platform, including GPUs, CPUs, networking and infrastructure software.  

The initial deployment is expected to provide 4.25 gigawatts of AI factory capacity. Each generation of NVIDIA AI factory systems deployed at PORTS-Pike could represent approximately 1.5 million NVIDIA GPUs, or approximately $150 billion to $200 billion in NVIDIA revenue. Over 20 years, the site can support multiple upgrade cycles. 

This is the essential economic point: the LPS commitment secures a long-lived AI factory site, while the NVIDIA compute inside can be upgraded repeatedly. Each new generation can deliver greater production, more intelligence and better economics. 

NVIDIA may also choose to extend the arrangement at PORTS-Pike beyond the initial 4.25 gigawatts to secure the remaining capacity of 3.75 gigawatts. 

OpenAI and NVIDIA Expanding Compute Opportunity

More broadly, OpenAI has committed to substantial deployments of NVIDIA AI infrastructure through 2030. OpenAI’s existing and planned commitments represent approximately 12 gigawatts of NVIDIA compute, with an opportunity to expand to approximately 16 gigawatts if NVIDIA extends the PORTS-Pike arrangement beyond the initial 4.25 gigawatts. 

At these levels, the opportunity represents roughly $600 billion of NVIDIA compute through 2030. 

The Important Questions

What is NVIDIA guaranteeing, and for how long?

NVIDIA is supporting the LPS infrastructure at PORTS-Pike for approximately 4 gigawatts over a 20-year term, securing a site on which NVIDIA compute will be exclusively deployed.  

Our support is limited to defined portions of lease and power payments, along with a specified residual-value commitment — not the full cost of the site or all of the tenant’s obligations.   

The guarantee will become effective in phases as data centers are placed in service between 2028 and 2030. As OpenAI makes lease payments and capacity comes online, NVIDIA’s remaining exposure declines.

Why is NVIDIA guaranteeing PORTS-Pike?

LPS has become a critical constraint on AI factory deployment. NVIDIA is selectively securing exceptional sites where we can host multiple generations of NVIDIA compute and serve durable customer demand. 

The productive life of the site extends through multiple generations of NVIDIA systems, each capable of producing more intelligence and more revenue than the generation before.

Is this circular financing?

No. OpenAI will pay the lease.

NVIDIA uses its scale and long-term visibility to secure PORTS-Pike to host NVIDIA compute. This is the same discipline we apply to supply-chain management: we secure critical inputs when we have visibility into customer demand and when doing so enables long-term productive capacity.

What happens to PORTS-Pike if OpenAI does not use the site in the future?

NVIDIA compute is versatile, fungible and broadly adopted. The capacity can be resold to another qualified tenant across NVIDIA’s global ecosystem of cloud service providers, enterprises, AI labs and startups. 

CUDA makes NVIDIA compute more than hardware. It gives developers and NVIDIA engineers a common platform to continually improve installed systems. 

CUDA makes NVIDIA compute versatile. Versatility makes it fungible. Fungibility drives utilization and durability — making NVIDIA compute a productive asset: rentable and financeable. 

The value of an exceptional site, like PORTS-Pike, is not limited to one customer or one generation of compute. NVIDIA’s standardized platform, broad developer ecosystem and large market of potential users support the ability to redeploy productive capacity over time.

How much LPS will NVIDIA secure?

It will be strategic and disciplined. 

Most NVIDIA customers will continue to secure their own LPS.  The vast majority of LPS hosting NVIDIA compute will continue to be secured directly by CSPs, enterprises, sovereign AI builders and other customers. 

NVIDIA will focus selectively on exceptional sites where visible, durable demand can support multiple generations of NVIDIA compute.

The Infrastructure of Intelligence

PORTS-Pike represents the next step in NVIDIA’s journey. 

We began by building accelerated computing chips. We then expanded to systems, networking, CUDA and full-stack AI factories. Today, we are helping secure the critical infrastructure required to build these factories. 

NVIDIA is the full-stack AI infrastructure platform. 

We are investing in the long-lived foundations of AI factories so our customers can deploy the most productive compute platform in the world, generation after generation. 

By securing the critical resources needed to host NVIDIA compute, we can help the world’s most innovative companies build the AI factories that will power the age of intelligence.