country_code

2023 Predictions: AI That Bends Reality, Unwinds the Golden Screw and Self-Replicates

15 NVIDIA AI experts predict digital twins and generative AI are set to advance enterprise goals and consumer needs even as the world enters a third year of planning uncertainty.
by

After three years of uncertainty caused by the pandemic and its post-lockdown hangover, enterprises in 2023 — even with recession looming and uncertainty abounding — face the same imperatives as before: lead, innovate and problem solve.

AI is becoming the common thread in accomplishing these goals. On average, 54% of enterprise AI projects made it from pilot to production, according to a recent Gartner survey of nearly 700 enterprises in the U.S., U.K. and Germany. A whopping 80% of executives in the survey said automation can be applied to any business decision, and that they’re shifting away from tactical to strategic uses of AI.

The mantra for 2023? Do more with less. Some of NVIDIA’s experts in AI predict businesses will prioritize scaling their AI projects amid layoffs and skilled worker shortages by using cloud-based integrated software and hardware offerings that can be purchased and customized to any enterprise, application or budget.

Cost-effective AI development also is a recurring theme among our expert predictions for 2023. With Moore’s law running up against the laws of physics, installing on-premises compute power is getting more expensive and less energy efficient. And the Golden Screw search for critical components is speeding the shift to the cloud for developing AI applications as well as for finding data-driven solutions to supply chain issues.

Here’s what our experts have to say about the year ahead in AI:

Anima Anandkumar headshot

ANIMA ANANDKUMAR
Director of ML Research, and Bren Professor at Caltech

Digital Twins Get Physical: We will see large-scale digital twins of physical processes that are complex and multi-scale, such as weather and climate models, seismic phenomena and material properties. This will accelerate current scientific simulations as much as a million-x, and enable new scientific insights and discoveries.

Generalist AI Agents: AI agents will solve open-ended tasks with natural language instructions and large-scale reinforcement learning, while harnessing foundation models — those large AI models trained on a vast quantity of unlabeled data at scale — to enable agents that can parse any type of request and adapt to new types of questions over time.

Manuvir Das headshot

MANUVIR DAS
Vice President, Enterprise Computing

Software Advances End AI Silos: Enterprises have long had to choose between cloud computing and hybrid architectures for AI research and development — a practice that can stifle developer productivity and slow innovation. In 2023, software will enable businesses to unify AI pipelines across all infrastructure types and deliver a single, connected experience for AI practitioners. This will allow enterprises to balance costs against strategic objectives, regardless of project size or complexity, and provide access to virtually unlimited capacity for flexible development.

Generative AI Transforms Enterprise Applications: The hype about generative AI becomes reality in 2023. That’s because the foundations for true generative AI are finally in place, with software that can transform large language models and recommender systems into production applications that go beyond images to intelligently answer questions, create content and even spark discoveries. This new creative era will fuel massive advances in personalized customer service, drive new business models and pave the way for breakthroughs in healthcare.

Kimberly Powell headshot

KIMBERLY POWELL
Vice President, Healthcare

Biology Becomes Information Science: Breakthroughs in large language models and the fortunate ability to describe biology in a sequence of characters are giving researchers the ability to train a new class of AI models for chemistry and biology. The capabilities of these new AI models give drug discovery teams the ability to generate, represent and predict the properties and interactions of molecules and proteins — all in silicon. This will accelerate our ability to explore the essentially infinite space of potential therapies.

Surgery 4.0 Is Here: Flight simulators serve to train pilots and research new aircraft control. The same is now true for surgeons and robotic surgery device makers. Digital twins that can simulate at every scale, from the operating room environment to the medical robot and patient anatomy, are breaking new ground in personalized surgical rehearsals and designing AI-driven human and machine interactions. Long residencies won’t be the only way to produce an experienced surgeon. Many will become expert operators when they perform their first robot-assisted surgery on a real patient.

DANNY SHAPIRO
Vice President, Automotive

Training Autonomous Vehicles in the Metaverse: The more than 250 auto and truck makers, startups, transportation and mobility-as-a-service providers developing autonomous vehicles are tackling one of the most complex AI challenges of our time. It’s simply not possible to encounter every scenario they must be able to handle by testing on the road, so much of the industry in 2023 will turn to the virtual world to help.

On-road data collection will be supplemented by virtual fleets that generate data for training and testing new features before deployment. High-fidelity simulation will run autonomous vehicles through a virtually infinite range of scenarios and environments. We’ll also see the continued deployment of digital twins for vehicle production to improve manufacturing efficiencies, streamline operations and improve worker ergonomics and safety.

Moving to the Cloud: 2023 will bring more software-as-a-service (SaaS) and infrastructure-as-a-service offerings to the transportation industry. Developers will be able to access a comprehensive suite of cloud services to design, deploy and experience metaverse applications anywhere. Teams will design and collaborate on 3D workflows — such as AV development simulation, in-vehicle experiences, cloud gaming and even car configurators delivered via the web or in showrooms.

Your In-Vehicle Concierge: Advances in conversational AI, natural language processing, gesture detection and avatar animation are making their way to next-generation vehicles in the form of digital assistants. This AI concierge can make reservations, access vehicle controls and provide alerts using natural language understanding. Using interior cameras, deep neural networks and multimodal interaction, vehicles will be able to ensure that driver attention is on the road and ensure no passenger or pet is left behind when the journey is complete.

Rev Lebaredian headshot

REV LEBAREDIAN
Vice President, Omniverse and Simulation Technology

The Metaverse Universal Translator: Just as HTML is the standard language of the 2D web, Universal Scene Description is set to become the most powerful, extensible, open language for the 3D web. As the 3D standard for describing virtual worlds in the metaverse, USD will allow enterprises and even consumers to move between different 3D worlds using various tools, viewers and browsers in the most seamless and consistent fashion.

Bending Reality With Digital Twins: A new class of true-to-reality digital twins of goods, services and locations is set to offer greater windfalls than their real-world counterparts. Imagine selling many virtual pairs of sneakers in partnership with a gaming company that are simply undergoing design testing — long before sending the pattern to manufacturing. Companies also stand to benefit by saving on waste, increasing operational efficiencies and boosting accuracy.

Ronnie Vasishta

RONNIE VASISHTA
Senior Vice President, Telecoms

Cutting the Cord on AR/VR Over 5G Networks: While many businesses will move to the cloud for hardware and software development, edge design and collaboration also will grow as 5G networks become more fully deployed around the world. Automotive designers, for instance, can don augmented reality headsets and stream the same content they see over wireless networks to colleagues around the world, speeding collaborative changes and developing innovative solutions at record speeds. 5G also will lead to accelerated deployments of connected robots across industries — used for restocking store shelves, cleaning floors, delivering pizzas and picking and packing goods in factories.

RAN in the Cloud: Network operators around the world are rolling out software-defined virtual radio access network 5G to save time and money as they seek faster returns on their multibillion-dollar investments. Now, they’re shifting away from bespoke L1 accelerators to 100% software-defined and full-stack, 5G-baseband acceleration that includes L2, RIC, Beamforming and FH offerings. This shift will lead to an increase in the utilization of RAN systems by enabling multi-tenancy between RAN and AI workloads.

BOB PETTE
Vice President, Professional Visualization 

An Industrial Revolution via Simulation: Everything built in the physical world will first be simulated in a virtual world that obeys the laws of physics. These digital twins — including large-scale environments such as factories, cities and even the entire planet — and the industrial metaverse are set to become critical components of digital transformation initiatives. Examples already abound: Siemens is taking industrial automation to a new level. BMW is simulating entire factory floors to optimally plan manufacturing processes. Lockheed Martin is simulating the behavior of forest fires to anticipate where and when to deploy resources. DNEG, SONY Pictures, WPP and others are boosting productivity through globally distributed art departments that enable creators, artists and designers to iterate on scenes virtually in real time.

Rethinking of Enterprise IT Architecture: Just as many businesses scrambled to adapt their culture and technologies to meet the challenges of hybrid work, the new year will bring a re-architecting of many companies’ entire IT infrastructure. Companies will seek powerful client devices capable of tackling the ever-increasing demands of applications and complex datasets. And they’ll embrace flexibility, moving to burst to the cloud for exponential scaling. The adoption of distributed computing software platforms will enable a globally dispersed workforce to collaborate and stay productive under the most disparate working environments.

Similarly, complex AI model development and training will require powerful compute infrastructure in the data center and the desktop. Businesses will look at curated AI software stacks for different industrial use cases to make it easy for them to bring AI into their workflows and deliver higher quality products and services to customers faster.

Azita Martin

AZITA MARTIN
Vice President, AI for Retail, Consumer Packaged Group and Quick-Service Restaurants

Tackling Shrinkage: Brick-and-mortar retailers perennially struggle with a commonplace problem: shrinkage, the industry parlance for theft. As more and more adopt AI-based services for contactless checkout, they’ll seek sophisticated software that combines computer vision with store analytics data to make sure what a shopper rings up is actually the item being purchased. The adoption of smart self-tracking technology will aid in the development of fully automated store experiences and help solve for labor shortages and lost income.

AI to Optimize Supply Chains: Even the most sophisticated retailers and e-commerce companies had trouble the past two years balancing supply with demand. Consumers embraced home shopping during the pandemic and then flocked back into brick-and-mortar stores after lockdowns were lifted. After inflation hit, they changed their buying habits once again, giving supply chain managers fits. AI will enable more frequent and more accurate forecasting, ensuring the right product is at the right store at the right time. Also, retailers will embrace route optimization software and simulation technology to provide a more holistic view of opportunities and pitfalls.

Malcolm DeMayo

MALCOLM DEMAYO
Vice President, Financial Services

Better Risk Management: Firms will look for opportunities like accelerated compute to drive efficiencies. The simulation techniques used to value risk in derivatives trading are computationally intensive and typically consume large swaths of data center space, power and cooling. What runs all night on traditional compute will run over a lunch break or faster on accelerated compute. A real-time value of sensitivities will enable firms to better manage risk and improve the value they deliver to their investors.

Cloud-First for Financial Services: Banks have a new imperative: get agile fast. Facing increasing competition from non-traditional financial institutions, changing customer expectations rising from their experiences in other industries and saddled with legacy infrastructure, banks and other institutions will embrace a cloud-first AI approach. But as a highly regulated industry that requires operational resiliency, an industry term that means your systems can absorb and survive shocks (like a pandemic), banks will look for open, portable, hardened, hybrid solutions. As a result, banks are obligated to purchase support agreements when available.

Charlie Boyle headshot

CHARLIE BOYLE
Vice President, DGX systems

AI Becomes Cost-Effective With Energy-Efficient Computing: In 2023, inefficient, x86-based legacy computing architectures that can’t support parallel processing will give way to accelerated computing solutions that deliver the computational performance, scale and efficiency needed to build language models, recommenders and more.

Amidst economic headwinds, enterprises will seek out AI solutions that can deliver on objectives, while streamlining IT costs and boosting efficiency. New platforms that use software to integrate workflows across infrastructure will deliver computing performance breakthroughs — with lower total cost of ownership, reduced carbon footprint and faster return on investment on transformative AI projects — displacing more wasteful, older architectures.

DAVID REBER
Chief Security Officer

Data Scientists Are Your New Cyber Asset: Traditional cyber professionals can no longer effectively defend against the most sophisticated threats because the speed and complexity of attacks and defense have effectively exceeded human capacities. Data scientists and other human analysts will use AI to look at all of the data objectively and discover threats. Breaches are going to happen, so data science techniques using AI and humans will help find the needle in the haystack and respond quickly.

AI Cybersecurity Gets Customized: Just like recommender systems serve every consumer on the planet, AI cybersecurity systems will accommodate every business. Tailored solutions will become the No. 1 need for enterprises’ security operations centers as identity-based attacks increase. Cybersecurity is everyone’s problem, so we’ll see more transparency and sharing of various types of cybersecurity architectures. Democratizing AI enables everyone to contribute to the solution. As a result, the collective defense of the ecosystem will move faster to counter threats.

Kari Briski headshot

KARI BRISKI
Vice President, AI and HPC Software

The Rise of LLM Applications: Research on large language models will lead to new types of practical applications that can transform languages, text and even images into useful insights that can be used across a multitude of diverse organizations by everyone from business executives to fine artists. We’ll also see rapid growth in demand for the ability to customize models so that LLM expertise spreads to languages and dialects far beyond English, as well as across business domains, from generating catalog descriptions to summarizing medical notes.

Unlabeled Data Finds Its Purpose: Large language models and structured data will also extend to the reams of photos, audio recordings, tweets and more to find hidden patterns and clues to support healthcare breakthroughs, advancements in science, better customer engagements and even major advances in self-driving transportation. In 2023, adding all this unstructured data to the mix will help develop neural networks that can, for instance, generate synthetic profiles to mimic the health records they’ve learned from. This type of unsupervised machine learning is set to become as important as supervised machine learning.

The New Call Center: Keep an eye on the call center in 2023, where adoption of more and more easily implemented speech AI workflows will provide business flexibility at every step of the customer interaction pipeline — from modifying model architectures to fine-tuning models on proprietary data and customizing pipelines. As the accessibility of speech AI workflows broadens, we’ll see a widening of enterprise adoption and giant increase in call center productivity by speeding time to resolution. AI will help agents pull the right information out of a massive knowledge base at the right time, minimizing wait times for customers.

Kevin Deierling

KEVIN DEIERLING
Senior Vice President, Networking

Moore’s Law on Life Support: As CPU design runs up against the laws of physics and struggles to keep up with Moore’s law — the postulation that roughly every two years the number of transistors on microchips would double and create faster, more efficient processing — enterprises increasingly will turn to accelerated computing. They’ll use custom combinations of CPUs, GPUs, DPUs and more in scalable data centers to innovate faster while becoming more cloud oriented and energy efficient.

The Network as the New Computing Platform: Just as personal computers combined software, hardware and storage into productivity-generating tools for everyone, the cloud is fast becoming the new computing tool for AI and the network is what enables the cloud. Enterprises will use third-party software, or bring their own, to develop AI applications and services that run both on-prem and in the cloud. They’ll use cloud services operators to purchase the capacity they need when they need it, working across CPUs, GPUs, DPUs and intelligent switches to optimize compute, storage and the network for their different workloads. What’s more, with zero-trust security being rapidly adopted by cloud service providers, the cloud will deliver computing as secure as on-prem solutions.

DEEPU TALLA
Vice President, Embedded and Edge Computing

Robots Get a Million Lives: More robots will be trained in virtual worlds as photorealistic rendering and accurate physics modeling combine with the ability to simulate in parallel millions of instances of a robot on GPUs in the cloud. Generative AI techniques will make it easier to create highly realistic 3D simulation scenarios and further accelerate the adoption of simulation and synthetic data for developing more capable robots.

Expanding the Horizon: Most robots operate in constrained environments where there is limited to no human activity. Advancements in edge computing and AI will enable robots to have multi-modal perception for better semantic understanding of their environment. This will drive increased adoption of robots operating in brownfield facilities and public spaces such as retail stores, hospitals and hotels.

Marc Spieler

MARC SPIELER
Senior Director, Energy

AI-Powered Energy Grid: As the grid becomes more complex due to the unprecedented rate of distributed energy resources being added, electric utility companies will require edge AI to improve operational efficiency, enhance functional safety, increase accuracy of load and demand forecasting, and accelerate the connection time of renewable energy, like solar and wind. AI at the edge will increase grid resiliency, while reducing energy waste and cost.

More Accurate Extreme Weather Forecasting: A combination of AI and physics can help better predict the world’s atmosphere using a technique called Fourier Neural Operator. The FourCastNet system is able to predict a precise path of a hurricane and can also make weather predictions in advance and provide real-time updates as climate conditions change. Using this information will allow energy companies to better plan for renewable energy expenditures, predict generation capacity and prepare for severe weather events.

NVIDIA Alpamayo 2 Super, the Frontier Open Model for Robotaxis and Autonomous Vehicles, Now Available for Commercial Use

Open commercial licensing, benchmark‑leading reasoning and inspectable decisions bring autonomous vehicles, including robotaxis, closer to production and widescale deployment.
by

For robotaxis and other autonomous vehicles (AVs), the hardest problems aren’t the everyday scenarios. They’re the rare, complex situations that are difficult to anticipate and train for.

Handling these long‑tail events takes more than just object detection and motion prediction. AVs must understand the situation, reason about cause and effect, choose the right action and turn that decision into a safe, comfortable path — all in real time and in a way developers can inspect, validate and trust.

NVIDIA Alpamayo 2 Super, available now for commercial use, is part of the Alpamayo family, the most-adopted open reasoning models for autonomous driving on Hugging Face, supporting a wide range of AV-relevant capabilities within a single foundation model. 

Built on NVIDIA Cosmos 3 Super Reasoner and post‑trained with reinforcement learning, the model advances the AV ecosystem on two fronts: open commercial licensing and leading multitask capabilities for autonomous driving. 

Alpamayo 2 Super is part of NVIDIA’s growing collection of open models, datasets and tools for autonomous driving, expanding access, strengthening competition, giving developers greater control and supporting safer, more transparent AV deployment. 

Open Licensing for Production AVs

Alpamayo 2 Super is available on Hugging Face under OpenMDW‑1.1, the Linux Foundation’s permissive license for open AI model distributions. The license covers fine‑tuning, derivative models and commercial redistribution, allowing AV developers, automakers, truckmakers and suppliers to adapt Alpamayo to their own data, driving policies and deployment strategies. 

This openness lets AV researchers and companies keep control of their own data and infrastructure, as well as own the value they create through specialized models and accumulated know‑how. Such control is essential for workflows involving proprietary fleets and safety. 

Earlier Alpamayo releases were initially introduced for R&D. The OpenMDW license is now being  applied across the entire Alpamayo model family so developers can deploy any of the models commercially without requiring additional permissions. This creates a direct path from adaptation to deployment.

Open weights make that path economically viable. Teams can build on advanced reasoning without re‑training every foundation capability from scratch or paying frontier‑model costs for every task, matching the right model to the right job at the right cost. 

Alpamayo 2 Super enables frontier-scale reasoning in cloud-based development workflows, where developers can generate high-quality reasoning traces, synthetic training data and teacher outputs for model distillation. Within the Alpamayo model family, Alpamayo 2 Super delivers the highest reasoning and driving performance for multimodal autonomous driving development, while Alpamayo 1.5 and Alpamayo 1 provide more cost-efficient options for cloud-based development and model distillation.

The resulting distilled models can then be optimized for efficient, real-time inference in production vehicles. Together, the Alpamayo model family provides a cloud-to-car workflow that combines frontier-scale reasoning with scalable deployment across commercial AV fleets.

For AV programs, that means frontier‑scale reasoning in the cloud and efficient, specialized models in the vehicle — a more sustainable way to scale safe autonomy into commercial fleets.

Benchmark-Leading Reasoning at Frontier Scale

Alpamayo 2 Super ranks first on LingoQA, an autonomous driving reasoning benchmark, among nearly 40 models evaluated. In NVIDIA testing using the Lingo‑Judge metric, it outperformed Qwen2.5‑VL 72B by 17.0 points, Gemini 2.5 Pro by 15.1 points and GPT‑4o by 23.2 points, demonstrating state‑of‑the‑art reasoning for driving‑centric scenarios. Alpamayo 2 Super also ranks first across all autonomous driving benchmarks evaluated by NVIDIA, underscoring its leading performance across a broad range of AV capabilities. 

Alpamayo 2 Super offers 3x the scale of the 10‑billion‑parameter NVIDIA Alpamayo 1.5 and Alpamayo 1 models. The added capacity helps the model better generalize reasoning from sparse examples — a critical capability for the rare, multi‑agent interactions where conventional systems often struggle. 

The model reasons over full‑surround camera coverage, fusing views from the vehicle’s front, sides and rear. This 360‑degree context enables richer understanding of lane changes, merges, unprotected turns and complex intersections, where risks commonly arise.

A Multitask Foundation Model for Robotaxis and Autonomous Driving

For each driving situation, Alpamayo 2 Super can produce five tightly coupled outputs:

  • A trajectory describing the vehicle’s planned path.

  • A chain‑of‑causation (CoC) trace that explains the reasoning behind the decision.

  • A meta‑action (e.g., yield, lane changes, stops) that captures the model’s intent.

  • Reasoning auto-labels that generate CoC annotations for training and validation data.

  • Visual question answering responses with 2D visual grounding that link the model’s answers to specific regions in camera images.

Together, these outputs offer insight into the model’s decision-making process. Developers can tie what the model observed to the action it selected, making decisions easier to understand, critique and validate. 

CoC traces integrate with NVIDIA Halos safety‑validation workflows and support AI safety aligned with ISO/PAS 8800 requirements, providing a stronger foundation for AV safety engineering. 

Alpamayo 2 Super can also be deployed as an autolabeler to generate CoC labels and perform visual question answering with 2D grounding on proprietary fleet data. By linking its reasoning to specific regions in camera images, the model can transform raw driving clips into richer training data, compressing annotation cycles from months to days.

Beyond planning and auto-labeling, Alpamayo 2 Super supports scene understanding, model critiquing and knowledge distillation. These multitask capabilities enable developers to use a single foundation model across more of the development stack, simplifying tooling and accelerating iteration.

An Open Ecosystem for Reasoning‑Based AVs

Alpamayo 2 Super is part of a broader family of open models, frameworks and datasets for AV development. 

Other tools in the family include: 

  • NVIDIA AlpaSim, which provides closed‑loop simulation.

  • NVIDIA AlpaGym, which enables high‑throughput reinforcement learning.

  • NVIDIA Physical AI Open Datasets, which supply data for training and testing.

  • Open training recipes and an autolabeling pipeline to accelerate model development, training and validation.

Alpamayo has already surpassed 500,000 downloads on Hugging Face, reinforcing its position as the most-adopted open reasoning model family for autonomous driving on the platform.

Download NVIDIA Alpamayo 2 Super on Hugging Face to explore the model, evaluate its reasoning capabilities and start building the next generation of robotaxis and autonomous vehicles.

For Robotaxis, Safety Must Be Built In, Not Bolted On

by

Editor’s note: The name of NVIDIA DRIVE Hyperion was changed to NVIDIA Hyperion in September 2026. All references to the name have been updated in this blog.

A car pulls up to the curb. The app says, “Your ride is here.” No one’s in the driver’s seat. For people who live in one of the dozens of cities now hosting robotaxi services, this is already a reality.

The robotaxi industry has moved from prototype milestones to commercial operations, with an expanding ecosystem accelerating the pace of deployment. New collaborations announced at NVIDIA GTC Taipei reflect robotaxi programs spinning up around the world:

  • Uber and Autobrains are launching a robotaxi program in Munich on the NVIDIA Hyperion platform, using Autobrains’ agentic AI to support scalable operations. 
  • Foxconn is expanding its collaboration with NVIDIA to deploy robotaxi fleets, combining its services with NVIDIA Hyperion for rapid integration and scaling in Taiwan.
  • VinFast is working with Autobrains to bring level 4 vehicles built on Hyperion to the Southeast Asia market.
  • HUMAIN is working to bring Hyperion-powered robotaxis to Saudi Arabia, expanding the platform’s global footprint into the Middle East.

Building a Safe Software Foundation

As the robotaxi industry scales, safety is paramount.

Regulators, certification bodies and developers are scrutinizing what safe deployment at scale requires. 

Industry discussion on level 4 autonomy often centers on what the vehicle can perceive and decide. 

That discussion is well-founded. Accurate perception, sound decision-making and handling the unexpected are difficult problems, and real progress toward solving them is being made.

But perception and decisions alone are not the whole story. Regulators require something more: proof that the overall system behaves reliably, isolates faults before they escalate and never operates outside the boundaries it was designed for. 

Robotaxi safety requires solving four distinct challenges simultaneously:

  • A safety-certifiable operating system
  • Safe, standardized hardware and software interfaces
  • AI that operates within verifiable guardrails
  • Validation at scale before vehicles touch public roads

To help solve these challenges, the recently introduced Halos Operating System (OS) — a component of the NVIDIA Halos full-stack, comprehensive safety system — offers a unified, production-ready safety foundation for AI-driven vehicles, built on NVIDIA Hyperion. It comprises: 

Halos Core: A Certified OS Foundation

At the foundation of NVIDIA Halos OS is Halos Core, which is the next generation of NVIDIA DriveOS and certified to automotive safety standards. It’s audited, documented and proven to behave predictably under fault conditions, with a hypervisor — a specialized software layer — that isolates safety-critical functions so failures can’t reach vehicle controls. 

Halos Core is compliant with ISO 26262 ASIL D, includes safety-certified support for NVIDIA CUDA and TensorRT, and provides the TensorRT Edge-LLM open source framework for high-performance large language model inference.

Halos SDK: Standardized and Safe Interfaces

A robotaxi integrates cameras, radar, lidar and other sensors, each streaming data in a different format at a different rate. Without a standardized middleware layer, every hardware change forces teams to manually rebuild those integrations. 

Halos SDK removes that burden. Its sensor abstraction layer decouples the autonomous driving stack from individual sensor drivers, so adding or swapping a sensor no longer causes ripples through application code, while a vehicle abstraction layer connects the autonomous driving stack to the rest of the vehicle through a single, consistent interface. 

On top, Halos SDK provides the runtime building blocks that safety-critical software demands: a deterministic application-level scheduler for predictable timing, zero-copy inter-process communication that moves data without added latency, a comprehensive system error-handling framework and a robust scenario data recorder — delivering the foundation for highly reliable and low-latency automotive applications.      

Halos Applications: Safety Guardrails for AI

AI models can match human driving behavior, but regulators require more than performance. 

The Halos Applications layer provides safety guardrails for AI through deterministic, rule-based functions, analyzed and designed to behave within defined bounds. It includes world model perception and the top-rated NVIDIA DRIVE active safety stack featuring automatic emergency braking, lane departure warning, blind spot monitoring, collision warning and more. 

In addition, in Halos Applications, Halos OS can be combined with end-to-end AI models for which explainability and transparency are essential. This includes the NVIDIA Alpamayo family of open models for autonomous vehicle development, which enables chain-of-thought reasoning, continuously evaluating the road, planning next steps and adapting to changing conditions.

The Halos Safety Evaluation Framework

Halos Infra is the cloud-side development infrastructure that enables autonomous vehicle training, simulation and validation at scale. It’s the foundation for the recently released NVIDIA Halos Safety Evaluation Framework (SEF).

SEF provides the tools and guidelines needed to build a credible safety case, from L2 driver assistance to L4 robotaxis. It draws on more than 330 research papers and 1,000 patents developed within NVIDIA Halos OS.

Halos Infra runs on NVIDIA’s three-computer autonomous driving solution: 

Halos OS spans the full development lifecycle — from training and simulation in Halos Infra to inference in the vehicle itself.

Learn more about NVIDIA Halos.

NVIDIA Research Unlocks Advanced Grasping, Smarter Autonomous Driving and Agent Training at Scale

New NVIDIA Research breakthroughs show how training at scale — across gripper types, driving scenarios and virtual worlds — creates AI that generalizes to diverse applications.
by

What makes a robot gripper useful isn’t that it can pick up one object — it’s that it can pick up the next one, and the one after that, with a tool it’s never held before. 

What makes an autonomous vehicle system safe isn’t just that it can reason through a situation — it’s that it can do so quickly enough on the hardware actually installed in the car. 

What makes a virtual agent capable is exposure to as many different environments as possible before it faces the real world. 

At this year’s Computer Vision and Pattern Recognition (CVPR) conference, NVIDIA Research is presenting three papers that address each of these challenges — and share a common theme: training at scale creates systems that generalize across diverse applications.

The three papers cover different challenges in physical AI research: 

  • GraspGen-X, the first foundation model for zero-shot grasping, was trained on billions of simulated grasps to work with any gripper it’s shown.
  • LCDrive introduces a model that replaces expensive text-based reasoning with compact latent representations, letting autonomous vehicles think faster on embedded hardware.
  • NitroGen is a generalized gameplay AI foundation model that harnesses the NVIDIA Isaac GR00T robot foundation model architecture to help train embodied agents in virtual environments across tens of thousands of hours of interaction.

NVIDIA also unveiled at CVPR new physical AI agent skills that help researchers and developers speed the development of autonomous vehicles, robots and vision AI systems.

NitroGen and another NVIDIA-authored paper, PixelDIT, were named best paper finalists at the conference — an accolade given to just 15 of over 4,000 accepted papers at CVPR.

The First Foundation Model for Grasping

Most AI systems for robotic grasping are specialists.

A vision-language-action policy trained for a two-finger gripper only learns to grasp with those two fingers. Similarly, a policy for dextrous grasping will only work for the bespoke multi-fingered gripper it’s trained on. For every new embodiment, the process typically needs to be repeated — requiring new training data, fine-tuning and validation. This constraint means most robotics companies pick a gripper, train for it and stick with it.

GraspGen-X is the first foundation model for grasping built to eliminate this bottleneck. 

Like a large language model that can apply its understanding of language to a new task without retraining, GraspGen-X applies its understanding of geometry and contact to any robotic gripper it encounters. Given the geometry of a new gripper and an unknown object it’s never seen before, the model generates reliable grasp pose proposals to enable the robot to grasp the object.

To get there, the researchers needed a dataset that’s impossible to collect in the real world at scale. They generated 2 billion simulated grasps across thousands of object shapes and synthetic gripper configurations, spanning the diversity of form factors a deployed robot might encounter. 

For robot developers, this foundation model eliminates the need for per-gripper training cycles and can be applied out of the box for several commonly used grippers. GraspGenX can be used in conjunction with curoboV2, a new CUDA-accelerated motion planning library, to achieve these grasp poses in unknown environments. 

Building on the GraspGen research foundation, another paper, Grasp-MPC — presented at ICRA 2026 — advances the next step in the pipeline: moving from grasp generation to closed-loop grasp execution.

Teaching Autonomous Vehicles to Think Faster

In recent years, researchers have found that letting an AI reason — generating intermediate thinking steps before committing to an answer — reliably improves its decision-making. 

For autonomous vehicles, the challenge is doing that reasoning on the hardware inside an actual vehicle. Text-based chain-of-thought reasoning generates words, and every word is a token that takes time to produce. On the processor running inside a car, token count is a real constraint on how fast the system can respond.

LCDrive tackles this problem by replacing words with compressed latent representations. 

Instead of generating human-readable reasoning steps, the system thinks in a compact latent space — states that capture spatial information rather than producing text. The architecture alternates between two kinds of thinking: proposing candidate actions, then predicting what the world will look like if those actions are taken. 

It uses that predicted world state to refine its next step. It’s the same reasoning loop — just in a more computationally efficient form than natural language.

The result: comparable output trajectory quality to text-based reasoning, using roughly half the tokens. 

The model was built on NVIDIA Alpamayo and trained using supervision derived from existing vehicle data.

Embodied Agents Trained in Virtual Worlds

Isaac GR00T — NVIDIA’s open foundation model for humanoid robots — is built on a simple principle: expose a model to enough diverse situations, and it will generalize to ones it hasn’t seen. 

NitroGen extends that principle to virtual environments, using the GR00T architecture to train a foundation model for embodied agents across a breadth of virtual worlds.

Video games offer something that’s hard to build from scratch: structured, varied worlds with defined goals and well-specified success conditions. They’re high-quality training environments, available at scale. 

NitroGen treats them that way — as a training ground for agents that will eventually be trained to handle novel real- or simulated-world situations, like powering a robot that helps with housework based on broad instructions such as, “Put these items away in the pantry.”  

Trained across more than 1,000 games and 40,000 hours of interaction using a model based on GR00T, the resulting agents learn to generalize across environments. The model was evaluated across a range of action role-playing games, platformers, roguelikes and open-world games, demonstrating gameplay behaviors spanning combat, navigation and exploration. 

The same techniques could eventually help enable more adaptive nonplayable characters, AI companions and gameplay systems inside games, as well as broader testing of complex game environments.

In low-data conditions — where an agent has seen only a handful of examples of a new environment — starting with NitroGen gives agents a huge head start, improving performance by up to 52% over previous state-of-the-art methods. 

The model is open source, available on GitHub and Hugging Face. 

Learn more about NVIDIA at CVPR and explore NVIDIA Research’s work in physical AI, computer vision and autonomous systems. Get started with Isaac GR00T and NVIDIA robotics tools.