country_code

Five Things You Always Wanted to Know About AI, But Weren’t Afraid to Ask

by
Looking for inferencing in the real world? Turn on your smartphone.

They say there’s no such thing as a dumb question. As someone who asks dumb questions for a living, I can tell you that’s a really stupid thing to say.

But the best questions are often the ones where someone smart explains something from the ground up to a total novice (read: me). The beauty of NVIDIA: there are a lot of smart people upon whom I can inflict my very dumbest questions.

Turns out I’m not alone. Month in and month out, tens of thousands of readers ask search engines these very questions. And they get connected to the answers through our blog.

What are they? Smart question. Here are five of our most popular in 2019.

What’s the Difference Between a CPU and a GPU?

This post is over a decade old, but the answer — thanks to the emergence of deep-learning driven AI, supercomputing, and self-driving cars — is more relevant than ever. That’s why we’ve updated our original post earlier this year, and why more readers are seeing this post than ever.

What’s the Difference Between Artificial Intelligence, Machine Learning and Deep Learning?

Visualize the fields of AI, machine learning and deep learning as concentric circles. AI — the idea that came first — is the largest circle. Then comes machine learning, which blossomed later. And finally deep learning — which is driving today’s AI explosion — fitting inside both. Click on the link, above, for more.

What’s the Difference Between Supervised, Unsupervised, Semi-Supervised and Reinforcement Learning?

This is one of the key questions in AI right now, which is why this post has become one of our most popular. Click on the link for a plain English answer to these questions, and a walk through the kinds of datasets and problems that lend themselves to each kind of learning.

What’s the Difference Between Deep Learning Training and Inference?

This is another question that’s drawn more readers over time. Training, in short, is the process of running data through a neural network to teach it a task. That’s taught computers to do things that, just a decade ago, most believed could only be done by humans. Inference, by contrast, is the process of putting that trained network to work, in everything from hyperscale data centers to autonomous machines.

What’s the Difference Between Ray Tracing and Rasterization?

Used to be if you wanted to see ray tracing, you went to the movies. If you wanted to see rasterization, you fired up a video game. Ray tracing models the way light moves around the real world beautifully but it’s computationally intensive. Rasterization, by contrast, can be done in a hurry. NVIDIA’s latest Turing architecture GPUs blur these lines, with hardware acceleration for real-time ray tracing, making truly cinematic games possible.

Have a question you want answered? Send us your idea.

 

NVIDIA Brings Real-Time AI to Broadcast, Sports and Global Streaming at IBC

by

At the IBC conference, running Sept. 11-14 in Amsterdam, the creative, technology and business communities are coming together to turn ideas into action and discuss innovations across the media and entertainment industries. More than 44,000 attendees from 170+ countries are gathering to explore 1,300+ exhibitions in 14+ halls and outdoor spaces, with over 600 speakers delivering insights.

Read on to learn more about what NVIDIA’s highlighting at the show.


NVIDIA AI for Media Brings Real-Time Intelligence to Broadcast, Sports and Production Workflows 🔗

Media companies are increasingly integrating AI into live production, sports, news and streaming workflows to unlock richer performance insights, verify video authenticity, and enhance and localize content — all without disrupting trusted broadcast environments.

At IBC 2026 in Amsterdam, NVIDIA announced a major expansion to NVIDIA AI for Media — a collection of GPU-accelerated software development kits (SDKs), NVIDIA NIM microservices, playbooks, and blueprints that enhance audio, video and augmented-reality effects for media and entertainment workflows — to unlock new ways to understand motion, verify and enhance video, localize programming and build AI-powered media applications. 

The NVIDIA Synthetic Video Detector (SVD) NIM microservice, announced earlier this year at SIGGRAPH, helps organizations assess the probability of whether footage is authentic or AI-generated, giving editorial, content-authentication, digital-forensics and media-integrity teams another point of analysis in their review process.

Since its initial release, SVD’s accuracy has reached 99.3% for text-to-video content and 97.7% for image-to-video content, with especially large gains on difficult image-to-video cases. 

Dalet is integrating SVD into a secure, cloud-hosted verification workflow for news organizations. This allows editorial teams to submit footage through SVD, inspect and review the resulting scores and metadata within a Dalet interface. 

TwelveLabs announced the general availability of Compliance by TwelveLabs, its first application built on the company’s video intelligence platform, helping media and broadcast teams rapidly screen content against regional and custom compliance standards. The solution integrates SVD to add frame-level authenticity signals and confidence scores, enabling media teams to identify potentially synthetic media within the same compliance workflow.

Wowza, whose Wowza Streaming Engine media server technology powers more than 35,000 video deployments across over 170 countries, will distribute SVD through the Wowza Video Intelligence Framework. The solution, powered by NVIDIA-accelerated infrastructure, will enable broadcasters, streaming providers and other organizations to analyze live video feeds and extract data around detected objects, scenes and signs of AI generation in real time. It can be deployed and run on premises, at the edge, in the cloud, across hybrid deployments or fully air-gapped, giving organizations greater control over critical media workflows.

NVIDIA 3D Body Pose estimates 2D and 3D human joint locations and angles from video captured by a single camera, helping turn motion into structured data without marker-based capture systems.

For sports organizations, that data can support player and athlete movement tracking, biomechanics and performance analysis, replay enhancement, officiating and adjudication workflows, player-safety applications, and virtual interaction and immersive experiences. 

The technology can also provide structured human-motion data for content-creation workflows. When mapped to a compatible character rig, joint and motion data can serve as input for animation blocking, digital doubles, character retargeting and virtual-production experiences.

Vizrt is using Body Pose technology in live virtual-studio environments, with tracked body movement driving real-time 3D lighting effects such as reflections, shadows and environmental rendering.

Video Frame Generation (VFG) makes video motion appear smoother by using generative AI to create new frames between the original frames of a video. It can increase frame rates by 2x or 4x while preserving visual quality and temporal consistency, enabling more fluid sports, slow-motion replays, live media and other high-motion video experiences. VFG also supports frame-rate conversion and frame boosting for generative AI video workflows.

Ross Video is integrating VFG into its Rio Replay platform to create AI-assisted slow-motion video for sports production.

The work supports 6x slow-motion generation for sports replay. Development is underway toward 8x interpolation, meaning generated intermediate frames can give replay teams smoother motion without requiring every frame to be captured by an ultrahigh-frame-rate source camera. 

NVIDIA Video Super Resolution (VSR) uses AI to upscale video while reducing noise, blur and compression artifacts. New streaming modes let developers choose between real-time performance and higher image quality, while adjustable controls help achieve the desired level of enhancement. VSR also adds 10-bit video support and improves overall performance and quality. VSR is available through the NVIDIA Video Effects SDK and a NIM microservice for use in streaming, broadcast, conferencing, video playback and content-creation applications. 

The technology can support video players, conferencing applications, creator tools, streaming services, transcoders and broadcast systems through a common interface.

NVIDIA TrueHDR converts standard-dynamic-range video into high-dynamic-range output in real time, reaching up to approximately 2,000 nits while preserving local contrast and adapting brightness to the content.

VSR, VFG and TrueHDR can be combined within a single video-effects pipeline — helping media companies enhance existing content libraries for streaming, transcoding, gaming and creator workflows.  

The NVIDIA LipSync and Active Speaker Detection NIM microservices help developers build localization systems for interviews, news, sports, entertainment and other programming where multiple people may appear on screen.

LipSync transforms mouth movement in an input video to match a target audio track while preserving natural head pose, blinking and body movement. The new release improves facial occlusion handling and better preserves teeth, lip and facial textures.

The new Active Speaker Detection NIM microservice no longer requires speaker diarization for multiple audio tracks, adds voice activity detection and expands NIM microservice deployment support through a gRPC interface and broader GPU compatibility.

NDI is using NVIDIA AI for Media, including the NVIDIA LipSync NIM microservice, to enable real-time translation, lip-synced dubbing and regional language adaptation within existing broadcast workflows. By generating multiple language experiences from a common media stream, the approach can help broadcasters reach global audiences while reducing the bandwidth, infrastructure and production complexity traditionally required for multilingual distribution.

Studio Voice includes new Microphone Profiles built on NVIDIA Studio Voice NIM microservices, giving users more control over the tonal character of enhanced speech.

The capability is designed to suppress background noise, reduce room reverberation and improve speech clarity, then shape the enhanced output into a selected microphone profile for more polished live communications, streaming, podcasting and content creation.

Try NVIDIA AI for Media NIM microservices. See the latest NVIDIA and partner workflows at IBC 2026.


NVIDIA Holoscan for Media Provides Open Media Exchange Layer to Build and Connect Live Media Applications 🔗

As broadcasters, streaming services and sports organizations adopt software and AI, the infrastructure behind live content is becoming more flexible, more connected and increasingly built on shared accelerated computing.

The integration of Media Exchange Layer (MXL) with NVIDIA Holoscan for Media accelerates that transition — providing the common exchange layer that helps media applications connect and operate together. 

Holoscan for Media is an open reference architecture and developer toolkit for building AI-powered media functions and applications for software-defined live production. MXL adds an open way for those software-based media functions to exchange live video, audio and data across a distributed environment. 

As production functions move into software, developers can build applications that share accelerated infrastructure, connect dynamically and evolve independently. That can help media companies use infrastructure more efficiently, introduce new capabilities faster and reduce the amount of custom integration required between applications.

The integration also creates a stronger foundation for AI in live media. AI processing, video applications and traditional media functions can increasingly operate on the same accelerated infrastructure and in the same software-defined environment.

For technology vendors, this expands the opportunity to build applications that can work across broader, multi-vendor ecosystems. For media companies, it creates a path toward infrastructure that can adapt as formats, applications and AI capabilities change.

See the demo at IBC in EBU Stand 10.D21. ​Learn more about Holoscan for Media.


NVIDIA Sports Intelligence Playbooks Chart a Path to Multimodal AI for Sports 🔗

Sports is becoming a proving ground for a broader shift in AI: from general-purpose models toward fine-tuned open models built on proprietary data.

NVIDIA Sports Intelligence Playbooks are designed to accelerate that transition. They give leagues, media companies and technology providers structured frameworks to fine-tune NVIDIA open models on their own sports footage and annotations, creating multimodal AI that can understand the rules, players, scoring, strategy and context unique to a sport.

Sports organizations hold large volumes of proprietary video, metadata and performance information that are difficult for competitors to replicate. The playbooks provide a practical blueprint for converting those assets into AI capabilities that can underpin new analytics products, media experiences, automation tools and revenue streams.

The playbooks span the AI lifecycle, including data preparation, fine-tuning, inference, evaluation, optimization and deployment, and bring together NVIDIA technologies including Nemotron, NeMo AutoModel, Megatron Bridge, NIM microservices and NVIDIA accelerated computing

By providing an integrated path from model customization to production, Sports Intelligence Playbooks can reduce the cost and complexity of building specialized sports AI while increasing demand across its compute, software and inference stack.

Early testing demonstrates the potential of domain specialization. When evaluated on previously unseen footage using question formats similar to those used in training,  multiple-choice accuracy increased from approximately 53% to 94% and open-ended evaluation from approximately 5.7% to 66%.

Machina Sports is integrating Sports Intelligence Playbooks with its sports-native data, evaluation and agent infrastructure, enabling rights holders to turn proprietary media and expertise into private, deployable intelligence for live production, content and fan experiences.

The opportunity also expands as agentic AI becomes increasingly adopted. With the NVIDIA AI-Q Blueprint, organizations can use their domain-specific sports models as expert intelligence within agents that reason across video, enterprise data and software systems, extending the playbook from sports understanding into decision-making and automation.

Wowza is integrating vision language models, including NVIDIA Cosmos 3 and Nemotron, into the Wowza Video Intelligence Framework, fine-tuned through NVIDIA Sports Intelligence Playbooks to detect sports-specific moments in live streams and reduce time to action. 

Explore NVIDIA Sports Intelligence Playbooks.


NVIDIA Brings Multilingual Content Localization to Live Broadcast 🔗

Reaching global audiences with live programming requires more than translating words. Language nuances, voice, timing, facial movement, captions and onscreen graphics must work together in real time, while preserving the editorial intent and production quality of the original program.

To help broadcasters, sports leagues, rights holders and streaming services bring these elements into a unified, software-defined, real-time localization workflow, NVIDIA is bringing its Content Localization technologies to the NVIDIA Holoscan for Media developer toolkit. Designed for broadcast and streaming developers, the reference workflow enables captions, translated audio, dubbing, synchronized video and localized graphics.

Content Localization with Holoscan for Media provides a reference for how localization technologies can work together in software-defined broadcast applications. Developers can select the capabilities needed for each program, market or distribution channel rather than deploying separate infrastructure for every localized version.

Content Localization with Holoscan for Media incorporates the latest advancements from NVIDIA AI for Media, including improved LipSync when faces are partially obscured and enhanced Active Speaker Detection to help applications identify who’s speaking in multi-person scenes.

Expanding the Reach of Live Programming

Localization can transform the reach and economics of live programming. A shared, composable workflow can help media companies introduce regional coverage faster, serve more audiences and tailor experiences for individual markets — while preserving the timing, visual context and editorial control required for live production.

Technologies from AI-Media, CAMB.AI, Chyron and Panjaya each address a specific part of content localization with Holoscan for Media, from adapting voice and onscreen delivery to creating multilingual captions and translated audio, localizing graphics, and preserving expression and identity across live and on-demand content.

The Content Localization technologies also support file-based, streaming and post-production applications. Developers can use application programming interfaces for on-demand workflows and the Holoscan for Media reference workflow when localization must run as part of a live media environment. Together, they provide a consistent foundation for building multilingual media services across production and distribution.

Learn more about NVIDIA Holoscan for Media and AI for Media.

Sparks Fly: NVIDIA Accelerates Local AI at IFA 2026

Local agents get easier to install, faster to run and able to tap multiple RTX PCs at home with NVIDIA PAIR — plus, new NVIDIA RTX Spark Windows PCs arriving in October.
by

Frontier intelligence is going local. At IFA 2026, NVIDIA, Microsoft and its partners are teaming up to provide faster inference and new tools that make agents easier to set up and run locally on NVIDIA hardware. New compact NVIDIA RTX Spark Windows PCs are also coming in October to give AI enthusiasts, developers and creators more ways to run capable agents locally and securely. 

Today’s announcements include:

  • Simplified local AI support for NVIDIA GPUs is coming in Hermes Agent, OpenClaw and Perplexity Portable Computer.
  • Up to 1.9x faster local inference — new llama.cpp and vLLM optimizations are available now directly and through LM Studio and Ollama. 
  • NVIDIA PAIR — a Personal AI Router tool that intelligently distributes AI inference across the PCs on a user’s local network.
  • NVIDIA RTX Spark arrives in October  — with new Windows PCs from Lenovo and Acer. Electronic Arts, Embark and Ubisoft are among the latest game publishers and developers bringing their blockbuster titles to NVIDIA RTX Spark.

Also, August was a busy month for local AI:

  • Nemotron 3.5 Lightning — which can run on NVIDIA RTX PCs, RTX PRO Workstations, DGX Spark and Jetson — is a 30-billion parameter model that has been launched. Get started with Nemotron 3.5 Lightning today. 
  • Z.ai’s GLM-5.3-Flash is a multimodal mixture-of-experts (MoE) model that’s bringing agentic AI to DGX Station.
  • Qwen has released Qwen3.8-Flash-Next, an open weight multimodal MoE model, which can run locally on DGX Spark and DGX Station, along with Qwen3.8-27B, a 27-billion-parameter open model optimized for local agentic and coding workloads on NVIDIA GPUs. 
  • LTX’s LTX 2.5 is an open-world video generation model optimized for NVIDIA RTX GPUs, DGX Spark and DGX Station, with new NVFP4, FastVideo and ComfyUI enhancements for faster, more memory-efficient local generation.
  • MiniMax-H3 is an open-weight video generation model with synchronized audio that can run locally on NVIDIA GPUs through ComfyUI. FastVideo teamed up with NVIDIA researchers to improve this further by releasing FastH3 — an open-weight, four-step distilled version that improves performance by 7x. Optimized FastVideo recipes for NVIDIA RTX GPUs and DGX Spark are coming soon.
  • Meta’s Muse Glimmer is a 30-billion-parameter open-weight model for coding and agentic workloads that can run locally on GeForce RTX PCs, DGX Spark, DGX Station and Jetson. NVIDIA has also released NVFP4 quantization with DGX Spark support for more memory-efficient local deployment.
  • DeepSeek v4 Flash is a 284-billion-parameter MoE model with 13 billion active parameters that can run locally on 2x DGX Spark cluster and DGX Station.

A Simpler Start for Local Agents

Getting a local agent up and running with local models required some effort — choosing a model, finding a compatible inference server, dialing in quantization settings and keeping everything updated. That friction is disappearing on RTX and DGX systems.

Three of the most widely used agent apps will offer simplified local model setup on Windows, each built on llama.cpp and incorporating NVIDIA’s latest inference optimizations. The new setup experiences are designed to reduce manual configuration and make it easier to get local agents up and running. 

Last month, Perplexity introduced its Portable Computer agent, giving users a simple way to run Perplexity locally on Linux systems like NVIDIA DGX Spark with the models, orchestration and tools packaged into a single app experience.

Perplexity Portable Computer is available on NVIDIA RTX GPUs with at least 24GB VRAM running on Linux, with support on Windows coming soon, bringing that same streamlined setup to a broader group of PC users. Users can run complete workflows locally without consuming credits, while selectively escalating parts of a task to one of 15+ frontier models in the cloud when additional research or reasoning is needed. Portable Computer asks for permission before sending content to the cloud, helping users keep sensitive information on their device. Here’s some example use-cases:

  • Engineering: Review open PRs in a connected GitHub repo and sort them into ready, blocked, stale, and needs review, each tagged with the next step. Docs that fell out of sync with the latest merge get caught and fixed, with a PR opened for the changes.
  • Finance: Point the agent at two years of brokerage summaries, consolidated 1099s and tax returns and have it trace the recurring holdings creating the most avoidable fees and tax drag, with every figure cited to the exact file and page — all without a document ever reaching a chatbot.
  • Startups: Ask why activation went flat, and the agent analyzes the funnel export locally to find where new signups drop off between install and first completed task, then posts the top insights straight to the team’s Slack channel.

Try Portable Computer today.

Hermes Agent — developed by Nous Researchis a general-purpose agent used by millions that excels at reliability and self-improvement. Model- and provider-agnostic, Hermes is built to run all day on local systems, making RTX PCs, RTX PRO workstations and DGX Spark a natural fit.

Configuring a local model in Hermes will provide users with one-click setup across RTX and DGX systems on Windows. The agent will automatically detect the NVIDIA GPU, select an appropriate model and configuration, and run it through integrated llama.cpp with NVIDIA inference optimizations already in place, eliminating manual model downloads and tuning. Support for Linux is coming soon.

Once it is running, Hermes works the way it does anywhere else. It uses tools, maintains context across tasks, remembers information between sessions and creates reusable skills over time, allowing the agent to become more capable with continued use. Running the model locally on a GPU keeps performance fast while keeping data on the system.

One-click local model setup is available now on Windows, with support coming soon to Linux. Learn more about Hermes Agent.

OpenClaw has become one of the defining projects of the open-agent movement — the largest AI project on GitHub, with more than 380K stars and a fast-growing community that’s building tools and skills across research, engineering, project management and everyday productivity.

NVIDIA, Microsoft and OpenClaw have been working together to make that experience easier to set up on Windows PCs. To reduce onboarding friction, the OpenClaw Windows App simplifies the process of setting up an optimized local model on any RTX GPU with at least 24GB of VRAM.

Learn more in the OpenClaw blog.

Faster Inference Gives Local Agents a Boost

Inference performance is critical to keeping local agents responsive. NVIDIA is continuing to collaborate with the open-source llama.cpp and vLLM communities to accelerate agentic workloads across local NVIDIA platforms.

llama.cpp delivers up to 1.9x higher throughput through kernel optimizations on a GeForce RTX 5090, enhanced speculative decoding techniques and faster prefill. ​

vLLM delivers 1.2x on RTX PRO 6000 Blackwell Workstation Edition and up to 1.4x on two DGX Spark clusters. New XQA attention kernels in FlashInfer and backend optimizations help to accelerate inference across both platforms.

These gains are available on the llama.cpp and vLLM inferencing backends. 

Users can also experience these via the LM Studio and Ollama applications.

Tap Idle PCs for More Local AI Compute With NVIDIA PAIR

More than half of U.S. households have two or more PCs, and much of that computing power sits idle throughout the day. NVIDIA Personal AI Router (PAIR) is a free, open source software tool that puts those systems to work together for local AI.

Agentic workflows often break complex tasks into smaller jobs that can run in parallel, but performance can slow when every request is competing for the same GPU. PAIR automatically discovers compatible PCs on a local network and routes independent inference requests to whichever system has capacity. It works with Ollama and LM Studio and can adapt as devices join or leave the network.

For example, a user could ask Hermes to create a “Sunday Reset” plan by sorting through a cluttered inbox and prioritizing what needs attention now, what can wait and what can be skipped. Hermes can split that work across multiple subagents, while PAIR distributes those jobs across available PCs instead of having them all wait on a single GPU.

The result is more compute for local agents, with more tasks running in parallel and the flexibility to move AI workloads to another PC while the main system is being used for gaming, creating or other work.

The NVIDIA PAIR beta is available for Windows, macOS and Linux through both graphical and terminal interfaces, supporting NVIDIA GeForce RTX 20 Series GPUs and newer, NVIDIA RTX PRO workstation GPUs (Turing architecture and newer), NVIDIA DGX Spark and Apple M4 or newer silicon.

Check out the NVIDIA tech blog to get started with NVIDIA PAIR. 

Powerful On Device Photo Editing With Cyberlink PhotoDirector AI PC Mode on RTX Spark

Open image and video models enable artists to experiment with Creative AI models on PCs.  This enables artists to iterate and explore concepts and ideas, without the dreaded token anxiety and keep more of their creative work private and on-device.

CyberLink’s new PhotoDirector AI PC Mode is one of the first applications to integrate these diffusion models directly into a creative software, and turn them into a creative tool at the finger tips of the artists. Coming to PhotoDirector 365 and optimized for NVIDIA RTX Spark when it launches, AI PC Mode users are getting AI-powered editing tools for generative editing, image enhancement, object and distraction removal, background removal and replacement, portrait refinement and the creation of entirely new visuals — with the flexibility to choose between local or cloud processing, depending on the task.

On NVIDIA GPUs, PhotoDirector uses TensorRT-RTX and FP8 to accelerate local AI.

Start using Cyberlink’s PhotoDirector 365 photo editing software and learn more about PhotoDirector AI PC Mode, launching with RTX Spark in October.

NVIDIA RTX Spark Windows PCs Arrive October 2026

NVIDIA RTX Spark is coming this October— and at IFA 2026, partners are showing off their hardware. At IFA, newly announced designs join the existing six OEMs shipping in October. Acer showed its compact desktop RTX Spark concept, and Lenovo announced its Yoga Pro  9n and Yoga 9n 2-in-1.  

RTX Spark is a new beginning for Windows PCs. One PC built for creators, gamers and AI agents. With a powerful 1 Petaflop RTX Blackwell GPU, up to 128GB of unified memory and a highly efficient 20-core Grace CPU, RTX Spark delivers incredible performance and efficiency. This superchip enables high performance thin laptops with all day battery life and compact desktops to power always-on agents. Paired with the new Windows Agent framework, it enables agents that run safely in the background under OS level control.

Last week at Gamescom, Electronic Arts, Embark and Ubisoft were among the latest game publishers and developers bringing their blockbuster titles to NVIDIA RTX Spark Windows PCs. They join the publishers that announced RTX Spark support at COMPUTEX in May, including KRAFTON, NetEase, Riot Games and XBOX. Read more.

Sign up to be notified when RTX Spark laptops and desktops are available.

#ICYMI: More Updates From NVIDIA Local AI 

🎮 NVIDIA Brings New RTX Tech and Games to Gamescom — NVIDIA released DLSS 4.5 Ray Reconstruction, featuring a new second-generation transformer model for improved image quality in ray-traced and path-traced games. Gamescom also brought new RTX announcements for titles including 007 First Light, CONTROL Resonant and Gears of War: E-Day, plus expanded game support for the upcoming NVIDIA RTX Spark.

🐋Introducing DeepSeek Harness — DeepSeek’s new open source harness pairs with DeepSeek-V4-Flash to power local agentic coding workflows on NVIDIA DGX Station and multi-DGX Spark setups.

📊MLPerf Client v2.0 Expands AI PC Benchmarking — MLCommons released MLPerf Client v2.0, developed in collaboration with NVIDIA and other industry leaders. The update adds new benchmarks for agentic AI and image generation, alongside expanded LLM testing for real-world local AI workloads.

Follow NVIDIA RTX Spark on X, Instagram, TikTok and Facebook — and stay informed by subscribing to the NVIDIA Local AI newsletter. Follow NVIDIA Workstation on LinkedIn and X

See notice regarding software product information.

  • Categories:
  • AI

NVIDIA to Acquire Hugging Face

by

I’m excited to announce that NVIDIA has agreed to acquire Hugging Face for $12,930,300,000. Together, we will scale Hugging Face’s platform, strengthen its infrastructure and expand access to AI for developers and institutions worldwide.

Over the past decade, Clem, Julien, Thomas and the team at Hugging Face have built something remarkable: a vibrant home for the open model developer community.

More than 18 million developers, researchers and creators use Hugging Face to share more than 3 million models, 500,000 datasets and 1 million applications. More than 200,000 companies use the platform to discover, evaluate, customize and deploy AI.

Hugging Face will remain an open platform for the entire AI ecosystem. Developers will choose the models they want, the frameworks they want, the clouds and inference service providers they want and the computing platforms they want. NVIDIA compute will not be required to build on or deploy through Hugging Face.

Hugging Face will continue to support open source and open weight models from across the ecosystem, from every model builder. It will continue to support multi-cloud and multi-accelerator development and deployment, so builders can use the hardware and infrastructure that best fit their work.

Recently, I coauthored an open letter on the importance of open weights to the AI economy. Joined by leaders from across the industry, we made a simple point: open weights broaden access to AI and help ensure that AI leadership is distributed across companies, institutions and communities.

Open models let startups, businesses, universities and public institutions build on advanced capabilities without training every model from scratch. They enable organizations to match the right model to the right job. That is how AI can advance safely, strengthen cybersecurity and sovereignty, accelerate innovation, and reach factories, hospitals, farms, classrooms and Main Street businesses around the world.

AI advances faster when people can build together.

NVIDIA has been committed to open weight models for years, demonstrated by multiyear investments and major contributions to open source platforms, including Hugging Face. NVIDIA has said that open models, data and tools broaden access to AI, and it has contributed hundreds of open models and datasets to Hugging Face as part of that effort.

  • NVIDIA is the largest contributor of open models and data to Hugging Face, and our contributions continue to grow.
  • NVIDIA has released more than 500 models on Hugging Face and more than 250 open datasets.
  • We build our own models, libraries and tools in the open so developers everywhere can use them, modify them and build on top of them.

As the opportunity for open models accelerates, Hugging Face can serve the global AI community at unprecedented scale. NVIDIA’s infrastructure, engineering and global reach can help improve platform reliability, safety, model evaluation, inference and deployment capabilities, while preserving the open ecosystem that made Hugging Face foundational.

I am honored that Clem came to me as he considered the next chapter of Hugging Face and believed NVIDIA would be a great home for the company, its community and the future of open models. We share this vision, and the Hugging Face team will now bring their passion and expertise to a much larger canvas, with their same iconic 🤗 brand.

To the millions of builders on Hugging Face: thank you for pushing the boundaries of what is possible. We can’t wait to build the future together with you. Together, we will make AI more open, more capable and more accessible to people and institutions around the world.