Open Agent Safety Platform
NVIDIA's Open Agent Safety Platform adds a hardware watchdog for rogue agents
NVIDIA's Open Agent Safety Platform includes a hardware watchdog that can quarantine rogue agents.
NVIDIA's AI coverage centers on the open Nemotron model family — reasoning, nano and speech variants — alongside research models and ecosystem news. ThursdAI — the weekly AI news podcast hosted by Alex Volkov — has covered 30 NVIDIA releases since Jan 2025, most recently Open Agent Safety Platform on Oct 1, 2026. Highlights include Nemotron 3 Super 120B, Hugging Face acquisition, Nemotron 3 Ultra, Nemotron 3.5 Lightning. 17 of them shipped with open weights. Every entry below has the episode segment where we covered it live, plus primary-source links and key numbers where we have them.
NVIDIA's Open Agent Safety Platform adds a hardware watchdog for rogue agents
NVIDIA's Open Agent Safety Platform includes a hardware watchdog that can quarantine rogue agents.
NVIDIA releases an open speaker diarization model
NVIDIA released an open model for speaker diarization, tracking who is speaking in audio.
NVIDIA reportedly agrees to acquire Hugging Face for $12.9B
The Information reports NVIDIA has agreed to acquire Hugging Face for $12.9 billion — roughly 3x the 2023 valuation, after a declined $500M offer during the $7B era. Neither company had confirmed at air time. Hugging Face reached about $100M ARR in 2026 with roughly 13 million accounts, and the ThursdAI panel leaned positive for open source given NVIDIA's open-source push and the cash infusion hosting requires.
NVIDIA ships Nemotron 3.5 Lightning: 30B MoE with 3B active
NVIDIA released Nemotron 3.5 Lightning, a 30B MoE with just 3B active parameters delivering up to 4x output speed and strong voice-agent results, with weights on Hugging Face in NVFP4. CoreWeave Inference picked it up with day-zero support.
NVIDIA launches the Open Secure AI Alliance: an open defensive stack, born from the Hugging Face hack
Jensen Huang's second letter of the week proposes an open defensive stack — identity, permissions, isolation, harnesses, logs, and evals — under Linux Foundation stewardship, with launch partners including Microsoft, Hugging Face, CrowdStrike, Mistral, Cloudflare, and Nous Research. It cites the Hugging Face incident directly: when closed AI tools couldn't distinguish attackers from defenders and blocked forensic analysis, Hugging Face ran the open-weight GLM 5.2 on its own infrastructure to contain the intrusion. OpenAI and Anthropic are absent.
Jensen Huang joins X and publishes the Open Weights and American AI Leadership letter
Jensen Huang's first-ever X post published a coalition letter arguing open-weight models are the path to AI diffusion and security, signed at launch by NVIDIA, Microsoft, Meta, Google, and OpenAI and growing from 25 to 230 signatories within a week — CoreWeave among them, announced first on ThursdAI. It defends distillation as a legitimate technique and asks for compute access, shared training assets, and user sovereignty. Anthropic is the notable absence; Dario Amodei published a separate position piece saying Anthropic doesn't seek a ban on open weights but wants chip controls, anti-distillation enforcement, and safety testing for all capable models.
NVIDIA ships Nemotron 3.5 ASR, a 600M streaming speech model
NVIDIA released Nemotron 3.5 ASR, a 600M-parameter open multilingual streaming speech-to-text model aimed at voice agents. It supports 40 languages and reportedly delivers 17x more throughput than Parakeet-style baselines at half the size, pushing the latency/accuracy frontier for open voice-agent infrastructure.
NVIDIA releases Nemotron 3 Ultra, a 550B open-weight MoE for agents
NVIDIA dropped Nemotron 3 Ultra the day of the show, a 550B-parameter sparse MoE with 55B active parameters built for long-running agentic harnesses like OpenCode, Hermes, and OpenClaw. Chris Alexiuk joined to explain the hybrid Mamba/Transformer architecture and the unusually complete open release: weights, training data, recipes, a GenRM reward model, and an NVFP4 quantized checkpoint.
NVIDIA announces RTX Spark Arm + Blackwell platform for local AI PCs
At Computex, NVIDIA unveiled RTX Spark, an Arm CPU plus Blackwell GPU PC platform with 128GB unified memory targeting local AI agents and 120B-class local inference. A wave of thin laptops with RTX 5070-class GPUs and roughly one petaflop of local AI compute raises the question of what agents should run locally versus in the cloud.
NVIDIA Nemotron 3 Nano Omni: hybrid Transformer-Mamba MoE
NVIDIA released Nemotron 3 Nano Omni, a 30B-total/3B-active hybrid Transformer-Mamba MoE with 256K context. It delivers 9x throughput on consumer hardware.
NVIDIA Lyra 2.0: single image to explorable 3D worlds, Apache 2.0
NVIDIA released Lyra 2.0 under Apache 2.0, generating persistent, explorable 3D worlds from a single image. Together with Baidu ERNIE-Image and Tencent HYWorld 2.0, it rounds out a week of open releases in the 3D-world-from-single-image race.
NVIDIA DLSS 5 adds a generative AI filter for photo-realistic lighting
Announced at GTC, NVIDIA's DLSS 5 introduces a new generative AI filter bringing photo-realistic lighting to RTX 50-series GPUs. It applies generative models to real-time game rendering, extending DLSS beyond upscaling and frame generation.
NVIDIA GTC: GR LPX pairs Rubin NVL72 servers with the new Groq 3 chip
NVIDIA's GTC hardware reveal integrates the new Groq 3 chip (gen 2 was never publicly seen) into Rubin NVL72 servers via the GR LPX system. Claims include 3x tokens-per-watt efficiency at baseline, up to 30x at higher throughput, and 1000+ tokens/sec on a 2T-parameter frontier model with 400K context — performance the current Blackwell generation can't reach at any price.
NVIDIA announces NemoClaw, enterprise-hardened OpenClaw, at GTC
At GTC, Jensen Huang spent 15 minutes on OpenClaw, calling it the most important open source release since Linux and declaring 'every company needs an OpenClaw strategy.' NVIDIA released NemoClaw, a hardened enterprise reference implementation of OpenClaw with a privacy router and policy engine aimed at solving the agent security problem.
NVIDIA releases Nemotron 3 Super 120B with $26B open-source bet
NVIDIA launched Nemotron 3 Super, a 120B Hybrid Mamba-Transformer MoE model with 12B active parameters, a 1M-token context window, and 450 tok/s throughput. It shipped with BF16/FP8/NVFP4 weights, a base checkpoint, SFT and pre-training data, and the full training recipe, alongside a $26B 5-year open-source commitment. It is available on W&B Inference at $0.20/M input and $0.80/M output.
NVIDIA releases PersonaPlex-7B voice model
NVIDIA released PersonaPlex-7B, an open voice/audio model published on Hugging Face with code on GitHub. Listed in the week's Voice & Audio releases.
NVIDIA Alpha Mayo: open source reasoning self-driving models
NVIDIA announced Alpha Mayo at CES, a family of open source reasoning-based self-driving AI models. The models perform end-to-end autonomous driving with explicit reasoning steps, like identifying jaywalkers and stopping accordingly, demoed in a Mercedes-Benz.
NVIDIA acquires Groq team and licenses its tech for ~$20B
NVIDIA entered an exclusive licensing deal with Groq and acquired most of its team for approximately $20B. Groq's inference-optimized chips, created by former Google TPU lead Jonathan Ross, complement NVIDIA's training dominance as inference demand grows exponentially across AI use cases.
Nemotron Speech ASR: 600M streaming model with 24ms latency
NVIDIA released Nemotron Speech ASR, a 600M parameter open source streaming speech recognition model with 24ms median latency and support for 900 concurrent streams on a single H100. Kwindla Hultman Kramer of Daily.co demoed sub-500ms voice-to-voice latency using a three-model pipeline of Nemotron ASR, Nemotron Nano LLM, and Magpie TTS.
NVIDIA Vera Rubin platform: 5x Blackwell inference at CES 2026
Jensen Huang unveiled the Vera Rubin platform at CES 2026, NVIDIA's next-gen AI computer delivering 50 PFLOPS and 5x inference performance over Blackwell while adding only ~200W of power draw. It needs 75% fewer GPUs for 10 trillion parameter MoE training, packs 72 GPUs per rack with 20.7TB memory and 13 TB/s bandwidth, is 100% liquid cooled, and entered full production just four months after the B300.
NVIDIA Project Digits: $3,000 desktop that runs 200B-param models
NVIDIA announced Project Digits in January, a $3,000 desktop supercomputer capable of running 200B parameter models locally. It brought serious local-inference hardware to individual developers and was one of January's standout hardware stories.
NVIDIA ships Nemotron 3 Nano, a 30B hybrid Mamba-MoE with full recipes
NVIDIA released Nemotron 3 Nano, a 30B-parameter hybrid Mamba-MoE model with only 3B active parameters for efficient inference. The panel called it the most consequential open release of the week because NVIDIA shipped not just weights but technical reports, training recipes, and details on the 25T-token training data.
NVIDIA releases ChronoEdit-14B Upscaler LoRA
NVIDIA released an Upscaler LoRA for its ChronoEdit-14B image editing model, available on Hugging Face with Diffusers pipeline support. It adds high-quality upscaling to the ChronoEdit physics-aware editing stack.
NVIDIA DGX Spark: a desktop personal supercomputer for local AI
NVIDIA started shipping DGX Spark, a desktop personal AI supercomputer aimed at prototyping and local inference. The show pointed to the LMSYS deep dive on its real-world performance, and Alex shared his own first impressions of the device.
Nvidia commits up to $100B to OpenAI for 10GW of compute
Nvidia and OpenAI announced a letter of intent under which Nvidia would invest up to $100 billion in OpenAI as the two deploy at least 10 gigawatts of Nvidia systems for OpenAI's next-generation infrastructure. The episode's big-company segment centered on this deal as evidence that money and infrastructure, not just models, now drive the AI race.
NVIDIA releases DAM-3B for region-based image and video captioning
NVIDIA dropped the Describe Anything Model (DAM-3B), a 3 billion parameter multimodal model for region-based image and video captioning. You can point it at a specific region of an image or video and it generates a detailed description of just that area. NVIDIA also published an accompanying DescribeAnything dataset and a Hugging Face demo.
NVIDIA ships Nemotron Ultra, a 253B pruned and distilled Llama 3.1-405B
NVIDIA released Nemotron Ultra, a pruned and distilled finetune of Llama 3.1-405B at roughly half the parameters (253B). Its benchmarks even included Llama 4 comparisons, showing the older finetuned Llama beating the new models on AIME, GPQA and more. It supports 128K context and fits on a single 8xH100 node for inference.
NVIDIA Canary Flash: Apache 2 speech recognition and translation
NVIDIA released Canary 1B Flash and 180M Flash, Apache 2.0 licensed speech recognition and translation models built as Llama finetunes. The permissive license makes them freely usable for commercial ASR and translation workloads.
NVIDIA drops Llama-Nemotron reasoning models plus training dataset
NVIDIA released the Llama-Nemotron family, including Super 49B and Nano 8B reasoning models, announced around GTC. Alongside the open weights, NVIDIA published the Llama-Nemotron post-training dataset, giving the community both the models and the data recipe behind them.
NVIDIA releases Eagle 2 open vision-language models
NVIDIA published Eagle 2, a family of open vision-language models with an accompanying paper, model weights on Hugging Face, and a live demo. It is a fully transparent VLM release covering training data strategy and recipes, competitive with much larger vision models.
Never miss a NVIDIA launch — we cover every release live, every Thursday.