RL Rollouts
CoreWeave RL Rollouts hot-load weights about 15x faster
CoreWeave shipped RL Rollouts, which hot-load model weights about 15x faster for reinforcement-learning workloads.
Compute, GPUs, hardware, serving, quantization, and efficiency for running models at scale. — 100 releases covered on the show.
CoreWeave RL Rollouts hot-load weights about 15x faster
CoreWeave shipped RL Rollouts, which hot-load model weights about 15x faster for reinforcement-learning workloads.
CoreWeave Serverless GPU Sandboxes are free during the preview
CoreWeave's Serverless GPU Sandboxes give you an isolated sandbox with a GPU, started from Python at forge.coreweave.com, with no salesperson in the middle. They are free during the preview; sign up through Deok's form.
Cloudflare goes agent-first with cf, one CLI for its whole dashboard
For its birthday week Cloudflare shipped cf, an agentic CLI covering 3,000+ API operations, so anything you can click in the Cloudflare dashboard an agent can do from one command-line tool.
CoreWeave Agent Lens brings human-friendly insights to agent observability
Agent Lens, in public preview, pairs Weave observability with insights that surface small trends across long agent runs, such as the 1% of conversations that share the same problem.
CoreWeave launches Forge, the whole agentic loop in one place
Forge combines run, observe, curate, improve and evaluate in one product, with Weights & Biases (W&B Models included), OpenPipe's post-training, marimo notebooks and CoreWeave Sandboxes. It has a free tier, and Pro starts at $60 a month with a 30-day trial.
CoreWeave to offer NVIDIA Vera, the first CPU built for AI agents
CoreWeave is bringing NVIDIA's Vera CPU to its cloud for agent workloads. Deok Filho's new metric is agent packing: one rack of Vera CPUs runs over 20,000 agents at once on about 11,000 cores.
CoreWeave Sandboxes go GA as part of Forge
CoreWeave Sandboxes reached general availability as part of Forge. Wolfram Ravenwolf's WolfBench evaluations run on them.
CoreWeave adds serverless model distillation to Forge
Forge's serverless side covers inference, RL and SFT, and now model distillation: train a small student on a frontier teacher for your task without managing Kubernetes or a cluster.
CoreWeave launches serverless GPUs: GPU sandboxes by the hour, no contract
Announced live on ThursdAI by Deok Filho: serverless GPUs on CoreWeave, GPU sandboxes with untrusted code execution. Sign up at forge.coreweave.com, ask to have it enabled, and pay per hour with no contract and no commitment. Private preview started September 30, with more GPU SKUs planned.
Cognition becomes the first production customer on Vera Rubin NVL72 at CoreWeave
Cognition now runs its SWE-2 models on NVIDIA Vera Rubin NVL72 at CoreWeave, the first production customer anywhere, with CoreWeave citing up to 4.8x more token throughput than on GB200.
OpenAI launches UltraFast mode on Cerebras with a $500 Pro tier
A new $500 Pro tier adds UltraFast mode, OpenAI's models served on Cerebras: up to 8x faster in Codex and around 300 tokens a second for GPT-6 Astra, at 6x the price. The $200 Pro plan returns with half the usage.
SGLang adds native support for Jev-style decision models
SGLang announced native support for turning any model into a Jev-style decision model, starting with Qwen.
CoreWeave earns Platinum on SemiAnalysis ClusterMax 3.0 again
CoreWeave received the Platinum rating on SemiAnalysis ClusterMax 3.0, its third report in a row at the top tier.
DeepSeek V4.1 Flash: 552B MoE, 8B/16B active, KV cache 400x smaller than V1, MIT
DeepSeek's V4.1 Flash is a 552B-parameter multimodal MoE that activates only 8B parameters on prefill and 16B on decode, with a 1M-token context window, trained from scratch on 45 trillion multimodal tokens and released under MIT. It returns to an encoder-decoder architecture at half-trillion scale and cuts the KV cache to about 890 bytes per token, down from 389,000 bytes in the first DeepSeek release (over 400x smaller), which is why it is so cheap to serve. DeepSeek's own evals put it at 90.6 on Terminal-Bench 2.1 and 74.2 on DeepSWE 1.1, above Opus 5 and GPT-5.6 Sol, and second behind GPT-6 Astra on an Open Design leaderboard at about two cents per task. TokenJuice.ai serves it free in the US for a limited time in exchange for training data.
Kimi K3 (2.8T) runs on CoreWeave Dedicated Inference on GB300 NVL72
CoreWeave added Moonshot's 2.8T-parameter Kimi K3 to Dedicated Inference, running on GB300 NVL72 systems. Alex says it 'purrs like a kitten' and a token promotion is planned on the CoreWeave and W&B accounts.
Apple announces Mac Studio with M5 Max and M5 Ultra plus a new Mac Mini
Apple announced a new Mac Studio with M5 Max and M5 Ultra chips, plus a refreshed Mac Mini with higher-spec options. Starting at $2,499 but configurable to roughly $22K, the panel framed it as home AI infrastructure: Wolfram's 'central heating' theory — financed over two years it costs about the same as a Pro AI subscription, with unlimited local tokens.
SemiAnalysis benchmarks OpenAI's Jalapeño chip: beats Vera Rubin on throughput per watt
SemiAnalysis published a benchmark report on OpenAI's upcoming Jalapeño inference chip, reporting it beats NVIDIA's Blackwell and Vera Rubin on throughput per watt. Caveats: the numbers were supplied by OpenAI, and the AgentX suite had not yet been run.
Chroma launches Foundation: unified memory for your agents
Launched live during the show (founder Jeff Huber hopped on within minutes of the announcement), Foundation is a research preview of a shared memory system between you and your agents — memory as infrastructure. It ingests sources natively from Codex, Claude Code, Cursor and Slack on day one, with Notion, GitHub and Google Drive connectors coming, and is built on ChromaDB plus the Context-1 agentic search model (a GPT-OSS 20B fine-tune running ~400 tokens/sec at 25x less cost than Opus). Each Foundation manages and improves its own system prompt from natural-language feedback. Part of Chroma Cloud starting at $30/mo.
Mojo goes fully open source under Apache 2.0
The Mojo language is now fully open source under Apache 2.0 with LLVM exceptions, three weeks after Qualcomm's $3.9B acquisition of Modular.
OpenAI joins PORTS-Pike: 8 GW Ohio data center on a 20-year lease
OpenAI joined the PORTS-Pike project: an 8 GW data center in Ohio on a 20-year lease, with NVIDIA backing $105B in credit support.
OpenAI previews ultrafast GPT 5.6 Sol on Cerebras at ~14x speed
OpenAI previewed an ultrafast serving mode for GPT 5.6 Sol running on Cerebras hardware at roughly 14x normal speed. Access starts behind a work-account waitlist. The news broke during the live show.
Weave ships BYOB: media stays in your own S3/GCS bucket
Weights & Biases shipped bring-your-own-bucket support in Weave, so trace media stays in your own S3 or GCS bucket instead of W&B-managed storage.
Cloudflare OS: Kenton Varda's Sandstorm reborn as open source agent infrastructure
Kenton Varda's 'secret 10-year master plan': a remake of Sandstorm on Workers and Durable Objects, Apache 2.0 with no open-core catch. Every app instance ('Gadget') is a sandboxed Dynamic Worker with zero default internet; 'Gatekeepers' are supercharged MCP servers holding credentials, enforcing per-resource policy, and logging every action, with pending approvals simulated locally so agents keep working while humans review in batch. Thousands of Cloudflare employees have used v1 internally since May. Landing the week of the sandbox-escape disclosures, the timing wrote its own headline.
Endpoint Accuracy Index: the same open weights score 52% to 100% depending on your provider
Artificial Analysis launched an index measuring whether inference providers actually serve the model they claim: GLM-5.2 scores 52% on one provider and 100% on three others with a 5.2x price spread, and some gpt-oss-120b endpoints hit 22% on tool calling versus the reference's 37%. Low scorers produce about half the output tokens per task, pointing at aggressive quantization. v1.0 covers GLM-5.2 (15 providers), gpt-oss-120b (20), and DeepSeek V4 Pro (9), with Kimi K3 next. CoreWeave came in cheapest on gpt-oss-120b at $0.04/M blended with 98% accuracy.
Breaking on the show: OpenAI cuts GPT-5.6 Luna prices 80% and Terra 20%, crediting Sol's self-optimization
Dropped live during the episode: Luna prices fall 80%, Terra 20%, and a faster GPT-5.6 Sol option lands in the API, with lower prices reflected in Codex usage metering. OpenAI explicitly credits efficiency work GPT-5.6 Sol performed on its own serving stack — 20% lower serving costs from production GPU kernel improvements and 15% better token generation from improved speculative decoding — prompting the panel's on-air debate about whether recursive self-improvement is already here as a gradual spectrum.
MCP's biggest update ever: fully stateless core, MCP Apps and Tasks extensions, OAuth 2.1
The 2026-07-28 spec makes MCP fully stateless — no handshakes, no sessions, every request self-describing — enabling serverless deployment behind plain round-robin load balancers (GitHub dropped its Redis session store). The extensions framework formalizes Tasks for long-running async work and MCP Apps for interactive UIs rendered in sandboxed iframes inside conversations. Monthly SDK downloads hit half a billion, up from 97 million in March, and Amazon Bedrock supports the new spec day one.
Ant Group's Ling-3.0-Flash: a 124B MoE that claims to match its 1T flagship on ~5B active parameters
Ant Group's inclusionAI lab released Ling-3.0-Flash, a 124B-parameter MoE activating only ~5.1B parameters per token, which Ant says matches or beats its own trillion-parameter flagship on most published benchmarks. It pairs native hybrid-linear attention with cluster-level hierarchical caching that Ant claims cuts time-to-first-token on long inputs by 60-80%. Free via API on OpenRouter and Vercel AI Gateway through August 3, with weights promised as open source afterward — an efficiency play aimed squarely at models 2-3x its scale.
CoreWeave posts first Vera Rubin results: up to 10x more tokens per megawatt than Blackwell
CoreWeave published the first customer results for NVIDIA's Vera Rubin NVL72 platform, claiming up to 10x more DeepSeek-R1 tokens per megawatt than GB200 at similar interactivity — an efficiency leap that lands directly on the industry's power-constrained bottleneck. Covered on the Jul 23 live show's infrastructure block.
Anthropic signs an AMD compute deal: up to 2GW of MI450/Helios capacity, with up to $5B in AMD equity
Covered on the Jul 23 live show: Anthropic and AMD struck a capacity agreement giving Anthropic access to up to 2 gigawatts of MI450/Helios-generation compute, a deal that includes up to $5 billion in AMD equity — landing the same week AMD's Advancing AI 2026 event launched the Helios/MI400 platform and reports surfaced of separate Meta/Anthropic compute-lease talks (~$10B over two years). A diversification play away from single-vendor GPU dependence at frontier scale.
Google quietly patches Gemma 4 with Flash Attention 4 and tool-calling fixes — no version bump
Google shipped a stealth update to the Gemma 4 family: Flash Attention 4 support on Hopper-class GPUs (a reported 25-70% prefill throughput speedup), tool-calling bug fixes, reduced model 'laziness,' and configurable vision resolution. The ThursdAI panel criticized shipping new weights under the same Gemma 4 name with no version bump, leaving users unsure which checkpoint they're actually running.
PyTorch 2.13 lands FlexAttention on Apple Silicon and big memory wins
3,328 commits from 526 contributors: FlexAttention on Apple Silicon at roughly 12x over SDPA for sparse patterns, a deterministic CUDA backward path, nn.LinearCrossEntropyLoss with up to 4x peak-memory reduction, torchcomms for large-cluster training, and expanded ROCm/Arm/XPU support.
Exo Labs launches local.ai to track the local-AI frontier
Announced live on ThursdAI at AI Engineer: local.ai tracks the best model for your hardware, the performance trade versus the cloud, and whether running local beats API-token pricing. Early access is live with signup codes, and the Exo CLI — 'vLLM for consumer devices, with the configs figured out for you' — ships in the coming weeks.
Together AI raises $800M Series C at an $8.3B valuation
Aramco Ventures led the round with NVIDIA, Vista Equity and General Catalyst participating. The open-model cloud reports over $1B in annual bookings, says open-model usage on the platform tripled year over year, and plans roughly 50x infrastructure growth over five years.
W&B Aria auto-research agent goes GA
Aria went generally available on Monday — an auto-research agent living in the W&B UI ('Just Ask Aria') that reads your traces and debugs your loss curves. In Zubin Aysola's AI Engineer talk, Aria read its own production traces and updated its own prompts.
OpenAI unveils Jalapeno custom inference chip with Broadcom
OpenAI unveiled Jalapeno, its first custom inference ASIC built with Broadcom, positioning it as part of a full-stack strategy to make ChatGPT, Codex, API, and agent workloads cheaper and faster at scale.
Kimi K2.7 Code goes live on W&B/CoreWeave Inference
Kimi K2.7 Code became available on W&B/CoreWeave Inference, with the episode notes calling out Blackwell NVFP4 serving, speculative decoding, and 289 tokens per second near the top of Artificial Analysis speed and price-performance charts.
Weights & Biases launches HiveMind for coding-agent observability
Weights & Biases launched HiveMind, a dashboard for tracking AI coding-agent sessions, spend, transcripts, ROI, and reusable organizational learning. Chris Van Pelt and Adrian Swanberg joined the show to explain why teams need observability for their growing fleet of coding agents.
NVIDIA announces RTX Spark Arm + Blackwell platform for local AI PCs
At Computex, NVIDIA unveiled RTX Spark, an Arm CPU plus Blackwell GPU PC platform with 128GB unified memory targeting local AI agents and 120B-class local inference. A wave of thin laptops with RTX 5070-class GPUs and roughly one petaflop of local AI compute raises the question of what agents should run locally versus in the cloud.
PrismML's 1-bit Bonsai Image 4B runs local image gen under 1GB
PrismML released 1-bit and ternary versions of Bonsai Image 4B, a sub-1GB diffusion transformer for local image generation. The quantized model even runs in-browser via WebGPU and ships with an iOS app and a Hugging Face demo.
Weights & Biases launches MCP server with 20 tools for agents
W&B officially launched its MCP server with 20 schema-first tools so coding agents can read experiments, monitor training, and run autonomous research loops. Agents can query metadata before pulling full 300-metric runs, keeping their context windows from blowing up.
SpaceX IPO filing reveals Anthropic pays $1.25B/month for Colossus compute
The SpaceX IPO filing revealed Anthropic is paying $1.25 billion per month for AI compute at the Memphis Colossus facility. The crew called it a bombastic deal that lets Anthropic serve far more inference at scale and feel less compute-constrained.
CoreWeave Sandboxes launch in preview via the W&B SDK
CoreWeave Sandboxes is now an official Harbor provider, letting teams run agentic workloads like Terminal-Bench safely at scale on CoreWeave infrastructure. It plugs CoreWeave's isolated execution environments directly into the Harbor eval/agent ecosystem.
AWS brings GPT-5.5 and Codex to Bedrock as Azure exclusivity ends
AWS announced GPT-5.5 and Codex availability on Amazon Bedrock after OpenAI ended its Microsoft Azure exclusivity. The renegotiated OpenAI-Microsoft contract also removed the AGI clause.
Stripe opens Projects.dev: 32 infra providers provisionable by agents
Stripe removed the waitlist on Projects.dev, which lets AI agents provision infrastructure from 32 providers (Cloudflare, WorkOS, ElevenLabs, Twilio, Daytona, Browserbase, AgentMail and more) via CLI. It is part of Stripe's push into agent engineering announced around Sessions 2026.
CoreWeave signs Anthropic, Meta ($21B), and Jane Street ($6B + $1B)
CoreWeave announced a multibillion-dollar deal with Anthropic, a $21B expansion with Meta (taking the relationship past $35B total), and a Jane Street deal worth $6B in cloud plus $1B in equity. CoreWeave now serves 9 of the top 10 AI labs, cementing its position as the neocloud backbone of frontier AI.
Gemma 4 goes live on W&B Inference with LoRA inference support
Weights & Biases put Gemma 4 live on W&B Inference, running on CoreWeave infrastructure with LoRA inference support. Replying to the W&B announcement post on X with the code 'Gem Drop' gets $20 in free inference credits.
Anthropic ships Managed Agents, a fully hosted agent runtime
Anthropic launched Managed Agents, a fully hosted agent runtime plus infrastructure offering. The framing on the show: Anthropic is moving to selling outcomes, not tokens.
W&B Automations launch: event triggers from training runs
Weights & Biases shipped Automations, event-triggered actions that pipe signals from your training runs into notifications (Slack), GitHub Actions, and deployments, pairing nicely with the new W&B iOS app. In the same Buzz segment: GLM-5.1 and Gemma 4 both went live on W&B Inference.
OpenAI closes $122B funding round at $852B valuation
OpenAI closed a reported $122 billion funding round, described as the largest in history, at an $852B valuation with an IPO said to be incoming. The panel discussed what that scale of capital implies for AI infrastructure spending, product velocity, and competitive pressure across the market.
PrismML releases Bonsai 1-bit models, an 8B model in 1.15 GB
PrismML released Bonsai, a family of 1-bit quantized open models fitting an 8B model into 1.15 GB and claiming 10x intelligence density, built on decades of compression research. The panel discussed one-bit quantization as a cost/performance lever for cheap local inference.
Google TurboQuant claims 6x KV-cache compression and 8x faster inference
Google Research published TurboQuant, a KV-cache quantization technique claiming 6x compression and 8x inference speedup with near-zero accuracy loss. The panel framed it as a potential unlock for LLM inference economics, while calling stock-market panic over the result premature without broader production validation.
Modular 26.2 runs FLUX.2 in under a second, 99% cheaper than Nano Banana
Modular shipped its 26.2 release with state-of-the-art image generation, running FLUX.2 in under one second (sub-300ms claims) at 99% lower cost than Nano Banana, plus upgraded AI coding with Mojo. Alex noted the surprise of an inference platform releasing model-level optimization and hoped the approach spreads to all image generation.
NVIDIA DLSS 5 adds a generative AI filter for photo-realistic lighting
Announced at GTC, NVIDIA's DLSS 5 introduces a new generative AI filter bringing photo-realistic lighting to RTX 50-series GPUs. It applies generative models to real-time game rendering, extending DLSS beyond upscaling and frame generation.
NVIDIA GTC: GR LPX pairs Rubin NVL72 servers with the new Groq 3 chip
NVIDIA's GTC hardware reveal integrates the new Groq 3 chip (gen 2 was never publicly seen) into Rubin NVL72 servers via the GR LPX system. Claims include 3x tokens-per-watt efficiency at baseline, up to 30x at higher throughput, and 1000+ tokens/sec on a 2T-parameter frontier model with 400K context — performance the current Blackwell generation can't reach at any price.
Weights & Biases launches native iOS app for monitoring training runs
W&B shipped its most-requested feature ever: a native iOS app for monitoring AI training runs with live metrics and push notifications for crash alerts. Practitioners can now keep an eye on long-running training jobs from their phone instead of staying glued to a dashboard.
Google launches Gemini 3.1 Flash-Lite with 1M context at 360 tok/s
Google launched Gemini 3.1 Flash-Lite, a fast and cheap model with 1M token context aimed at the instant/fast tier, running around 360 tokens per second. The panel flagged a material pricing jump versus the prior Flash-Lite generation but saw it as well suited for judge, guardrail, and orchestration workloads in agent systems.
Taalas demos 15,000+ tokens/sec with model weights baked into silicon
Taalas published a live demo (chatjimmy.ai) showing Llama 3 8B running at 15,691 tokens per second on a chip with weights baked directly into the hardware. The panel called it a 10x speed-class jump that points at chip-level innovation compressing inference costs and iteration cycles.
W&B Inference adds MiniMax 2.5 and Kimi K2.5
Weights & Biases added MiniMax M2.5 and Kimi K2.5 to its CoreWeave-backed Inference service. The panel emphasized price/performance, with MiniMax 2.5 presented as roughly 10x cheaper than premium alternatives in some tiers and Kimi K2.5 praised for practical function calling and image-in-loop use cases.
W&B adds Kimi K2.5 to its inference service
Weights & Biases launched Kimi K2.5 on its inference service, making Moonshot AI's model available to W&B users. In Wolfram's Terminal Bench deep dive for W&B, Kimi K2.5 achieved a 67.4% ceiling score across multiple runs, among the strongest open-model results he measured.
OpenAI ships GPT 5.3 Codex Spark on Cerebras for real-time coding
OpenAI released GPT 5.3 Codex Spark, a smaller Codex variant built for real-time coding, served on Cerebras hardware — OpenAI's first model on Cerebras — with reported speeds of over 1000 tokens/sec. Available to ChatGPT Pro users in the Codex app, CLI, and IDE extension. It broke during the show as the second breaking-news drop of the episode.
W&B Inference adds day-zero GLM-5 and Kimi K2.5 support
Weights & Biases launched day-zero GLM-5 support on its CoreWeave-powered W&B Inference service, alongside Kimi K2.5, with MiniMax 2.5 coming soon. Alex announced $50 in free credits for listeners to test the new open-weights models.
OpenAI inks $10B deal with Cerebras for 750MW of high-speed compute
OpenAI announced a $10 billion partnership with Cerebras for 750 megawatts of high-speed inference compute, with capacity starting in 2028. It extends OpenAI's pattern of locking in massive compute supply deals beyond its existing cloud partners.
NVIDIA acquires Groq team and licenses its tech for ~$20B
NVIDIA entered an exclusive licensing deal with Groq and acquired most of its team for approximately $20B. Groq's inference-optimized chips, created by former Google TPU lead Jonathan Ross, complement NVIDIA's training dominance as inference demand grows exponentially across AI use cases.
NVIDIA Vera Rubin platform: 5x Blackwell inference at CES 2026
Jensen Huang unveiled the Vera Rubin platform at CES 2026, NVIDIA's next-gen AI computer delivering 50 PFLOPS and 5x inference performance over Blackwell while adding only ~200W of power draw. It needs 75% fewer GPUs for 10 trillion parameter MoE training, packs 72 GPUs per rack with 20.7TB memory and 13 TB/s bandwidth, is 100% liquid cooled, and entered full production just four months after the B300.
NVIDIA Project Digits: $3,000 desktop that runs 200B-param models
NVIDIA announced Project Digits in January, a $3,000 desktop supercomputer capable of running 200B parameter models locally. It brought serious local-inference hardware to individual developers and was one of January's standout hardware stories.
Project Stargate: $500B AI infrastructure commitment announced
Announced in January, Project Stargate committed $500 billion to AI infrastructure in the US — described on the show as the Manhattan Project for AI. It set the tone for a year in which investment numbers stopped making sense.
GLM 4.5 runs on Cerebras fast enough to win hackathons
Zhipu's GLM 4.5 came out in July and was the first open model that ran on Cerebras hardware fast enough that hackathon competitors were winning with it. It set up GLM's quiet rise as a business workhorse later in the year.
Gemini 3 Flash delivers frontier intelligence at $0.50/1M input tokens
Google launched Gemini 3 Flash, offering frontier-tier capability at flash-tier pricing of $0.50 per million input tokens. It scores 78% on SWE-bench Verified, beating larger models on some agentic tasks, and supports tool-calling at scale with up to 100 simultaneous function calls.
NVIDIA ships Nemotron 3 Nano, a 30B hybrid Mamba-MoE with full recipes
NVIDIA released Nemotron 3 Nano, a 30B-parameter hybrid Mamba-MoE model with only 3B active parameters for efficient inference. The panel called it the most consequential open release of the week because NVIDIA shipped not just weights but technical reports, training recipes, and details on the 25T-token training data.
Pruna P-Image promises sub-second image generation at $0.005
Pruna AI promoted P-Image, an image generation offering with sub-second generation times at roughly $0.005 per image. The release fit the week's diffusion theme of competing on speed and cost efficiency rather than just quality.
W&B launches LLM Evaluation Jobs for OpenAI-compatible APIs
Weights & Biases launched LLM Evaluation Jobs, letting teams run evaluations against any OpenAI-compatible API during training cycles instead of only at the end. The show framed it as a practical workflow upgrade for getting earlier model quality signals without blindly burning compute.
W&B launches Serverless LoRA Inference on CoreWeave
Weights & Biases launched Serverless LoRA Inference on CoreWeave: upload a LoRA adapter to W&B Artifacts and serve it instantly on top of any supported base model with no cold starts and no dedicated GPU instances. Alex demoed a 'Mocking SpongeBob' LoRA he trained in 25 minutes, served on a Qwen 2.5 base.
W&B ships LEET, an open-source terminal UI for monitoring ML runs
Weights & Biases released LEET (Lightweight Experiment Exploration Tool), an open-source terminal-native dashboard for tracking ML runs, demoed live by Dima Duev of the SDK team. It works fully offline for air-gapped HPC clusters and brings real-time metrics, system stats, and zoomable interactive charts to the terminal.
AWS announces multi-year strategic infrastructure partnership with OpenAI
AWS announced a multi-year strategic infrastructure partnership with OpenAI to power ChatGPT inference, training, and agentic AI workloads. It is another sign of OpenAI spreading its compute needs across every major cloud provider, and a notable win for AWS in the frontier-AI infrastructure race.
Sandbar launches Stream voice assistant and Stream Ring wearable
Sandbar launched Stream, a voice-first personal assistant, alongside Stream Ring, a wearable described as a 'mouse for voice' that is now available for preorder. The pairing pushes always-available voice interaction into dedicated hardware rather than the phone.
Claude Haiku 4.5: fast, cheap model rivals Sonnet 4 accuracy
Anthropic released Claude Haiku 4.5, its smallest and fastest current-generation model. The show highlighted that it approaches Sonnet 4 level accuracy at a fraction of the cost and latency, making it attractive for high-volume agentic and production workloads.
Apple announces M5 chip with double the AI performance
Apple unveiled the M5 chip, claiming roughly double the AI performance of the previous generation for Apple Silicon. For local-model enthusiasts on the show, it means more on-device headroom for running and fine-tuning models on Macs.
NVIDIA DGX Spark: a desktop personal supercomputer for local AI
NVIDIA started shipping DGX Spark, a desktop personal AI supercomputer aimed at prototyping and local inference. The show pointed to the LMSYS deep dive on its real-world performance, and Alex shared his own first impressions of the device.
OpenAI and Broadcom to deploy 10 gigawatts of custom AI accelerators
OpenAI announced a strategic collaboration with Broadcom to co-develop and deploy 10 gigawatts of custom AI accelerators. It is another massive compute commitment in OpenAI's infrastructure buildout, this time with chips designed in-house.
OpenPipe Qwen3 14B Instruct lands on W&B Inference
OpenPipe, now part of Weights & Biases / CoreWeave, released a Qwen3 14B instruct model available through W&B Inference. Co-founder Kyle Corbitt joined the show to talk RL, Serverless RL, and practical agent evaluation and deployment.
Nvidia commits up to $100B to OpenAI for 10GW of compute
Nvidia and OpenAI announced a letter of intent under which Nvidia would invest up to $100 billion in OpenAI as the two deploy at least 10 gigawatts of Nvidia systems for OpenAI's next-generation infrastructure. The episode's big-company segment centered on this deal as evidence that money and infrastructure, not just models, now drive the AI race.
Meta Connect: new AI glasses with a display and neural control interface
At Meta Connect, Meta unveiled new AI glasses featuring a built-in display, a neural wristband control interface, and a new AI mode. The panel treats the glasses as an interface milestone, arguing the product surface for AI is shifting from apps to display-equipped wearables.
W&B brings Weave traces into Models workspaces for RL runs
Weights & Biases shipped Weave inside W&B Models workspaces, so reinforcement learning runs can now be logged and inspected with Weave trace tooling alongside training metrics. The show frames it as giving RL training 'x-ray vision' into what the model is actually doing.
Cloudflare launches one-click AI bot blocking for the web
Cloudflare announced a one-click feature letting site owners block AI scraping bots, a direct response to the economics of perpetual web scraping by AI labs. The move puts a default-off switch in front of a large share of the internet and highlights the tension between open research norms and commercial scraping.
Huawei's Pangu Pro MoE: 72B model trained entirely on Ascend NPUs
Huawei released Pangu Pro, a 72B-parameter MoE trained on its own Ascend NPUs rather than Nvidia or AMD hardware, hitting 1,528 tokens/sec and pretrained on 13T tokens. The panel framed it as the geopolitical open-model story of the week, showing how far Chinese compute stacks have advanced under sanctions.
Nous Research launches Psyche, a decentralized cooperative-training network
Psyche is Nous Research's decentralized cooperative-training network that lets distributed participants jointly train large models over the internet. The launch includes open code on GitHub and a live dashboard tracking the first run, a 40B model called Consilience. COO Dillon Rolnick joined the show to explain the decentralized training push.
Falcon-Edge: ternary BitNet LLMs for edge deployment under 1GB VRAM
TII's Falcon-Edge project releases ternary BitNet LLMs (1B and 3B base models) that slash memory and compute requirements, enabling inference on less than 1GB of VRAM. Fine-tuners get pre-quantized checkpoints and a clear path to 1-bit LLMs.
Meta announces the Llama API at LlamaCon, powered by Groq
At LlamaCon, Meta unveiled an official Llama API for developers, with fast inference powered by Groq hardware. Zuckerberg also confirmed Llama thinking models are coming, along with a new meta.ai app with a social feed and a full-duplex voice model in the works.
Google ships Quantization-Aware Trained Gemma 3 models for consumer GPUs
Google released Quantization-Aware Training (QAT) versions of the Gemma 3 family, dramatically cutting memory requirements while preserving quality. The 27B model drops from a hefty 54GB to just 14.1GB, and even the 1B model goes from 2GB to about half a gig, making state-of-the-art open models runnable on consumer GPUs. Wolfram took the 4B QAT model for a spin in LM Studio on the show.
Microsoft releases BitNet 1.58-bit model weights on Hugging Face
Microsoft published BitNet (listed in the show notes as BitNet v1.5), its native 1.58-bit quantized LLM, as open weights on Hugging Face. The ternary-weight approach targets extremely efficient CPU inference at a fraction of the memory of standard models.
CoreWeave hits 800 tok/s on Llama 405B with NVIDIA GB200 Blackwell
CoreWeave announced record-breaking AI inference benchmarks using NVIDIA's new GB200 Grace Blackwell superchips: 800 tokens/sec on Llama 3.1 405B, plus 33,000 tokens/sec on Llama 2 70B with H200s. It is a marker of how fast inference hardware is accelerating.
Arcee AI announces Conductor, an intelligent model router
Arcee AI's Lucas Atkins joined the show to announce Conductor, a model router that picks the best model (including Arcee's small specialized models) for each query. It targets cost and quality optimization by routing requests instead of sending everything to one large model.
Nous Research opens Portal, an inference API for Hermes models
Nous Research launched Portal, its new inference API service offering access to models like Hermes 3 Llama 70B and DeepHermes 3 8B directly via API. It marks another open-source lab standing up hosted API access to make its models more accessible.
Weights & Biases is acquired by CoreWeave
CoreWeave announced it is acquiring Weights & Biases, the AI developer platform and ThursdAI's home company. The deal pairs W&B's experiment tracking, Weave, and models tooling with CoreWeave's AI cloud infrastructure.
DeepSeek open-sources its infra stack during Open Source Week
DeepSeek ran its Open Source Week, releasing a series of production infrastructure repos (including FlashMLA, DeepEP, and DeepGEMM) that power its training and inference stack. The drops gave the open-source community a rare look at the low-level kernels and communication libraries behind DeepSeek's efficient frontier models.
Inception Labs debuts Mercury, a commercial diffusion LLM
Inception Labs announced Mercury, billed as the first commercial-scale diffusion large language model, generating text via diffusion rather than autoregressive decoding. The approach promises dramatically faster token throughput, demoed first with the Mercury Coder playground.
Hao AI Lab's FastVideo makes HunyuanVideo 3x faster with no extra training
Hao AI Lab released FastVideo, a method that makes HunyuanVideo (HY-Video) three times faster with no additional training, using a technique called Sliding Tile Attention that outperforms even flash attention for this workload. Faster inference makes open-source video models far more practical, and it supports HY-Video LoRAs for fine-tuned applications.
Hugging Face publishes the Ultra Scale Playbook for training on GPU clusters
Hugging Face released the Ultra Scale Playbook, a guide to building and scaling AI models on large GPU clusters. The team ran 4,000 scaling experiments on up to 512 GPUs to distill practical guidance for labs training big models.
Microsoft unveils Majorana 1 quantum chip and a new state of matter
Microsoft announced the Majorana 1 quantum chip alongside a claimed new state of matter called topological superconductivity, carving a new path for quantum computing. Alex called the announcement 'absolutely mind blowing' as a potential big deal for the future of computing.
Stargate Project: $500B AI infrastructure investment announced
OpenAI, SoftBank (Masayoshi Son's Vision Fund), and Oracle (Larry Ellison) announced the Stargate Project, a planned $500 billion investment in US AI infrastructure. The announcement, made alongside the White House, was framed on the show as an AI 'Manhattan Project'-scale buildout of datacenters and compute.
Follow Infrastructure & Inference and everything else in AI — live every Thursday.