EmbeddingGemma 2
EmbeddingGemma 2, an open multimodal embedding model
Google released EmbeddingGemma 2, an open multimodal embedding model, the same week Perplexity opened pplx-embed-v2.
Search products, deep research, RAG, embeddings, and retrieval systems. — 31 releases covered on the show.
EmbeddingGemma 2, an open multimodal embedding model
Google released EmbeddingGemma 2, an open multimodal embedding model, the same week Perplexity opened pplx-embed-v2.
Perplexity opens pplx-embed-v2 multimodal embedders
Perplexity released pplx-embed-v2, open multimodal embedding models on Hugging Face, with a write-up on multimodal embeddings beyond a single vector.
Chroma launches Foundation: unified memory for your agents
Launched live during the show (founder Jeff Huber hopped on within minutes of the announcement), Foundation is a research preview of a shared memory system between you and your agents — memory as infrastructure. It ingests sources natively from Codex, Claude Code, Cursor and Slack on day one, with Notion, GitHub and Google Drive connectors coming, and is built on ChromaDB plus the Context-1 agentic search model (a GPT-OSS 20B fine-tune running ~400 tokens/sec at 25x less cost than Opus). Each Foundation manages and improves its own system prompt from natural-language feedback. Part of Chroma Cloud starting at $30/mo.
Google Search adds Gemini 3.5 Flash-powered agentic capabilities
Google Search is getting new Gemini 3.5 Flash-powered agentic capabilities, including a new AI-powered Search box and background information agents. The crew framed the rollout as a massive intelligence uplift across one of Google's largest surfaces, with billions of Search users getting frontier-model capabilities.
Google launches Gemini Embedding 2, a natively multimodal embedder
Google launched Gemini Embedding 2, a natively multimodal embedding model that supports text, image, video, and audio in a single unified embedding space. It is available through the Gemini Embeddings API.
Karpathy open-sources AutoResearcher for autonomous ML experiments
Andrej Karpathy open-sourced AutoResearch, a framework that runs AI-driven ML experiments autonomously. Over two days it ran 700 experiments on nanochat GPT-2, stacked 20 improvements, and achieved an 11% training speedup. Tobi Lütke adapted it overnight for Shopify's Liquid templating engine for a 51% render-time improvement, and the repo hit 26K GitHub stars quickly.
Mixbread embed-large-v3 beats Gemini Embedding 2
mixbread.ai dropped embed-large-v3, an embedding model that beats Gemini Embedding 2 on nearly every benchmark, including a jaw-dropping 98% vs 6.9% on structured-data tasks. Benjamin Clavie announced it live during the show.
Perplexity launches pplx-embed SOTA embedding models
Perplexity released pplx-embed, a family of state-of-the-art embedding models built for web-scale retrieval. The models are available on Hugging Face and through Perplexity's API with quickstart docs.
xAI silently drops Grok 4.20 with four 500B-param collaborating agents
xAI released Grok 4.20, a multi-agent system where four 500B-parameter agents collaborate in a multi-agent UI, with a $300/month Heavy tier scaling to 16 agents. No benchmarks or evals were released with the drop. The panel found it underwhelming for coding and day-to-day agent work but still top tier for deep research thanks to xAI's RAG over X data; Grok 4.1 Fast remains #8 on OpenRouter by API usage.
MiroThinker 1.5: 30B search agent beats trillion-param models
MiroMind AI released MiroThinker 1.5, a 30B parameter open source search agent that achieves 56.1% on BrowseComp and 66.8% on BrowseComp Chinese, outperforming trillion-parameter models. It introduces 'interactive scaling' as a third scaling dimension beyond parameters and context, and is a fine-tune of Qwen 3 Thinking with 147K open training samples.
DeepSeek-OCR turns text into compressed vision tokens for massive contexts
DeepSeek open-sourced DeepSeek-OCR, a 3B model (~570M active parameters) that is less an OCR model and more a context-compression breakthrough: it renders text as images, compresses it up to 10x while retaining 97% decoding accuracy (60% even at 20x), and reads it back with a tiny vision decoder. The approach suggests text tokenization is far from optimal and points at vastly cheaper long-context processing; alphaXiv reportedly OCR'd all of arXiv for $1000 versus $7500 with MistralOCR, and a single H100 can process up to 200K pages.
PokeeResearch-7B: open-source SOTA deep research agent model
Pokee AI released PokeeResearch-7B, an open-source 7B deep research agent model claiming state-of-the-art results for its size. Weights, code, a paper, and a hosted deep-research preview all shipped together.
Tongyi DeepResearch: open-source A3B web agent rivals OpenAI Deep Research
Alibaba's Tongyi Lab open-sourced Tongyi DeepResearch, a 30B mixture-of-experts web research agent with only 3B active parameters. The lab claims parity with OpenAI's Deep Research on agentic search and report-writing tasks, and the weights are available on Hugging Face.
Alibaba's Tongyi Lab open-sources WebWatcher vision-language research agent
Alibaba's Tongyi Lab open-sourced WebWatcher, a vision-language deep research agent that sets new state-of-the-art results on agentic browsing and research tasks. The 32B model combines visual understanding with web research capabilities and is available on Hugging Face.
Google releases EmbeddingGemma, a 300M-param SOTA embedding model for RAG
Google released EmbeddingGemma, a 300M-parameter open embedding model that achieves state-of-the-art results for its size, aimed at RAG and on-device semantic search. It dropped as breaking news during the show, with browser-based demos like Semantic Galaxy showing it running fully client-side.
Mistral ships new state-of-the-art embedding API
Mistral announced a new state-of-the-art embedding API. The release gives developers a SOTA option for retrieval and semantic search workloads served through Mistral's platform.
Anthropic launches Web Search API for real-time retrieval in Claude
Anthropic released a Web Search API that gives Claude models real-time web retrieval, letting developers ground responses in current information directly through the API. It was covered among the week's big-company API updates.
ChatGPT adds shopping capabilities
OpenAI rolled out shopping features in ChatGPT, letting the assistant find and recommend products for users. Mentioned briefly in the big-companies roundup amid the week's OpenAI sycophancy drama.
Cohere Embed 4: multimodal embeddings for enterprise search
Cohere released Embed 4, a multimodal embedding model aimed at enterprise search and retrieval over mixed text and image documents. It is available through Cohere's API.
Jina Reranker M0: SOTA multilingual, multimodal document reranker
Jina AI released Jina Reranker M0, a state-of-the-art multimodal and multilingual document reranker model. It reranks documents that include both text and images, targeting retrieval and RAG pipelines, with weights available on Hugging Face.
Nomic Embed Multimodal: SOTA embeddings for visual documents
Nomic AI released Nomic Embed Multimodal, new 3B and 7B parameter embedding models built on Alibaba's Qwen2.5-VL. They achieve SOTA on visual document retrieval by embedding interleaved text-image sequences, ideal for PDFs and complex webpages. The 7B model ships under Apache 2.0 with open weights, code, and data; guest Zach Nussbaum discussed the release on the show.
Google makes Deep Research free, adds Canvas and Live Previews to Gemini
Google made its Deep Research agent free for Gemini users and shipped Canvas, a collaborative workspace with live previews for code and documents. Demos on the show included a playable Tetris game and a markdown word counter built and previewed directly inside Gemini.
EuroBERT: multilingual encoder models from 210M to 2.1B parameters
EuroBERT is a new family of multilingual encoder models ranging from 210M to 2.1B parameters, trained on a 5 trillion-token dataset across 15 languages with 8K context support. It targets European and global language NLP tasks like retrieval and RAG, where properly encoding non-English character sets matters.
Google makes Deep Research free in the Gemini app, powered by Gemini Thinking
Google made its Deep Research agent free for everyone in the Gemini app and upgraded it to run on Gemini Thinking. In a live test on the show it browsed over 150 websites to compile a comprehensive answer, with a polished interface and export to Google Docs.
Manus AI research agent has everyone talking
Manus is a new AI research agent (manus.im) that creates a to-do list, browses the web in a real Chrome browser, and generates files, described on the show as 'Operator on steroids' and seemingly powered by Claude 3.7 behind the scenes. The crew tested it live on a research task and praised its slick UI.
Google announces AI Mode in Search powered by Gemini 2.0
Google announced AI Mode, a new conversational search experience in Google Search, alongside Gemini 2.0-powered upgrades to AI Overviews. Robby Stein, VP of Product for Google Search, joined the show for an exclusive interview about the launch, which brings full AI chat-style answers with follow-ups directly into Search.
xAI launches DeepSearch, an agentic research feature with live X access
Alongside Grok 3, xAI launched DeepSearch, an agentic deep-research feature comparable to Perplexity or OpenAI's Deep Research, with a leg up on real-time information thanks to native access to X search. Alex's initial tests were underwhelming, nicknaming it 'Shallow Search' after it spent 34 seconds on a query where OpenAI's Deep Research took 11 minutes and cited 17 sources.
Exa ships free DeepSeek R1 chat demo with web search
Exa integrated DeepSeek R1 into a free hosted chat demo that combines the reasoning model with Exa's web search. Mentioned in the tools section as a no-cost way to try R1 grounded with live search results.
Perplexity adds DeepSeek R1 as a Pro reasoning model option
Perplexity integrated DeepSeek R1 into its Pro search product, letting subscribers choose R1 as the reasoning model behind answers. It was one of several tools that raced to host R1 on Western infrastructure within days of the model's release.
Anthropic adds Citations to the Claude API
Anthropic launched a Citations capability in the Claude API, letting Claude ground its answers in provided source documents and return precise citations. It targets RAG and document-QA use cases where verifiable sourcing matters.
Perplexity ships Sonar Pro search API and an Android AI assistant
Perplexity released its Sonar Pro search-grounded API, giving developers programmatic access to Perplexity-style web-grounded answers, and also launched an AI assistant for Android. Two shipping moves that push Perplexity beyond its consumer answer engine.
Follow Search & Retrieval and everything else in AI — live every Thursday.