Agentica AI Releases: ARC-AGI-3 public set result & DeepSWE-Preview

agentica-project.com ↗

ThursdAI — the weekly AI news podcast hosted by Alex Volkov — has covered 2 Agentica releases since Jul 2025, most recently ARC-AGI-3 public set result on Feb 26, 2026. Highlights include DeepSWE-Preview. 1 of them shipped with open weights. Every entry below has primary-source links and the episode segment where we covered it live, plus key numbers where we have them.

2 releases1 open weight2 episodesJul 2025 – Feb 2026

February 2026 1

Agentica
Benchmarks & Evals

ARC-AGI-3 public set result

Agentica claims to solve all public ARC-AGI-3 tasks

Agentica published a claim of solving all public ARC-AGI-3 tasks, adding to the week's theme of benchmark saturation. The panel discussed it alongside METR and ARC-AGI-2 results as part of weighing signal versus noise in headline benchmark leaps.

July 2025 1

Agentica
New ModelsOpen weights

DeepSWE-Preview

DeepSWE-Preview hits 59% SWE-Bench Verified with pure RL on Qwen3-32B

Agentica and collaborators (with guest Michael Luo of UC Berkeley) released DeepSWE-Preview, a fully open-sourced RL-trained coding agent built on Qwen3-32B that reached 59% on SWE-Bench Verified, a top open result in a benchmark dominated by closed systems. The team published training methodology and weights, emphasizing reproducible reward design and verification over sealed benchmark numbers.

59% SWE-Bench Verified