Zhipu AI (GLM) Releases: GLM-5.3-Flash, GLM-5.3 & ZCode

z.ai ↗

Zhipu AI's coverage is anchored by the GLM model line — mostly open weights — plus CogView image models. ThursdAI — the weekly AI news podcast hosted by Alex Volkov — has covered 13 Zhipu AI (GLM) releases since Mar 2025, most recently GLM-5.3-Flash on Aug 27, 2026. Highlights include GLM-5.3, GLM-5.2, GLM-4.7-Flash, CogView 4 (6B). 10 of them shipped with open weights. Every entry below has the episode segment where we covered it live, plus primary-source links and key numbers where we have them.

13 releases10 open weights12 episodesMar 2025 – Aug 2026

August 2026 2

Z.ai
New ModelsOpen weights

GLM-5.3-Flash

GLM-5.3-Flash: the OX Alpha mystery model, open-sourced under MIT

Z.ai open-sourced GLM-5.3-Flash, a 320B-parameter MoE with 18B active under MIT license, after stealth-testing it for about six days as 'OX Alpha' with effectively unlimited free traffic on OpenRouter — all served on Chinese chips. Company-reported DeepSWE is 63.4 with Claude Opus 4.8-level coding claims, it's natively multimodal, and a hybrid sparse/linear attention architecture cuts KV cache size 4x versus GLM 5.3 with 3x serving performance.

320B-A18B parameters (total / active), MIT license63.4 DeepSWE (company-reported)4x smaller KV cache vs GLM 5.3
Z.ai
New Models

GLM-5.3

GLM-5.3: post-training alone delivers a 6x Terminal-Bench jump

Z.ai announced GLM-5.3, keeping the same 743B base as GLM 5.2 but jumping from 4.6 to 28.3 on Terminal-Bench 3 and gaining almost 20% on DeepSWE from post-training alone, at unchanged pricing with a 1M context window. It scores 60 on the Artificial Analysis Intelligence Index, roughly Kimi K3 level with about a third of the parameters, and shows emergent cybersecurity capabilities: 84% on CyberGym and 54.5 on ExploitGym, beating GPT-5.6 Sol. API-only for now — weights (and license terms) expected later.

4.6→28.3 Terminal-Bench 3, a 6x jump from post-training alone60 Artificial Analysis Intelligence Index84% CyberGym

July 2026 1

Z.ai
Dev Tools

ZCode

Z.ai launches ZCode, a GLM-5.2 agentic coding environment

ZCode is an agentic coding environment built on GLM-5.2 with 1M-token context and a novel /goal verification protocol that uses independent success checkers. Output reaches 173 tokens/second with 1.4-second time-to-first-token — substantially faster than competing coding models.

173 tokens/second output1M token context

June 2026 1

Z.ai (Zhipu AI)
New ModelsOpen weights

GLM-5.2

Z.ai releases GLM-5.2, a 753B open MoE with 1M context

Z.ai released GLM-5.2 as a major open-source coding and agentic model: a 753B-parameter MoE, MIT-licensed, with a one-million-token context window. The episode treated it as the open-source model that arrived exactly as Fable access disappeared, with strong coding and agentic performance close to the frontier.

753B parameters1M context windowMIT license

April 2026 1

Z.ai (Zhipu AI)
New ModelsOpen weights

GLM-5.1

GLM-5.1 takes #1 open-source spot on SWE-Bench Pro at 58.4%

Z.ai released GLM-5.1, now the #1 open-source model on SWE-Bench Pro at 58.4%. It can run autonomously for 8 hours with 1,700+ agent steps, and is already live on W&B Inference. Open weights are up on Hugging Face alongside an arXiv paper.

February 2026 2

Zhipu AI (Z.ai)
New ModelsOpen weights

GLM-5

Z.ai launches GLM-5, the open-weights agentic coding crown

Z.ai released GLM-5, a 744B-parameter MoE model (40B active) trained on 28.5 trillion tokens that takes the #1 open-source ranking for agentic coding with 77.8% SWE-bench Verified. It introduces the SLIM asynchronous RL framework for post-training, adopts DeepSeek's sparse attention to cut deployment cost, and was trained on Huawei chips rather than NVIDIA. Lou from Z.ai joined the show live and summed it up as bigger, faster, better, and cheaper.

744B GLM-5 Parameters28.5T Training tokens

January 2026 2

Z.AI (Zhipu)
New ModelsOpen weights

GLM-4.7-Flash

GLM-4.7-Flash: 30B MoE local coding agent with only 3B active params

Z.AI released GLM-4.7-Flash, a 30B parameter MoE model with only 3B active parameters, designed as the ultimate local coding and agent assistant. It hits 59% on SWE-Bench Verified (approaching Sonnet 4's 64%) and runs at 120 tokens/sec on a stock Mac Studio M3 Ultra, fast enough to run RALF autonomous coding loops even on CPU.

59% SWE-Bench Verified120 tps Speed on Mac Studio M3 Ultra

December 2025 2

Zhipu AI (GLM)
New ModelsOpen weights

GLM 4.5

GLM 4.5 runs on Cerebras fast enough to win hackathons

Zhipu's GLM 4.5 came out in July and was the first open model that ran on Cerebras hardware fast enough that hackathon competitors were winning with it. It set up GLM's quiet rise as a business workhorse later in the year.

Zhipu AI (GLM)
New ModelsOpen weights

GLM 4.6

GLM 4.6 quietly becomes the model businesses actually use

Zhipu's GLM 4.6 arrived in October and, per Nisten, quietly became a go-to model that many businesses still run today. It continued GLM's trajectory from hackathon favorite to production workhorse.

April 2025 1

March 2025 1