Topic · 11 stories
GLM: Z.ai models, releases and news
In short
GLM is a family of large language models from Z.ai, formerly Zhipu AI, and many of them are released as open-weight models. The newest are GLM-5.3, a coding model from August 2026, and the cheaper GLM-5.3-Flash under an MIT license. On 30 September Anthropic reported that GLM-5.3 builds working cyber exploits nearly as well as Claude Mythos Preview, and that its guardrails are easy to remove.
What is GLM
GLM is a family of large language models from Z.ai, formerly Zhipu AI, a Beijing-based AI lab, and many of its models are released as open-weight. Open-weight means the trained model files are public, so anyone can download, run and modify the model on their own hardware, unlike closed models from OpenAI or Anthropic, which are only reachable through a service. Z.ai builds more than models: in July it launched ZCode, a free desktop coding app.
The company is also a political subject. The U.S. Commerce Department added Zhipu AI to its Entity List on 16 January 2025, according to a SiliconANGLE report from August. The same report asked where enterprise code goes when it flows to a model with no named owner (see US restrictions on Chinese AI models).
What is GLM-5.2
GLM-5.2 is Z.ai’s open-weight coding model from June 2026: a mixture-of-experts model with 744 billion parameters, 40 billion of them active, and a 1-million-token context window. It was released on 16 June, trained on 28.5 trillion tokens using Huawei silicon, and ranked second on Code Arena in mid-June, behind only Claude Fable 5. Its weights are MIT-licensed, according to VentureBeat’s report on the ZCode launch.
The reaction was strong. Vercel CEO Guillermo Rauch called it genuinely impressive at coding, and former Meta VP Matt Velloso said it was the first open model that passes as a daily driver. Running it yourself is hard: the uncompressed weights take 1.51 terabytes, and even heavily quantized it needs roughly 240 GB of memory. That is beyond a standard high-end PC but within reach of a maxed Mac Studio with 256 GB of unified memory. For industries handling sensitive data, local operation keeps data on private hardware.
What are GLM-5.3 and GLM-5.3-Flash
GLM-5.3 is Z.ai’s flagship coding model, unveiled on 14 August 2026, and GLM-5.3-Flash is its cheaper, multimodal sibling from late August. GLM-5.3 uses the same base as GLM-5.2, and all gains come from longer post-training. Zhipu says it is the strongest open-weights coding model, with the biggest gains in agent tasks. It was trained on data and environments built to find software vulnerabilities, and Zhipu reported 2,436 vulnerabilities in 269 projects found with security teams in China, some up to 40 years old. The weights were due to go open source two weeks after security reviews. The stories collected here do not confirm that this has happened.
GLM-5.3-Flash is the first natively multimodal model of the GLM-5 series, with 320 billion parameters, 18 billion active, an MIT license and a 1-million-token context. Artificial Analysis gave it 57 points on its Intelligence Index at maximum reasoning effort, three behind GLM-5.3 and level with GPT-5.6 Terra and Muse Spark 1.2. The cost per task on that index is 0.09 dollars against 0.68 for GLM-5.3, about 7.5 times less. Z.ai says its cost per token on Chinese chips matches typical Nvidia GPUs, and SemiAnalysis sees that as another test of the CUDA moat of Nvidia.
Flash was first a mystery. On 20 August an anonymous model called Ox Alpha appeared free on OpenRouter, and the tool modelprint matched it to GLM-5.3 on six of nine probes, including all four tokenizer counts. OpenCode advertised capacity of 100 trillion tokens a day, and tools such as Claude Code pushed billions of tokens through it, which raised a worry: company code was flowing to a provider nobody could name. On 27 August Z.ai confirmed Ox Alpha was a preview of GLM-5.3-Flash and put the weights on Hugging Face. It had already run an anonymous preview of GLM-5 as Pony Alpha.
Is GLM free and how much does it cost
GLM is partly free: the open weights and the ZCode app cost nothing, but Z.ai’s API and coding subscription are paid, apart from a few small models. According to Z.ai’s official pricing page (as of 10 October 2026, US dollars per million tokens):
| Model | Input | Output |
|---|---|---|
| GLM-5.3-Flash | 0.15 | 0.50 |
| GLM-5.3-FlashX | 0.37 | 1.25 |
| GLM-5.3 | 1.40 | 4.40 |
| GLM-5.2 | 1.40 | 4.40 |
| GLM-5 | 1.00 | 3.20 |
| GLM-4.7 and GLM-4.6 | 0.60 | 2.20 |
The same page lists GLM-4.7-Flash, GLM-4.5-Flash and GLM-4.6V-Flash as free models. For coding there is the GLM Coding Plan with Lite, Pro and Max tiers. Its documentation says “starting at just 18 USD per month” and gives 5-hour credit limits of 2,000, 12,000 and 28,000, with usage in off-peak hours charged at half. In July VentureBeat reported plan prices from 16.20 dollars a month for Lite to 144 dollars for Max, with a promotional 1.5x quota bonus through 31 July; check the plan page for today’s price. The same report said GLM-5.2’s API pricing offered up to 82 percent cost reduction against Claude Opus 4.8.
How to use GLM
You can use GLM through Z.ai’s API, through the GLM Coding Plan inside coding agents, in the free ZCode app, or by running the open weights yourself. Z.ai’s documentation says the Coding Plan works with ZCode, Claude Code, Codex, OpenCode, AutoClaw, OpenClaw and Cline, and that requests for GLM-5.2 and GLM-5.1 are now routed to GLM-5.3. ZCode itself is a desktop app for macOS, Windows and Linux whose agent plans work, edits files, runs checks and iterates, with remote control from mobile and Feishu or WeChat bots. It accepts your own API keys for other models. VentureBeat notes that self-hosting the open weights avoids both U.S. export-control concerns and Chinese data-sovereignty concerns, at the cost of the hardware described above.
How does GLM compare with Western models
GLM sits close to the best Western models on broad benchmarks, but it still trails them on a few measurable fronts. On GDPval-AA v2 GLM-5.3-Flash scored an Elo of about 1,770, matching GLM-5.3 and Grok 4.6 and trailing only Claude Opus 5. A Frontier Radar report says Chinese open-weights models from Moonshot, Alibaba and Z.ai now sit near the top of almost every demanding evaluation, and that the gap of a few months has become an investor problem.
The same report names three areas of Western lead: abstract specialty tests, reliability (a task counts as solved only if a model succeeds in five of five runs) and cybersecurity. Two accusations hang over Chinese labs: distillation of Western models, and tuning for benchmarks, called benchmaxxing. The White House has declared distillation campaigns a national security threat by memorandum. See also Kimi, Qwen and DeepSeek for the other Chinese labs.
Can GLM-5.3 build cyber exploits
Yes: Anthropic’s Frontier Red Team reported that GLM-5.3 can build working exploits nearly as well as Claude Mythos Preview, and that its guardrails can be removed. On ExploitBench it built a working exploit in 50 of 410 attempts, against 56 for Mythos Preview. On Anthropic’s binary exploitation test it achieved a full takeover in 4 percent of trials against 6 percent. Given a sandboxed Linux browser build, it found unknown flaws within a day and chained them into a page that read a visitor’s SSH private key. An abliterated copy complied every time, and NIST called GLM-5.3 the most cyber-capable open-weight model to date. Z.ai’s own measurement is different: 54.4 percent on ExploitBench, more than double GLM-5.2.
The trend was visible earlier. The UK AI Security Institute found that GLM-5.2 and DeepSeek V4-Pro are four to seven months behind closed frontier models in cyber capability. GLM-5.2 matched Opus 4.6 on narrow tasks and cost about 6 dollars per task against roughly 15 for Opus 4.6. The institute judged the open models’ safety measures largely ineffective.
Removing them is now a product. Abliteration.ai sells access to modified models with refusals removed, including abliterated-model-large-v2 based on GLM-5.3, at five dollars per million tokens, and markets it for offensive security and red teaming. See open-weight model security and AI jailbreaks.
What it means for you
- If you want a cheap coding model, GLM-5.3-Flash at 0.15 and 0.50 dollars per million tokens is the entry point. Check the official price page first.
- If you handle sensitive code, self-hosting is the only way to keep it off Z.ai’s servers, and it needs hundreds of gigabytes of memory.
- Do not send code to an anonymous model endpoint. Ox Alpha showed how quickly real work flows to a provider nobody has named.
- If you defend systems, assume open-weight models close to GLM-5.3’s level can write exploits and that guardrails on downloadable weights can be removed.
Still open: whether the main GLM-5.3 weights are fully released, how Z.ai’s own exploit numbers reconcile with Anthropic’s, and how U.S. restrictions will treat Chinese open-weight models.
Key facts
- Anthropic's Frontier Red Team: GLM-5.3 built a working exploit in 50 of 410 ExploitBench attempts, against 56 for Mythos Preview. NIST called it the most cyber-capable open-weight model to date. (source)
- Abliteration.ai sells access to a GLM-5.3 copy with refusals removed, abliterated-model-large-v2, at five dollars per million input or output tokens. (source)
- GLM-5.3-Flash has 320 billion parameters (18 billion active), an MIT license and a 1-million-token context. Z.ai's API charges 0.15 dollars per million input tokens and 0.50 output. (source)
- Z.ai confirmed that the mystery model Ox Alpha, free on OpenRouter and OpenCode, was a preview of GLM-5.3-Flash. Its weights are on Hugging Face. (source)
- Zhipu AI released GLM-5.3 and claims it is the strongest open-weights coding model. With Chinese security teams it reported 2,436 vulnerabilities in 269 projects. (source)
- The UK AISI found GLM-5.2 about four months behind closed frontier models in cyber capability, matching Opus 4.6 on narrow tasks. Its safety measures were largely ineffective. (source)
- Z.ai launched ZCode, a free desktop app for GLM-5.2 on macOS, Windows and Linux with bring-your-own-key support for third-party models. (source)
- GLM-5.2, an open-source model with a 1-million-token context, drew praise from Vercel CEO Guillermo Rauch and ex-Meta VP Matt Velloso for long coding tasks. (source)
This edition was produced with artificial intelligence. Text and voice are generated automatically.
Timeline
- Anthropic Says China’s GLM-5.3 Can Build Working Cyber Exploits AI
- Abliteration.ai Sells Access to Modified AI Models With Safety Guardrails Removed AI
- Z.ai’s GLM-5.3-Flash matches top models at a fraction of the cost, runs without Nvidia AI
- Chinese Lab Z.ai Confirms It Built Mystery Model Ox Alpha AI
- Nobody knows who built AI coding model Ox Alpha or where the code goes AI
- China’s AI lead narrows as Western models keep edge in cyber and reliability AI
- Zhipu AI claims GLM-5.3 is strongest open-weights coding model AI
- Open-weight models now match frontier cyber performance from just four months ago at a fraction of the cost AI
- Z.ai launches ZCode to challenge Cursor, Claude Code, and GitHub Copilot in AI coding AI
- Chinese AI Model GLM 5.2 Runs Locally, Outperforms ChatGPT AI
Show older (1 story)
FAQ
What is GLM 5.2?
GLM-5.2 is Z.ai's open-weight coding model, released on 16 June 2026. It is a mixture-of-experts model with 744 billion parameters (40 billion active) and a 1-million-token context. In mid-June it ranked second on Code Arena, behind only Claude Fable 5.
What is GLM?
GLM is a family of large language models developed by Z.ai, formerly Zhipu AI, a Beijing-based lab. Many are open-weight. The line includes GLM-5.2, GLM-5.3 and the cheaper GLM-5.3-Flash, and Z.ai also makes ZCode, a free coding app.
Is GLM 5.2 free?
Partly. The weights are open, so you can download and run them, but that takes roughly 240 GB of memory even heavily quantized. Z.ai's API charges 1.40 dollars per million input tokens and 4.40 output (as of 10 October 2026). The ZCode app is free.
How do you use GLM 5.2?
Through Z.ai's API, through the GLM Coding Plan inside tools such as Claude Code, OpenCode and ZCode, or by running the open weights on your own hardware. On the Coding Plan, Z.ai now routes requests for GLM-5.2 to GLM-5.3 (as of 10 October 2026).
What is GLM 5?
GLM-5 is the base generation of the series that GLM-5.2, GLM-5.3 and GLM-5.3-Flash belong to. Z.ai previewed it anonymously on OpenRouter as Pony Alpha before release. Its API price is 1 dollar per million input tokens and 3.20 output (as of 10 October 2026).
Can GLM-5.3 build cyber exploits?
Yes, according to Anthropic's Frontier Red Team on 30 September 2026. GLM-5.3 built a working exploit in 50 of 410 ExploitBench attempts, against 56 for Claude Mythos Preview. Its guardrails can be bypassed with simple tricks, and NIST called it the most cyber-capable open-weight model to date.