August 27th.ZhipuLast night on line andOpen Source GLM-5.3-Flash (320B-A18B), the first birth of the GLM-5 seriesMultimodal ModelI DON'T KNOW. THE TOTAL SIZE OF THE MODEL IS 320 B (320 BILLION) AND ONLY 18 B (18 BILLION) IS ACTIVATED。

Before it was officially released, a large-scale test was carried out on OpenCode and OpenRouter with an anonymous model, Ox-Alpha. Ox-Alpha quickly became the most popular model of the week, creating a new high number of double platform calls。And all of these requests are supported by national chips.

The GLM-5.3-Flash capacity exceeds the GLM-5.2, achieving 57 points in the global authority Artificial Analysis Inteligence Index (AA-Composite Smart Index) and entering the global front-line model capability range, which is the same as the most popular model of Anthropic, Claude Opus 4.8. In the self-study Z.ai Code Bench sensitisation assessment, the programming performance was comparable to Claude Opus 4.8。
At the same time, the GLM-5.3-Flash price is 1/10 for GLM-5.3, 1/20 for GLM-5.3 and 1/40 for Opus 4.8。
Officially, the GLM-5.3-Flash price is only 1/10 of the GLM-5.2, which is twice as high as the parameter in all baseline tests and actual use。
It was described as benefiting from its entirely new model structure. The GLM-5.3-Flash architecture is designed for very low cost, and the total parameter volume corresponds to the GLM-4.5 (355B vs. 320B), but the activated parameter volume (32B → 18B) and the layers (92 → 45) are almost halved. GLM-5.3-Flash, combined with the latest 30T Token multi-modular pre-training language, can achieve greater performance with fewer computing resources。
The GLM-5.3-Flash has been accompanied by a number of structural upgrades, the first open-source front-line model to introduce a thin and linear mix of attention, which, while maintaining its ability to keep a precise long context, significantly reduces the cost of long context services; and the use of the Fluid-Constrained Hyper-Connects, mHC to upgrade the Scaling capacity。
Officially, the GLM-5.3-Flash visual abilities are incorporated into the Coding cycle, enabling the model to observe the interface, render results and interactive feedback, and to use them for continuous testing and improvement. Whether it's a front-end development, game construction, a Blender 3D scenario, or a real environmental operation driven by BUA, CUA, the model works in tandem with the code, browser and graphical interface。
GLM-5.3-Flash further expands to work areas such as Office, financial research and professional documentation. It is able to break down complex targets, call appropriate tools, check and optimize outputs and complete the full workflow from research analysis, modelling to PPTX, PDF, DOCX, XLSX product delivery。
GLM-5.3-Flash is officially open to the global open source, has access to coding platforms such as ZCode and incorporates the GLM Coding Plan (a limited 10,000 experience cards are issued daily) to synchronize the API call。
In addition, the entire GLM Code Plan post has been replaced, and all user backstages are fully utilized。
1AI WITH RELEVANT LINKS:
- BigModel Open Platform: https://docs.bigmodel.cn/cn/guide/models/vl/glm-5.3-flash
- Z.ai: https://docs.z.ai/guides/vlm/glm-5.3-flash
- GMCing Plan: https://bigmodel.cn/glm-coding
- HuggingFace: https://huggingface.co/zai-org/GLM-53-Flash