GLM-5.3-FlashFrontier Intelligence, Open Weights
GLM-5.3-Flash is the open-weights frontier model from Z.ai, released August 26, 2026 under the MIT license. It matches Claude Opus 4.8 on Artificial Analysis' Intelligence Index with a 1,310,720-token context window, native text, image, and video input, and a 320B-parameter MoE that activates only 18B. Try it on Felo AI Search, or call it through the Felo API.
Available on Felo AI Search and the Felo API
Model card
GLM-5.3-Flash at a glance
- License
- MIT · Open weights
- Parameters
- 320B total · 18B active
- Context window
- 1.31M tokens
- Max output
- 48K tokens
- Input
- Text · Image · Video
57
Intelligence Index
Artificial Analysis · level with Claude Opus 4.8
1M
Context Window
1,310,720 tokens input
48K
Max Output
48,000 tokens per response
MIT
License
MIT · open weights
What Sets the Architecture Apart
Z.ai rebuilt the GLM-5 series around a hybrid attention design, making front-tier intelligence practical to serve.
Hybrid Sparse + Linear Attention
The first open frontier model to combine sparse and linear attention. A lightweight IndexPool keeps the index cache at a quarter of the size and Manifold-Constrained Hyper-Connections (mHC) boost scaling — roughly 3× less attention compute and 4.4× smaller KV cache than GLM-5.3, while staying precise on long context.
Open Weights, MIT License
Full weights published under MIT. Self-host it, fine-tune it, or deploy it through Z.ai or any of the 10+ providers that host it on OpenRouter.
First Natively Multimodal GLM-5
Trained on a 30T-token multimodal corpus, it reads text, images, and video directly — no separate vision pipeline, no stitching several models together.
Frontier Work, Flash Performance
A 320B-parameter MoE that activates only 18B per token gets frontier-level work done at agent speed.
Level With Claude Opus 4.8
Artificial Analysis Intelligence Index: 57 for GLM-5.3-Flash, 57 for Claude Opus 4.8. Agentic Index: 58, above Claude Opus 4.8, Claude Fable 5, and GPT-5.6 Sol.
Built for Coding and Agents
Z.ai Code Bench (max effort): 29.0, versus Claude Opus 4.8's 29.5. DeepSWE v1.1: 63.4 — 17.2 points above GLM-5.2.
One Request for Text, Images, Video
A 1,310,720-token context window with native multimodal input. Chartography: 78.0 versus DeepSeek-V4-Flash-Vision-Exp's 64.3; MVBench: 77.8 versus 69.4.
Sub-Second Responses
Providers on OpenRouter measured a 0.55s median first-token time and peak throughput around 134 tokens/s, with 10+ providers to choose from.
GLM-5.3-Flash vs. Top Frontier Models
Official Z.ai model-card results: GLM-5.3-Flash against GLM-5.2, DeepSeek-V4-Vision-Exp, Opus 4.8, GPT-5.6 Terra, and Gemini 3.7 Flash, August 2026. Higher is better — bold marks the best score in each row.
Benchmark
GLM-5.3-Flash
GLM-5.2
DeepSeek-V4-Vision-Exp
Opus 4.8
GPT-5.6 Terra
Gemini 3.7 Flash
Coding
Terminal Bench 2.1
84.3
81.0
83.9
85.0
87.4
85.8
DeepSWE v1.1
63.4
46.2
59.3
58.0
69.6
65.3
NL2Repo
56.3
48.9
57.7
69.7
-
-
Agentic
Toolathlon Verified
78.4
59.9
75.9
76.2
74.9
-
AutomationBench v1.0.6
48.8
26.2
38.8
41.0
37.2
52.3
Agents' Last Exam
26.3
20.4
27.3
27.0
28.0
-
HLE w/ Tools
55.3
54.7
55.1
57.9
-
-
GDPval-AA v2
1773
1504
1675
1582
1571
1527
Vision
OfficeQA Pro
62.4
-
57.9
48.9
-
-
CharXiv Reasoning w/ Tools
89.4
-
80.4
89.9
88.0
88.7
Chartography w/ Tools
78.0
-
64.3
75.0
68.0
65.0
BabyVision
53.4
-
35.1
46.8
61.6
70.9
MVbench
77.8
-
69.4
67.1
75.0
82.2
MMVU
80.5
-
72.7
67.4
75.8
82.3
Source: Z.ai official model card, August 2026. Self-reported; results may vary by evaluation setup.
Z.ai Code Bench (max effort): 29.0 vs. Claude Opus 4.8's 29.5 · Artificial Analysis Agentic Index: 58, above Claude Opus 4.8
Read the Z.ai announcementGLM-5.3-Flash in Third-Party Rankings
Frontend design skill on DesignArena and value on the Artificial Analysis Pareto frontier.

DesignArena — Overall Frontend (Non-Agentic)
GLM-5.3-Flash lands at 1348 and 1345 on DesignArena's frontend Elo — 5th and 6th overall, behind Kimi K3, GPT-5.6 Sol, and Claude Opus 5, and ahead of GLM-5.2, Gemini 3.7 Flash, and Claude Sonnet 5.
Source: DesignArena by Intelligence · Elo Rating · All Models

Artificial Analysis — Cost-Performance Pareto Frontier
57 points on the Intelligence Index at $0.045 per task — GLM-5.3-Flash sits on the Pareto frontier with GPT-5.6 Sol (max) and Claude Opus 5 (max), at a fraction of their cost.
Source: Artificial Analysis, updated August 26, 2026
Technical Specifications
The details that matter when you use GLM-5.3-Flash in your own workflows.
Context & Output
1,310,720 tokens input (1M+ context)
48,000 tokens max output
Architecture
320B total / 18B active Mixture-of-Experts, 45 layers. Hybrid sparse + linear attention with IndexPool and Manifold-Constrained Hyper-Connections (mHC).
License & Weights
MIT-licensed open weights, trained on a 30T-token multimodal corpus.
Speed
Median first-token latency of 0.55s and peak throughput around 134 tokens/s across OpenRouter providers (measured August 2026).
Availability
Live on Felo AI Search and in the Felo API, plus the Z.ai platform and 10+ OpenRouter providers.
Modalities
Native text, image, and video input with text output — the first natively multimodal model in the GLM-5 series.
One Model, Every Input Type
No separate pipelines or stitched models — GLM-5.3-Flash reads text, images, and video natively.
Text & Code
Parses long documents, codebases, and structured data in a single pass — 63.4 on DeepSWE v1.1, 84.3 on Terminal-Bench 2.1.
Images & Charts
Feed screenshots, charts, and diagrams to get grounded answers: Chartography 78.0 vs. DeepSeek-V4-Flash-Vision-Exp's 64.3; OfficeQA-Pro 62.4% vs. Claude Opus 4.8's 48.9%.
Video
Analyze video alongside text and images: MVBench 77.8%, ahead of DeepSeek-V4-Flash-Vision-Exp's 69.4%.
Who Uses GLM-5.3-Flash
From solo developers to teams running their own inference, GLM-5.3-Flash covers coding, research, and multimodal work on one open model.
Agentic Coding
63.4 on DeepSWE v1.1 and 48.8 on AutomationBench v1.0.6 make it a strong pick for coding agents, multi-file edits, and long call chains.
Long-Horizon Agent Work
Toolathlon 78.4 plus a 1.31M-token window keep a single session going across many steps without dropping context.
Multimodal Research
Combine documents, screenshots, charts, and video frames in one request to extract structured answers with sources.
Video & Visual Analysis
MVBench 77.8 and Chartography 78.0 cover video understanding, chart reading, and document layout analysis.
Self-Hosted Deployment
MIT license and 18B active parameters mean you can run the weights yourself for data-sensitive or fine-tuned workflows.
Felo Agents & Pipelines
Call it through the Felo API to build agents, automations, and machine-readable tool pipelines.
Two Ways to Use GLM-5.3-Flash
Use It on Felo AI Search
Open Felo AI Search to ask questions, search the web with AI, and get cited answers from GLM-5.3-Flash.
Open Felo AI SearchCall the Felo API
Get a Felo API key, set the base URL, and call the model from Claude Code, Codex, Hermes Agent, or your own OpenAI-compatible agent.
View the Felo APIFrequently Asked Questions
GLM-5.3-Flash is Z.ai's open-weights frontier model, released on August 26, 2026. It is a 320B-parameter Mixture-of-Experts model with 18B active parameters, a 1,310,720-token context window, native text, image, and video input, and hybrid sparse-plus-linear attention. It is licensed under MIT.
Start With GLM-5.3-Flash on Felo
The open frontier model from Z.ai — 1.31M-token context, native multimodal input, and agent-class coding in one place.
Open GLM-5.3-Flash on Felo AI SearchAlso on the Z.ai platform and the Felo API