MIT License · Open Weights · Released Aug 26, 2026

GLM-5.3-FlashFrontier Intelligence, Open Weights

GLM-5.3-Flash is the open-weights frontier model from Z.ai, released August 26, 2026 under the MIT license. It matches Claude Opus 4.8 on Artificial Analysis' Intelligence Index with a 1,310,720-token context window, native text, image, and video input, and a 320B-parameter MoE that activates only 18B. Try it on Felo AI Search, or call it through the Felo API.

Available on Felo AI Search and the Felo API

Model card

GLM-5.3-Flash at a glance

License
MIT · Open weights
Parameters
320B total · 18B active
Context window
1.31M tokens
Max output
48K tokens
Input
Text · Image · Video
Hybrid sparse + linear attention: about 3× less attention compute and 4.4× smaller KV cache than GLM-5.3.

57

Intelligence Index

Artificial Analysis · level with Claude Opus 4.8

1M

Context Window

1,310,720 tokens input

48K

Max Output

48,000 tokens per response

MIT

License

MIT · open weights

What Sets the Architecture Apart

Z.ai rebuilt the GLM-5 series around a hybrid attention design, making front-tier intelligence practical to serve.

Hybrid Sparse + Linear Attention

The first open frontier model to combine sparse and linear attention. A lightweight IndexPool keeps the index cache at a quarter of the size and Manifold-Constrained Hyper-Connections (mHC) boost scaling — roughly 3× less attention compute and 4.4× smaller KV cache than GLM-5.3, while staying precise on long context.

Open Weights, MIT License

Full weights published under MIT. Self-host it, fine-tune it, or deploy it through Z.ai or any of the 10+ providers that host it on OpenRouter.

First Natively Multimodal GLM-5

Trained on a 30T-token multimodal corpus, it reads text, images, and video directly — no separate vision pipeline, no stitching several models together.

Frontier Work, Flash Performance

A 320B-parameter MoE that activates only 18B per token gets frontier-level work done at agent speed.

Level With Claude Opus 4.8

Artificial Analysis Intelligence Index: 57 for GLM-5.3-Flash, 57 for Claude Opus 4.8. Agentic Index: 58, above Claude Opus 4.8, Claude Fable 5, and GPT-5.6 Sol.

Built for Coding and Agents

Z.ai Code Bench (max effort): 29.0, versus Claude Opus 4.8's 29.5. DeepSWE v1.1: 63.4 — 17.2 points above GLM-5.2.

One Request for Text, Images, Video

A 1,310,720-token context window with native multimodal input. Chartography: 78.0 versus DeepSeek-V4-Flash-Vision-Exp's 64.3; MVBench: 77.8 versus 69.4.

Sub-Second Responses

Providers on OpenRouter measured a 0.55s median first-token time and peak throughput around 134 tokens/s, with 10+ providers to choose from.

Official Model Card Benchmarks

GLM-5.3-Flash vs. Top Frontier Models

Official Z.ai model-card results: GLM-5.3-Flash against GLM-5.2, DeepSeek-V4-Vision-Exp, Opus 4.8, GPT-5.6 Terra, and Gemini 3.7 Flash, August 2026. Higher is better — bold marks the best score in each row.

Benchmark

GLM-5.3-Flash

GLM-5.2

DeepSeek-V4-Vision-Exp

Opus 4.8

GPT-5.6 Terra

Gemini 3.7 Flash

Coding

Terminal Bench 2.1

84.3

81.0

83.9

85.0

87.4

85.8

DeepSWE v1.1

63.4

46.2

59.3

58.0

69.6

65.3

NL2Repo

56.3

48.9

57.7

69.7

-

-

Agentic

Toolathlon Verified

78.4

59.9

75.9

76.2

74.9

-

AutomationBench v1.0.6

48.8

26.2

38.8

41.0

37.2

52.3

Agents' Last Exam

26.3

20.4

27.3

27.0

28.0

-

HLE w/ Tools

55.3

54.7

55.1

57.9

-

-

GDPval-AA v2

1773

1504

1675

1582

1571

1527

Vision

OfficeQA Pro

62.4

-

57.9

48.9

-

-

CharXiv Reasoning w/ Tools

89.4

-

80.4

89.9

88.0

88.7

Chartography w/ Tools

78.0

-

64.3

75.0

68.0

65.0

BabyVision

53.4

-

35.1

46.8

61.6

70.9

MVbench

77.8

-

69.4

67.1

75.0

82.2

MMVU

80.5

-

72.7

67.4

75.8

82.3

Source: Z.ai official model card, August 2026. Self-reported; results may vary by evaluation setup.

Z.ai Code Bench (max effort): 29.0 vs. Claude Opus 4.8's 29.5 · Artificial Analysis Agentic Index: 58, above Claude Opus 4.8

Read the Z.ai announcement
Independent Leaderboards

GLM-5.3-Flash in Third-Party Rankings

Frontend design skill on DesignArena and value on the Artificial Analysis Pareto frontier.

DesignArena — Overall Frontend (Non-Agentic)

DesignArena — Overall Frontend (Non-Agentic)

GLM-5.3-Flash lands at 1348 and 1345 on DesignArena's frontend Elo — 5th and 6th overall, behind Kimi K3, GPT-5.6 Sol, and Claude Opus 5, and ahead of GLM-5.2, Gemini 3.7 Flash, and Claude Sonnet 5.

Source: DesignArena by Intelligence · Elo Rating · All Models

Artificial Analysis — Cost-Performance Pareto Frontier

Artificial Analysis — Cost-Performance Pareto Frontier

57 points on the Intelligence Index at $0.045 per task — GLM-5.3-Flash sits on the Pareto frontier with GPT-5.6 Sol (max) and Claude Opus 5 (max), at a fraction of their cost.

Source: Artificial Analysis, updated August 26, 2026

Technical Specifications

The details that matter when you use GLM-5.3-Flash in your own workflows.

Context & Output

1,310,720 tokens input (1M+ context)

48,000 tokens max output

Architecture

320B total / 18B active Mixture-of-Experts, 45 layers. Hybrid sparse + linear attention with IndexPool and Manifold-Constrained Hyper-Connections (mHC).

License & Weights

MIT-licensed open weights, trained on a 30T-token multimodal corpus.

Speed

Median first-token latency of 0.55s and peak throughput around 134 tokens/s across OpenRouter providers (measured August 2026).

Availability

Live on Felo AI Search and in the Felo API, plus the Z.ai platform and 10+ OpenRouter providers.

Modalities

Native text, image, and video input with text output — the first natively multimodal model in the GLM-5 series.

One Model, Every Input Type

No separate pipelines or stitched models — GLM-5.3-Flash reads text, images, and video natively.

Text & Code

Parses long documents, codebases, and structured data in a single pass — 63.4 on DeepSWE v1.1, 84.3 on Terminal-Bench 2.1.

Images & Charts

Feed screenshots, charts, and diagrams to get grounded answers: Chartography 78.0 vs. DeepSeek-V4-Flash-Vision-Exp's 64.3; OfficeQA-Pro 62.4% vs. Claude Opus 4.8's 48.9%.

Video

Analyze video alongside text and images: MVBench 77.8%, ahead of DeepSeek-V4-Flash-Vision-Exp's 69.4%.

Who Uses GLM-5.3-Flash

From solo developers to teams running their own inference, GLM-5.3-Flash covers coding, research, and multimodal work on one open model.

Agentic Coding

63.4 on DeepSWE v1.1 and 48.8 on AutomationBench v1.0.6 make it a strong pick for coding agents, multi-file edits, and long call chains.

Long-Horizon Agent Work

Toolathlon 78.4 plus a 1.31M-token window keep a single session going across many steps without dropping context.

Multimodal Research

Combine documents, screenshots, charts, and video frames in one request to extract structured answers with sources.

Video & Visual Analysis

MVBench 77.8 and Chartography 78.0 cover video understanding, chart reading, and document layout analysis.

Self-Hosted Deployment

MIT license and 18B active parameters mean you can run the weights yourself for data-sensitive or fine-tuned workflows.

Felo Agents & Pipelines

Call it through the Felo API to build agents, automations, and machine-readable tool pipelines.

Two Ways to Use GLM-5.3-Flash

Use It on Felo AI Search

Open Felo AI Search to ask questions, search the web with AI, and get cited answers from GLM-5.3-Flash.

Open Felo AI Search

Call the Felo API

Get a Felo API key, set the base URL, and call the model from Claude Code, Codex, Hermes Agent, or your own OpenAI-compatible agent.

View the Felo API

Frequently Asked Questions

GLM-5.3-Flash is Z.ai's open-weights frontier model, released on August 26, 2026. It is a 320B-parameter Mixture-of-Experts model with 18B active parameters, a 1,310,720-token context window, native text, image, and video input, and hybrid sparse-plus-linear attention. It is licensed under MIT.

Start With GLM-5.3-Flash on Felo

The open frontier model from Z.ai — 1.31M-token context, native multimodal input, and agent-class coding in one place.

Open GLM-5.3-Flash on Felo AI Search

Also on the Z.ai platform and the Felo API