Agents / Z.ai

GLM-5.3-Flash

A multimodal model for coding and agent tasks, with 18B active parameters from a 320B total model.

Not yet testedRepository opened Aug 25, 2026
LOCALCLOUDOPEN WEIGHTS
This is a source-based starting point. We haven’t independently tested this model yet. Source-based estimates are separate from our test results. Measured speed and our verdict will appear after a reproducible test.

At a glance

Parameters
320B total · 18B active
Architecture
Mixture of experts
Context length
See publisher documentation
License
MIT
Recommended VRAM
Not yet tested
Minimum tested VRAM
Not yet tested
Disk space
Not yet tested
Software
vLLM, SGLang, Transformers
Read the official model documentation
HARDWARE ESTIMATE

What will it need?

Our test notes

Not yet tested. We’ll publish the hardware, quantization, context, speed, peak memory, and load time together. A parameter count alone is not a hardware requirement.

Our verdict

Not yet tested. Check the publisher’s documentation for capabilities and limitations; we’ll add an independent verdict after testing.

Explore its uses

AGENTSCODINGMULTIMODALREASONING

Keep exploring

GLM-5.3-Flash: hardware, VRAM & setup | YouRunAI