← all models
Umans GLM 5.3 Flash (lab) Playground Retired
umans-glm-5.3-flash-lab · GLM
playground experiment; the GLM-5.3-Flash pre-release window ran Aug 26 to Sep 8, 2026; its metrics carry on the released model's status page as its pre-release period
retired Sep 8, 2026 · replaced by umans-glm-5.3-flash · weights ↗
Retired
120.3tok/s
throughput · p50 · whole period
1.59s
TTFT · p50 · whole period
91.67%
uptime · whole period

GLM-5.3-Flash as a Labs experiment, open for a short test window: temporary, not a permanent id. Served from the native weights (Z.ai's GLM-5.3-Flash): a 320B mixture-of-experts model with 18B active parameters per token, the first natively multimodal release in the GLM-5 series, built for fast coding and agentic tasks on a 1M context window. It always thinks at max effort by default; dial reasoning to low or high (thinking cannot be turned off). Access is seat-gated through the Labs page while an experiment is live. It is offered at limited capacity and low availability, so expect it to be flaky and to go down under load: crash it, give it a moment, and try again. For production work we recommend umans-coder or umans-deepseek-v4-pro-0813.

Aug 26, 2026retired Sep 8, 2026Sep 10, 2026
Trends

Speed over its final 90 days

daily medians · dashed line = target
throughput p50 · output tokens per second, higher is better
Aug 26, 2026retired Sep 8, 2026Sep 10, 2026
TTFT p50 · time to first token, lower is better
Aug 26, 2026retired Sep 8, 2026Sep 10, 2026
Changelog

Events for Umans GLM 5.3 Flash (lab)

incl. gateway-wide announcements
Sep 92026
Playground closed: Umans GLM 5.3 Flash (pre-release) Testing
The GLM-5.3-Flash pre-release window closed on September 8, 2026 (opened August 26 and extended past the announced August 28 close). Thanks to everyone who pushed it - text, images, and the thinking dials - and shared findings. The experiment answered its framing question: the model earns a permanent slot, so the seat-gated experiment on umans-glm-5.3-flash-lab ended and the model keeps serving as the pay-per-token umans-glm-5.3-flash: the model stays.
Aug 262026
New Labs experiment: Umans GLM 5.3 Flash Testing
A new lab opened on umans-glm-5.3-flash-lab: GLM-5.3-Flash served from its native weights, a 320B mixture-of-experts model (18B active per token) and the first natively multimodal release in the GLM-5 series, with native image and video understanding and a 1M context window. It always thinks at max effort by default (dial to low or high; thinking cannot be turned off). Free and seat-gated while the experiment runs, served at limited capacity and low availability: it is a lab, so expect it to be flaky and to go down under load. Crash it, give it a moment, and try again. The window closes August 28, 2026.