umans/status/umans-glm-5.3-flash
Live · refreshes every 30s
← all models
Umans GLM 5.3 Flash
umans-glm-5.3-flash · GLM-5.3-Flash · Z.ai
also served as umans-coder
Operational
110.3tok/s
throughput · p50 · last 5 min
1.42s
TTFT · p50 · last 5 min
100.00%
uptime · 24h

GLM-5.3-Flash is Z.ai's fast open-weights model for coding and agentic work: a 320B mixture-of-experts with 18B active parameters per token, the first natively multimodal release in the GLM-5 series with native image and video understanding, on a 1M-token context window. It always thinks at max effort by default; dial reasoning to low or high (thinking cannot be turned off). Billed per token ($0.15 / $0.50 / $0.03 per 1M; input / output / cache read).

90 days agoin production since Sep 9, 2026today
Context
1049K
Max output
131K
Recommended
131K
Vision
Yes
Tools
Yes
Reasoning
Always on
Weights
Trends

Speed over the last 90 days

daily medians · dashed line = target
throughput p50 · output tokens per second, higher is better
90 days agopre-release before Sep 9, 2026today
TTFT p50 · time to first token, lower is better
90 days agopre-release before Sep 9, 2026today
Changelog

Events for Umans GLM 5.3 Flash

incl. gateway-wide announcements
Sep 92026
Released pay-per-token: Umans GLM 5.3 Flash Released
umans-glm-5.3-flash joins the lineup: GLM-5.3-Flash, the checkpoint that served here as the seat-gated pre-release lab since August 26: a 320B mixture-of-experts (18B active per token), the first natively multimodal release in the GLM-5 series with native image and video understanding, on a 1M context window. It always thinks at max effort by default; dial reasoning to low or high (thinking cannot be turned off). Billed per token: $0.15 / $0.50 / $0.03 per 1M (input / output / cache read). The pre-release window's metrics stay on the model's status page as its pre-release period (before Sep 9).