aisumate

Z.ai GLM-5.3-Flash

★ 4/5 · Open Source Model

Z.ai GLM-5.3-Flash homepage

GLM-5.3-Flash pairs native multimodal input with a one-million-token context window and MIT-licensed open weights, aimed at coding and long-running agent work at flash-tier prices.

Pros

Cons

Cost

Promotional launch pricing of $0.075/M input, $0.015/M cached input and $0.25/M output runs until 24:00 on 9 September 2026 (UTC+8), after which Z.ai's list price of $0.15/M input, $0.03/M cached input and $0.50/M output applies, with cached-input storage free for a limited period. OpenRouter resells the same model at $0.07125/M input, $0.01425/M cached input and $0.2375/M output with a 1,310,720-token context and 131,072-token maximum output, and also lists a discounted batch endpoint. The weights are free under the MIT licence, so self-hosting costs are hardware only, and Z.ai offers GLM-4.7-Flash and GLM-4.5-Flash at no charge for lighter text-only work.

Verdict

Best for teams that want near-frontier coding and agent performance at roughly a twentieth of flagship API prices, and for anyone who values the fallback of pulling MIT-licensed weights in-house rather than being locked to a hosted endpoint. Price in the September promotional expiry, treat the vendor benchmarks as vendor benchmarks, and read Z.ai's own API data terms before sending anything confidential rather than assuming the retention position that applied while the model was being served anonymously as Ox Alpha.

Visit Z.ai GLM-5.3-Flash