Z.ai GLM-5.3-Flash
★ 4/5 · Open Source Model

GLM-5.3-Flash pairs native multimodal input with a one-million-token context window and MIT-licensed open weights, aimed at coding and long-running agent work at flash-tier prices.
Pros
- Near-frontier capability at flash-tier cost: a 320B-parameter mixture-of-experts that activates only 18B per token, accepting text, images, video and files across a one-million-token context window and returning up to 128K output tokens, with function calling, thinking mode, streaming, context caching and structured output all supported. The weights are published on Hugging Face under the MIT licence, so the model can be self-hosted, fine-tuned or shipped inside a commercial product without negotiating a licence, and it runs on SGLang, vLLM, Transformers and other common inference stacks. Z.ai's published evaluations put it at 63.4 on DeepSWE v1.1 against 46.2 for GLM-5.2 and 48.8 on AutomationBench against 26.2, with 84.3 on Terminal-Bench 2.1 on the model card, and its hybrid sparse-and-linear attention design is built specifically to hold long-context accuracy while cutting the cost of serving those long contexts.
Cons
- Provenance is the thing to know before you trust it with anything sensitive. The model spent its preview period on OpenRouter under the anonymous alias "Ox Alpha", where the operator was described only as a third-party provider who had chosen to remain anonymous, and OpenRouter's stealth terms state plainly that prompts and completions are retained by the provider, so anything sent during that window is held by a company users had not knowingly picked. That company has since been confirmed as Z.ai (Zhipu), a Chinese lab, though Z.ai states it generally provides the services from Singapore. The position today is materially better but worth checking yourself rather than assuming: Z.ai's privacy policy says API information is processed in real time and is not saved on its servers, and that the policy does not cover content processed on behalf of business customers, while the same document lists training and improving its models among the legitimate interests for personal data it does hold, and the Data Processing Addendum that actually governs API customers is not published at a public URL. Beyond that, the launch price is promotional and doubles when it lapses in September 2026, the headline benchmark figures are vendor-published rather than independently audited, output is text-only despite multimodal input, and self-hosting 320B weights is a serious hardware commitment even at 18B active.
Cost
Promotional launch pricing of $0.075/M input, $0.015/M cached input and $0.25/M output runs until 24:00 on 9 September 2026 (UTC+8), after which Z.ai's list price of $0.15/M input, $0.03/M cached input and $0.50/M output applies, with cached-input storage free for a limited period. OpenRouter resells the same model at $0.07125/M input, $0.01425/M cached input and $0.2375/M output with a 1,310,720-token context and 131,072-token maximum output, and also lists a discounted batch endpoint. The weights are free under the MIT licence, so self-hosting costs are hardware only, and Z.ai offers GLM-4.7-Flash and GLM-4.5-Flash at no charge for lighter text-only work.
Verdict
Best for teams that want near-frontier coding and agent performance at roughly a twentieth of flagship API prices, and for anyone who values the fallback of pulling MIT-licensed weights in-house rather than being locked to a hosted endpoint. Price in the September promotional expiry, treat the vendor benchmarks as vendor benchmarks, and read Z.ai's own API data terms before sending anything confidential rather than assuming the retention position that applied while the model was being served anonymously as Ox Alpha.
Visit Z.ai GLM-5.3-Flash