On 26 August 2026, Z.ai claimed authorship of the model that had appeared anonymously on OpenRouter days earlier as Ox Alpha. GLM-5.3-Flash carries 320B parameters, 18B of them active, and ships with open weights under an MIT licence.
Key takeaways
- Revealed 26 August 2026 after anonymous testing as Ox Alpha
- 320B total parameters, 18B active per request
- Context window up to 300K tokens
- Priced at $0.15 per million input tokens and $0.50 per million output
- Open weights under the MIT licence
The anonymous debut had a second layer
A few days ago a coding model turned up on OpenRouter with no stated author. We covered Ox Alpha then as an industry puzzle. The puzzle is now solved: it was a public dress rehearsal for GLM-5.3-Flash.
There is practical sense to that launch mode. An anonymous model gathers feedback without brand baggage and without the expectations attached to a particular lab.
Two points lower, seven times cheaper
On the Artificial Analysis intelligence index the model scores 57 at roughly 9 cents per task. GPT-5.6 Sol scores 59 but costs 67 cents, and Grok 4.6 reaches 61 for 94 cents.
| Model | Score | Cost per task |
|---|---|---|
| GLM-5.3-Flash | 57 | ≈ 9 cents |
| GPT-5.6 Sol | 59 | 67 cents |
| Grok 4.6 | 61 | 94 cents |
The quality gap is small, the price gap sevenfold. Among task results, terminal work stands out most, with 84.3 on Terminal Bench 2.1.
Sparse and linear attention in one
The architecture combines sparse attention with linear attention and adds an mHC?mHC: Manifold-Constrained Hyper-Connections — a mechanism that lowers the cost of serving long context. mechanism that cuts the cost of serving long context. It is the first natively multimodal model in the GLM-5 line, trained on a 30T-token corpus. According to Z.ai, the whole thing was served on Chinese chips and Chinese infrastructure.
Why it matters
Open weights under MIT at this level of performance change the maths for teams that have been paying for closed-model APIs. The point is not catching the frontier on benchmarks — it is that the quality gap stops justifying the price gap. Once a near-frontier model can run in-house, the case for an external vendor shrinks to convenience.
What’s next?
- An OpenRouter promotion halves the rates until 9 September 2026
- MIT open weights allow self-hosting with no licence restrictions
- Z.ai, GMI Cloud and Cloudflare are already serving the model
Sources
- VentureBeat — GLM-5.3-Flash will likely handle 45% of your AI workloads
- Hugging Face — zai-org/GLM-5.3-Flash





