Models
Z.ai reveals stealth model Ox Alpha as GLM-5.3-Flash and releases MIT-licensed weights
The 320B-parameter mixture-of-experts model activates 18B per token and mixes linear with sparse attention. How it works, what it takes to run, and what to check first.
HackHoster Team · · 11 min read

At a glance
- Z.ai confirmed on August 26, 2026, that the stealth model Ox Alpha is GLM-5.3-Flash and published its weights on Hugging Face under the MIT license.
- The mixture-of-experts model has 320 billion parameters but activates about 18 billion per token, routing each token to 8 of 288 experts plus one shared expert.
- Of its 45 layers, 34 use linear attention with a fixed-size state and 11 use sparse attention, a hybrid Z.ai says is a first for the GLM series.
- Z.ai says the design cuts attention compute by 3.0 times and KV cache by 4.4 times compared with its larger GLM-5.3.
- The FP8 weights take about 306 GiB; vLLM lists 386 GB of GPU memory as the minimum, while Unsloth's smallest 1-bit GGUF is 93 GB.
Z.ai confirmed on Wednesday that Ox Alpha, the anonymous model that OpenRouter had been serving for free since the previous week, is its new GLM-5.3-Flash. The Beijing company, known as Zhipu AI until it rebranded internationally in 2025, published the weights on Hugging Face the same day under the MIT license, which allows commercial use and modification with little more than keeping the copyright notice.
The stealth run worked as a public test. Z.ai's developer guide says it served the model anonymously as ox-alpha on OpenRouter and the coding tool OpenCode to gather feedback, that it quickly became the most popular model of the week, and that all of that traffic ran on Chinese AI chips. SiliconANGLE reported that the free, unattributed listing drew a lot of industry attention and that users soon guessed Z.ai was behind it.
The headline specs, from the model card, the released config file and Z.ai's guide:
| Spec | GLM-5.3-Flash |
|---|---|
| Total parameters | 320 billion (vLLM's recipe counts about 321 billion) |
| Active parameters per token | about 18 billion |
| Layers | 45: 34 linear attention, 11 sparse attention |
| Experts | 288 routed plus 1 shared, 8 routed per token; first 3 layers dense |
| Context window | 1,048,576 tokens |
| Maximum output on Z.ai's API | 128K tokens |
| Inputs on Z.ai's API | text, images, video and files |
| Pre-training data | 30 trillion multimodal tokens |
| Released weights | FP8 (about 328 GB of files) and a separate BF16 copy |
| License | MIT |
Z.ai calls it the first natively multimodal model in the GLM-5 series and claims it beats its own GLM-5.2 across benchmarks at one-tenth the price, while approaching Anthropic's Claude Opus 4.8 on coding and agentic tests. SiliconANGLE summarized the pitch as 10 times more cost-efficient than the previous generation.
Where GLM-5.3-Flash comes from
Z.ai began in Beijing in 2019 as a spin-out of Tsinghua University, with Jie Tang and Li Juanzi as founders, according to Wikipedia. The US Commerce Department added it to its Entity List in January 2025, it began releasing models under the MIT license in July 2025, and it listed on the Hong Kong Stock Exchange in January 2026.

Its recent model line has moved quickly. The GLM-5 repository on GitHub records the steps:
- GLM-4.5 (July 2025) had 355 billion parameters, 32 billion active.
- GLM-5 (February 2026) grew to 744 billion parameters with 40 billion active, trained on 28.5 trillion tokens, and adopted DeepSeek Sparse Attention (DSA) to cut deployment cost at long context.
- GLM-5.2 (June 2026) added what Z.ai calls a solid 1-million-token context and IndexShare, which reuses the same indexer across every four sparse-attention layers and, Z.ai says, cuts per-token compute 2.9 times at a million tokens.
- GLM-5.3 (August 14) kept GLM-5.2's base model and gained everything from post-training, according to the repository.
GLM-5.3-Flash breaks that pattern. It starts from a newly trained base model, with the architecture and training recipe redesigned for efficiency. Z.ai's guide frames the comparison with GLM-4.5: a similar total size (320 billion against 355 billion parameters), but roughly half the active parameters (18 billion against 32 billion) and half the layers (45 against 92).
How the architecture cuts long-context cost
The problem with standard attention
In a standard transformer, every new token is compared with every earlier token. Compute for the whole sequence grows with the square of its length, and the model keeps a key-value (KV) cache entry for every token it has seen. At a million tokens, that cache, not the weights, is often what limits how many requests a GPU can serve at once.

GLM-5.3-Flash attacks the problem from two directions at once, which the model card describes as a first for the GLM series.
Linear attention layers
Thirty-four of the 45 layers use linear attention. The released config names these layers KDA, short for Kimi Delta Attention, a design the team behind the Kimi models published in October 2025. Instead of storing every past token, a linear-attention layer carries a running state of fixed size and updates it as each token arrives, much like a recurrent network. The Kimi Linear paper describes KDA as an extension of Gated DeltaNet with finer-grained gating, so the model can make better use of that limited memory.

The trade-off is precision. A fixed-size state cannot remember everything, which is why the Kimi team, and now Z.ai, mix linear layers with layers that can still look back exactly. In the Kimi team's own 48-billion-parameter test model, the paper reports that a KDA hybrid cut KV cache use by up to 75% and decoded up to six times faster at a million tokens than full attention.
Sparse attention layers
Every fourth layer (11 in total) uses DeepSeek-style sparse attention with a compressed, latent KV cache. These layers keep a cache for every token, but a small indexer picks which earlier positions each query actually reads; the config's top-k setting for that selection is 2,048. To make the indexer itself cheaper at a million tokens, Z.ai's guide describes IndexPool, which merges four indexer key vectors into one by weighted pooling. The guide's division of labor is simple: linear attention handles local dependencies through its state, and sparse attention retrieves the relevant global context.
What the cache costs, roughly. By our arithmetic from config.json, each sparse layer caches a 512-number compressed vector per token. Across 11 layers that is about 11 KB per token in BF16, or around 12 GB for a full million-token context. Each linear layer holds a fixed state of about a million numbers no matter how long the prompt gets. Treat these as estimates; real serving stacks add their own overhead.
Z.ai puts the combined effect at 3.0 times less attention compute and 4.4 times less KV cache than GLM-5.3, measured per head and per layer. It also concedes a limit: compared with Kimi-K3 and DeepSeek-V4-Flash, GLM-5.3-Flash has the lowest attention compute but a slightly larger KV cache.
Experts and residual connections
The feed-forward side is a mixture of experts. Each token goes to 8 of 288 routed experts plus one shared expert, after three dense layers at the bottom of the stack. About 18 billion of the 320 billion parameters do work for any given token. The model also includes one multi-token-prediction layer, which serving stacks use as a built-in draft model for speculative decoding.
Z.ai also adopted Manifold-Constrained Hyper-Connections (mHC), a change to the residual connections between layers. Z.ai says it improves scaling efficiency; SiliconANGLE describes it as reducing the risk that training gradients get distorted on their way through the network.
Vision built in
"Natively multimodal" shows up in the config as a separate vision encoder: 24 layers with a hidden size of 1,024, cutting images into 14-pixel patches. Its output is projected into the language model's 4,096-wide hidden space, and the config groups video frames in pairs along the time axis. The config also defines dedicated start and end tokens for images and for video clips. That matters for builders because a video's cost in tokens depends on how many frames your serving stack samples, so long clips can eat the context window quickly.
The numbers Z.ai reports
These are vendor-reported results from the chart on the model card. Z.ai's footnotes say GDPval-AA v2 scores come from Artificial Analysis, and the model card recommends the default maximum reasoning effort for reproducing them.
| Benchmark | GLM-5.3-Flash | GLM-5.2 | Claude Opus 4.8 | GPT-5.6 Terra | Gemini 3.7 Flash |
|---|---|---|---|---|---|
| Terminal-Bench 2.1 | 84.3 | 81.0 | 85.0 | 87.4 | 85.8 |
| DeepSWE v1.1 | 63.4 | 46.2 | 58.0 | 69.6 | 65.3 |
| Agents' Last Exam | 26.3 | 20.4 | 27.0 | 28.0 | not shown |
| AutomationBench | 48.8 | 26.2 | 41.0 | 37.2 | 52.3 |
| HLE with tools | 55.3 | 54.7 | 57.9 | not shown | not shown |
| GDPval-AA v2 | 1773 | 1504 | 1582 | 1571 | 1527 |
Read the table the way SiliconANGLE did. GLM-5.3-Flash has the top GDPval-AA v2 score in Z.ai's comparison and the second-best AutomationBench score. It beats GLM-5.2 everywhere, by a wide margin on DeepSWE and AutomationBench. On Terminal-Bench, Agents' Last Exam and HLE it trails the closed models by a few points. Z.ai's guide adds an in-house result, 29.0 against 29.5 for Claude Opus 4.8 on its Code Bench at maximum effort, and says the model scores 57 on version 4.1.1 of Artificial Analysis's Intelligence Index at $0.045 per task with discounted pricing.
A caveat on all of these numbers. Every score above was chosen, run or reported by the vendor, sometimes with its own harness and sampling settings. Independent evaluations take time. Until they arrive, treat the table as a claim about what the model can do, and test it on your own tasks.
What self-hosting takes
"18B active" describes compute per token, not memory. Every token can route to different experts, so all 320 billion parameters need to sit in fast memory, or be offloaded to slower system memory at a speed cost. The main checkpoint ships in FP8 and totals about 328 GB of files on Hugging Face; vLLM's recipe counts about 306 GiB of weights before runtime and cache overhead, with a BF16 copy needing roughly twice that.

vLLM's recipe, as of August 27, lists these minimums and options:
| Variant | GPU memory minimum (vLLM) | Notes |
|---|---|---|
| FP8 (default) | 386 GB | verified on H100, B200, GB200 and MI355X |
| NVFP4 (Red Hat AI) | 229 GB | experts in 4-bit; Blackwell GPUs only |
| BF16 | 772 GB | official full-precision copy |
As rough arithmetic, four 80 GB cards come to 320 GB and fall short of the FP8 minimum, so on older hardware plan for a larger node. The recipe's example runs the FP8 model with tensor parallelism of 4 on one GB200 tray, with multi-token-prediction drafting turned on. Two details matter for Hopper owners: the recipe says the FP8 KV cache works only on Blackwell for this model, so H100s must keep the cache in BF16, and at launch it required a dedicated Docker image rather than a standard vLLM release.
A short checklist from the same recipe, for anyone deploying in the first weeks:
- Use FlashInfer 0.6.17 or newer, which the recipe requires for the sparse attention layers.
- Pass the GLM tool-call parser (
glm47) with automatic tool choice, plus the GLM reasoning parser, so tool calls and reasoning text come back as structured fields. - Turn on multi-token-prediction drafting with the recipe's speculative-decoding setting; it is opt-in.
- For heavier traffic, the recipe shows one 8-GPU node split into a prefill pool and a decode pool of four GPUs each, with the linear-attention state layouts pinned identically on both sides.
- On AMD's MI355X, the recipe notes that multi-token-prediction drafting was not supported in the launch image. It also reports what the small cache buys: at tensor parallelism of 4, the FP8 model had room for a KV pool of about 14.9 million tokens, roughly 114 concurrent requests at a 128K context.
The model card also lists SGLang, TokenSpeed, Transformers, KTransformers and Unsloth as supported.

For workstations, Unsloth published GGUF quantizations the day the weights went up. As of August 27 its repository listed a 1-bit file of about 93 GB, a 3-bit version of about 120 GB and a 4-bit version of about 200 GB, plus a separate 1.1 GB vision projector. Running them required Unsloth's own llama.cpp pull request or its desktop app. Expect quality to drop at the lowest bit widths.

The hosted route and its two gotchas
For a weekend project, the API is simpler. Z.ai serves the model as glm-5.3-flash with the full 1-million-token window. Two details in the docs matter:
- Thinking is always on. On Z.ai's API,
thinking.typeonly acceptsenabled. You control the budget withreasoning_effortset tolow,highormax, and it defaults tomax, which spends the most tokens. vLLM's recipe confirms the same behavior for self-hosted serving: any value other thanloworhighfalls back tomax. - Clear old reasoning in chat. The chat template's
clear_thinkingflag defaults tofalse, and the model card says to passtruefor chat use so earlier turns' reasoning is dropped.
Open questions and caveats
- Independent results. Every benchmark above is vendor-reported. Z.ai cites Artificial Analysis for one test and for its Intelligence Index score, but a fuller independent picture will take time.
- Cost claims are relative. "Ten times cheaper" and "one-tenth the price" compare against Z.ai's own previous model and pricing. Your costs depend on context lengths, batch sizes, reasoning effort and hardware, and a model that always thinks at maximum effort by default can spend far more tokens than you expect.
- Long-context quality. Hybrid linear attention saves memory by design. Z.ai says long-context precision is preserved, but builders who rely on exact recall deep in very long documents should run their own needle-in-a-haystack and multi-document tests.
- Multimodal support varies by stack. Z.ai's API accepts video, and the released config defines video tokens, but image and video input in local tools depends on your serving framework and version. Check before you design around it.
- The serving story is partly Chinese hardware. Z.ai says it served the stealth traffic on tens of thousands of domestic accelerators with a custom engine built on SGLang, reaching per-token cost comparable to mainstream Nvidia GPUs after a threefold improvement over its first attempt. That is a notable claim, and it is Z.ai's.
What builders can try
- Long-context agents. The cheap million-token window is the main reason to try this model. Feed it a full repository or a large document set and compare answer quality and cost against a smaller window plus retrieval.
- Vision in the loop. Z.ai pitches the model for visual coding, such as turning screenshots into working front ends and checking rendered output against the reference. That is a natural hackathon demo, and it exercises the multimodal input.
- Effort tuning. Run the same task at
low,highandmaxand log tokens, latency and pass rate. For many simple calls,lowmay be the right default once you have measured it. - Quantization trade-offs. If you have a 128 GB machine, compare the 3-bit GGUF against the hosted model on your own prompts to see what the quantization costs you.
What to watch
As of August 27, the model had been public for one day. The things to watch are concrete: independent benchmark runs, whether the hybrid attention holds up on long-context tests that Z.ai did not choose, how quickly mainstream vLLM, SGLang and llama.cpp releases absorb the new layer types, and whether other labs follow Z.ai and Moonshot in replacing most attention layers with linear ones. Z.ai's guide says it is already scaling the same recipe to larger models.
Sources
- GLM-5.3-Flash model card, config and weights (Hugging Face)
- GLM-5.3-Flash developer guide (Z.ai)
- Z.ai open-sources Ox Alpha model as GLM-5.3-Flash (SiliconANGLE, August 26, 2026)
- GLM-5 series repository and release notes (GitHub)
- GLM-5.3-Flash serving recipe (vLLM)
- GLM-5.3-Flash GGUF quantizations (Unsloth, Hugging Face)
- Kimi Linear, an expressive, efficient attention architecture (arXiv, October 2025)
- Z.ai (Wikipedia)
More from the blog

Models ·
OpenAI begins a staged rollout of GPT-6 Astra, starting with its security program
GPT-6 Astra went to OpenAI's Daybreak security customers first. It is OpenAI's first model rated Critical for cyber capability, and its reasoning is harder to monitor.
12 min read

Models ·
NVIDIA releases Nemotron 3 Ultra, a 550B open-weights model built for long-running agents
NVIDIA published weights, data and recipes for a 550-billion-parameter hybrid Mamba-attention model. It computes like a 55B model but needs memory for all 550B.
11 min read

Models ·
DeepSeek V4 preview ships open weights with a 1M-token context and Huawei Ascend support
DeepSeek released V4-Pro (1.6T parameters, 49B active) and V4-Flash (284B, 13B active) under the MIT license, both with 1M-token context. Huawei says its chips run them.
9 min read