Z.ai revealed that the stealth Ox Alpha preview was GLM 5.3 Flash and released its weights under the MIT license. The vendor describes a natively multimodal mixture-of-experts model with 320 billion total parameters, 18 billion active per token and a one-million-token context window.
The host's eight-test King Bench retest scored the official API version at 63 out of 80, or 78.75 percent, compared with 70 out of 80 during the stealth preview. Reasoning, math and an end-to-end fine-tuning pipeline remained strong, while several visual and front-end generations were less polished. The host notes that run-to-run variance or serving differences could explain the gap.
For local use, total parameters determine memory capacity while active parameters dominate per-token compute. A four-bit quantization may require about 180 GB, while more aggressive dynamic quantizations could approach 100 GB. That puts the model within reach of high-memory unified systems but not ordinary consumer machines.
The video also repeats vendor claims about the scale of the stealth beta, low API pricing and deployment on Chinese AI accelerators. Those claims are not independently verified here. The practical case is narrower: GLM 5.3 Flash offers open weights, low active compute and competitive agentic performance for users who can provide enough memory.
Watch on YouTube


