GLM 5.3 Flash has 320 billion total parameters but activates 18 billion for each token. Its mixture-of-experts architecture aims to provide broad model capacity without applying every parameter during each inference step, and it supports a context window of up to one million tokens.
The released weights use the MIT license, which permits local deployment subject to the license terms. Running the model still requires substantial memory, with practical requirements depending on quantization, context length, runtime overhead and serving software.


