Tag
2 articles
Zhipu AI's GLM-5.3-Flash matches top models in performance at a fraction of the cost and runs without Nvidia hardware.
Z.ai has released GLM-5.3-Flash, a 320B-parameter, 18B-active MoE model with a 1M-token context window and native multimodal capabilities. It features a 3x reduction in attention compute and 4.4x in KV cache usage compared to previous versions.