Zhipu AI has officially released GLM-5.3-Flash, a highly efficient 320-billion-parameter multimodal model that previously went viral during an anonymous stealth trial as Ox Alpha. Zhipu AI claims the model's massive inference workload was served entirely on a cluster of 100,000 domestic Chinese chips. While this claim of hardware independence remains unverified by third parties, Zhipu AI shares surged 12% in Hong Kong following the announcement. The model features architectural optimizations to run on memory-constrained hardware and is priced aggressively at $0.15 per million input tokens.
Aug 31, 2026 · 3 sources
Aug 27, 2026 · 2 sources
Aug 26, 2026 · 2 sources
Aug 26, 2026 · 4 sources
Story comments
Loading comments…