Geo News
Community curated by people like you
LatestAICryptoHealthWorld AffairsUS Politics
Zhipu AI reveals GLM-5.3-Flash model ran on 100,000 Chinese chips
00

Zhipu AI reveals GLM-5.3-Flash model ran on 100,000 Chinese chips

Aug 26, 2026

Zhipu AI has officially released GLM-5.3-Flash, a highly efficient 320-billion-parameter multimodal model that previously went viral during an anonymous stealth trial as Ox Alpha. Zhipu AI claims the model's massive inference workload was served entirely on a cluster of 100,000 domestic Chinese chips. While this claim of hardware independence remains unverified by third parties, Zhipu AI shares surged 12% in Hong Kong following the announcement. The model features architectural optimizations to run on memory-constrained hardware and is priced aggressively at $0.15 per million input tokens.

GLM-5.3-Flash model release

  • ▪Zhipu AI officially released GLM-5.3-Flash, a natively multimodal mixture-of-experts model with 320 billion total parameters and 18 billion active parameters per token.
  • ▪GLM-5.3-Flash features a 1,048,576-token context window and natively processes text, image, and video inputs.
  • ▪Zhipu AI released the weights for GLM-5.3-Flash globally on Hugging Face under an MIT license on August 26, 2026.

Ox Alpha stealth launch

  • ▪An anonymous model labeled stealth/ox-alpha launched on OpenRouter and OpenCode on August 20, 2026, quickly rising to the top of usage charts.
  • ▪During its stealth preview, the model code-named Ox Alpha processed 62 trillion total tokens, including over 11 trillion tokens in its first three days on OpenRouter.

Chinese chip infrastructure claim

  • ▪Zhipu AI claimed that GLM-5.3-Flash ran entirely on a cluster of 100,000 domestically produced Chinese chips during its stealth trial.
  • ▪Third-party organizations, including CNBC, reported they were unable to independently verify Zhipu AI's claim of running entirely on domestic Chinese hardware.
  • ▪To bypass the lower memory capacity of Chinese chips, Zhipu AI built a custom inference engine that split processing into independent pools for encoding, prefill, and decoding.

Model pricing strategy

  • ▪Zhipu AI offered a 50% promotional discount through September 9, 2026, reducing GLM-5.3-Flash API rates to $0.075 per million input tokens and $0.25 per million output tokens.
  • ▪Zhipu AI priced GLM-5.3-Flash at $0.15 per million input tokens and $0.50 per million output tokens, which is approximately one-tenth the rate of standard GLM-5.3.

Hybrid attention architecture

  • ▪GLM-5.3-Flash utilizes a hybrid attention mechanism that interleaves KDA linear-attention layers with NoPE sparse MLA layers to handle local and global context.
  • ▪GLM-5.3-Flash implements an IndexPool optimization that compresses key vectors, resulting in a 4.4 times smaller KV cache and three times less attention compute.

Community forensic identification

  • ▪By August 22, 2026, AI researchers matched Ox Alpha's tokenizer to Zhipu AI's GLM-5.3 family across 30 probe strings with a constant 75-token offset.
  • ▪A deliberately malformed API call during the preview period exposed a Java stack trace containing Zhipu AI's internal class path, confirming the model's origin.

4 sources

Scmp
Zhipu’s viral Ox Alpha AI model runs entirely on Chinese chips
View source article
Marktechpost
Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context
View source article
Zerohedge
Viral Sensation Ox Alpha Model Revealed As GLM-5.3-Flash, Running Entirely On Chinese Chips
View source article
Techtimes
Ox Alpha Was GLM-5.3-Flash: China Inference Chip Claim Stands Unverified
View source article

Featured stories

View more in AI foundation models

Nvidia unveils NVHBM memory architecture, moving controller into HBM stack to free 25% more compute

Aug 31, 2026 · 3 sources

Anthropic plans IPO with prospectus unveiling after Labor Day

Aug 27, 2026 · 2 sources

Amazon triples Nvidia GPU order to 3 million chips amid surging AI demand

Aug 26, 2026 · 2 sources

Anthropic signs $45 billion AI computing deal with Nscale ahead of IPO

Aug 26, 2026 · 4 sources

Story comments

Loading comments…

Related entities

China

Topics

AI foundation modelsAIMixture of experts (MoE)AGI arms race & geopoliticsCompute, chips & AI infrastructureChina tech & industrial strategy

Featured stories

View more in AI foundation models

Nvidia unveils NVHBM memory architecture, moving controller into HBM stack to free 25% more compute

Aug 31, 2026 · 3 sources

Anthropic plans IPO with prospectus unveiling after Labor Day

Aug 27, 2026 · 2 sources

Amazon triples Nvidia GPU order to 3 million chips amid surging AI demand

Aug 26, 2026 · 2 sources

Anthropic signs $45 billion AI computing deal with Nscale ahead of IPO

Aug 26, 2026 · 4 sources