Z.ai's GLM-5.3-Flash ran entirely on Chinese chips during preview
The anonymous Ox Alpha model that topped OpenRouter usage charts was served on a 100,000-chip domestic cluster before its official release.
Chinese chipmaker reveals infrastructure behind viral AI model
Z.ai confirmed Wednesday that Ox Alpha, an anonymous AI model that dominated OpenRouter usage charts in recent weeks, was its own creation—and that every inference request during the preview period ran on domestically produced Chinese chips.
The company released the model publicly as GLM-5.3-Flash and published its weights on HuggingFace. According to details first reported by Bloomberg and CNBC, Z.ai deployed a cluster of 100,000 Chinese-made AI processors to serve the model, though the company did not identify specific chipmakers. Counterpoint Senior Research Analyst Ivan Lam told CNBC the hardware likely includes Huawei Ascend processors combined with components from other domestic vendors.
GLM-5.3-Flash carries 320 billion total parameters with 18 billion active at any given time. Z.ai priced the model at $0.15 per million input tokens and $0.50 per million output tokens, positioning it alongside DeepSeek in the low-cost tier. The company also listed a per-task price of $0.045, which it said represents roughly one-tenth the cost of comparable models.
Performance and positioning
The model is the first in Z.ai's GLM-5 series to handle multimodal input, including text, images, and video. On coding and agentic benchmarks, Z.ai said GLM-5.3-Flash approaches the performance of Anthropic's Claude Opus 4.8. The Artificial Analysis Intelligence Index placed it 10th overall, ahead of DeepSeek V4 Pro Max.
Z.ai built a custom inference engine on top of the SGLang framework and achieved what it described as a 3x improvement in serving performance over its initial baseline on the same hardware. The company said the optimization brought efficiency and per-token cost to levels comparable with mainstream Nvidia GPUs.
Why it matters
The disclosure demonstrates that Chinese AI companies can deploy competitive large language models at scale using domestic chip infrastructure, despite U.S. export restrictions on advanced semiconductors. Z.ai's decision to run the entire Ox Alpha preview on Chinese processors—and publicize that fact—signals growing confidence in homegrown hardware capabilities. For enterprises evaluating AI vendors, the announcement raises questions about supply chain resilience and the durability of technology export controls as a strategic lever.
Anonymous launch strategy
Z.ai named the preview version Ox Alpha after "Niu Lai," a Chinese film whose title translates to "Ox Comes," Bloomberg reported. The anonymous rollout mirrored tactics Alibaba and Xiaomi each used this year, releasing models without attribution to gather unbiased user feedback before claiming ownership.
The strategy worked. Ox Alpha became the most popular model of the week on OpenRouter before Z.ai stepped forward to identify itself as the developer. Z.ai stock climbed more than 8% in Hong Kong trading Thursday following the announcement. Shares have risen more than 800% since the company's Hong Kong listing in January.
Z.ai made GLM-5.3-Flash available to all GLM Coding Plan subscribers and released the model weights publicly. The launch precedes the company's first detailed earnings report covering the first half of the year, scheduled for Monday.
Details were first reported by Bloomberg and CNBC.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call
