Connect with us

NEWS

Ox Alpha Was a Chinese Chip Stress Test All Along

Ox Alpha was GLM-5.3-Flash, a cheap MIT model Z.ai served on Chinese chips for more than 20 trillion OpenRouter tokens.

Published

on

GLM-5.3-Flash is Z.ai’s 320B open model that scores 57 on an independent index at $0.09 a task. A week ago the lab named the free stealth listing Ox Alpha, and the MIT weights are already on 22 clouds.

The score is close to paid US mid-tier models. The second effect is the one that moves money: you can buy that intelligence from Western hosts, or run the weights yourself, after a preview that Z.ai says ran on Chinese chips.

Ox Alpha Burned More Than 20 Trillion Tokens in Six Days

On August 20, 2026 a model with no lab name appeared on OpenRouter as stealth/ox-alpha. It was free. It advertised a 1M-token window, 131K max output, and image and video input. Hobbyists pushed it hard because the price was zero and the answers were good enough to keep using.

Six days later, on August 26, Z.ai put its name on it. OpenRouter listed z-ai/glm-5.3-flash the same afternoon. The stealth slug vanished with no alias, so anything still pointing at ox-alpha had to move, self-host, or stop.

OpenRouter called it the biggest model the router had ever carried and said the six-day run processed over 20 trillion tokens.

That is not a lab demo in a closed cluster. It is public traffic from coding agents, wrappers, and people who will try any free endpoint that does not immediately fail.

THE STEALTH WEEK

  1. August 20, 2026: stealth/ox-alpha goes live on OpenRouter, free, with a 1M-token window.
  2. August 26, 2026: Z.ai names GLM-5.3-Flash, posts MIT weights, and OpenRouter switches on paid IDs.
  3. September 9, 2026: the launch promo ends at 16:00 UTC, and list rates take over.

Guesses during the week ran from a US lab sitting on spare capacity to a Chinese lab testing silicon in public. Tokenizer traces pointed at the GLM family before the blog post did. The reveal ended the mystery and started the price argument.

Chinese Accelerators Under a Week of Free Traffic

Z.ai did not only claim a new model. It claimed the preview was a hardware run. The launch post says the lab all of this traffic served on Chinese AI chips after testing the model anonymously as ox-alpha on OpenCode and OpenRouter.

Compared with our initial baseline on the same hardware, we achieved a 3× improvement in end-to-end serving performance, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs.

Z.ai, GLM-5.3-Flash launch post

The company has not named the vendor. Treat every specific card count you see as unconfirmed. What Z.ai did publish is the design that made a 1M-token model cheap enough to give away for a week.

GLM-5.3-Flash is a mixture of experts with 320B total parameters and 18B active per token. It starts from a new base, not a trim of GLM-5.3. Versus the GLM-4.5 series it cuts layers from 92 to 45 and active parameters from 32B to 18B. Hybrid sparse and linear attention, plus an IndexPool compressor, cut attention compute by 3.0× and KV cache size by 4.4× against GLM-5.3.

The lab also trained on a 30T-token multimodal corpus and built a dedicated inference engine on SGLang for chips it describes as tight on memory and bandwidth. Encode, prefill, and decode run as separate pools. That is a serving story as much as a model story, and it is why the free week was possible at all.

On Z.ai’s own tables, Flash scores 63.4 on DeepSWE v1.1 against 46.2 for GLM-5.2, and 84.3 on Terminal Bench 2.1 against 85.0 for Claude Opus 4.8. Those are vendor numbers. The independent read is the 57 index score, three points behind full GLM-5.3 at a small fraction of the token price.

Twenty-Two Hosts Now Serve the Same Weights

List price on Z.ai is $0.15 per million input tokens and $0.50 per million output, with cached input at $0.03. A 50 percent discount through September 9 cuts that to $0.075 and $0.25 ($0.015 cached). Full GLM-5.3 stays at $1.40 / $4.40. Flash users on the Coding Plan get 3× the quota of GLM-5.3.

The weights are not locked to Beijing. They are MIT-licensed weights on Hugging Face, with day-one paths through SGLang, vLLM, and TokenSpeed. OpenRouter told developers it expected many providers to onboard during the first week. By September 2 the router’s own table showed 22 inference hosts on OpenRouter, including GMI Cloud, Cloudflare, Fireworks, Together, DigitalOcean, and Baseten.

WHO IS ACTUALLY SERVING FLASH

Host Input / output per 1M 1-day token share
Z.ai $0.075 / $0.25 (promo) 63.3%
GMI Cloud $0.075 / $0.25 (promo) 11.5%
Parasail $0.15 / $0.50 6.1%
Novita $0.075 / $0.25 (promo) 5.2%
Fireworks $0.15 / $0.50 3.6%
Cloudflare $0.15 / $0.50 0.9%

Z.ai still takes most of the routed volume. That is not the same as lock-in. A 36.7% slice already lands somewhere else, and the MIT file means a new host can appear without a contract with Zhipu. For a US buyer, the legal path is now “Chinese weights on a US or EU endpoint,” not “every token must hit a Beijing API.”

That split is the part a 45% workload recipe skips. The model is Chinese. The serving map is not.

How Much Extra Intelligence Costs at the Mid Tier

Artificial Analysis puts GLM-5.3-Flash at 57 on the Intelligence Index and $0.09 per index task at list rates, with a 1M context window. Full GLM-5.3 scores 60 at $0.68 a task. Grok 4.6 scores 61 at $0.84. Later index reads put GPT-5.6 Sol on the same 61-point rung as Grok 4.6.

INTELLIGENCE AGAINST COST PER TASK

Model Index Cost per task Input / output per 1M
GLM-5.3-Flash 57 $0.09 $0.15 / $0.50
GLM-5.3 60 $0.68 $1.40 / $4.40
Grok 4.6 61 $0.84 $2 / $6

Three extra index points on GLM-5.3 cost more than seven times Flash’s $0.09. Four extra points on Grok 4.6 cost more than nine times. Z.ai’s own line is that 57 was a level of intelligence that used to cost about 10× as much per task.

The top of the curve has flattened. Paying for 61 instead of 57 can still be right for a plan you cannot unwind. It is a hard sell for the millionth boilerplate test, the nightly agent loop, or the marketing draft that a junior can rewrite.

This is also not a one-lab accident. OpenRouter rankings in June 2026 had combined US lab token share near 30%, down from about 70% a year earlier, with Chinese-origin labs around 46% of identified volume. DeepSeek, Qwen, MiniMax, Kimi, and now GLM are the default cheap pipe for people who are not already trapped in a seat contract.

Hermes, Cline, and the New Cheap Default

Post-launch traffic on the named ID is already agent traffic, not chat-window curiosity. OpenRouter’s public app board for GLM-5.3-Flash is a list of tools that spend tokens the way a factory spends power.

WHERE THE NAMED TOKENS ARE GOING

  • Hermes Agent: 1.48 trillion tokens, a persistent open agent with tools, memory, and subagents.
  • Claude Code: 1.08 trillion tokens, Anthropic’s coding agent pointed at a cheaper brain.
  • Cline: 752 billion tokens, an in-editor agent that edits, runs commands, and drives a browser.
  • omp: 425 billion tokens on a new app climbing the same board.
  • pi: 270 billion tokens from another personal coding agent.

Those five apps alone account for about 4 trillion tokens on the paid ID. Claude Code routing through Flash is the tell. Teams are already taking an expensive harness and swapping the model underneath it.

Seven days after the unmasking, that swap is showing up in daily use. People bounce overflow off Fable 5.1 agents and onto Flash so the premium model stops spawning copies of itself. Others run a task on DeepSeek V4 Flash, then rerun it on GLM because the first pass did not stick. Flash is becoming the default cheap bar, the model you have to beat, not a curiosity on a leaderboard.

Forty-Two Tokens a Second Is the Hidden Bill

The independent page is blunt: Flash is among the leading models in intelligence for its class, and it is notably slow and somewhat verbose. Output speed is 42.5 tokens per second, 52nd of 111 in that ranking. While running the index it wrote 150 million tokens against a 110 million median.

THE SPEED AND TALK TAX

  • Output speed: 42.5 tokens per second on the independent run, versus 114 tokens per second at Baseten’s OpenRouter route and 24 at Z.ai’s own.
  • Verbosity: 150 million index tokens against a 110 million median, so a “cheap” token can still multiply if the model rambles.
  • Latency spread: OpenRouter shows first-token times from about 0.41 seconds on the fastest host to more than 3 seconds on slower ones.
  • Uptime: 100% host uptime over three days and 99.87% availability, so the drag is speed, not outage.

A slow model with a 1M window is a different product from a fast 8B distill. Long agent jobs wait. Interactive editors feel sticky. If your harness already burns extra reasoning tokens, verbosity taxes you twice: you pay for the chatter, then you wait for it.

That is the honest limit on the “put 45% of volume here” pitch. Flash can take the boring mass of tokens. It should not take the call where a human is staring at the cursor, or the job where a four-point index gap is the difference between a merge and a rollback.

Company Seats Now Have a Price They Can Be Measured Against

American firms were already hitting the wall before Ox Alpha had a name. Uber CTO Praveen Neppalli Naga said in April he was going “back to the drawing board because the budget I thought I would need is blown away already,” after the company’s full-year 2026 coding budget went in four months and he personally burned $1,200 in a two-hour demo. By June, Uber had a $1,500-per-person-per-tool cap. COO Andrew Macdonald still could not draw a line from those dashboards to “25% more useful consumer features.”

A 2026 McKinsey AI survey found 80% of people say they are faster, 37% of companies see some EBIT, and 32% skipped at least one software purchase because agents could build the feature in-house. The tools work. The bill is the fight. Finance will not keep paying frontier rates for volume that a 57-index open model can chew, hosted in Virginia or self-hosted in a rack you already own.

Seat contracts with OpenAI, Anthropic, or xAI are now a sunk cost that has a public comparable. If pay-as-you-go intelligence at $0.09 a task is good enough for the fat part of the distribution, the renewal meeting gets shorter. The remaining spend belongs on the 5% of tasks that are strategy, legal-grade wording, or an irreversible production change, where 61 or 63 still earns the premium.

Labs that cannot get serving cost down will lose that fat part, and they will lose the users who only show up when tokens are cheap. GLM-5.3-Flash is one more Chinese open model on that path, with a twist the last wave did not have: the silicon story and the Western host map arrived on the same day.

The promo clock runs to September 9 at 16:00 UTC. List prices after that are still $0.15 and $0.50. The weights stay MIT. The 22 hosts will not all remain, and some will get faster. The comparison that is not going away is 57 intelligence at $0.09 a task, served on someone else’s chip bill.

Frequently Asked Questions

Can You Run GLM-5.3-Flash on a Laptop?

No. It is a 320B mixture-of-experts checkpoint meant for datacenter GPUs, with official paths through SGLang, vLLM, and TokenSpeed, plus community setups such as KTransformers. A laptop build is the wrong shape; budget for a multi-GPU box or a hosted route.

Is GLM-5.3-Flash a Distilled Copy of GLM-5.3?

No. Z.ai trained a new base and cut the stack for cheap serving, 45 layers and 18B active against 92 layers and 32B active in the GLM-4.5 series. Full GLM-5.3 is a separate 753B / 40B reasoning model at $1.40 / $4.40 whose open weights were delayed for safety work.

Did Z.ai Name the Chinese Chips Behind Ox Alpha?

No. The company says a large cluster of Chinese AI accelerators carried the preview and that per-token cost is now comparable to mainstream NVIDIA GPUs, but it has not published a vendor, a SKU, or an official card count.

Does the MIT License Allow Commercial Use of the Weights?

Yes. The Hugging Face card lists MIT with no extra commercial rider, which allows use, modification, and redistribution, including in paid products, without a separate license from Z.ai, subject only to MIT’s notice rules.

Harry is the editor of COVER 365, an independent publication he owns and runs, and a journalist of ten years who moved from reporting into editing. Anything the site reviews has been used before it is judged. A phone, a car, a game or a piece of travel gear is tested in ordinary conditions, its measured results are set against the maker's specification sheet, and where the two disagree the article says which one to trust and why. No product gets a verdict Harry has not earned by using it. Off the test bench, the same rule of primary evidence applies: business stories come from filings and results, science from the published paper, sports from the governing body's records, and news from statements and transcripts rather than second hand accounts. Coverage runs across technology, auto, gaming, lifestyle and travel as well as news, business, science, sports and entertainment, for readers in every part of the world. Every figure is checked before publication and corrected publicly under a stated policy when wrong. Reader mail is answered at support@cover365.in.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending