
Ox Alpha hits OpenRouter as a free stealth reasoning model with a 100T tokens/day claim
Attribution remains unresolved as observers split between a Z.ai/GLM-5 lineage and a Microsoft MAI-family hypothesis.
An anonymous provider has listed a “stealth” reasoning model called Ox Alpha on OpenRouter and is offering it free for a limited window. The rollout is drawing attention for a claimed 100 trillion tokens per day of capacity even as the model’s origin remains unverified.
Ox Alpha appeared on OpenRouter on Thursday as a “stealth model” from an anonymous third-party provider. OpenRouter is a model-routing platform that lets developers send requests to multiple models through one interface, which makes a new listing immediately usable across existing apps without a bespoke integration.
OpenRouter describes Ox Alpha as a “reasoning model designed for coding, sustained agentic work, and production workloads,” and says it is suited for “long-horizon software engineering, complex reasoning, and workflows that combine text with visual context.” “Agentic work” here means multi-step workflows where a system plans and executes actions toward a goal, often by calling tools, rather than answering a single prompt.
The model launched into the market with a distribution hook that matters more than the benchmark chatter: it is free, at least temporarily. OpenCode, an open-source AI coding agent, said on X that Ox Alpha would be free for a week with “near unlimited usage,” framing the rollout as a time-limited access window rather than a permanent price point.
Early validation arrived fast. Stripe CEO Patrick Collison posted on X after trying the model that “it's very impressive.” That kind of endorsement does not verify lineage, but it does increase the odds developers keep routing traffic to it long enough for clearer fingerprints to emerge.
Attribution Whiplash: Z.ai/GLM-5 vs Microsoft MAI, and What Evidence Is There
The identity debate has moved faster than the evidence. Early speculation centered on a Chinese AI lab, with Z.ai, the lab behind GLM-5, floated as a candidate. The mechanism behind that theory is familiar: developers look at response patterns and tokenizer behavior, where a tokenizer is the text-splitting component that can sometimes leak family resemblance across model lines.
That Z.ai theory leaned on precedent as much as on technical tells. GLM-5 was previously tested anonymously under the name “Pony Alpha,” and observers pointed to perceived similarities between Ox Alpha and GLM-family behavior as a reason to suspect another stealth rollout.
The catch is that the same kind of tokenizer-based attribution can cut the other way. A competing analysis cited in the discussion suggested Ox Alpha’s tokenizer could instead point toward Microsoft’s MAI family, which would flip the narrative from “China lab stealth drop” to “US incumbent testing distribution.” As of Saturday, the public conversation was already cooling on any single answer. AI analyst Andrew Curran wrote on X that GLM had been the leading theory on Friday night, but by Saturday morning “people seem less sure of anything.”
Without a provider disclosure, weights release, or a hosting partner putting its name on the bill, both theories remain inference from side effects. That is useful for narrowing possibilities, but it is not the same thing as attribution.
Why the 100T Tokens/Day Claim and Week-Long Free Access Matter to Inference Economics
The market-relevant part of Ox Alpha is the distribution shock: a reasoning model positioned for coding and long-horizon agentic workloads, offered free with “near unlimited usage,” inside a router that can redirect developer traffic quickly. When routing is one config change, “free for a week” can be enough to reset defaults, at least temporarily, especially for teams benchmarking agents in production-like loops.
OpenCode also said the provider had capacity for 100 trillion tokens per day, described as roughly 100x the number of AI tokens Visa said it uses in an entire month. Tokens are the small text units models process, and token counts are how inference vendors talk about both usage and throughput. If that capacity claim is even directionally accurate, it implies unusually large inference availability for a stealth, time-limited free rollout.
That matters for AI-crypto traders less because it proves a specific lab is “winning,” and more because it pressures the pricing narrative around inference. A high-capacity free window can pull demand away from paid endpoints, change what developers consider an acceptable latency and rate-limit profile, and create a short-lived expectation that “reasoning-grade” models can be routed like commodities.
The broader backdrop is that Chinese labs have been narrowing the gap on performance while competing hard on price and openness. The same conversation around Ox Alpha referenced Chinese labs including Zhipu, DeepSeek, and Moonshot AI as increasingly challenging US rivals, and pointed to Moonshot’s Kimi K3, released in July, as a 2.8 trillion-parameter open-weight model built for coding, reasoning, and agentic tasks that drew attention for performance and lower price.
The forward signals are straightforward and mostly operational. The first is whether OpenRouter, OpenCode, or the provider discloses identity, weights status, or a hosting partner before the stated free-week window ends. The second is what happens to pricing and limits after the week, including throttling, queueing, or a shift to paid tiers. The third is whether further technical attribution work, including tokenizer fingerprints, materially strengthens either the Z.ai/GLM-5 lineage theory or the Microsoft MAI-family hypothesis. The last is whether sustained usage on OpenRouter, including availability and latency under load, corroborates or contradicts the 100T tokens/day capacity claim.
My Take: Treat Ox Alpha as a Capacity-and-Distribution Signal Until the Provider Is Verified
The threshold that matters here is not which logo ends up attached to Ox Alpha, it is whether the “free for a week” promise translates into sustained, high-throughput availability that developers can actually lean on for agentic workloads. If the model stays responsive under real traffic and then moves into a credible paid tier, that is a real inference-market datapoint. If it degrades into queues, throttles hard, or disappears when the free window closes, it was a marketing spike with a mystery wrapper.
Attribution remains too uncertain to trade directly because the evidence being passed around is mostly tokenizer and behavior inference, and even the loudest observers have gotten less confident over time. What would make this matter in practical terms is a verified provider plus post-free pricing that forces the rest of the inference stack to reprice or re-route to compete.