Close-up of a computer motherboard with
AI

OpenAI and Anthropic tee up flagship price cuts as DeepSeek usage jumps on OpenRouter

GPT-6 Sol/Luna are set for a 50% cost cut and Opus 5.5 is pitched at ~40% cheaper, while DeepSeek V4.1 Flash logged a 172% weekly usage spike.

By Elliot Marsh6 min read

A new week-of release cycle is pushing flagship AI models down the cost curve, with OpenAI and Anthropic marketing steep price reductions rather than just discounting older tiers. Early usage data from OpenRouter, where DeepSeek’s V4.1 Flash sits at No. 1 with a 172% spike “this week,” is consistent with cheaper per-token pricing expanding demand even as labs face margin pressure.

Key Takeaways

  • OpenAI planned to ship GPT-6 Sol and GPT-6 Luna on Tuesday with costs for top business customers cut 50% versus prior iterations of the same models.
  • Anthropic also targeted Tuesday for Opus 5.5, saying it costs around 40% less to run than Opus 5 while maintaining “top intelligence.”
  • xAI released Grok 4.7 on Monday and marketed it on price-performance, but competing launches narrowed much of that pricing edge within 24 hours.
  • DeepSeek’s V4.1 Flash ranked No. 1 on OpenRouter’s leaderboard and recorded a 172% spike in usage “this week.”

Flagship Models Get Cheaper, Fast: GPT-6 Sol/Luna, Opus 5.5, and Grok 4.7

This week’s AI release cadence is being marketed less like a pure capability race and more like a price sheet rewrite. OpenAI planned to release GPT-6 Sol and GPT-6 Luna on Tuesday, cutting costs for top business customers by 50% versus previous iterations of the same models.

Anthropic planned to release Opus 5.5 on the same day and said it costs around 40% less to run than Opus 5 while maintaining top intelligence. The common thread is explicit cost compression on the newest “flagship” tier, not just the usual pattern of lowering prices on older models after a new launch.

xAI’s move set the tempo. Elon Musk’s xAI released Grok 4.7 on Monday, positioning the model around price-performance, and the competitive response was immediate. Within 24 hours, new options had reduced much of Grok 4.7’s price advantage, a fast reminder that pricing power in frontier models can evaporate on a single news cycle.

Per-Token Costs, Explained: Why Price Cuts Can Increase Total AI Spend

The unit that matters in this fight is the token. Per-token costs are the price a provider charges to process or generate a unit of text, and it is the simplest way to compare operating costs across models when everything else is moving.

Cutting per-token prices does not automatically mean the market spends less on AI. The mechanism is straightforward: when a model becomes cheaper to call, developers ship more features that were previously too expensive, users run longer sessions, and companies push more workflows through the model instead of reserving it for “high value” prompts. That can raise total token volume enough to keep aggregate spend climbing even as the price per token falls.

That dynamic is the intuition behind Jevons paradox, the economics idea that when a resource becomes cheaper to use, total consumption can increase rather than decrease. In this case, the “resource” is inference, and the bet is that lower prices unlock enough new usage to keep the compute meter running.

Sell-side framing in the packet leans into that volume story. Citadel Securities told clients that falling per-token costs are fueling additional usage and that overall AI spending is increasing, which it argued points to higher eventual profits for AI labs or companies providing the computing firepower behind AI usage. Morgan Stanley made a similar demand-side case, framing competition from Chinese providers as bullish for AI because cheaper access can capture more customers and increase overall demand for computing.

Adoption Proof Point: DeepSeek V4.1 Flash Tops OpenRouter as Usage Jumps 172%

The cleanest datapoint in the packet is not a benchmark score. It is a leaderboard and a usage spike.

DeepSeek’s V4.1 Flash ranked No. 1 on OpenRouter’s leaderboard and recorded a 172% spike in usage “this week.” OpenRouter is an aggregator that routes requests across multiple models and publishes rankings that traders and builders use as a rough proxy for adoption momentum.

DeepSeek is also the example the packet uses to define the open-weight pressure on incumbents. An open-weight model is one where the trained parameters are made available so others can run or fine-tune it, which often enables cheaper or more flexible deployment than fully closed offerings. The packet notes DeepSeek said V4.1 Flash beats its prior flagship on a number of benchmarks via technical updates that allow it to charge less, though the benchmark suite and results are not provided here.

The broader point is competitive, not ideological. When open-weight providers climb widely watched rankings while undercutting price, they force the closed labs to respond on cost, and this week’s flagship pricing language reads like that response.

Confirmation Checklist for Traders: Usage-to-Revenue, Pricing Floors, and Leaderboard Share

The first confirmation point is whether the price cuts are real in the only place that matters: post-release pricing pages and enterprise terms. OpenAI’s stated 50% reduction for top business customers and Anthropic’s ~40% lower run cost for Opus 5.5 are large numbers, but the packet does not include absolute $/token figures or the exact scope of which customers and usage tiers qualify.

Second is whether OpenRouter’s rankings and usage trends keep DeepSeek V4.1 Flash at or near No. 1 over the next one to two weeks, or whether the leaderboard reverts after the Tuesday flagship releases land. A one-week spike can be novelty, incentives, or routing defaults. Sustained share is a different claim.

Third is the usage-to-revenue bridge. The packet flags margin pressure on OpenAI and Anthropic as the price war accelerates, and it is still unresolved whether higher token volumes translate into sustained revenue growth when per-token pricing is compressing.

Finally, watch how fast repricing spreads. Grok 4.7’s advantage being competed away within 24 hours is the template. If the same rapid repricing hits more “flash” or small variants, the market is likely moving toward a pricing floor set by whoever has the cheapest compute and the best distribution, not whoever wins a single flagship launch.

My Take: The Price War Looks Like a Margin Story for Labs—and a Volume Story for Compute

The part that decides this isn’t whether one model is “best,” it’s whether cheaper inference expands the workload enough to keep total spend rising. The packet’s hard evidence is narrow but directionally consistent: DeepSeek’s V4.1 Flash sitting No. 1 on OpenRouter with a 172% weekly usage spike is what Jevons-paradox demand looks like in the wild, even before the Tuesday flagship cuts hit.

The real test is whether the post-launch pricing pages and enterprise terms confirm the stated 50% and ~40% reductions and whether leaderboard share holds after the new flagships land. If those two things line up, the setup starts to look structural: labs fight a margin war while compute demand stays elevated on volume.

Sources