Dimly lit server room with blue and orange lights
AI

OpenAI and Anthropic roll out cheaper GPT-6 tiers and Claude Opus 5.5

OpenAI cited a 50% API price cut versus GPT-5.6 promo pricing, while Anthropic pitched ~40% lower run costs via token efficiency.

By Elliot Marsh4 min read

OpenAI and Anthropic each introduced lower-cost model options on Sept. 22, framing the releases around cutting customer inference spend as cheaper open-weight rivals squeeze hosted API pricing. The launches also land as the first new models from either lab since recent calls for an industrywide slowdown tied to the AI safety debate.

OpenAI’s GPT-6 Sol/Luna and Anthropic’s Opus 5.5 Put Cost Cuts Front and Center

OpenAI added two new tiers to its GPT-6 family on Sept. 22: GPT-6 Sol and GPT-6 Luna. The company said the new tiers cut API pricing by 50% versus GPT-5.6 promotional pricing, positioning the move as a direct reduction in what developers pay to call the model programmatically.

The segmentation is explicit. OpenAI positioned Sol as a tier below Astra and aimed it at more complex workloads like coding. Luna is framed for “high-volume tasks” such as extracting information or summarizing documents, the kind of throughput-heavy work where token-based API bills can dominate the unit economics.

Anthropic’s release hit the same lever from a different angle. The company unveiled Claude Opus 5.5 as the latest model in its Claude 5.5 family and described it as more token-efficient, estimating it will cost around 40% less to run than Opus 5. Tokens are the chunks of text models process, and token efficiency matters because most API pricing is metered by tokens in and out.

Anthropic product leader Dianne Penn described the mechanism as reducing how many tokens the model needs to do the same job at a given “effort setting.” “One of the things we’re continuing to innovate on is how to make that thinking, how to make the answering more efficient, so it uses less tokens depending on your effort setting,” Penn said.

Both launches were framed against the same competitive backdrop: pressure from cheaper, open-weight models, meaning models whose weights are available for others to run or fine-tune rather than being locked behind a hosted API. The rivals named in that context were Alibaba, Moonshot AI, and DeepSeek.

The timing also matters. These are described as the first releases from either lab since the recent safety debate intensified and calls for an industrywide slowdown on developing advanced AI gained traction. Anthropic CEO Dario Amodei called for a slowdown in recent weeks, and OpenAI CEO Sam Altman and Elon Musk also joined that call. The safety debate was catalyzed in part by a Sept. 8 X post from former Anthropic researcher Jacob Coxon, who said he quit and warned both labs were “gambling with our lives.”

What Traders Can Monitor Next: Pricing Sheets, Benchmarks, and Open-Weight Spillover

The immediate limitation is that the disclosures are directional, not fully priced. OpenAI and Anthropic gave percentage reductions, but the provided text does not include absolute API pricing tables, such as $/1M tokens for input and output, or any details on whether the cuts apply uniformly across context windows, tool use, or higher-effort reasoning modes.

Benchmarks are the next missing primitive. Without third-party comparisons of GPT-6 Sol/Luna and Claude Opus 5.5 against prior versions and against the open-weight competitors cited, the market is left with a cost-curve signal but not a quantified performance-per-dollar delta.

Adoption is the other leg. If Luna is genuinely tuned for “high-volume tasks,” the validation would show up as measurable throughput growth, developer migration, or usage concentration in extraction and summarization workloads where open-weight deployments can undercut hosted APIs. The excerpt provides no usage metrics, so any downstream read-through to compute demand remains speculative.

Finally, the releases set up a near-term credibility test for the slowdown rhetoric. Follow-on statements from leadership on whether the recent calls change release cadence or impose deployment constraints will matter more than the headline discount, because cadence and gating determine how quickly cheaper frontier inference actually propagates into production workloads.

My Read: Cheaper Frontier APIs Raise the Bar for ‘AI Token’ Narratives Without New Proof Points

The mechanism here is simple: both labs are trying to defend inference volume by lowering effective run costs, either through explicit API price cuts (OpenAI’s 50% versus GPT-5.6 promo pricing) or by marketing token efficiency as “same work, fewer tokens” (Anthropic’s ~40% cheaper-to-run claim). That reads less like a one-off promo and more like segmentation for cost-sensitive use cases, especially the high-throughput extraction and summarization lane where open-weight models can be run cheaply if you can operate the stack.

The threshold that matters is whether absolute pricing sheets and independent benchmarks land quickly enough to turn this from narrative into a measurable performance-per-dollar reset. If the numbers and adoption data confirm that cheaper tiers pull real workload back from open-weight deployments, the practical impact is a lower floor for inference pricing that forces every AI-adjacent token story to justify itself with usage, not slogans.

Sources