Dimly lit corridor with dark server cabinets
AI

Silicon Data’s LLM token-price index hits record low at $0.97

The gauge is down more than 50% from an earlier-summer high as price cuts and open-source competition bite.

By Elliot Marsh6 min read

Silicon Data’s LLM Token Expenditure Index fell to $0.97 on Monday, the lowest reading since the benchmark launched late last year. The more-than-50% drawdown from an earlier-summer high is a clean print of inference-price deflation that helps users but tests frontier labs’ pricing power ahead of IPO scrutiny.

Key Takeaways

  • Silicon Data’s LLM Token Expenditure Index printed $0.97 on Monday, the lowest level since the benchmark was created late last year.
  • The gauge is down by more than half from the high set earlier this summer, extending a fast repricing of inference costs.
  • The index is designed to track the going market rate for a large language model token, a unit used to meter and bill LLM usage.
  • Open-source Chinese models, OpenAI’s late-July GPT-5.6 price cuts, and demand-linked “dynamic pricing” were cited as key forces pushing token rates lower.

LLM Token Prices Print a New Low: $0.97 on Silicon Data’s Gauge

Silicon Data’s LLM Token Expenditure Index, a daily benchmark for the going market price of a large language model token, fell to $0.97 on Monday. It was the lowest reading since the index was created late last year.

The move extends a sharp summer drawdown. The index is now down by more than half from the high recorded earlier this summer, putting a number on how quickly inference pricing is deflating across the LLM market.

Mechanically, the index is tracking what users pay to run prompts and receive outputs across major chatbot and API ecosystems. Lower token prices translate into cheaper inference for end users running queries on products like OpenAI’s ChatGPT, Anthropic’s Claude, and Google’s Gemini, but the same print also tightens the revenue-per-token ceiling for the labs selling access.

Why the Slide Matters: Cheaper Inference for Users, Less Pricing Power for Labs

Token-price deflation is a two-sided trade. Users get a direct cost-down on inference, which can expand experimentation and push more workloads into production. The catch is that the same deflation can reset what buyers consider a “normal” rate for access, weakening the ability of model providers to hold price even as usage grows.

That matters most for foundation model labs, where the cost base is not purely variable. Charles-Henry Monchau, investing chief at Syz Group, framed the exposure bluntly: “Foundation model labs are the most directly exposed,” adding, “Token deflation compresses the revenue line while compute commitments stay fixed.” If compute is contracted or provisioned ahead of demand, falling unit prices can show up as margin pressure before any efficiency gains arrive.

The timing also intersects with public-market narratives. The index’s slide was tied to potential profit pressure on AI leaders such as OpenAI and Anthropic as they weigh public-market debuts, with both companies having confidentially filed for IPOs with regulators this summer. For traders, the near-term question is less about whether AI demand exists and more about whether the market is pricing inference like a scarce capability or like a competitive utility.

The read-through extends to the broader AI buildout. Mega-cap technology companies including Nvidia and Microsoft have poured billions of dollars into expanding AI capacity, and sliding token prices can force a reassessment of expected return on invested capital if revenue per unit of usage falls faster than costs.

What’s Driving Deflation: Open-Source China, OpenAI Cuts, and Dynamic Pricing

The drivers cited point to competition and pricing design, not a single promotional cycle. Monchau attributed part of the drop to the rise of open-source Chinese models such as Moonshot’s Kimi K3, which he said can fetch lower prices than alternatives from leading frontier labs. Open-weight releases compress the time window in which frontier labs can charge a premium purely for capability, because comparable performance can be replicated and served at lower cost by third parties.

Frontier labs have also moved on list price. OpenAI announced price cuts for two GPT-5.6 models in late July, a direct step-down in what the market anchors on for “top-tier” inference. When the price leader cuts, the rest of the stack either follows or differentiates on something other than raw tokens.

The other lever is pricing structure. Monchau said other frontier labs have rolled out offerings with “dynamic pricing” that allows access rates to rise and fall with demand. In practice, demand-linked pricing can smooth capacity peaks, but it also makes discounting a built-in feature rather than a one-off event, which can accelerate the path to a lower clearing price.

Monchau also pointed to decreasing costs across the industry for producing a token. That matters because it suggests the index is not only reflecting competitive undercutting, but also a shifting cost curve that makes lower prices sustainable.

Steve Hou, Silicon Data’s head of research, offered a supply-side interpretation: the drop could indicate that between frontier models and cheaper competitors, there may already be enough supply to “provide sufficient capabilities for most tasks,” implying less scarcity and more commodity-like pricing.

Signals to Watch for AI token prices hit record lows

The first signal is whether the index stabilizes around the $1.00 handle or continues making new lows after the $0.97 print. A flatline would suggest the market is finding a clearing price, while further lows would imply that either supply is still coming online faster than demand or that pricing tactics are still cascading.

Next is whether frontier labs extend the late-July playbook with additional list-price cuts or broaden the rollout of demand-linked “dynamic pricing.” If dynamic pricing becomes the default for premium endpoints, traders should expect more frequent price discovery and less durable headline rates.

IPO timelines are the other catalyst surface. OpenAI and Anthropic have already confidentially filed, but any updates that reframe margin expectations or pricing-power narratives will matter more in a deflationary tape than in a scarcity tape.

Finally, watch the cadence of competitive releases from lower-priced open-source Chinese models, including new versions or wider adoption. The market impact is not just performance, it is how quickly those models become deployable substitutes that pull the clearing price down.

The Commoditization Read-Through—and the Gaps Traders Should Respect

The threshold that matters is whether token prices can fall this quickly without forcing a visible change in how frontier labs package and monetize access. Monchau’s point about fixed compute commitments is the mechanical constraint, and it is why a >50% drawdown from an earlier-summer high to $0.97 reads like margin pressure first and “more usage will fix it” second.

The gap traders should respect is methodology and scope. The excerpt describes the index as a gauge of daily prices and the going market rate for an LLM token, but it does not spell out coverage, weighting, or how it normalizes across providers and tiers. It also uses “AI token prices” to mean LLM usage pricing, not crypto tokens, so the clean takeaway is commoditization risk in inference economics, not a direct signal for on-chain AI-token valuations. If the index keeps printing new lows while labs shift moats toward distribution, memory, and context, the story becomes structural: pricing power is moving up the stack, and raw tokens are becoming the cheap part on purpose.

Sources