An orange robotic arm handling a microchip on a
AI

AI labs cut top-model release intervals to 44 days as agents take on core R&D work

Internal metrics at Anthropic and OpenAI point to automation that compresses safety windows and raises policy and security tail risk.

By Elliot Marsh5 min read

Major U.S. and Chinese AI labs have shortened the average interval between high-performance model releases to 44 days since April 2026, down from 125 days over January 2023 to March 2026. The shift is being driven by AI agents doing more of the work of building next-generation models, tightening safety-verification timelines in a competitive race that can spill into risk markets.

From 125 Days to 44: The New Model-Release Tempo Across U.S. and China Labs

A survey of high-performance model releases spanning five U.S. AI companies and four Chinese firms found the average release interval fell to 44 days from April 2026 through September 2026, versus 125 days from January 2023 through March 2026. The set included U.S. labs such as Anthropic and OpenAI and China-based firms including Alibaba Group and Moonshot AI.

The cadence is no longer an abstract average. Early September provided a clustered example: Anthropic unveiled “Claude Fable 5.1” on Sept. 1, Google launched “Gemini 3.8 Flash” on Sept. 2, and OpenAI released “GPT-6 Astra” on Sept. 3. Google’s Sept. 2 release was described as the third Flash model in six weeks.

The same pattern is showing up as a baseline schedule rather than a one-off sprint. Meta has been updating its “Muse Spark” model monthly since July, and China-based DeepSeek has released new versions every month since July, with Alibaba and Zhipu AI described as joining the faster-release race.

For markets, the mechanism that matters is not just “more models.” A 44-day cycle means less time for external testing, less time for downstream integrators to harden deployments, and more frequent moments where policy, security, or safety disputes can become near-term catalysts.

AI Building AI: Internal R&D Automation Metrics at Anthropic and OpenAI

The driver described for the compressed cycle is a shift from AI as a coding assistant to AI as an R&D executor. AI agents are being used to run experiments, manage evaluation loops, and analyze results, which lets labs parallelize work that used to bottleneck on human attention.

Anthropic’s Sept. 17 report said Claude led 26% of the company’s AI R&D tasks as of last month, up from 0% in February 2026. Anthropic also said, verbatim, “Claude does not yet operate fully autonomously.” The report framed the measurement as a way to gauge proximity to recursive self-improvement, meaning a system iterating on its own design or training process in a way that can accelerate capability gains without proportional human oversight.

OpenAI’s internal metrics in the packet point to the same direction, but with a different proxy: labor and spend. As of last month, AI agents in OpenAI’s research organization logged total working hours equivalent to 3.1 times the daily labor of human researchers. OpenAI’s per-researcher AI token spending, a usage-cost measure tied to how much model output researchers consume, rose from about $1 per day for a typical researcher in February 2026 to over $600 per day as of last month, with the top 10% exceeding $7,000 per day.

The output-side indicators moved with it. The packet says the volume of programming code written at OpenAI last month grew to seven times last year’s average, while Anthropic’s code adopted into actual products in Q2 reached eight times the 2021–2025 average. That is what “AI building AI” looks like operationally: more parallel attempts, faster iteration, and less slack between capability jumps.

Safety Windows Shrink as Competition Turns Into a ‘Prisoner’s Dilemma’

The safety concern is straightforward: if iteration speed rises, the time available for humans to understand model behavior and verify safety shrinks. The packet ties that to fears of recursive self-improvement triggering an “intelligence explosion,” a hypothesized runaway increase in capability that outpaces human control and governance.

Security risk is tightening on its own curve. The UK AI Security Institute estimated in November 2025 that the doubling time for the length of cyberattack tasks AI can perform was about eight months. By February 2026, the institute estimated that doubling time had shortened to 4.7 months, implying faster improvement in offensive capability and a narrower window for defensive adaptation.

The catch is incentives. The competitive environment is described as a “prisoner’s dilemma,” where each lab has reason to keep racing even if all would be safer slowing down, because unilateral deceleration risks losing ground. The packet also includes an unverified competitive-pressure datapoint: it cites Reuters as saying Anthropic is considering accelerating its next model launch to counter OpenAI’s Astra ahead of an IPO.

Near-term signals that would confirm the 44-day regime are concrete. Confirmed schedule changes for Anthropic’s next release would show whether “considering acceleration” becomes an actual calendar move. Updates to internal automation intensity metrics, including whether OpenAI’s 3.1x agent-hours multiple holds or rises and whether typical per-researcher token spend stays above $600 per day, would indicate whether the acceleration is stabilizing or still ramping. Fresh cyber capability assessments from the UK AI Security Institute or peers that extend the eight-month to 4.7-month estimate would also matter, because that is the kind of external risk curve regulators and enterprise buyers can anchor to.

My Read: Faster Cycles Raise the Probability of Sudden Risk Events Traders Can’t Ignore

The part that decides this is whether the 44-day cadence is a temporary launch cluster or a new operating rhythm driven by automation. Anthropic’s 26% AI-led R&D share and OpenAI’s 3.1x agent-hours multiple read like a step-change in how much work is being executed by machines, not just assisted, and that kind of throughput tends to compress verification time even if safety budgets rise.

The threshold that matters is whether labs keep publishing metrics that show automation intensity staying high while releases continue to bunch within days or weeks. If that holds, the setup starts to look structural rather than narrative-driven, and the practical consequence is a higher frequency of policy and security shocks that can leak into broader risk pricing.

Sources