A computer setup with a dark background
AI

Google launches Gemini 4 Argon with $2/$10 token pricing as agent race heats up

Argon’s intro rates later double, while Meta’s free Muse tops 5M downloads and Google’s Spark stays paywalled.

By Elliot Marsh7 min read

Google unveiled Gemini 4 Argon with aggressive introductory token pricing and benchmark claims in coding and cybersecurity. The release lands as investors increasingly price AI leadership through consumer personal agents, where Meta’s free Muse is scaling faster than Google’s paywalled Spark.

Key Takeaways

  • Google introduced Gemini 4 Argon, positioning the frontier model as a step up for coding, cybersecurity, and complex multi-step tasks.
  • Argon launched at $2 per million input tokens and $10 per million output tokens, with Google saying its published rates will later double to $4 and $20.
  • Google says Argon ties OpenAI on a key cybersecurity test and leads software engineering benchmarks, but the benchmark names and independent replication are not available in the provided excerpt.
  • Meta’s free Muse personal agent app passed 5 million downloads as of Sept. 30 (Sensor Tower), while Google’s Spark remains subscriber-only and cannot place outbound calls or complete purchases.

Gemini 4 Argon Lands as Wall Street Shifts From Models to Agents

Google unveiled Gemini 4 Argon on Sept. 30, pitching the model as a frontier-tier upgrade for coding, cybersecurity, and complex tasks. The company framed Argon as capable of handling assignments that require multiple steps and can run for extended periods, a workload profile that matters for agent-style products that do more than chat.

The timing matters because the investor narrative is moving away from “who has the best model this week” toward “who has an agent people actually use.” Personal agents are the consumer-facing layer that turns model capability into distribution, retention, and eventually monetization. In that framing, a model launch is table stakes, while an agent with real adoption is the scoreboard.

That shift is visible in the competitive set described around Argon. Meta’s Muse launched “last month” and climbed to the top of Apple’s App Store, topping ChatGPT in the process, while OpenAI rolled out its own personal agent product, Dots, earlier in the same week. The market is treating agent products as the new surface area for mindshare, and mindshare is what tends to leak into risk-on positioning across tech and crypto beta.

Argon’s Token Pricing Play: Intro Discounts Now, Higher Rates Later

Argon’s introductory API pricing was set at $2 per million input tokens and $10 per million output tokens, matching the discounted pricing cited for OpenAI’s GPT-6.1 Sol model. Tokens are the billing unit for model usage. Input tokens are the text a user or application sends into the model, while output tokens are the text the model generates back.

Google also made the second half of the pricing story explicit: it said Argon’s published rates will eventually double to $4 per million input tokens and $20 per million output tokens. That structure reads like a land-grab for early workloads, followed by a normalization phase once demand and cost curves are better understood.

For agent workflows, the input-output split is not cosmetic. Agents tend to run longer, call tools, and iterate, which can push output-heavy usage and make per-token economics show up quickly in gross margin. Intro pricing can win pilots and headlines, but the published-rate schedule is the part that determines whether “always-on” agents are cheap enough to be default.

Google’s own product framing reinforces that tradeoff. Tulsee Doshi, head of product for Gemini, described a portfolio approach where everyday agents do not require the most powerful frontier models, and lighter models remain essential for routine requests at scale. Frontier models like Argon are positioned as the reasoning layer for harder assignments, which is useful, but also the expensive part to run continuously.

Benchmarks vs Reality: What Google Claims, What Can’t Yet Be Verified

Google’s benchmark pitch for Argon is straightforward: it ties OpenAI on a key cybersecurity test and posts leading results in software engineering, based on industry benchmarks. The catch is that the specific benchmark names were not provided in the excerpt, and Argon is not broadly available for independent testing.

Access is constrained by design. Google is participating in a voluntary U.S. government safety-testing process and is initially restricting Argon to select cybersecurity defenders and enterprise cloud customers. The company said it plans to expand availability to developers, enterprise customers, and consumers, but it has not announced a public release date. Until that changes, the market is left trading on claims rather than replication.

The consumer-agent gap is even more concrete. Google’s personal agent Spark launched in May at Google I/O and can work across Gmail and Calendar, navigate websites through Chrome, and complete tasks such as filling out online forms. Spark is available through Google’s mobile and desktop apps, but it remains limited to paying subscribers.

Spark also has action limits that matter in practice. It cannot make outbound phone calls or complete purchases on a user’s behalf. It instead guides users through purchases and hands control back for final approval. That design reduces risk, but it also caps the “agent” moment where users feel the system actually took the task off their plate.

Meta’s Muse is taking the opposite distribution posture. It is free with usage caps and had reached more than 5 million downloads as of Sept. 30, according to Sensor Tower. Meta’s approach has also introduced complications. The company experimented with a human concierge system where contractors handled some phone calls when Muse could not complete them autonomously, before suspending the experiment amid internal privacy concerns, while its automated phone-calling feature remains in beta. A Meta spokesperson said internal testing “is core to the product development process” and added: “We’re working with merchants to continue improving this potential calling feature, and will only roll it out when it’s ready and with the proper disclosures.”

There is also a credibility dispute around Argon’s real-world coding performance. A Bloomberg characterization referenced in the source material said some Google employees questioned the model’s practical coding quality despite strong benchmarks. Google disputed that characterization and said employees across the company have been testing versions of Gemini 4 for weeks, with some receiving unlimited access. Without underlying evidence in the excerpt and without broad external access, that disagreement remains unresolved.

Catalysts That Could Reprice the Narrative

The next repricing catalyst is simple: a public release date and broader developer access for Gemini 4 Argon. Until independent users can run the model across real workloads, benchmark leadership is a marketing claim, not a market-clearing fact.

The second catalyst is distribution. If Google loosens Spark’s paywall, or expands Spark’s action capabilities beyond its current limits on outbound calls and purchases, it changes the adoption slope more than another benchmark chart does.

Meta’s Muse metrics are the other side of that trade. Updated adoption figures beyond the Sept. 30 Sensor Tower snapshot, and any change to usage caps as Meta scales, will tell traders whether Muse is a novelty spike or a durable consumer habit.

Wall Street’s framing is already leaning this way. Over the past three months, Alphabet stock was down about 6% while Meta was up 19%. Follow-on commentary that explicitly ties agent traction to Alphabet versus Meta positioning is likely to be the cleanest narrative driver for how this theme trades.

My Read: Why Agent Distribution, Not Frontier Benchmarks, Is the Near-Term Trade

The threshold that matters is whether Argon becomes something developers can actually touch at scale, because restricted access keeps the story stuck at “trust us.” Google can plausibly compete on frontier-model economics in the near term, but the planned move from $2/$10 to $4/$20 per million tokens is the reminder that cost pressure comes back fast once agents run continuously.

If Spark stays paywalled and capped on high-value actions like calls and purchases, Argon’s benchmark narrative will struggle to translate into consumer mindshare. If Google pairs broader Argon access with a real Spark distribution shift, the agent story stops being aspirational and starts being measurable in adoption and usage curves.

Sources