
Ramp launches Router, a U.S.-only LLM routing API free through 2026
The service routes requests across multiple model providers and defaults to one-year, opt-out data retention.
Ramp has launched Router, an API that lets companies route AI requests across multiple large language model providers with built-in routing strategies and a spend-and-latency dashboard. The service is U.S.-only at launch, free to use through the rest of 2026 excluding inference costs, and ships with a default one-year, opt-out retention policy for prompts and outputs.
Key Takeaways
- Ramp launched Router, an API that lets teams switch between multiple large language model providers with routing strategies and a spend/performance dashboard.
- Router is limited to U.S. users at launch and is free to use through the rest of 2026 excluding underlying model inference costs, with a $26 launch credit.
- Ramp says the router has been running internally for three years, positioning the release as an externalized production tool rather than a new experiment.
- The default policy retains model inputs, outputs, and tool calls for one year on an opt-out basis, and Ramp says it removes personally identifiable information before using content to improve the product.
Ramp Launches Router: A U.S.-Only LLM Routing API, Free Through 2026
Ramp, best known for corporate expense management, has pushed further into AI infrastructure with Router, a model-routing service that sits between an application and multiple large language model providers. Mechanically, it is a single API integration that can send each request to a chosen model vendor based on rules. The consequence is less vendor lock-in for teams running AI-heavy workflows, and a clearer path for Ramp to charge for “traffic direction” in inference.
Router is only available in the United States at launch. Ramp is also making the service free to use for the remainder of 2026, while still requiring customers to pay the underlying model inference costs. A $26 launch credit is included, and Ramp has not disclosed what Router will cost in 2027.
Ramp is framing Router as an external productization of internal tooling, saying it has used its router for its own AI usage needs over the past three years. That matters because routing layers tend to fail in boring ways first, like edge-case latency spikes, provider outages, and billing mismatches. “We’ve already lived with it in production” is the credibility pitch.
How Router Works: One Integration Across OpenAI, Anthropic, DeepSeek, xAI and More
A model router is a traffic cop for AI inference, not a new model. Developers integrate once, then choose which provider handles each request without rewriting their app for every vendor’s API. Router’s pitch is that switching costs move from “engineering project” to “configuration change,” which is what makes vendor optionality real.
At launch, Router provides access to models from OpenAI, Anthropic, DeepSeek, Moonshot, Minimax, Nvidia, xAI, and Z.ai. Ramp is not trying to win on catalog breadth yet. The product is positioned as similar to OpenRouter, but with fewer model options than OpenRouter’s lineup.
Router also ships with routing “strategies,” which are pre-baked rules for where requests go. One strategy lets users set a preference for model providers’ flex usage tiers. Another lets Router choose which model to route queries to based on up to three user-specified benchmarks. Ramp also says customers can route only difficult problems to more expensive models, or test models without manually switching integrations.
The gating detail for some teams is data handling. Router has an opt-out retention policy that records model inputs, outputs, and tool calls for one year by default. Ramp says it will remove personally identifiable information before using that content to improve the product, but it has not detailed how that PII removal is implemented or what controls enterprises get beyond opting out.
Cost, Latency, and Fallbacks: The Spend Dashboard Ramp Is Betting On
Router’s second product surface is observability: a dashboard that reports token spend, cost, latency, fallback attempts, and other details. Tokens are the billing unit for most text models, latency is the time-to-response, and fallback attempts are retries or provider switches when the first choice fails or cannot serve the request. Put together, it is a view of both the bill and the failure modes.
That dashboard is the bridge back to Ramp’s core business. If a team is routing across multiple providers, the messy part is not just picking the cheapest model. It is understanding when “cheap” becomes “slow,” when a provider’s capacity constraints trigger fallbacks, and when a routing rule quietly shifts spend from one vendor to another. Ramp is betting that the spend layer is sticky, and that once a company trusts the measurement, it will tolerate Ramp sitting in the request path.
The free-through-2026 structure reinforces that adoption strategy. Ramp is not waiving the actual inference bill, which still accrues at the model providers. It is waiving its own routing fee for now, lowering the friction to try Router as a control plane while Ramp learns what customers optimize for in practice.
Pricing Is the Open Question After 2026—and Ramp Isn’t Saying Yet
The forward curve on Router is mostly about what Ramp chooses to charge for once the free period ends. Ramp has not said how much the service will cost in 2027, leaving open whether pricing lands as a take rate on inference, a subscription, a per-request fee, or enterprise tiers tied to controls and reporting.
Geography is another tell. Router is U.S.-only today, and expansion beyond the U.S. would signal broader enterprise readiness and distribution ambition, especially for teams that need consistent policy and data handling across regions.
Data posture is the third lever. A default one-year, opt-out retention policy is a high-friction default for some operators, even with PII removal claims. Shorter retention, opt-in storage, or clearer enterprise controls would be a concrete signal that Ramp is optimizing for regulated and security-sensitive workloads.
Finally, the provider list is likely to move. Router currently supports OpenAI, Anthropic, DeepSeek, Moonshot, Minimax, Nvidia, xAI, and Z.ai. Additions beyond that set would clarify whether Ramp is building a broad gateway or a curated set of relationships with specific labs and inference providers.
My Read: Ramp Is Trying to Own the AI ‘Spend Layer’ Beyond Corporate Cards
The mechanism here is simple: put a measurement-and-control plane in front of inference, then make switching cheap enough that teams stop treating model choice as a one-way door. Free Router access through 2026 does that, because the only immediate cost is the underlying inference bill, and the dashboard makes the tradeoffs legible in tokens, latency, and fallbacks.
The threshold that matters is whether Ramp can turn that control plane into an enterprise default without getting blocked by its own retention posture. If the product keeps its one-year, opt-out logging while trying to sit in the middle of sensitive prompts and tool calls, adoption will skew toward lower-sensitivity workloads and experimentation. If Ramp tightens retention controls and lands a credible 2027 pricing model, Router starts to look like a durable toll layer for inference rather than a promotional feature bolted onto spend management.