A recent report from Vercel reveals that open-weight AI models now process over half of all tokens passing through its AI Gateway, marking a sharp rise in adoption since late 2025. Despite this traffic dominance, Anthropic's proprietary models continue to command nearly two-thirds of total spending, illustrating a complex cost and infrastructure dynamic for cloud AI operations.

  • Open-weight models handle 56% of Vercel AI Gateway tokens in August 2026.
  • Anthropic still accounts for 64% of total AI Gateway spending despite lower token share.
  • Average token price across the gateway fell 23% in August, reflecting growing cost efficiency.

Infrastructure signal

The substantial rise of open-weight AI models to processing a majority of tokens through Vercel’s AI Gateway signals a major shift in AI inference workloads. Since December 2025, their share surged from single digits to over half of all tokens by August 2026, driven largely by models originating from Chinese developers and smaller open-weight vendors like DeepSeek, Moonshot AI, and Z.ai. This trend points to increasing confidence in open-weight models’ reliability and scalability on production cloud infrastructure.

Cost dynamics also shift with this evolution. Open-weight models, while handling the majority of token volume, only represent about 14% of total spend, underscoring their significantly lower price points compared to US-based proprietary alternatives. This pricing gap suggests that cloud cost optimization strategies may increasingly favor workload routing to open-weight models, alongside traditional providers, to balance performance with budget constraints.

Developer impact

For developers utilizing Vercel’s AI Gateway, the rise of open-weight models offers greater choice and cost efficiency, but also increases complexity in managing multi-model environments. Developers must adjust tooling—such as CLIs, SDKs, and IDEs—to be model-agnostic to seamlessly switch between or combine open-weight and proprietary AI models. This adaptation period reflects the nascent stage of enterprise adoption and indicates opportunity for improved workflow integrations that abstract away provider-specific details.

Observability and cost tracking become more critical as token volume and spend diverge across model types. Teams need enhanced monitoring to accurately attribute usage and forecast expenses, especially given the dominant spending role of Anthropic’s models, which remain essential for premium inference tasks despite their higher costs. Clear visibility into model selection and switching patterns will help teams optimize performance-cost tradeoffs in deployment pipelines.

What teams should watch

Teams should closely monitor pricing trends and token volume distribution across AI models on gateways like Vercel’s. The average token price fell by over 23% in August 2026, indicating improving cost efficiency and increasing adoption of cheaper open-weight inference options. This price evolution may impact budgeting and cloud cost allocation models, prompting reassessment of AI infrastructure spend strategies.

Anthropic’s models continue to dominate AI Gateway spending, accounting for around 64% share, with consistent holding of the highest-spend positions. Developers and infrastructure managers should track shifts within Anthropic’s portfolio—such as the decline of Fable 5 and rise of Opus 5—to align application deployment choices with optimal cost-performance balance. Additionally, enterprises preparing for broader AI adoption must prioritize model-agnostic deployment architectures that facilitate flexible routing and integration across an expanding ecosystem of AI providers.

Source assisted: This briefing began from a discovered source item from The New Stack. Open the original source.
How SignalDesk reports: feeds and outside sources are used for discovery. Public briefings are edited to add context, buyer relevance and attribution before they are published. Read the standards

Related briefings