Agentic AI models can increase token usage per task by 10 to 100 times compared to simpler inferences, straining the per-token pricing model that enabled early AI experimentation but now risks escalating operational costs.

  • Agentic AI increases token usage per task 10x to 100x, inflating costs.
  • Majority of AI compute now on reserved or owned infrastructure, not public cloud.
  • Unpredictable variable pricing risks undermining successful AI initiatives.

Market signal

The traditional per-token pricing model for AI usage, once praised for enabling rapid experimentation, is under significant strain as agentic AI systems drive token consumption massively higher. Unlike simple inference calls, agentic agents plan, execute multi-step workflows, check their work, and iterate, consuming tokens at every step. This complexity naturally leads to costs that grow disproportionately relative to simpler workloads.

Futurum's recent report quantifies this phenomenon, forecasting a 219% increase in agentic and reasoning inference usage within the year, with total AI inference spending ballooning from $120 billion in 2025 to $885 billion by 2030. This pricing and consumption dynamic signals an urgent need for enterprises to rethink how AI usage is metered and budgeted, especially as successful projects routinely overshoot initial cost estimates.

Operator impact

Chief Information Officers and Chief Financial Officers are increasingly challenged by the unpredictable, often rapidly escalating costs associated with agentic AI deployments. Use cases that initially appeared cost-effective in pilot phases can become budget risks in scaled production, with at least one example cited where a $1 million annual budget was burned through in only three months due to unanticipated token demands.

Beyond the raw cost impact, this unpredictability threatens to stall AI adoption by creating governance and financial management challenges. Some organizations have reportedly abandoned useful, internally developed automation tools because they could not justify or forecast ongoing costs, highlighting that this issue extends beyond pricing structures to enterprise governance and operational planning.

What to watch next

Enterprises appear to be moving beyond the assumption that AI workloads must run fully on-demand in hyperscale public clouds. A significant majority—66%—of AI compute consumption is now on reserved or owned infrastructure, and 59% of surveyed decision-makers run AI workloads primarily outside hyperscaler cloud environments. This trend reflects a maturing market where capacity commitments and hybrid infrastructure models become critical to controlling skyrocketing token costs.

Innovative deployment models such as partitioning high-density GPU servers across multiple customers with tiered service levels illustrate how AI infrastructure providers and operators can optimize utilization and profitability in a cost-effective manner. Industry stakeholders should watch how these hybrid and containerized approaches to AI compute evolve alongside pricing model experiments to better accommodate agentic AI’s heavy token consumption.

Source assisted: This briefing began from a discovered source item from SiliconANGLE Business. Open the original source.
How SignalDesk reports: feeds and outside sources are used for discovery. Public briefings are edited to add context, buyer relevance and attribution before they are published. Read the standards

Related briefings