When a Siri‑style chatbot first steps into the world, its cost looks trivial: free. Yet behind that price lies a labyrinth of micro‑units called tokens, the building blocks that large language models (LLMs) parse to deliver answers. Token count scales with prompt length and complexity, and because each token consumes a tiny slice of the model’s computational budget, price can spike unexpectedly.
Firms such as Microsoft, Google, Anthropic and Google’s Gemini arm have poured hundreds of billions into creating these models. To recover that outlay, they roll out paid tiers offering more capabilities—code generation, business analytics, higher throughput, and future features locked behind a paywall. However, when a startup or enterprise employs the model via a third‑party API, the calculation of costs becomes a nightmare.
“Trying to tie someone into a cost model for the next twelve, two or three years, it doesn’t make any sense,” says Simon Gooch of Saviynt, whose identity‑management platform is piling on agentic AI. Token consumption battles the “non‑deterministic” nature of what the model outputs: the same question can yield a dozen different answers depending on wording, and multi‑agent pipelines amplify token churn.
Goldman Sachs’ latest research forecasts that token usage will burgeon 24‑fold between 2026 and 2030, reaching 120 quadrillion tokens per month, driven by companies migrating from ad‑hoc requests to integrated AI agents. That scale means even a modest per‑token price translates into multi‑million‑pound expense cycles that businesses struggle to predict.
Will Venters of the London School of Economics points out that companies can roam perilously in hidden, flat‑fee accounts while vendors notice the “bath” they’re creating. He notes that “it will start clamping down” once the vendors feel shareholder pressure.
Answering the cost dilemma, Oliver King‑Smith of smartR AI suggests smaller firms should lean on personal accounts, but stresses companies must assess which LLM they use and keep prompts tight. “If you weren’t precise, you could buy a million tokens and use them on unseen demand,” he says.
Rob Steele, CFO at iplicit, adds that clearer prompts are equivalent to a grocery list: it informs the AI what you need in the shopping basket of tokens. In a citizen‑AI era, a single ambiguous prompt can push costs to a sold‑out billing month, threatening an entire project.
Adam PT figures that while the cost fluctuations look like a cracked calculator, the value derived often outweighs the expense: more tokens can mean higher quality results, yet translating that value to customers remains an open issue.
Bill Peterson, product marketing director at Sumo Logic, underscores the need for flexible pricing frameworks, ranging from price‑per-incident bundles to result‑based models. Yet the churn of token prices—shifting each couple of months on the LLM provider side—means that budgeting for AI-rich projects will remain a challenge in the near future.
In short, token cost is only the tip of an iceberg: as AI services climb in complexity, they demand new financial models that can handle unpredictability while providing clear value to users and investors alike.


















