The budget pieces for AI Adoption

This article was orginally post in LinkedIn by Indira Hayams on 9/10/2026. The core idea was to break down the elements to consider when starting an AI adoption project. As all about AI, concepts may have evolved or their weights have changed. But it is a topic that is more than worthy to follow up these days.

What is Tokenomics?

“AI Tokenomics is the discipline of converting energy and capital into AI, then efficiently consuming AI services to enable intelligent outcomes and drive business value.” – Tokenomics Draft Definition V0.5.2. September 1, 2026.

This article originally started as an inquiry into the blockers slowing successful AI implementation, especially in Latin America.  The initial thought was to focus on the models’ pricing perspective, a natural bias from years at Microsoft, my previous employer.  Much of my tenure there was spent trying to make software pricing and enterprise agreements compelling to customers, since in the SaaS era, the price per user was usually the dominant variable in the value equation.

And I was headed in the right direction but, then this term “Tokenomics” knocked my door.   The fact that the Linux Foundation and FinOps experts around the world have joined forces to create the Tokenomics Foundation, to define concepts, vocabulary, build frameworks and engage in a long-term effort to measure the true economic value of AI got me interested in understanding their mandate and what operational friction they plan to solve.  

Let’s learn more about the status of Tokenomics in 2026.

AI Budget & ROI Reality

·         According to McKinsey (State of AI) as many as 89% of surveyed business organizations report any use of AI in any business function. However, only 6% of the survey respondents are companies who attribute an EBIT impact of 5% or more to AI use, those are considered AI high performers.

·         A survey of 500 US/UK finance leaders by Sapio Research revealed that 79% of enterprises missed their AI infrastructure forecast by more than 25% in the past 12 months, overruns were caused mostly by token’s usage-based pricing, computing power for agents and data platforms associated with the operation.

·         PwC’s 29th Global CEO Survey reported that “despite widespread experimentation, only 12% CEOs say AI has delivered both cost and revenue benefits. Overall, 33% report gains in either cost or revenue, while 55% say they have seen no significant financial impact to date”.

·         And about Latin America, the region is adopting AI at a faster rhythm (47% over global mean of 45%) while capturing only 1.1% of global investment. The region faces higher costs of acquisition versus a higher demand for AI services.  

Tokenomics

Data gives the context; we are facing the biggest disruption humanity has experienced since the iPhone. Everyone wants it but not everyone can take advantage of it, because it is so disruptive that not everyone is prepared to use it.

We see how difficult is to identify its value (as impact in the EBIT), when even companies with mature FinOps practices overran budgets by a mean of 30.9%, as the survey by Sapio Research revealed.

Of course, the next classical questions come in order: why to buy AI? Why now? What will be changed for good? And most probably, because resources are not unlimited, what will be left for later, replaced or eliminated to get the AI on board?

The imminent adoption of AI as part of any business ecosystem requires leaders to understand well how it works, and to monetize its impact on the P&L, what are its elements of value. 

Tokenomics, obviously, comes from Tokens, the smallest unit of text that Large Language Models (LLM) read and generate; not the same as a digital asset in blockchain. Although AI’s cost has many components, tokens are the most visible, probably because they are metered or usage-based billed. Other known components of the cost equation are orchestration, agent loops, memory, retrieval, routing, evaluation, governance and people. With all these parts playing at the same time, then we can imagine why it can be easy to overrun AI budgets.

The Architecture of the Leak: Why Budgets Break

In enterprise SaaS, scaling was predictable: you added headcount, you provisioned user subscription licenses (USL) at a negotiated flat monthly rate, and finance could forecast spend three years out with confidence. AI completely shatters that predictability.

As Jared Spataro, Microsoft’s CMO of AI at Work, recently pointed out, we are witnessing the emergence of two fundamentally different bills inside enterprise AI spend. On one side sits the traditional USL, standard subscription access for everyday knowledge work (like email writing or basic summarization), where declining frontier model prices accrue directly to the business as higher capability at a fixed cost. On the other side sits Usage-Based Billing (UBB) driven by long-running agentic workloads. This spending cannot be judged by seat economics because there is no seat to divide the cost by. It is a variable operational expense scaling with the volume of work delegated.

When organizations attempt to manage usage-based intelligence with SaaS budgeting mindsets, the leaks begin immediately. Let’s see through the lens of the new Tokenomics Foundation, which maps token economics into three distinct domains: Production (energy and capital), Consumption (allocation, forecasting and optimization), and Value (revenue, productivity, quality, etc.), and figure out if we can pinpoint exactly where and why the financial model breaks down:

1. The Multiplier Effect of Autonomous Agents

Conversational AI operates on human-to-model turns; you ask a question, the model responds, and the interaction pauses. But Agentic AI operates on model-to-model and model-to-software loops.

A single high-level user prompt (i.e. "evaluate this vendor contract against our regional compliance guidelines") triggers an autonomous cascade. The orchestrator calls a reasoning model, executes tool calls against external software, spins up specialized sub-agents, queries vector databases, tests code in sandboxes, and re-evaluates the output. As Shruti Koparkar from NVIDIA’s Accelerated Computing Team notes, the number of turns and LLM calls in agentic workflows multiplies token consumption exponentially compared to standard chat. If an agent enters an inefficient reasoning loop or hits an edge case, it burns through thousands of unconstrained, billable operations in seconds before returning a single outcome.

2. Context Window Creep and the Asymmetry Trap

Modern foundational models are stateless—they carry no memory of prior interactions. To maintain conversational continuity, an application must resend the entire conversation history, system instructions, and retrieved context on every turn.

Engineers call this Context Window Creep. On turn one, you might consume 34 tokens. By turn three or four, resending transcripts, tool descriptions, and grounding documents can easily push the input payload past several thousand tokens per call.

Compounding this is the fundamental pricing asymmetry of AI providers: Output tokens cost 3x to 5x more than input tokens because generating text requires real-time, sequential GPU computation rather than parallel reading. When teams deploy reasoning models that generate thousands of internal "thinking tokens" or fail to constrain generation limits, the resulting invoice can explode by 700% without adding any measurable business insight.

3. The "Frontier Overkill" and Jevons Paradox

There is a widespread executive misconception that "falling token prices" will naturally rescue AI margins. J.R. Storment from the Linux Foundation pointed out that we face Jevons Paradox in enterprise computing: as models become more capable and cost per turn drops, organizations do not spend less; they invent exponentially more complex agentic workflows, completely overwhelming efficiency gains.

Compounding the problem is model misallocation. Accenture’s research reveals that only 10% to 20% of enterprise tasks justify frontier-grade reasoning models. The remaining 80% to 90% consist of structured, repeatable steps that can be handled just as accurately by small language models (SLMs), fine-tuned open-weight models, or simple deterministic code. Deploying a flagship reasoning engine to process standard customer tickets or summarize emails is the organizational equivalent of hiring a specialized senior partner to do basic data entry.

The Way Forward: Moving from Experimentation to Token Discipline

If your organization is going to join the 6% of high performers attributing genuine EBIT impact to AI, leadership must start managing it as an operational supply chain.

Here are five foundational disciplines for leaders to consider before scaling further:

  1. Focus on Growth Through Innovation: Redesign workflows rather than simply layering AI onto legacy processes. High performers manage AI investment as a portfolio which covers everyday productivity, functional workflows that improves repeatable tasks and strategic bets that build market differentiation, all while tracking net impact the P&L.

  2. Implement Intelligent Model Routing: Eliminate single-model dependency. Adopt routing architectures (i.e. open-source dynamic routers, or custom gateways) that evaluate incoming prompt complexity. Direct bounded, high-volume tasks to lightweight models or cached data, and reserve frontier reasoning engines strictly for multi-layered, ambiguous, and high-risk decisions.

  3. Architect for Token Efficiency: Enforce engineering guardrails early. Leverage prompt caching (which slashes repeated input costs by up to 80–90%), dynamic tool filtering, and state summarization to kill context creep.

  4. Governance is a key priority: High performers do not treat governance as a post-launch compliance; they treat it as financial control infrastructure that allows them to see overruns in time to act on them, it becomes the operating layer that determines which AI work can scale.

  5. Measure Accepted Work per Dollar (AWpD): Vanity metrics like "number of prompts executed" or "active monthly users" tell you nothing about financial health. High performers benchmark against real business output: What is the cost per resolved support ticket? What is the cost per verified compliance review? What is the accepted work per dollar spent?

Tokenomics looks at all AI costs, not just token costs. It aligns total technology and labor investment against business value. But, as with its main object of study - artificial intelligence – we should expect rapid changes and updates from the knowledge it produces.

References

·         Tokenomics Foundation. The Tokenomics Foundation - AI Value

·         McKinsey. State of AI in 2026: On the road to ROI.

·         Doit/Sapio Research. The AI Spending Data Your Board Is About to Ask You About.

·         PwC. PwC’s 29th Global CEO Survey. Leading through uncertainty in the age of AI.

·         LinkedIn. Jared Spataro. The two very different bills inside your AI spend.

·         Accenture. AI is on your P&L. Most companies are only reading half of it.

·         OpenAI. How to manage AI investments in the agentic era.

·         Youtube. Microsoft Mechanics | Tokenomics | The new AI currency & your options explained.

·         Youtube. Nvidia | Inside AI Tokenomics: How to profitably turn tokens into Business Value.

·         Youtube. Tokenomics Foundation. What is Tokenomics? Production, Consumption, Value and Open Collaboration on Emerging Answers.

 

Next
Next

From Vision to Execution: Welcome to Via MotuS