Home Platform Implementation Transformation Blog About Get started
Back

Tokenomics: why your AI bill is rising faster than your AI value

Tokenomics — AI cost rising faster than AI value

The cost of a given level of AI capability has been falling for two years. Providers keep cutting rates, shipping cheaper models, competing on who can serve a million tokens for less. Like for like, AI has been getting cheaper to run. And yet the line on the CFO's dashboard goes the other way.

That gap now has a name. The FinOps Foundation named it formally this year: tokenomics. The discipline of treating AI consumption as a cost to be understood and managed rather than a bill to be discovered. It exists because enough finance and engineering leaders hit the same wall in the same quarter. The unit got cheaper, the usage exploded, and nobody could explain the difference.

The mechanics are not mysterious. A chat query costs a fraction of a cent. An agent running a multi-step workflow, reading a repository, calling three internal APIs, retrying when it fails, can consume a thousand times more for a single task. Multiply that by every developer with a coding agent, every custom agent wired into the ERP, every workflow somebody built and forgot to switch off. The unit price falls and the bill still triples, because consumption is growing faster than price is dropping. The bill is where those two lines cross, and right now they are crossing in the wrong place.

The unit price falls and the bill still triples, because consumption is growing faster than price is dropping.

Then there is the other half of the phrase, the part people hurry past: AI value. And here it is worth being honest, because the market mostly isn't.

The boards have noticed the spend. In the space of a quarter they went from asking how to expedite AI adoption to demanding the return on it. The uncomfortable answer, almost everywhere, is that the value side cannot be cleanly measured. Not because the value isn't there, but because isolating what a given workflow actually returned — in revenue, in hours, in risk avoided — is genuinely hard, and it is hard for everyone. Anyone handing you a tidy productivity ROI figure is handing you a model's guess dressed up as an accounts entry. The honest version is that the value half of tokenomics is unsolved across the industry, and pretending otherwise is how you end up with a leaderboard that rewards burning tokens for their own sake.

So the two halves are not equally tractable. The value half is hard for everyone. The cost half is completely knowable, and almost nobody knows it. That asymmetry is the whole opportunity, because it tells you where to start.

Knowing the cost half looks like four things, none of which require solving the value problem first.

  1. Attribution. Spend broken down by user, team, model and workflow, not arriving as one undifferentiated invoice at month end. The moment you can attribute, you can separate the coding agent paying for itself ten times over from the loop that has been quietly burning tokens since March. A number you cannot break down is a number you can only flinch at.
  2. Right-sizing. Most requests do not need the most expensive model. The default sends them there anyway. Routing each request to the cheapest model that can actually do the job is the single largest lever on the bill. It is the lever Marc Benioff was reaching for when he said out loud that he wished for a smart router to decide which queries genuinely needed the frontier model and which did not.
  3. A savings ledger. This is the part that turns cost control into something you can defend. Every time a request that would have defaulted to the frontier model gets routed to a cheaper one that does the job, you book the difference. Actual cost against default cost, logged, auditable, added up. That ledger is the closest thing to a clean ROI number that exists in AI today, because it is cost avoided rather than value imagined. It is the platform proving its own spend, in your own figures.
  4. Caps. Budgets enforced ahead of the surprise instead of discovered behind it. The first signal of a runaway workflow should be an alert, not an invoice.

One honest boundary, because it matters. This works on the AI traffic that routes through your gateway. Someone on a personal account in a browser tab is a different control problem, one a CASB handles, and we will say so rather than pretend the meter sees everything. But for the governed path — the agents, the coding tools, the workflows the company actually runs on — every token can be attributed, routed and capped.

None of this proves the value half, and it is not meant to. It is meant to do the thing you can do this quarter while the industry works out the thing nobody can do yet. There is a quieter point underneath it too. You cannot measure what a workflow returned until you can say what it cost and who ran it. Cost attribution is not the answer to the ROI question. It is the precondition for ever asking it properly.

The bill will keep rising, because consumption pricing does not forgive inattention. The organisations that come through this well will not be the ones that used the most AI or the least. They will be the ones who can say, with evidence, what each token cost, who spent it, and what the cheaper alternative would have been.

MisaLabs builds the enterprise AI control plane: token-level cost attribution, model routing and a savings ledger you own, governed at one gateway and deployed entirely in your environment.

Talk to us →