JRJérôme RaguilletFinOps · Cloud & AI
← All insights
FINOPS · CLOUD & AI

Tokenomicon and FinOps X Amsterdam: what exploding AI costs change for FinOps

In his post about the Amsterdam events, Jérôme Raguillet offers a reading of the emergence of Tokenomics: dramatic token consumption growth, hidden costs missed by most dashboards, and early tangible standards. A framing to help FinOps teams decide where to focus their efforts in 2026.

Two days in Amsterdam, one massive finding

The Tokenomicon and FinOps X were held in Amsterdam on 22-23 September, bringing together more than 500 practitioners. Beyond the variety of sessions, Jérôme Raguillet notes in his original post a shared finding: 2026 would be the year when AI budgets started to burn, faster than expected.

The idea deserves to be taken seriously, even though it remains a field impression attributed to the post. When several hundred practitioners from different organizations converge on the same diagnosis, it usually signals a gap between the planned budget trajectory and the one actually observed. If not addressed early, that gap is later paid for through painful mid-year arbitrations.

That is precisely FinOps work: detecting these deviations as early as possible, before they become budget crises. But measuring the right thing is a prerequisite, and that is where the topic becomes complex.

Growth figures that force a rethink of allocation

Two data points anchor the post's argument. The first: per-developer token spend was multiplied by 18.6 over 9 months, according to Jellyfish data cited in the post. The second: Goldman Sachs projects 24x growth by 2030, reaching 120 quadrillions of tokens per month. These figures are presented by the author and are not verified here; but their order of magnitude is enough to frame the problem.

A growth rate of that scale makes budget allocation mechanisms designed for linear spending growth obsolete. Allocating an AI bill proportionally to headcount or revenue, then applying annual inflation, cannot work when the unit cost per employee moves by a factor above ten in under a year.

Suggestion to discuss with your teams: rather than allocating the past bill, project future consumption per use case, and negotiate budgets on that basis. It is a classic FinOps perspective inversion, but it becomes urgent when the curve is that steep.

The visible bill tells only part of the story

The most structural point in the post: the notion of nine layers of AI cost, presented based on a KostKompass analysis. These layers cover tokens, retrieval, orchestration, GPU infrastructure, cache, evaluation, governance, labor and waste.

In one real case analyzed and reported in the post, the visible layers accounted for only 62% of the total cost. This case illustrates the risk of blind spots in AI cost monitoring, without establishing that the proportion applies to all organizations.

The 62% figure deserves methodological reflection. It points to a simple, prudent exercise: map, for each AI use case, the nine layers, and document for each whether it is measured, estimated, or simply ignored. The exercise often reveals blind spots not out of negligence, but because today's tools were designed for classic cloud billing, where these layers did not exist or were bundled in.

The first deliverables of the Tokenomics Foundation

An important institutional signal reported in the post: the Tokenomics Foundation was launched in August under the Linux Foundation, with 30 members including SAP, JPMorgan and Microsoft. It is already delivering two concrete contributions.

The first is Big-T Notation: a way to measure how the cost of an AI task grows with its complexity, echoing the Big-O notation familiar to engineers. The value is shifting the conversation from micro to macro: instead of optimizing the price of a token, analyze the complexity class of the usage. As a design hypothesis, two functionally equivalent implementations may exhibit different cost trajectories; the difference needs to be measured for the use case concerned.

The second is version 1.5 of the FOCUS standard, which distinguishes input, output and cache tokens in bills, and attaches each usage to a team or an agent. Combined with a common standard to compare OpenAI, Anthropic, Google and self-hosted models, this finally gives teams a vendor-agnostic allocation foundation.

FinOps and Tokenomics: two complementary questions

The post condenses the topic into a sentence worth pondering: "FinOps asks whether the AI bill matches the plan. Tokenomics asks whether the plan can cost less without losing quality." The author adds that these two questions are complementary, and that most teams staff only one.

This is a useful lens to assess your own maturity. The first question belongs to control: budget tracking, allocation, forecasting. The second belongs to design: architecture, model selection, task granularity. Both require different skills and mandates, and a team covering only one axis will systematically leave a blind spot.

Pragmatic suggestion: map who in your organization owns each question. If nobody owns the second, AI usage design is probably happening without a cost constraint, which partly explains observed waste levels.

The agentic coding blind spot and a checklist to act

A final salient point from the post: agentic coding remains the industry's blind spot, described as the fastest-growing and least-tooled cost line. That, the author concludes, is where the work lies. This qualification is presented as an event observation, to be checked against your own context.

To turn these takeaways into action, here is a suggested checklist, to be adapted to your organization:

1. Check whether your AI bill is allocated by use case rather than as a global line; if not, that is the priority workstream.

2. Map your nine AI cost layers and document each layer's degree of visibility.

3. Test the Big-T Notation reading on a representative use case to measure how cost grows with complexity.

4. Assess your billing tools' FOCUS 1.5 coverage, notably input/output/cache distinction and per-team or per-agent attribution.

5. Explicitly identify who staffs the Tokenomics question, distinct from the FinOps question.

6. Specifically review your agentic coding usage, as the fastest-moving line.

7. Revise your 2026-2030 budget projections in light of the orders of magnitude cited in the post.

The question closing the post remains the best way to open an internal discussion: is your AI bill allocated by use case or as a global line?

The post behind this insight

Expanded from the LinkedIn post. The links below come from the original post; listing them does not imply independent verification.

LinkedIn ↗

Links included in the post