JRJérôme RaguilletFinOps · Cloud & AI
← All insights
FINOPS · CLOUD & AI

Tokens are billed, value is measured elsewhere: structuring generative AI cost tracking

Dashboards that only display millions of tokens cannot explain spending or support trade-offs. This article details a three-level approach — visibility, attribution, decision — before any optimization, and offers a practical checklist inspired by Jérôme Raguillet's original post.

The meter is not the balance sheet

The starting point of the post is a simple observation: tokens are billed, but value is measured elsewhere. A dashboard that only displays token volumes, even expressed in millions, does not say which team spent, for which product, or whether the work was completed. In other words, the meter is running, but the balance sheet remains silent.

This situation is far from trivial. With growing use of coding assistants and agents, volumes consumed rise quickly, and the associated billing follows. Without an analytical framework, these amounts show up as an opaque line item that is hard to defend in budget reviews, because no one can tie the spend to a deliverable or a business outcome.

One point deserves emphasis: consumption is not a problem in itself. The problem is the absence of a link between that consumption and the value produced. A high volume may be perfectly justified if it corresponds to completed, useful tasks; it may also hide silent waste. Without visibility and attribution, it is impossible to tell the two apart.

Level 1: visibility, or knowing where every token comes from

The first level proposed in the post is visibility: identifying the provider, the model, the account, the key, and the user or service behind every call. It is the foundation of everything else, because spending that cannot be tied to an actor is spending that can be neither explained nor controlled.

In practice, this means systematically capturing these dimensions at the level of calls to models: which provider and which model were used, along with their pricing differences; which account and which technical key served as the pivot; and above all, which user or which service is responsible. The distinction between a human user and an automated service matters, because consumption patterns and optimization levers differ radically.

This level answers seemingly trivial questions that often go unanswered in organizations: how many providers are actually in use? Which models are being called, at what unit prices? Are there shared keys that make charge-out impossible? These are checkpoints to verify before even talking about optimization.

Level 2: attribution, or linking spend to product and task

The second level is attribution: tying each unit of consumption to a product, a session, and a task. This is the move from description to explanation. Knowing that a team spent a certain amount is useful; understanding that this spend corresponds to a given product, to specific work sessions, and to specific tasks is what enables action.

For AI-assisted coding, attribution by session and by task is particularly relevant. A work session can involve many calls: rephrasing, iterations, corrections. Without that breakdown, spend remains an indigestible aggregate; with it, spend becomes a series of events that can be qualified: this task cost so much, it required so many iterations, and it did or did not complete.

The post notes that the FinOps Foundation itself places attribution and governance before advanced chargeback. In other words, this is not a marginal idea but a recognized maturity principle: costs should only be charged back to teams once they can be reliably attributed and understood by everyone. Chargeback without reliable attribution produces only disputes, not trade-offs.

Attribution also raises design choices: task identifier granularity, naming conventions for products and projects, instrumentation of agents and services. These choices must be made early, because retrofitting an untracked history is far more costly than instrumenting it correctly from the start.

Level 3: decision, metrics that drive choices

The third level is decision-making: cost per validated task, quality, lead time, and alert thresholds. This is the level that turns reporting into a steering tool. Cost per validated task is arguably the central metric: it measures what completed work actually costs, not merely a quantity of computation.

Quality and lead time complete the picture. A low cost achieved at the price of degraded quality is not a saving; a high cost with a sharply reduced lead time can be an excellent investment. That is precisely because value is measured elsewhere than in token volume that these dimensions must be tracked together, over the same scope.

Alert thresholds, finally, make it possible to detect drift before it sets in: a session whose cost spikes, a task that accumulates iterations without completing, a service whose consumption becomes atypical. Properly calibrated, they trigger targeted investigations rather than generalized alarms.

These metrics must serve concrete decisions: continue, adjust a prompt, switch models for a given use case, or drop a use case that costs more than it returns. A metric that triggers no decision is just an additional measurement cost.

Optimize later, measure after: order matters

The post is explicit about the order of operations: only after the three levels do you optimize caching, routing, prompts, and model size. This sequencing is not arbitrary. Optimizing without visibility or attribution means hunting for savings without being able to verify where they materialize, or what they cost in terms of quality.

Each lever has its own logic. Caching avoids paying twice for identical content; routing sends each request to the model with the right level of capability; prompt work reduces iterations and superfluous output; model size choices match power to actual need. But all these levers share one requirement: being able to measure the effect obtained.

Hence the stated rule: gains must be measured after implementation, on the same scope. Comparing before and after requires a stable reference frame — the same tasks, the same scope, the same validation criteria. Without that discipline, variations driven by context get attributed to an optimization, and conclusions go wrong.

Checklist and closing question

The post ends with a question for readers: can you explain the spend of an AI coding session end to end? It is an excellent maturity test. If the answer is no, the following checklist — offered here as a suggestion, to be adapted to your context — can serve as a guide.

Suggested checklist: first, verify that every call is traceable by provider, model, account, key, and user or service; second, implement attribution by product, session, and task; third, define decision metrics, notably cost per validated task, along with quality, lead time, and alert thresholds; fourth, only then activate optimization levers — caching, routing, prompts, model size; fifth, measure gains after implementation, on a strictly identical scope; sixth, follow the sequencing recalled in the post, which places attribution and governance before advanced chargeback, as the FinOps Foundation indicates in the original post.

In short, the approach comes down to one sentence: make the spend explainable before reducing it. Tokens will keep being billed; but with visibility, attribution, and decision metrics, the value produced will stop being measured elsewhere.

The post behind this insight

Expanded from the LinkedIn post. The links below come from the original post; listing them does not imply independent verification.

LinkedIn ↗

Links included in the post