The FinOps black hole of AI consumption
In the original post, Jérôme Raguillet recalls an observation that has long shaped how technical teams relate to cost: for years, engineers only saw cloud costs once in production. The bill arrived after the fact, with no ability to act on the architecture decisions that had determined it. This is what he calls the FinOps black hole.
With generative AI, that pattern is repeating itself in a new form. Consumption is measured in tokens, and depends on the model chosen, the routing between models, and how usage is put into production. These parameters are decided very early on, by teams that often have no visibility into their financial translation. The gap between technical decisions and economic impact therefore tends to recur, even though consumption granularity is finer than in traditional cloud.
The question raised by the author is simple to state, harder to answer: what is the full cost of your AI stack, and what real gain can better routing deliver?
What Infracost announces with Infracost AI
According to the original post, Infracost has now presented Infracost AI, currently in early access. It is worth stressing that everything that follows reflects capabilities announced by the vendor, not observed or independently verified results.
The stated goal is to bring together AI consumption — tokens, models and routing — and then simulate the impact of changing a model or a routing strategy. In other words, the idea is to move cost analysis upstream in the development cycle, at the point where design choices are still reversible.
This logic extends what FinOps has sought to do from the start: connect architecture decisions to their economic expression, and give the teams that build the means to understand what they spend, and why. Applying this to AI adds a specific twist: routing — the ability to direct each request to the most suitable model — becomes itself an optimization lever to be simulated.
The promise: four expected benefits
The original post lists four expected benefits, described as promising but still to be confirmed in practice:
Visibility before the bill: knowing the financial impact of a model or routing choice before consumption is committed, rather than after receiving the bill.
Comparison of model scenarios: being able to place several models side by side and assess their cost in a given usage context, rather than comparing unit prices out of context.
Routing simulation: estimating what a change in routing strategy — for instance directing certain request types to a cheaper model — would yield in savings or in overspend.
Reconciliation between architecture and cost: making technical designs and financial data talk to each other, so that every component of the AI stack can be linked to its contribution to total cost.
What will need to be verified in practice
Perhaps the most important point in the post is this one: these capabilities are, at this stage, announced. The author therefore calls for caution and lists five verification points, which form a useful reading grid for any solution of this kind.
First, the providers and models covered. A solution that covers only part of the ecosystem will not allow an honest comparison of scenarios: actual coverage determines the real scope of the simulation.
Second, the handling of contractual discounts. Public prices rarely reflect what an organization actually pays. If negotiated discounts are not factored in, simulations risk diverging significantly from the final bill.
Third, the quality of usage data. A simulation is only as good as the consumption data it relies on. Poorly categorized volumes or inaccurately captured call patterns will produce misleading projections.
Fourth, the calculation of cost per useful outcome. Counting tokens is one thing; measuring what they produce is another. This is probably the most discriminating point, since it determines the ability to compare models on a basis that makes sense to the business.
Fifth, integration with gateways and observability tooling. A solution that does not fit into existing flows risks remaining a parallel analysis layer, with no effect on real decisions.
Open weight models: cost is not quality
The original post issues an important warning: an open weight model may cost less, but its quality must be evaluated on the real use case.
In other words, the trade-off between models cannot be made on the sole criterion of unit token price. A cheaper model that produces insufficient results can trigger retries, manual rework, fixes, or even functional failures whose cost far exceeds the initial saving. Conversely, a more expensive but better-suited model can reduce the number of calls or the need for human supervision.
This is exactly where the notion of cost per useful outcome comes into its own: it requires relating spend not to the quantity of tokens consumed, but to what that consumption actually delivers for the organization. Any scenario comparison that overlooks this dimension risks favoring choices that look cheap and are structurally expensive.
Questions to ask and a checklist before getting started
Pending validation under real conditions, here are methodological suggestions — clearly distinct from the capabilities announced by the vendor.
Preliminary questions: do you currently have a consolidated view of your AI consumption, across all models and all use cases? Do you know which teams and which use cases drive that consumption? Is routing currently a documented decision or an implicit setting?
A suggested checklist before evaluating an AI cost simulation solution:
1. Map your providers and models in actual use, and check they are covered by the solution under consideration.
2. Inventory your contractual discounts and verify they can be factored into simulations.
3. Audit the quality of your usage data: completeness, granularity, ability to link consumption to use cases.
4. Define what a useful outcome means for your major use cases, and how you will measure it.
5. Check compatibility with your existing gateways and observability tools.
6. Define a limited pilot scope, with documented assumptions, before any generalization.
The final question from the post remains the guiding compass: what is the full cost of your AI stack, and what real gain can better routing deliver? Until that question has a quantified, contextualized answer, every optimization decision rests on untested assumptions.
The post behind this insight
Expanded from the LinkedIn post. The links below come from the original post; listing them does not imply independent verification.
LinkedIn ↗