JRJérôme RaguilletFinOps · Cloud & AIFR
← All insights
FINOPS · CLOUD & AI

Cosmos DB: moving from 30,000 RU/s around the clock to autoscale, and what hourly billing reveals

A reported case where 30,000 RU/s provisioned continuously only served two hours a day: switching to autoscale brings the bill from $5,700 to $2,280 per month. The reasoning rests less on the mode chosen than on understanding the hourly billing rule and matching that mode to the actual load profile.

Throughput sized for the peak, not for actual usage

The starting point is easy to describe and common in distributed databases: throughput provisioned at 30,000 RU/s at all times, while load peaks lasted only two hours a day. Capacity is therefore sized for the worst moment of the day, but paid for around the clock.

RU/s is the unit of throughput: it is the capacity reserved for the container, regardless of how many requests are actually executed. Provisioning continuously on the basis of the peak is like booking a full-capacity meeting room for a one-hour meeting, all day long. The 30,000 RU/s figure and the two-hour duration are reported in a specific case; they describe a situation, not a standard.

What autoscale changes, and what it does not

Autoscale adjusts provisioned throughput between 10% and 100% of the configured maximum. With a maximum of 30,000 RU/s, the theoretical operating range therefore runs from 3,000 to 30,000 RU/s — arithmetic derived directly from the stated rule, to be checked in Microsoft documentation for each configuration.

The decisive nuance lies elsewhere: this is not purely per-request billing. For each hour, billing depends notably on the maximum RU/s level reached by autoscale during that hour. In other words, an hour in which load briefly climbs near the maximum is billed at that level, even if the rest of the hour is quiet. Hourly smoothing therefore plays a central role: the relevant question is not only how long the peak lasts, but at what level consumption peaks within each hour.

This rule also explains the limit of the reverse reasoning: an active container does not fall to zero. The 10% floor of the configured maximum keeps generating a baseline cost, which argues for setting that maximum as tight as possible.

The four-step migration method

The reported approach has four movements, and their order matters: 1) a 30-day analysis of actual traffic per container; 2) keeping the Autoscale maximum at 30,000 RU/s to absorb peaks; 3) a comparison with Serverless for genuinely sporadic containers; 4) checking partitions, regions and consistency levels.

The 30-day window separates real variability from perceived variability: many teams overestimate how often peaks occur. Keeping the same maximum avoids changing two variables at once — the effect of the billing mode is isolated without cutting peak headroom. The Serverless comparison targets containers whose traffic is genuinely episodic, where a model without provisioned throughput may be more relevant, but that choice remains a container-by-container trade-off, depending on expected latency and access patterns.

The fourth point is the one most often forgotten: changing throughput mode without revisiting the partitioning strategy, geographic distribution or consistency level can move the bill… and performance. Consistency has a cost in RU, and regions multiply throughput: these are cost levers just as much as the billing mode.

The arithmetic: $5,700 to $2,280 per month

The reported amounts are as follows: $5,700 per month before, $2,280 per month after, a difference of $3,420 per month, or 60%. Over twelve months, 3,420 × 12 = $41,040, rounded to $41,000 a year.

These figures come from a single case and do not constitute a general rule: they depend on the number of containers, the actual hourly profile, replicated regions and the consistency level selected — elements that are not part of the calculation. The 60% figure should therefore be read as an order of magnitude from one specific situation, to be replayed against your own metrics before any decision. It also assumes constant performance between the two configurations: that is a hypothesis to be validated by measurement, not a given.

Autoscale, Serverless or provisioned throughput: the selection criteria

Three modes, three profiles. Autoscale is presented as suited to variable loads: it absorbs peaks without paying the maximum continuously. Serverless addresses genuinely sporadic loads, with the latency and ceiling constraints that model implies. Provisioned throughput, by contrast, can remain more economical for a stable load close to the maximum: in that case, automatic adjustment adds nothing and hourly billing at the level reached does not favour a change.

Useful questions to ask: what share of hours actually exceeds a high consumption threshold? Are the peaks predictable or erratic? Is per-container traffic uniform or widely dispersed? Does each container justify its own arbitration? The answers determine the mode far more than overall volume does.

The KPI that matters: cost per million requests or per business transaction

RU/s is an intermediate indicator, not an indicator of value. The proposed KPI — cost per million requests or per business transaction, at constant performance — puts the bill back in context: lowering cost by degrading latency or consistency is not an optimisation, it is a transfer of cost onto users.

The phrase "at constant performance" is essential: it rules out comparing two configurations where one has been silently degraded. It is this guardrail that makes the 60% gap interpretable in the reported case.

Suggested checklist before switching

The following list is a methodological suggestion, to be adapted to your context; it does not follow from the case cited. 1) Measure traffic per container hourly over at least 30 days, to identify the maximum levels reached and not just averages. 2) Calculate the theoretical bill under both modes, including the 10% floor of the configured maximum. 3) Set the Autoscale maximum at the observed real peak, not at the existing provisioned level out of habit. 4) Identify genuinely sporadic containers and test the Serverless option on those only. 5) Check partitioning, regions and consistency level before and after the change. 6) Define a unit-cost KPI at constant performance and track it over time. 7) Plan an exit point: if load stabilises near the maximum, provisioned throughput may become the right choice again.

The real lesson is methodological: a billing mode does not reduce a bill by itself; the fit between the mode and the load profile does. The calculation is only worth anything if the load profile was measured before it was assumed.

The post behind this insight

Expanded from the LinkedIn post. The links below come from the original post; listing them does not imply independent verification.

LinkedIn ↗

Links included in the post

Continue reading

My FinOps approach · My Cloud and AI consulting services