JRJérôme RaguilletFinOps · Cloud & AI
← All insights
FINOPS · CLOUD & AI

R9g and R9gd on AWS: what the new Graviton5 generation really changes for your FinOps strategy

AWS has announced R9g and R9gd instances, built on Graviton5, with up to 25% more compute performance than R8g/R8gd. Beyond the technical announcement, the real issue for a FinOps team is the impact on existing commitments: Compute Savings Plans, EC2 Instance Savings Plans and Reserved Instances do not react the same way to a generation change. This article walks through the reasoning, the decision criteria and a phased migration approach.

A new line on the price-performance curve

According to the original post, AWS has made the R9g and R9gd instances available, built on the Graviton5 chip. The announced characteristics are as follows: up to 25% more compute performance compared with the R8g and R8gd generations, DDR5 memory running at 8,800 MT/s, and five times more L3 cache.

These instances come in 11 sizes, from 1 to 192 vCPU, and the family is presented as a fit for databases, in-memory caches and real-time analytics. It is therefore an offering aimed at so-called "memory-bound" workloads, that is, those whose performance depends heavily on memory bandwidth and capacity.

These figures are those stated by the vendor in the original post; they are not independently verified here. The useful question for a FinOps practitioner is not whether the announcement is accurate, but what it concretely implies for the pricing commitments already in place.

Why the FinOps impact is more nuanced than it looks

A new instance generation rarely changes a cloud bill overnight. What it does change, however, is the trade-off calculation between three things: the performance available per dollar spent, the coverage provided by existing commitments, and the cost of a technical migration.

The key point made in the post is that the different AWS discount mechanisms do not react uniformly to a change of instance family. Some automatically follow the new generation, others do not. Serious FinOps work therefore begins by classifying commitments into three categories according to how they behave with respect to R9g.

In other words, the new generation can improve cost per transaction for some workloads, but it does not render all existing commitments obsolete. It is this nuance, rather than the raw announcement, that should guide the decision.

Compute Savings Plans, Instance Savings Plans and Reserved Instances: three distinct behaviors

First case: the Compute Savings Plan. According to the post, it can continue to cover R9g instances, because it applies regardless of family, size and region. This is the most flexible mechanism: a commitment taken on instances of one family remains valid if you switch to R9g, as long as the committed hourly amount is being used.

Second case: the EC2 Instance Savings Plan. It is tied to a family and to a region. Consequently, a commitment taken on R8g instances does not automatically cover R9g instances. Switching a fleet without anticipating this point can leave part of the commitment unused while off-commitment consumption is billed at on-demand prices, which mechanically degrades the coverage rate.

Third case: Reserved Instances. A simple size modification stays within the same family and generation; it does not convert an R8g reservation to R9g. A Convertible reservation can be exchanged into another family, subject to AWS conditions and an exchange quote. A Standard reservation does not offer that exchange. Before switching, check the exact type, compatibility and cost of the commitment.

Decision criteria: who really gains from the change?

Not all memory-bound workloads benefit equally from R9g. The first criterion is the nature of the constraint: if performance is limited by memory bandwidth or cache size, a five-times-larger L3 cache and DDR5 at 8,800 MT/s can translate into substantial gains. If the constraint lies elsewhere, in the network or in an application dependency, the price-performance improvement will probably be marginal.

The second criterion is the granularity of the sizes: with 11 sizes from 1 to 192 vCPU, the family covers a wide range of profiles. This can make it possible to right-size oversized instances, but only if usage data justifies it.

The third criterion, often underestimated, is the cost of the migration itself: qualification effort, compatibility testing, regression risk. A 25% unit gain on paper can be wiped out by an expensive migration project. As a hypothetical example, a database whose unit cost drops by 10% after migration does not always justify the effort if the existing commitment is already optimized. This figure is illustrative and is not a forecast.

A four-step approach to avoid breaking anything

The post proposes a four-step FinOps routine, which deserves to be spelled out. Step one: identify the memory-bound workloads that genuinely gain from the change. This requires analyzing usage metrics, understanding which services are truly memory-limited, and ruling out candidates for whom the gain would be purely theoretical.

Step two: identify the exact type of each commitment. Compute Savings Plans, EC2 Instance Savings Plans and Reserved Instances do not behave the same way when the family changes. This mapping is the prerequisite for any reliable simulation.

Step three: simulate coverage and utilization after the migration. The goal is to verify, before switching, that the existing commitment will still be fully consumed, and that the off-commitment share will not explode during the transition phase.

Step four: migrate progressively, then right-size. A wave-based migration makes it possible to validate performance assumptions at small scale, and right-sizing at the end of the process avoids reproducing on R9g the over-provisioning inherited from R8g.

A preparation checklist before any switch

Here is a suggested checklist, to be adapted to your context. It is provided for guidance and does not replace an analysis specific to your environment.

1. Inventory commitments by type: Compute Savings Plans, EC2 Instance Savings Plans, Reserved Instances, with their family, region, size and expiry date. 2. Identify candidate workloads among databases, in-memory caches and real-time analytics, checking that the constraint is genuinely memory-related. 3. Check the compatibility of each Reserved Instance and determine whether a modification or an exchange is required. 4. Estimate the impact of a migration on the coverage rate of Instance Savings Plans, family by family and region by region. 5. Plan a progressive migration with measurement points before and after each wave. 6. Right-size once the migration has stabilized, relying on actual usage data rather than on theoretical vCPU equivalences.

In conclusion, a new generation such as R9g improves the potential cost per transaction, but it does not by itself disrupt your commitments. The right FinOps response is neither to ignore the announcement nor to switch everything at once: it is to map your commitments, target the workloads that genuinely gain, and migrate in a controlled manner. The technical claims and figures quoted in this article come from the original post and should be verified against the vendor's official documentation.

The post behind this insight

Expanded from the LinkedIn post. The links below come from the original post; listing them does not imply independent verification.

LinkedIn ↗

Links included in the post