JRJérôme RaguilletFinOps · Cloud & AI
← All insights
FINOPS · CLOUD & AI

EBS snapshots: when archiving increases the bill instead of reducing it

An AWS case with 94,000 snapshots and 1.9 PB of referenced data shows that the decision criterion is not a snapshot's age, but its role in the restore chain and its full cost after archiving. An analysis of the mechanics, the triage reasoning, and the safeguards to put in place.

An inventory finding that says it all

Before even discussing savings, it is worth setting the scene as the original post describes it: every virtual machine backed up its disk, every environment recreated snapshots, and no one was governing their lifecycle. The reported AWS result speaks for itself: 94,000 snapshots and 1.9 PB of referenced data. This kind of situation does not stem from a one-off mistake but from a governance gap: as long as a snapshot exists, it is billed, whether it still serves a purpose or not.

The key lesson is that snapshot storage growth is often silent. Unlike an oversized instance that stands out in a dashboard, snapshots accumulate by the thousands without alerting anyone. This is precisely why an initial inventory is indispensable before any optimization action.

The archiving trap: full snapshots

The original post raises an important and often misunderstood technical point. Snapshots in the Standard tier are incremental: each snapshot only stores the blocks changed since the previous one. When a snapshot is moved to EBS Snapshots Archive, AWS converts it into a full snapshot. In other words, the apparent benefit of incremental storage disappears at exactly the moment you expected to save.

Another constraint is a 90-day minimum billing period for archived snapshots. Early deletion or permanent restoration remains possible, but the remaining days are charged. Combined with full-snapshot conversion, this rule can increase the cost of systematically archiving old restore points. Archiving is therefore a case-by-case decision, not a technical requirement to retain data for 90 days.

The right criterion: the role in the restore chain

The central idea of the original post deserves emphasis: the right criterion is not the snapshot's age, but its role in the restore chain and its full cost after archiving. A recent snapshot can be useless if it is redundant, while a monthly or annual restore point can be valuable even if it is old.

The resulting triage reasoning is as follows. First, identify orphan snapshots, but only after validating the owner and the retention rules: a snapshot that looks orphaned on paper may be the only restore point for a critical system documented nowhere. Second, keep the still-active incremental chains in Standard, because that is where the incremental format remains most cost-effective. Third, archive only certain monthly, annual, or regulatory restore points, that is, the ones whose value justifies the extra cost of a full snapshot. In the post, this approach was complemented by automating retention and deletion with Data Lifecycle Manager, so that the problem does not recur.

Reported figures and a financial reading

The results indicated in the original post are as follows: a bill of $16,200 per month before optimization, reduced to $4,700 per month after, a difference of $11,500 per month. Over a year, that represents $11,500 multiplied by 12, i.e. $138,000 per year. These figures come from the author of the post; they are not independently verified and obviously depend on context: starting volume, applicable retention policy, and the pricing in effect at the time of the operation.

These amounts illustrate the overall result of the reported operation, without quantifying the contribution of each lever. The underlying logic combines validated deletion of snapshots no longer needed and selective archiving of restore points worth retaining. Savings must be recalculated for each environment.

Trade-offs and the questions to ask internally

Every triage decision involves trade-offs. Archiving too early exposes you to full-snapshot costs and incurs a minimum billing period of 90 days. Deleting too fast risks losing a regulatory restore point. Keeping everything in Standard avoids these risks but maintains the bill. The right level depends on your legal obligations, your disaster recovery objectives, and the actual frequency of restores.

A few practical questions to ask before acting: who is the functional owner of each snapshot chain? What retention rules apply, and are they written down somewhere? Which monthly, annual, or regulatory restore points are actually required? Is there a validation process before any deletion? These questions are not purely technical: they involve compliance, application teams, and operations.

A checklist to govern the lifecycle

To turn all of this into actions, here is a suggested checklist, to be adapted to your context. It is presented as a proposal, not as an independently validated method.

1. Build a complete inventory of snapshots and referenced volume. 2. Assign an owner and a retention policy to each chain, validated by the relevant teams. 3. Classify snapshots: active in an incremental chain, monthly, annual, or regulatory restore points, or orphans to delete. 4. Assess the full cost of each scenario before archiving, factoring in the conversion to a full snapshot and the 90-day minimum billing period. 5. Archive only the points whose value justifies it. 6. Automate retention and deletion, for example with Data Lifecycle Manager as mentioned in the original post. 7. Review the inventory periodically, because a snapshot's lifecycle is never settled once and for all.

The closing question in the original post sums up the challenge well: how many of your old snapshots still have an owner and a retention policy? As long as the answer is unclear, the bill will keep growing.

The post behind this insight

Expanded from the LinkedIn post. The links below come from the original post; listing them does not imply independent verification.

LinkedIn ↗

Links included in the post