Archiving files in Amazon S3: Glacier, lifecycle rules and recovery, the complete guide
How to keep years of business files in Amazon S3 cheaply and safely: choosing storage classes, moving files with lifecycle rules, restoring from Glacier, and protecting against deletes.
Price the archive by object count as well as terabytes, send only large and rarely opened files to Glacier, let lifecycle rules do the moving, and turn on Versioning before you need it. Decide in advance who can restore files and who can delete them, because those are the two actions an archive gets wrong.
Most business buckets are mostly archive: last year's projects, signed contracts, raw footage, scans nobody has opened since they were uploaded. Keeping all of it in S3 Standard is simple and expensive. Moving it to Glacier is cheap until someone needs a file back the same day, or until millions of small files turn the per-object overhead into the biggest line on the bill.
This guide collects our articles on running an archive in S3, in the order the decisions usually come up: what it costs, where files should live, how they get moved and brought back, and how to make sure a delete is never final by accident.
Know what the archive costs
The S3 bill rarely matches the per-gigabyte price. Object count, minimum storage durations, the 128 KB minimum billable size on the infrequent-access classes, retrieval fees and forgotten multipart uploads all change the arithmetic. Start here before moving anything.
Amazon S3 pricing explained: what a business file archive really costs
Choose where cold files live
S3 Glacier Deep Archive is the cheapest storage in AWS, at about $1 per terabyte per month, but it has a 180-day minimum and about 40 KB of overhead per object. It suits large files kept for years and rarely opened, and costs more than it saves for small files or anything you might delete soon.
S3 Glacier Deep Archive: when it saves money and when it costs more
Move files automatically
Lifecycle rules move objects between storage classes and expire old versions without anyone running a script. In a versioned bucket they also decide how long deleted and overwritten files stay recoverable, so the rule is a retention policy as much as a cost one.
Bring files back
A file in Glacier Flexible Retrieval or Deep Archive can't be opened until it is restored, which takes minutes to hours depending on the tier and creates a temporary copy that expires. Plan who can request a restore and who pays for it.
How to restore files from S3 Glacier Flexible Retrieval and Deep Archive
Make deletes recoverable
With Versioning on, a delete only adds a delete marker and an overwrite keeps the old version, so both can be undone. Object Lock goes further and makes versions impossible to delete for a set period. Replication keeps a second copy in another bucket or Region, with rules about what it skips.
How to recover deleted or overwritten files in Amazon S3Amazon S3 Object Lock: governance vs compliance mode, legal holds and retentionAmazon S3 replication explained: same-Region, cross-Region, and what never gets copied
Control who can read it
Every S3 object is already encrypted at rest. The choice between SSE-S3 and SSE-KMS changes who can decrypt it and which permissions a download needs, which matters once tools other than the AWS Console read the archive.
Amazon S3 encryption at rest: SSE-S3, SSE-KMS, DSSE-KMS or SSE-CNext guide: Previewing and sharing files in Amazon S3
Starter is free. Deploy a scoped role with CloudFormation, sign in, and browse, without handing anyone an access key.
Primary sources
Discussion
0 comments · open to guests · moderatedLiked this? Get the next article by email. No schedule, no filler, one click to leave.