aws s3 sync explained: --delete, --exclude, --dryrun and the files it skips
How aws s3 sync decides what to copy, how --exclude and --include filters are evaluated, what --delete removes (and what it never touches), and the flags that stop a sync from doing something you did not expect.
aws s3 sync copies a file when it is missing, a different size, or newer at the source, and it never compares contents. Filters are evaluated in order and the last match wins, so write --exclude "*" before --include. --delete removes whatever is missing from the source listing, except excluded files, so run --dryrun first and keep versioning on for any bucket you mirror into.
aws s3 sync copies a directory to a bucket prefix, a prefix to a directory, or one prefix to another, and on the next run it only sends what changed. That makes it the default tool for backups, website deploys and "mirror this folder to S3" jobs.
It is also easy to misread. Sync does not compare file contents, filters are evaluated in an order that surprises people, and --delete removes files from the destination based on what the source listing contains. This guide covers what the command actually compares, how to write filters that do what you mean, and how to run --delete without losing work.
What sync compares
For each file, sync copies it when one of three things is true: the file does not exist at the destination, its size differs, or the source's last-modified time is newer than the destination's. Nothing else is compared. There is no checksum comparison, so a file edited without changing its size, and then given an older timestamp, is skipped.
Downloads from S3 to a local directory have one more wrinkle. By default a same-sized file is skipped unless the local copy is newer, which means a local file that is older but the same size as the S3 object stays as it is. Add --exact-timestamps to copy same-sized files whenever the timestamps differ at all.
--size-only goes the other way and ignores timestamps completely. It suits a destination whose timestamps are meaningless, such as a build directory regenerated from scratch on every run, where every file looks newer than the S3 copy.
# Local folder to a bucket prefix: uploads new and changed files only.
aws s3 sync ./reports s3://amzn-s3-demo-bucket/reports/
# Bucket to a local folder, copying same-sized files whenever the timestamps differ.
aws s3 sync s3://amzn-s3-demo-bucket/reports/ ./reports --exact-timestamps
# Prefix to prefix, including across buckets, without downloading anything locally.
aws s3 sync s3://amzn-s3-demo-bucket/reports/ s3://amzn-s3-demo-archive/reports/--exclude and --include: later filters win
Every file is included by default. Filters are then evaluated in the order they appear on the command line, and when several match the same file, the last one wins. That is why the common "only these files" pattern is --exclude "*" first and --include afterwards: the reverse order excludes everything.
Patterns are matched against the path relative to the source directory, not against the full local path or the bucket name. So --exclude ".git/*" works for a local source, while --exclude "/home/me/project/.git/*" matches nothing. The pattern language is small: * for anything (including slashes), ? for one character, and [abc] or [!abc] for character sets. An --include on its own changes nothing, because it only re-includes files that an earlier --exclude removed.
# Only PDFs and spreadsheets: exclude everything, then re-include.
aws s3 sync ./finance s3://amzn-s3-demo-bucket/finance/ \
--exclude "*" --include "*.pdf" --include "*.xlsx"
# Everything except temp files and a cache folder.
aws s3 sync ./finance s3://amzn-s3-demo-bucket/finance/ \
--exclude "*.tmp" --exclude "cache/*"What --delete removes, and what it leaves
With --delete, files that exist in the destination but not in the source are deleted. It is what turns sync from "copy new things" into a true mirror, and it is the flag behind most "sync wiped my bucket" stories.
- The source and destination are the whole comparison. Point the source at an empty or wrong directory, or swap source and destination, and --delete will remove everything under the destination prefix.
- Excluded files are excluded from deletion as well. A file matching --exclude is neither copied nor deleted, so filters also protect destination files you want to keep.
- On a versioned bucket, the deletes are simple deletes. S3 adds a delete marker and keeps the previous version, so files removed by a mistaken sync can be restored. On an unversioned bucket, they are gone.
- Sync only creates folders that contain at least one file, so empty local directories never appear in S3, and --delete never has empty folders to remove.
Run --dryrun first
--dryrun prints every upload, download and delete the command would perform without doing any of them. Run it before the first --delete against a new destination, before changing filters, and in any script where the source path comes from a variable that could be empty.
The output is one line per operation, so it can be counted or grepped. A dry run that reports thousands of deletes when you expected a few is the cheapest warning you will ever get.
# See what a mirror would do, then count the deletes before running it for real.
aws s3 sync ./site s3://amzn-s3-demo-bucket/ --delete --dryrun
aws s3 sync ./site s3://amzn-s3-demo-bucket/ --delete --dryrun | grep -c "^(dryrun) delete:"Other flags worth knowing
Most sync flags match aws s3 cp, and apply to the files the run actually copies. Files that are skipped keep whatever storage class, encryption and metadata they already had, so changing a flag does not rewrite existing objects.
- --storage-class: write new copies straight to STANDARD_IA, INTELLIGENT_TIERING, GLACIER_IR or another class instead of transitioning them later with a lifecycle rule.
- --sse aws:kms with --sse-kms-key-id: encrypt the copies with a specific KMS key instead of the bucket default.
- --no-follow-symlinks: skip symbolic links. By default the target's contents are uploaded under the link's name, because S3 has no symlinks.
- --case-conflict: on a case-insensitive file system, decide what happens when two keys differ only by case (error, warn, skip or ignore).
- --only-show-errors: quiet the per-file output in scheduled jobs so the log only shows failures.
When sync is the wrong tool
Sync is a client-side job. It lists both sides, compares them and copies the difference each time it runs, from wherever it runs. That is fine for a folder of reports or a static site; it is slow and chatty for millions of objects, and it does nothing between runs.
For a bucket that should always have a copy in another Region or account, S3 Replication does the copying inside S3 as objects are written. For moving large existing datasets between buckets, S3 Batch Operations or AWS DataSync scale better than a laptop running sync. And for people who just need to put files into the bucket, a browser upload is simpler than handing out CLI credentials.
Replicating a bucket to another Region or accountThe other ways to upload files to S3Scripting S3 with Python and boto3Recovering files a sync deleted
Starter is free. Deploy a scoped role with CloudFormation, sign in, and browse, without handing anyone an access key.
Primary sources
Discussion
0 comments · open to guests · moderatedLiked this? Get the next article by email. No schedule, no filler, one click to leave.