How to see inside a zip file in Amazon S3 without downloading it
S3 can't unzip anything, but a zip keeps its table of contents at the end of the file. Here's how to read that list with a few range requests, and what to check before you extract.
A zip in S3 doesn't have to be downloaded to be understood: its central directory sits at the end of the file and can be fetched with a couple of range requests. Read that list first, check expanded size, file count, collisions and paths, and only then extract, ideally inside AWS so no copy ends up on a laptop.
Somebody drops a 300 MB zip called handover-final.zip into the bucket and asks you to "put the files in the right folders". Before you can do that you need to know what's inside, how big it gets when unpacked, and whether any of it will overwrite files that are already there.
Amazon S3 can't help with that directly. It stores the zip as one object, and the console's only option is to download it. This guide shows how to read the list of files from S3 without downloading the archive, what to check before extracting, and the usual ways to extract once you're ready.
Why S3 can't open a zip for you
S3 stores objects as opaque bytes. It has no unzip operation. The console can download the object or open it in a new tab, and a browser opening a zip just saves it to disk.
The usual workaround is to download the archive, unpack it on a laptop, look around, and then upload whatever you need. For a large archive that wastes time, uses disk space, and puts a copy of the files on someone's machine, which may be exactly what your data policy was trying to avoid.
A zip keeps its table of contents at the end
A zip file isn't one compressed blob. Each file inside is compressed separately, and at the very end of the archive is the central directory: a list of every entry with its name, compressed size, uncompressed size, and where its data starts. The last record in the file, the end-of-central-directory record, says where that list begins.
That layout is what makes remote inspection possible. S3's GetObject accepts a Range header, so you can ask for only the last few kilobytes, find where the central directory starts, then fetch just the directory. For a typical archive of documents or photos, the directory is a tiny fraction of the file.
This only works for zip. A .tar.gz has no index: tar headers are spread through the archive, and gzip compresses the whole stream, so listing its contents means reading the whole object from start to end.
# Fetch only the last 64 KB of the archive, where the zip directory lives
aws s3api get-object --bucket my-bucket --key Archives/handover.zip \
--range bytes=-65536 tail.binList a zip's contents with Python and range requests
Python's zipfile module needs a file object it can seek and read. The class below gives it one by turning each read into a ranged GetObject call. zipfile seeks to the end, reads the end record, then reads the central directory. That's usually two or three requests, however many files the archive holds.
We tested this against a 212 KB demo archive holding 1,500 small files. It listed every entry in 3 GET requests and never downloaded the file data.
import io
import sys
import zipfile
import boto3
class S3File(io.RawIOBase):
"""A read-only, seekable file that fetches bytes from S3 with range requests."""
def __init__(self, s3, bucket, key):
self.s3, self.bucket, self.key, self.pos = s3, bucket, key, 0
self.size = s3.head_object(Bucket=bucket, Key=key)["ContentLength"]
def readable(self):
return True
def seekable(self):
return True
def tell(self):
return self.pos
def seek(self, offset, whence=io.SEEK_SET):
base = {io.SEEK_SET: 0, io.SEEK_CUR: self.pos, io.SEEK_END: self.size}[whence]
self.pos = base + offset
return self.pos
def read(self, n=-1):
if n is None or n < 0:
n = self.size - self.pos
if n == 0 or self.pos >= self.size:
return b""
end = min(self.pos + n, self.size) - 1
resp = self.s3.get_object(
Bucket=self.bucket, Key=self.key, Range=f"bytes={self.pos}-{end}"
)
data = resp["Body"].read()
self.pos += len(data)
return data
bucket, key = sys.argv[1], sys.argv[2]
f = S3File(boto3.client("s3"), bucket, key)
with zipfile.ZipFile(f) as archive:
entries = archive.infolist()
for e in entries:
print(f"{e.file_size:>12,} {e.compress_size:>12,} {e.filename}")
total = sum(e.file_size for e in entries)
print(f"{len(entries)} files, {total:,} bytes expanded "
f"({total / f.size:.1f}x the {f.size:,}-byte archive)")What to check before you extract
The list of files is only useful if you act on it. Before anything gets written to the bucket, check the following. The sizes in the central directory are what the archive claims, so an extraction job should also stop if the real output goes past your limits.
- Expanded size: a small archive can expand far larger than its compressed size. An unusually high ratio is how zip bombs work.
- File count: 50,000 tiny files means 50,000 PUT requests and 50,000 objects to list later.
- Key collisions: an entry that maps to a key that already exists will silently overwrite it unless versioning is on.
- Unsafe paths: entry names like ../../config or absolute paths have to be rejected or normalised before they become S3 keys.
- Cost: S3 Standard charges per PUT request, and the extracted files are billed as new storage on top of the zip.
- Destination: decide the target prefix up front, so the files don't land at the bucket root.
Ways to extract once you're happy
For a one-off, downloading, unzipping and running aws s3 cp --recursive is fine if the archive fits on your disk. You give up the "no local copy" benefit, so delete the local files afterwards.
To keep the work in AWS, a Lambda function can stream entries from the archive and upload each one. Lambda caps a run at 15 minutes and temporary storage at 10 GB, so very large archives or very many entries belong in an AWS Batch or Fargate job instead. Whichever you choose, the checks above belong in the job itself, not only in the person who starts it.
Automating S3 file workflows with Python and boto3Copying folders with aws s3 cp --recursiveHow we're moving archive processing from PHP to Go
Doing it in BucketDesk
BucketDesk is a browser workspace on top of your own bucket. Click a zip and it lists the contents by reading the archive index by range, so nothing is downloaded and no copy lands on anyone's laptop.
When you choose to extract, BucketDesk shows the destination, file count, expansion ratio, any key collisions and an estimated cost before it writes anything. The extraction then runs as a job you can watch, and the files are written to your bucket under the role you connected. Archive inspection and extraction are part of the Pro and Business plans.
See archive operations in BucketDeskPreview other file types in S3 without downloading
Starter is free. Deploy a scoped role with CloudFormation, sign in, and browse, without handing anyone an access key.
Primary sources
Discussion
0 comments · open to guests · moderatedLiked this? Get the next article by email. No schedule, no filler, one click to leave.