The essentials
Quick reference
One focused task per row. Jump to the related section for complete, working examples.
| Use | Syntax | Examples |
|---|---|---|
| Create a ZIP archive | with ZipFile(path, 'w', compression=ZIP_DEFLATED) as archive: archive.write(source, arcname='data.txt') | View examples |
| Write bytes to ZIP | archive.writestr('manifest.json', payload) | View examples |
| Read one ZIP member | with archive.open('manifest.json') as member: data = member.read(limit + 1) | View examples |
| Inspect ZIP metadata | for info in archive.infolist(): print(info.filename, info.file_size, info.compress_size) | View examples |
| Check ZIP CRC values | bad_member = archive.testzip() | View examples |
| Create a compressed tar | with tarfile.open(path, 'w:gz') as archive: archive.add(source, arcname='data') | View examples |
| Detect tar compression | with tarfile.open(path, 'r:*') as archive: members = archive.getmembers() | View examples |
| Add generated bytes to tar | archive.addfile(info, io.BytesIO(payload)) | View examples |
| Use the tar data filter | archive.extractall(destination, filter='data') | View examples |
| Validate ZIP destinations | target = (destination / info.filename).resolve() | View examples |
| Compress with gzip | compressed = gzip.compress(data, compresslevel=6, mtime=0) | View examples |
| Compress with bzip2 | compressed = bz2.compress(data, compresslevel=9) | View examples |
| Compress with XZ | compressed = lzma.compress(data, format=lzma.FORMAT_XZ) | View examples |
| Incremental zlib compression | compressor = zlib.compressobj(level=6) | View examples |
| Create an archive by format | archive_path = shutil.make_archive(base_name, 'zip', root_dir=source_dir) | View examples |
The standard library handles interoperable ZIP and tar containers and several compression codecs. Archive extraction is a security boundary: inspect names and types, set resource limits, use tar extraction filters, extract into a new controlled directory, and never assume a valid archive is a safe archive.
Step by step
Detailed examples
Create and read ZIP archives
Use a context manager so central-directory records are finalized. ZIP_DEFLATED is broadly compatible; BZIP2, LZMA, and the Python 3.14 ZIP_ZSTANDARD option may not be readable by older tools. arcname prevents accidental disclosure of an absolute source path. Duplicate names are allowed by the format, so decide whether to reject them.
import io
from zipfile import ZIP_DEFLATED, ZipFile
buffer = io.BytesIO()
with ZipFile(buffer, 'w', compression=ZIP_DEFLATED) as archive:
archive.writestr('notes/ready.txt', 'ready\n')
with ZipFile(io.BytesIO(buffer.getvalue())) as archive:
print(archive.namelist())
print(archive.read('notes/ready.txt').decode().rstrip()) ['notes/ready.txt']
readyRepresent filesystem-oriented data with tar
tar preserves richer Unix metadata and member types than ZIP. Modes such as w:gz, w:bz2, and w:xz select compression; r:* detects it. When generating a member, set TarInfo.size exactly. Streaming modes use a vertical bar and avoid random access, which helps process pipelines and archives larger than memory.
import io
import tarfile
payload = b'alpha\nbeta\n'
buffer = io.BytesIO()
with tarfile.open(fileobj=buffer, mode='w:gz') as archive:
info = tarfile.TarInfo('data/items.txt')
info.size = len(payload)
archive.addfile(info, io.BytesIO(payload))
with tarfile.open(fileobj=io.BytesIO(buffer.getvalue()), mode='r:*') as archive:
member = archive.extractfile('data/items.txt')
print(archive.getnames())
print(member.read().decode().splitlines()) ['data/items.txt']
['alpha', 'beta']Inspect structure before consuming contents
is_zipfile and tarfile.is_tarfile recognize formats, not trustworthiness. Inspect every normalized name, member type, link target, declared size, total expanded size, compression ratio, and count. Enforce limits while reading because metadata can lie. Avoid loading huge member lists when a streaming approach can enforce limits incrementally.
import io
from zipfile import ZipFile
buffer = io.BytesIO()
with ZipFile(buffer, 'w') as archive:
archive.writestr('a.txt', b'abc')
archive.writestr('b.txt', b'12345')
with ZipFile(io.BytesIO(buffer.getvalue())) as archive:
facts = [(item.filename, item.file_size) for item in archive.infolist()]
print(facts)
print(archive.testzip()) [('a.txt', 3), ('b.txt', 5)]
NoneTreat extraction as untrusted file creation
Malicious archives can escape through absolute paths, .. traversal, links, device nodes, duplicate targets, case collisions, or resource exhaustion. For tar on Python 3.12+, pass filter='data'; it became the default in Python 3.14, but explicit policy documents intent and supports both versions. The filter mitigates many attacks but not denial of service. ZIP extraction sanitizes some names yet still requires your own containment, type, duplicate, and quota policy.
from pathlib import PurePosixPath
def portable_member(name: str) -> bool:
path = PurePosixPath(name.replace('\\', '/'))
windows_drive = bool(path.parts) and path.parts[0].endswith(':')
return not path.is_absolute() and '..' not in path.parts and not windows_drive and '\0' not in name
for name in ['docs/readme.txt', '../secret', '/etc/passwd', r'C:\Windows\win.ini']:
print(name, portable_member(name)) docs/readme.txt True
../secret False
/etc/passwd False
C:\Windows\win.ini FalseCompress bytes or streams without an archive
gzip, bz2, and lzma offer one-shot functions and file-like streaming APIs; zlib exposes raw DEFLATE-oriented primitives. Compression is not encryption or integrity authentication. Decompression can expand dramatically, so stream into bounded storage and enforce output, CPU, and nesting limits for untrusted data.
import bz2
import gzip
import lzma
data = b'cmdmemo:' * 20
codecs = [
('gzip', lambda value: gzip.compress(value, mtime=0), gzip.decompress),
('bz2', bz2.compress, bz2.decompress),
('xz', lzma.compress, lzma.decompress),
]
for name, compress, decompress in codecs:
packed = compress(data)
print(name, len(packed) < len(data), decompress(packed) == data) gzip True True
bz2 True True
xz True TrueChoose compatibility and resource policy deliberately
shutil.make_archive and unpack_archive are convenient for trusted directory trees. Prefer common codecs when consumers are unknown, and pin format choices in interfaces. Python 3.14 added ZIP_ZSTANDARD and the compression.zstd module; guard those APIs when supporting older runtimes. Password-protected ZIP decryption is slow and legacy-oriented, while stdlib zipfile cannot create encrypted archives—use modern authenticated encryption outside the archive format when confidentiality matters.
import shutil
formats = {name for name, _description in shutil.get_archive_formats()}
print('zip' in formats)
print('gztar' in formats) True
TrueLocal code tester
Create and inspect an in-memory ZIP
Write two generated members, inspect their declared sizes, and read one member without touching the filesystem.
Press Run to load Python locally.
Sources and further reading
References
Authoritative documentation used to verify and expand this cheat sheet.
- Python Software Foundationzipfile — Work with ZIP archivesdocs.python.org
- Python Software Foundationtarfile — Read and write tar archive filesdocs.python.org
- Python Software FoundationData Compression and Archivingdocs.python.org
- Python Software Foundationshutil — High-level file operationsdocs.python.org
Help us improve
Found a typo or missing example?
Tell us what would make this cheat sheet clearer, more complete, or more useful.



