The essentials

Quick reference

One focused task per row. Jump to the related section for complete, working examples.

UseSyntaxExamples
Create a ZIP archivewith ZipFile(path, 'w', compression=ZIP_DEFLATED) as archive: archive.write(source, arcname='data.txt')View examples
Write bytes to ZIParchive.writestr('manifest.json', payload)View examples
Read one ZIP memberwith archive.open('manifest.json') as member: data = member.read(limit + 1)View examples
Inspect ZIP metadatafor info in archive.infolist(): print(info.filename, info.file_size, info.compress_size)View examples
Check ZIP CRC valuesbad_member = archive.testzip()View examples
Create a compressed tarwith tarfile.open(path, 'w:gz') as archive: archive.add(source, arcname='data')View examples
Detect tar compressionwith tarfile.open(path, 'r:*') as archive: members = archive.getmembers()View examples
Add generated bytes to tararchive.addfile(info, io.BytesIO(payload))View examples
Use the tar data filterarchive.extractall(destination, filter='data')View examples
Validate ZIP destinationstarget = (destination / info.filename).resolve()View examples
Compress with gzipcompressed = gzip.compress(data, compresslevel=6, mtime=0)View examples
Compress with bzip2compressed = bz2.compress(data, compresslevel=9)View examples
Compress with XZcompressed = lzma.compress(data, format=lzma.FORMAT_XZ)View examples
Incremental zlib compressioncompressor = zlib.compressobj(level=6)View examples
Create an archive by formatarchive_path = shutil.make_archive(base_name, 'zip', root_dir=source_dir)View examples

The standard library handles interoperable ZIP and tar containers and several compression codecs. Archive extraction is a security boundary: inspect names and types, set resource limits, use tar extraction filters, extract into a new controlled directory, and never assume a valid archive is a safe archive.

Step by step

Detailed examples

01

Create and read ZIP archives

Use a context manager so central-directory records are finalized. ZIP_DEFLATED is broadly compatible; BZIP2, LZMA, and the Python 3.14 ZIP_ZSTANDARD option may not be readable by older tools. arcname prevents accidental disclosure of an absolute source path. Duplicate names are allowed by the format, so decide whether to reject them.

Build an in-memory ZIP
import io
from zipfile import ZIP_DEFLATED, ZipFile

buffer = io.BytesIO()
with ZipFile(buffer, 'w', compression=ZIP_DEFLATED) as archive:
    archive.writestr('notes/ready.txt', 'ready\n')
with ZipFile(io.BytesIO(buffer.getvalue())) as archive:
    print(archive.namelist())
    print(archive.read('notes/ready.txt').decode().rstrip())
Output
['notes/ready.txt']
ready
Back to quick reference ↑
02

Represent filesystem-oriented data with tar

tar preserves richer Unix metadata and member types than ZIP. Modes such as w:gz, w:bz2, and w:xz select compression; r:* detects it. When generating a member, set TarInfo.size exactly. Streaming modes use a vertical bar and avoid random access, which helps process pipelines and archives larger than memory.

Build an in-memory tar archive
import io
import tarfile

payload = b'alpha\nbeta\n'
buffer = io.BytesIO()
with tarfile.open(fileobj=buffer, mode='w:gz') as archive:
    info = tarfile.TarInfo('data/items.txt')
    info.size = len(payload)
    archive.addfile(info, io.BytesIO(payload))
with tarfile.open(fileobj=io.BytesIO(buffer.getvalue()), mode='r:*') as archive:
    member = archive.extractfile('data/items.txt')
    print(archive.getnames())
    print(member.read().decode().splitlines())
Output
['data/items.txt']
['alpha', 'beta']
Back to quick reference ↑
03

Inspect structure before consuming contents

is_zipfile and tarfile.is_tarfile recognize formats, not trustworthiness. Inspect every normalized name, member type, link target, declared size, total expanded size, compression ratio, and count. Enforce limits while reading because metadata can lie. Avoid loading huge member lists when a streaming approach can enforce limits incrementally.

Inspect ZIP declarations
import io
from zipfile import ZipFile

buffer = io.BytesIO()
with ZipFile(buffer, 'w') as archive:
    archive.writestr('a.txt', b'abc')
    archive.writestr('b.txt', b'12345')
with ZipFile(io.BytesIO(buffer.getvalue())) as archive:
    facts = [(item.filename, item.file_size) for item in archive.infolist()]
    print(facts)
    print(archive.testzip())
Output
[('a.txt', 3), ('b.txt', 5)]
None
Back to quick reference ↑
04

Treat extraction as untrusted file creation

Malicious archives can escape through absolute paths, .. traversal, links, device nodes, duplicate targets, case collisions, or resource exhaustion. For tar on Python 3.12+, pass filter='data'; it became the default in Python 3.14, but explicit policy documents intent and supports both versions. The filter mitigates many attacks but not denial of service. ZIP extraction sanitizes some names yet still requires your own containment, type, duplicate, and quota policy.

Reject escaping member names before extraction
from pathlib import PurePosixPath

def portable_member(name: str) -> bool:
    path = PurePosixPath(name.replace('\\', '/'))
    windows_drive = bool(path.parts) and path.parts[0].endswith(':')
    return not path.is_absolute() and '..' not in path.parts and not windows_drive and '\0' not in name

for name in ['docs/readme.txt', '../secret', '/etc/passwd', r'C:\Windows\win.ini']:
    print(name, portable_member(name))
Output
docs/readme.txt True
../secret False
/etc/passwd False
C:\Windows\win.ini False
Back to quick reference ↑
05

Compress bytes or streams without an archive

gzip, bz2, and lzma offer one-shot functions and file-like streaming APIs; zlib exposes raw DEFLATE-oriented primitives. Compression is not encryption or integrity authentication. Decompression can expand dramatically, so stream into bounded storage and enforce output, CPU, and nesting limits for untrusted data.

Round-trip three compression formats
import bz2
import gzip
import lzma

data = b'cmdmemo:' * 20
codecs = [
    ('gzip', lambda value: gzip.compress(value, mtime=0), gzip.decompress),
    ('bz2', bz2.compress, bz2.decompress),
    ('xz', lzma.compress, lzma.decompress),
]
for name, compress, decompress in codecs:
    packed = compress(data)
    print(name, len(packed) < len(data), decompress(packed) == data)
Output
gzip True True
bz2 True True
xz True True
Back to quick reference ↑
06

Choose compatibility and resource policy deliberately

shutil.make_archive and unpack_archive are convenient for trusted directory trees. Prefer common codecs when consumers are unknown, and pin format choices in interfaces. Python 3.14 added ZIP_ZSTANDARD and the compression.zstd module; guard those APIs when supporting older runtimes. Password-protected ZIP decryption is slow and legacy-oriented, while stdlib zipfile cannot create encrypted archives—use modern authenticated encryption outside the archive format when confidentiality matters.

Discover archive formats available at runtime
import shutil

formats = {name for name, _description in shutil.get_archive_formats()}
print('zip' in formats)
print('gztar' in formats)
Output
True
True
Back to quick reference ↑

Local code tester

Create and inspect an in-memory ZIP

Write two generated members, inspect their declared sizes, and read one member without touching the filesystem.

Runs in your browser
Output
Press Run to load Python locally.

Sources and further reading

References

Authoritative documentation used to verify and expand this cheat sheet.

  1. Python Software Foundationzipfile — Work with ZIP archivesdocs.python.org
  2. Python Software Foundationtarfile — Read and write tar archive filesdocs.python.org
  3. Python Software FoundationData Compression and Archivingdocs.python.org
  4. Python Software Foundationshutil — High-level file operationsdocs.python.org

Help us improve

Found a typo or missing example?

Tell us what would make this cheat sheet clearer, more complete, or more useful.

Share feedback