The essentials

Quick reference

One focused task per row. Jump to the related section for complete, working examples.

UseSyntaxExamples
Encode UTF-8payload = text.encode('utf-8')View examples
Decode UTF-8text = payload.decode('utf-8')View examples
Use strict decodingtext = payload.decode('utf-8', errors='strict')View examples
Open UTF-8 textwith open(path, 'r', encoding='utf-8') as stream: text = stream.read()View examples
Write stable LF newlineswith open(path, 'w', encoding='utf-8', newline='\n') as stream: stream.write(text)View examples
Wrap a binary streamtext_stream = io.TextIOWrapper(binary_stream, encoding='utf-8', newline='')View examples
Normalize canonical compositionnormalized = unicodedata.normalize('NFC', text)View examples
Check normalizationok = unicodedata.is_normalized('NFC', text)View examples
Build a caseless keykey = unicodedata.normalize('NFC', text).casefold()View examples
Lowercase for presentationlowered = text.lower()View examples
Read a code pointnumber = ord(character)View examples
Create a charactercharacter = chr(number)View examples
Read a Unicode namename = unicodedata.name(character, 'UNNAMED')View examples
Replace malformed inputtext = payload.decode('utf-8', errors='replace')View examples
Preserve OS bytestext = payload.decode(filesystem_encoding, errors='surrogateescape')View examples
Create a streaming decoderdecoder = codecs.getincrementaldecoder('utf-8')(errors='strict')View examples
Finish a decode streamtail = decoder.decode(b'', final=True)View examples

Python str stores Unicode code points; files, sockets, and binary formats store bytes. Correct software chooses an encoding at each boundary, preserves undecodable operating-system data when necessary, normalizes only for a defined comparison policy, and avoids assuming that code points equal visible characters. Locale defaults and permissive error handlers are data-model decisions, not conveniences.

Step by step

Detailed examples

01

Cross the text and byte boundary explicitly

encode transforms str into bytes and decode transforms bytes into str. UTF-8 is the default for many modern protocols but must still be named in durable interfaces. A byte order mark is not required for UTF-8. Do not repeatedly encode already-binary data or decode merely to apply byte-oriented framing.

Round-trip non-ASCII text
text = "café ☕"
payload = text.encode("utf-8")
print(len(text), len(payload))
print(payload.decode("utf-8"))
Output
6 9
café ☕
Back to quick reference ↑
02

Specify encoding and newline policy for files

TextIOWrapper performs incremental decoding and newline translation. Pass encoding for durable data rather than inheriting locale-sensitive defaults; use newline='' when a parser such as csv must control newlines. UTF-8 mode and Python UTF-8 defaults continue evolving, so explicit formats remain the clearest contract.

Use an in-memory text stream with explicit newlines
import io

buffer = io.BytesIO()
with io.TextIOWrapper(buffer, encoding="utf-8", newline="\n") as stream:
    stream.write("olá\n")
    stream.flush()
    print(buffer.getvalue().hex())
Output
6f6cc3a10a
Back to quick reference ↑
03

Normalize only for a defined equivalence policy

Unicode may represent visually identical text with different code-point sequences. NFC is commonly appropriate for stored general text, while NFKC performs compatibility folding that can erase meaningful distinctions. Normalize both operands at a comparison boundary, not blindly across source code, cryptographic input, identifiers, or forensic data.

Compare canonically equivalent text
import unicodedata

composed = "é"
decomposed = "e\u0301"
print(composed == decomposed)
print(unicodedata.normalize("NFC", composed) == unicodedata.normalize("NFC", decomposed))
print(len(decomposed))
Output
False
True
2
Back to quick reference ↑
04

Use casefold for caseless matching

casefold is more aggressive and language-neutral than lower for caseless comparison, but it is not locale-aware collation. Normalize along with case folding when the application's identity policy calls for it. Display the user's original spelling even when comparing a derived key, and account for collisions introduced by the policy.

Show why casefold differs from lower
word = "Straße"
print(word.lower())
print(word.casefold())
print(word.casefold() == "STRASSE".casefold())
Output
straße
strasse
True
Back to quick reference ↑
05

Distinguish code points from visible characters

len(str) counts code points, not bytes or user-perceived grapheme clusters. Combining marks, variation selectors, emoji modifiers, and zero-width joiners can form one displayed unit from many code points. The standard library exposes Unicode properties but not full grapheme segmentation; use a conforming segmentation implementation for cursor movement or length limits presented to users.

Inspect a combining sequence
import unicodedata

text = "e\u0301"
print(len(text))
for character in text:
    print(f"U+{ord(character):04X}", unicodedata.category(character))
Output
2
U+0065 Ll
U+0301 Mn
Back to quick reference ↑
06

Choose codec error behavior deliberately

strict preserves correctness by failing. replace and ignore are lossy and can create identifier, delimiter, or security ambiguities. surrogateescape is intended for round-tripping undecodable operating-system bytes through text APIs on supported paths; it is not valid interchange Unicode. backslashreplace is useful for diagnostic display, not reversible storage in general.

Compare strict and replacement decoding
payload = b"ok\xff"
try:
    payload.decode("utf-8")
except UnicodeDecodeError as error:
    print(type(error).__name__, error.start)
print(payload.decode("utf-8", errors="replace"))
Output
UnicodeDecodeError 2
ok�
Back to quick reference ↑
07

Decode streams incrementally and defend boundaries

Multibyte characters can span chunks, so use an incremental decoder or TextIOWrapper instead of decoding each chunk independently. Flush with final=True to detect incomplete tails. Validate protocol-specific controls, bidirectional text, confusable identifiers, and normalization separately; Unicode validity alone does not make text safe for HTML, SQL, shells, paths, or terminals.

Decode a character split across chunks
import codecs

decoder = codecs.getincrementaldecoder("utf-8")()
parts = [decoder.decode(b"caf\xc3"), decoder.decode(b"\xa9", final=True)]
print("".join(parts))
Output
café
Back to quick reference ↑

Local code tester

Build normalized caseless keys

Compare canonically equivalent spelling and Unicode case variants while preserving original display text.

Runs in your browser
Output
Press Run to load Python locally.

Sources and further reading

References

Authoritative documentation used to verify and expand this cheat sheet.

  1. Python Software FoundationUnicode HOWTOdocs.python.org
  2. Python Software FoundationText Sequence Type — strdocs.python.org
  3. Python Software Foundationcodecs — Codec registry and base classesdocs.python.org
  4. Python Software Foundationunicodedata — Unicode Databasedocs.python.org
  5. Python Software Foundationio — Core tools for working with streamsdocs.python.org

Help us improve

Found a typo or missing example?

Tell us what would make this cheat sheet clearer, more complete, or more useful.

Share feedback