The essentials
Quick reference
One focused task per row. Jump to the related section for complete, working examples.
| Use | Syntax | Examples |
|---|---|---|
| Encode UTF-8 | payload = text.encode('utf-8') | View examples |
| Decode UTF-8 | text = payload.decode('utf-8') | View examples |
| Use strict decoding | text = payload.decode('utf-8', errors='strict') | View examples |
| Open UTF-8 text | with open(path, 'r', encoding='utf-8') as stream: text = stream.read() | View examples |
| Write stable LF newlines | with open(path, 'w', encoding='utf-8', newline='\n') as stream: stream.write(text) | View examples |
| Wrap a binary stream | text_stream = io.TextIOWrapper(binary_stream, encoding='utf-8', newline='') | View examples |
| Normalize canonical composition | normalized = unicodedata.normalize('NFC', text) | View examples |
| Check normalization | ok = unicodedata.is_normalized('NFC', text) | View examples |
| Build a caseless key | key = unicodedata.normalize('NFC', text).casefold() | View examples |
| Lowercase for presentation | lowered = text.lower() | View examples |
| Read a code point | number = ord(character) | View examples |
| Create a character | character = chr(number) | View examples |
| Read a Unicode name | name = unicodedata.name(character, 'UNNAMED') | View examples |
| Replace malformed input | text = payload.decode('utf-8', errors='replace') | View examples |
| Preserve OS bytes | text = payload.decode(filesystem_encoding, errors='surrogateescape') | View examples |
| Create a streaming decoder | decoder = codecs.getincrementaldecoder('utf-8')(errors='strict') | View examples |
| Finish a decode stream | tail = decoder.decode(b'', final=True) | View examples |
Python str stores Unicode code points; files, sockets, and binary formats store bytes. Correct software chooses an encoding at each boundary, preserves undecodable operating-system data when necessary, normalizes only for a defined comparison policy, and avoids assuming that code points equal visible characters. Locale defaults and permissive error handlers are data-model decisions, not conveniences.
Step by step
Detailed examples
Cross the text and byte boundary explicitly
encode transforms str into bytes and decode transforms bytes into str. UTF-8 is the default for many modern protocols but must still be named in durable interfaces. A byte order mark is not required for UTF-8. Do not repeatedly encode already-binary data or decode merely to apply byte-oriented framing.
text = "café ☕"
payload = text.encode("utf-8")
print(len(text), len(payload))
print(payload.decode("utf-8")) 6 9
café ☕Specify encoding and newline policy for files
TextIOWrapper performs incremental decoding and newline translation. Pass encoding for durable data rather than inheriting locale-sensitive defaults; use newline='' when a parser such as csv must control newlines. UTF-8 mode and Python UTF-8 defaults continue evolving, so explicit formats remain the clearest contract.
import io
buffer = io.BytesIO()
with io.TextIOWrapper(buffer, encoding="utf-8", newline="\n") as stream:
stream.write("olá\n")
stream.flush()
print(buffer.getvalue().hex()) 6f6cc3a10aNormalize only for a defined equivalence policy
Unicode may represent visually identical text with different code-point sequences. NFC is commonly appropriate for stored general text, while NFKC performs compatibility folding that can erase meaningful distinctions. Normalize both operands at a comparison boundary, not blindly across source code, cryptographic input, identifiers, or forensic data.
import unicodedata
composed = "é"
decomposed = "e\u0301"
print(composed == decomposed)
print(unicodedata.normalize("NFC", composed) == unicodedata.normalize("NFC", decomposed))
print(len(decomposed)) False
True
2Use casefold for caseless matching
casefold is more aggressive and language-neutral than lower for caseless comparison, but it is not locale-aware collation. Normalize along with case folding when the application's identity policy calls for it. Display the user's original spelling even when comparing a derived key, and account for collisions introduced by the policy.
word = "Straße"
print(word.lower())
print(word.casefold())
print(word.casefold() == "STRASSE".casefold()) straße
strasse
TrueDistinguish code points from visible characters
len(str) counts code points, not bytes or user-perceived grapheme clusters. Combining marks, variation selectors, emoji modifiers, and zero-width joiners can form one displayed unit from many code points. The standard library exposes Unicode properties but not full grapheme segmentation; use a conforming segmentation implementation for cursor movement or length limits presented to users.
import unicodedata
text = "e\u0301"
print(len(text))
for character in text:
print(f"U+{ord(character):04X}", unicodedata.category(character)) 2
U+0065 Ll
U+0301 MnChoose codec error behavior deliberately
strict preserves correctness by failing. replace and ignore are lossy and can create identifier, delimiter, or security ambiguities. surrogateescape is intended for round-tripping undecodable operating-system bytes through text APIs on supported paths; it is not valid interchange Unicode. backslashreplace is useful for diagnostic display, not reversible storage in general.
payload = b"ok\xff"
try:
payload.decode("utf-8")
except UnicodeDecodeError as error:
print(type(error).__name__, error.start)
print(payload.decode("utf-8", errors="replace")) UnicodeDecodeError 2
ok�Decode streams incrementally and defend boundaries
Multibyte characters can span chunks, so use an incremental decoder or TextIOWrapper instead of decoding each chunk independently. Flush with final=True to detect incomplete tails. Validate protocol-specific controls, bidirectional text, confusable identifiers, and normalization separately; Unicode validity alone does not make text safe for HTML, SQL, shells, paths, or terminals.
import codecs
decoder = codecs.getincrementaldecoder("utf-8")()
parts = [decoder.decode(b"caf\xc3"), decoder.decode(b"\xa9", final=True)]
print("".join(parts)) caféLocal code tester
Build normalized caseless keys
Compare canonically equivalent spelling and Unicode case variants while preserving original display text.
Press Run to load Python locally.
Sources and further reading
References
Authoritative documentation used to verify and expand this cheat sheet.
- Python Software FoundationUnicode HOWTOdocs.python.org
- Python Software FoundationText Sequence Type — strdocs.python.org
- Python Software Foundationcodecs — Codec registry and base classesdocs.python.org
- Python Software Foundationunicodedata — Unicode Databasedocs.python.org
- Python Software Foundationio — Core tools for working with streamsdocs.python.org
Help us improve
Found a typo or missing example?
Tell us what would make this cheat sheet clearer, more complete, or more useful.



