The essentials
Quick reference
One focused task per row. Jump to the related section for complete, working examples.
| Use | Syntax | Examples |
|---|---|---|
| Parse an XML value | root = ET.fromstring(xml_bytes) | View examples |
| Parse an XML file | tree = ET.parse(path) | View examples |
| Catch malformed XML | except ET.ParseError as error: report(error.position) | View examples |
| Create an expanded name | tag = '{urn:books}title' | View examples |
| Query with a prefix map | items = root.findall('b:item', {'b': 'urn:books'}) | View examples |
| Choose a serialization prefix | ET.register_namespace('b', 'urn:books') | View examples |
| Find one matching element | item = root.find('./item[@id="a"]') | View examples |
| Find matching elements | items = root.findall('.//item') | View examples |
| Iterate matching descendants | for item in root.iter('item'): consume(item) | View examples |
| Append a child element | child = ET.SubElement(parent, 'item', {'id': item_id}) | View examples |
| Set element text | child.text = value | View examples |
| Remove a child | parent.remove(child) | View examples |
| Serialize to text | xml_text = ET.tostring(root, encoding='unicode') | View examples |
| Write UTF-8 XML | tree.write(path, encoding='utf-8', xml_declaration=True) | View examples |
| Canonicalize XML | canonical = ET.canonicalize(xml_data=xml_text) | View examples |
| Iterate completed elements | events = ET.iterparse(source, events=('end',)) | View examples |
| Release processed content | element.clear() | View examples |
| Create a pull parser | parser = ET.XMLPullParser(events=('start', 'end')) | View examples |
| Inspect the Expat version | version = pyexpat.EXPAT_VERSION | View examples |
| Bound XML input bytes | if len(xml_bytes) > MAX_XML_BYTES: raise ValueError('XML too large') | View examples |
XML is a tree-shaped syntax with namespaces, mixed content, processing instructions, and an external-entity history that demands defensive parsing. ElementTree covers common in-process tasks, but its XPath subset is intentionally limited. Preserve namespace meaning, escape through APIs rather than string concatenation, and apply explicit resource limits to untrusted documents.
Step by step
Detailed examples
Parse trusted XML into an element tree
fromstring parses one XML document from text or bytes and returns its root; parse reads a file or file-like object. ParseError exposes a position for malformed syntax. Passing bytes lets the XML declaration determine encoding, while passing str means decoding already occurred and an encoding declaration may be misleading.
import xml.etree.ElementTree as ET
root = ET.fromstring('<catalog version="2"><item id="a">Pen</item></catalog>')
item = root.find("item")
print(root.tag, root.get("version"))
print(item.get("id"), item.text) catalog 2
a PenAddress expanded names instead of prefixes
Namespace prefixes are aliases local to a document; the semantic element name is the namespace URI plus local name. ElementTree represents expanded names as {uri}local. Supply a prefix mapping to find/findall for readable queries, and do not assume the source document will reuse a preferred prefix.
import xml.etree.ElementTree as ET
root = ET.fromstring('<books xmlns="urn:books"><book><title>Python</title></book></books>')
namespaces = {"b": "urn:books"}
title = root.find("b:book/b:title", namespaces)
print(title.text) PythonUse ElementTree's supported XPath subset
find returns the first matching child or None, findall returns a list, and iter walks descendants. Element truth testing has historically been ambiguous and emits a deprecation warning in recent Python versions; test element is None explicitly. For full XPath, schema validation, or XSLT, choose a library that implements those features and define its security posture.
import xml.etree.ElementTree as ET
root = ET.fromstring('<items><item active="yes">A</item><item active="no">B</item></items>')
values = [item.text for item in root.findall('./item[@active="yes"]')]
print(values) ['A']Construct XML through element APIs
Element and SubElement escape text and attribute values during serialization, preventing malformed markup from ordinary data. text belongs before the first child and tail belongs after an element's end tag, so mixed-content edits require care. Removal works by element identity, and mutation during direct iteration should use a snapshot such as list(parent).
import xml.etree.ElementTree as ET
root = ET.Element("message", {"kind": "a&b"})
ET.SubElement(root, "text").text = "1 < 2"
print(ET.tostring(root, encoding="unicode")) <message kind="a&b"><text>1 < 2</text></message>Control declaration, encoding, and empty elements
tostring and ElementTree.write can emit bytes in a named encoding or Unicode text with encoding='unicode'. Request an XML declaration when writing a standalone byte document. Attribute order is preserved from creation in current Python, but semantically equivalent XML can serialize differently; use canonicalize for byte-stable comparisons or signatures and follow the signature profile's exact rules.
import io
import xml.etree.ElementTree as ET
tree = ET.ElementTree(ET.Element("status", {"ok": "yes"}))
target = io.BytesIO()
tree.write(target, encoding="utf-8", xml_declaration=True)
print(target.getvalue().decode("utf-8")) <?xml version='1.0' encoding='utf-8'?>
<status ok="yes" />Stream large documents and clear processed nodes
iterparse incrementally reports start and end events but still performs blocking reads and builds elements unless they are cleared. Inspect complete text and children only on end events. For non-blocking feeds use XMLPullParser. Clearing an element after processing releases children and attributes, but parent references or accumulated result lists can still retain memory.
import xml.etree.ElementTree as ET
parser = ET.XMLPullParser(events=("end",))
parser.feed("<items><item>A</item>")
parser.feed("<item>B</item></items>")
values = [element.text for event, element in parser.read_events() if element.tag == "item"]
parser.close()
print(values) ['A', 'B']Constrain untrusted XML
The standard XML modules document denial-of-service and entity-related risks, and Python builds can differ with their Expat version. Do not accept DTDs, entity expansion, or external resources unless a reviewed requirement demands them. Enforce input byte limits, nesting and element limits where possible, processing deadlines, and post-parse schema/business validation; defusedxml can provide hardened replacements for common untrusted-input cases.
import xml.etree.ElementTree as ET
def parse_bounded(xml_bytes: bytes, maximum: int):
if len(xml_bytes) > maximum:
raise ValueError("XML too large")
return ET.fromstring(xml_bytes)
try:
parse_bounded(b"<root><item /></root>", 10)
except ValueError as error:
print(error) XML too largeLocal code tester
Parse a namespace-aware XML inventory
Query namespaced elements and calculate an aggregate without external files.
Press Run to load Python locally.
Sources and further reading
References
Authoritative documentation used to verify and expand this cheat sheet.
- Python Software Foundationxml.etree.ElementTree — The ElementTree XML APIdocs.python.org
- Python Software FoundationXML Processing Modulesdocs.python.org
- Python Software FoundationXML vulnerabilitiesdocs.python.org
- Python Software Foundationpyexpat — Fast XML parsing using Expatdocs.python.org
- Python Software FoundationSAX — Support for SAX2 parsersdocs.python.org
Help us improve
Found a typo or missing example?
Tell us what would make this cheat sheet clearer, more complete, or more useful.



