Skip to content

Python XML 处理

XML (Extensible Markup Language,可扩展标记语言) 是一种标记语言,旨在以人类可读和机器可读的格式存储和传输数据。它广泛用于配置文件、数据交换和文档格式。

XML 定义了一套编码文档的规则。它使用标签(类似于 HTML)来定义元素,并使用属性来提供关于这些元素的元数据。与 HTML 不同,XML 不预定义标签;用户根据其数据定义自己的标签。

主要特点:

  • 可扩展:您可以定义自己的标签。
  • 层次化:数据以树形结构组织。
  • 文本化:使用 Unicode 字符。
  • 自描述:标签通常描述它们包含的数据。
  • 平台/语言无关:一个开放标准。

尽管 JSON 因其简单性以及与 JavaScript 对象的直接映射而在 Web API 中变得更受欢迎,但 XML 在许多企业系统、文档标准(如 DocBook、Office Open XML)和配置方面仍然很重要。

Python 的标准库提供了几个用于处理 XML 的模块,支持不同的解析策略:

  • xml.etree.ElementTree: (推荐用于大多数常见任务) 提供了一种简单且符合 Python 风格的方式来解析和创建 XML。它将整个文档加载到内存中的树状结构(类似于 DOM),但提供了一个更轻量级且通常更易于使用的 API。
  • xml.dom.minidom: 实现了 Document Object Model (DOM) API,一个 W3C 标准。它将整个 XML 文档解析到内存中的对象树中,允许完全访问和修改。对于非常大的文件可能占用大量内存。
  • xml.sax: 实现了 Simple API for XML (SAX),一个事件驱动的解析器。它按顺序读取 XML 文件,在解析器遇到特定 XML 结构(如开始标签、结束标签和字符数据)时触发回调函数(处理程序)。它不加载整个文档,因此内存效率高,适用于非常大的文件或流处理,但对于随机访问可能更复杂。

选择 API:

  • 对于一般的 XML 解析和创建,首先使用 ElementTree。
  • 如果需要严格遵循 DOM 标准或进行复杂的树修改,使用 minidom。
  • 对于处理非常大且内存有限的 XML 文件,使用 SAX。

我们将使用以下 XML 文件作为示例:

<?xml version="1.0" encoding="UTF-8"?>
<collection shelf="New Arrivals">
<movie title="Enemy Behind">
<type>War, Thriller</type>
<format>DVD</format>
<year>2003</year>
<rating>PG</rating>
<stars>10</stars>
<description>Talk about a US-Japan war</description>
</movie>
<movie title="Transformers">
<type>Anime, Science Fiction</type>
<format>DVD</format>
<year>1989</year>
<rating>R</rating>
<stars>8</stars>
<description>A scientific fiction</description>
</movie>
<movie title="Trigun">
<type>Anime, Action</type>
<format>DVD</nFormat>
<episodes>4</episodes>
<rating>PG</rating>
<stars>10</stars>
<description>Vash the Stampede!</description>
</movie>
<movie title="Ishtar">
<type>Comedy</type>
<format>VHS</format>
<rating>PG</rating>
<stars>2</stars>
<description>Viewable boredom</description>
</movie>
</collection>

ElementTree 将 XML 文档表示为 Element 对象的树。

#!/usr/bin/env python3
import xml.etree.ElementTree as ET
xml_file = 'movies.xml'
try:
tree = ET.parse(xml_file)
root = tree.getroot() # Get the root element ('collection')
# 获取根元素 ('collection')
print(f"Root element: {root.tag}, Shelf: {root.get('shelf')}")
# Iterate through 'movie' elements under the root
# 遍历根元素下的 'movie' 元素
for movie in root.findall('movie'): # findall finds direct children
# findall 查找直接子元素
title = movie.get('title') # Get attribute value
# 获取属性值
year_element = movie.find('year') # find gets the first matching child
# find 获取第一个匹配的子元素
type_element = movie.find('type')
# Element text content is in .text attribute
# 元素的文本内容在 .text 属性中
movie_type = type_element.text if type_element is not None else 'N/A'
year = year_element.text if year_element is not None else 'N/A'
print(f"\nMovie Title: {title}")
print(f" Type: {movie_type}")
print(f" Year: {year}")
# Accessing potentially missing elements safely
# 安全地访问可能缺失的元素
episodes_element = movie.find('episodes')
if episodes_element is not None:
print(f" Episodes: {episodes_element.text}")
except ET.ParseError as e:
print(f"Error parsing XML: {e}")
except FileNotFoundError:
print(f"Error: File '{xml_file}' not found.")

输出:

Root element: collection, Shelf: New Arrivals
Movie Title: Enemy Behind
Type: War, Thriller
Year: 2003
Movie Title: Transformers
Type: Anime, Science Fiction
Year: 1989
Movie Title: Trigun
Type: Anime, Action
Year: N/A
Episodes: 4
Movie Title: Ishtar
Type: Comedy
Year: N/A

ElementTree 关键概念:

  • ET.parse(source): 从文件路径或类文件对象解析 XML。
  • ET.fromstring(text): 从字符串解析 XML。
  • tree.getroot(): 返回解析树的根元素。
  • element.tag: 元素的标签名(字符串)。
  • element.attrib: 元素的属性字典。
  • element.text: 元素内直接的文本内容。
  • element.get(attr_name): 获取属性值。
  • element.find(match): 查找与标签名或路径匹配的第一个直接子元素。
  • element.findall(match): 查找与标签名或路径匹配的所有直接子元素。
  • element.iter(tag): 创建一个迭代器,遍历所有匹配标签的子元素(直接和嵌套)。

minidom 将 XML 解析为标准的 DOM 树结构。

#!/usr/bin/env python3
from xml.dom import minidom
import xml.parsers.expat # Often needed to catch specific parse errors
# 通常需要捕捉特定的解析错误
xml_file = 'movies.xml'
try:
# Parse the XML file into a DOM tree
# 将 XML 文件解析成 DOM 树
dom_tree = minidom.parse(xml_file)
# Get the collection element (root)
# 获取 collection 元素(根元素)
collection = dom_tree.documentElement # The root <collection> element
# 根元素 <collection>
if collection.hasAttribute("shelf"):
print(f"Root element: <{collection.tagName}>, Shelf: {collection.getAttribute('shelf')}")
# Get all 'movie' elements
# 获取所有 'movie' 元素
movies = collection.getElementsByTagName("movie")
# Print detail of each movie
# 打印每部电影的详细信息
for movie in movies:
print("\n***** Movie *****")
if movie.hasAttribute("title"):
print(f"Title: {movie.getAttribute('title')}")
# Helper function to get text from first child element
# 获取第一个子元素的文本内容的辅助函数
def get_element_text(parent, tag_name):
elements = parent.getElementsByTagName(tag_name)
if elements and elements[0].firstChild:
# Check if firstChild exists and is a text node
# 检查 firstChild 是否存在且是文本节点
if elements[0].firstChild.nodeType == elements[0].TEXT_NODE:
return elements[0].firstChild.data.strip()
return 'N/A'
print(f"Type: {get_element_text(movie, 'type')}")
print(f"Format: {get_element_text(movie, 'format')}")
print(f"Rating: {get_element_text(movie, 'rating')}")
print(f"Description: {get_element_text(movie, 'description')}")
except xml.parsers.expat.ExpatError as e:
print(f"Error parsing XML: {e}")
except FileNotFoundError:
print(f"Error: File '{xml_file}' not found.")
except Exception as e:
print(f"An unexpected error occurred: {e}")

输出(类似于 ElementTree,演示 DOM 访问):

Root element: <collection>, Shelf: New Arrivals
***** Movie *****
Title: Enemy Behind
Type: War, Thriller
Format: DVD
Rating: PG
Description: Talk about a US-Japan war
***** Movie *****
Title: Transformers
Type: Anime, Science Fiction
Format: DVD
Rating: R
Description: A scientific fiction
...

DOM 关键概念:

  • minidom.parse(source): 解析 XML。
  • dom.documentElement: 根元素节点。
  • element.tagName: 元素的标签名。
  • element.getAttribute(name): 获取属性值。
  • element.hasAttribute(name): 检查属性是否存在。
  • element.getElementsByTagName(name): 获取所有具有给定标签名的后代元素的列表。
  • node.firstChild、node.childNodes、node.parentNode: 遍历树。
  • node.nodeType: 节点的类型 (ELEMENT_NODE 元素节点, TEXT_NODE 文本节点等)。
  • node.data: TEXT_NODE 的文本内容。

SAX 解析需要创建一个 ‘Content Handler’ 类,该类定义了当解析器遇到特定 XML 结构(开始/结束标签、字符数据)时要调用的方法。

#!/usr/bin/env python3
import xml.sax
# Define the Content Handler class
# 定义内容处理程序类
class MovieHandler(xml.sax.ContentHandler):
def __init__(self):
super().__init__()
self.current_data = "" # Stores the tag name being processed
# 存储正在处理的标签名
self.type = ""
self.format = ""
self.year = ""
self.rating = ""
self.stars = ""
self.description = ""
self.episodes = "" # Added for Trigun
# 为 Trigun 添加
self.current_movie_title = ""
# Called when an element starts
# 当元素开始时调用
def startElement(self, tag, attributes):
self.current_data = tag
if tag == "movie":
print("\n***** Movie *****")
self.current_movie_title = attributes.get("title", "N/A")
print(f"Title: {self.current_movie_title}")
# Reset movie-specific fields
# 重置电影特定字段
self.type = self.format = self.year = self.rating = ""
self.stars = self.description = self.episodes = ""
# Called when an element ends
# 当元素结束时调用
def endElement(self, tag):
if tag == "type":
print(f" Type: {self.type}")
elif tag == "format":
print(f" Format: {self.format}")
elif tag == "year":
print(f" Year: {self.year}")
elif tag == "rating":
print(f" Rating: {self.rating}")
elif tag == "stars":
print(f" Stars: {self.stars}")
elif tag == "description":
print(f" Description: {self.description}")
elif tag == "episodes":
print(f" Episodes: {self.episodes}")
# Reset current tag tracker
# 重置当前标签追踪器
self.current_data = ""
# Called with character data between tags
# 在标签之间的字符数据时调用
def characters(self, content):
# Only store content if inside a relevant tag
# 只有在相关标签内时才存储内容
if self.current_data == "type":
self.type += content.strip()
elif self.current_data == "format":
self.format += content.strip()
elif self.current_data == "year":
self.year += content.strip()
elif self.current_data == "rating":
self.rating += content.strip()
elif self.current_data == "stars":
self.stars += content.strip()
elif self.current_data == "description":
self.description += content.strip()
elif self.current_data == "episodes":
self.episodes += content.strip()
# --- Main execution ---
# --- 主执行流程 ---
xml_file = 'movies.xml'
# Create a SAX parser
# 创建一个 SAX 解析器
parser = xml.sax.make_parser()
# Turn off namespaces (optional, simplifies example)
# 关闭命名空间 (可选,简化示例)
parser.setFeature(xml.sax.handler.feature_namespaces, 0)
# Create an instance of our handler
# 创建处理程序实例
handler = MovieHandler()
# Set the handler for the parser
# 设置解析器的处理程序
parser.setContentHandler(handler)
print("--- Starting SAX Parse ---")
# --- SAX 解析开始 ---
try:
parser.parse(xml_file)
print("\n--- SAX Parse Finished ---")
# --- SAX 解析完成 ---
except xml.sax.SAXParseException as e:
print(f"SAX Parsing Error: {e}")
# SAX 解析错误
except FileNotFoundError:
print(f"Error: File '{xml_file}' not found.")
# 错误:文件未找到

输出(内容相似,处理流程不同):

--- Starting SAX Parse ---
***** Movie *****
Title: Enemy Behind
Type: War, Thriller
Format: DVD
Year: 2003
Rating: PG
Stars: 10
Description: Talk about a US-Japan war
***** Movie *****
Title: Transformers
...

SAX 关键概念:

  • xml.sax.make_parser(): 创建一个 SAX 解析器对象。
  • xml.sax.ContentHandler: 处理解析事件的基类。
  • startElement(tag, attributes): 遇到每个开始标签时调用。
  • endElement(tag): 遇到每个结束标签时调用。
  • characters(content): 在标签之间的字符数据块时调用。
  • parser.setContentHandler(handler): 将您的处理程序与解析器关联。
  • parser.parse(source): 启动解析过程。

ElementTree 也方便用于创建 XML 文档。

import xml.etree.ElementTree as ET
# Create root element
# 创建根元素
root = ET.Element("data")
# Create child elements
# 创建子元素
item1 = ET.SubElement(root, "item")
item1.set("name", "item1_name") # Set attribute
# 设置属性
item1.text = "This is the first item."
item2 = ET.SubElement(root, "item", attrib={"name": "item2_name", "type": "testing"})
item2.text = "Second item content."
# Create a tree and write to file
# 创建一个树并写入文件
tree = ET.ElementTree(root)
# Pretty print (optional, adds indentation)
# 漂亮打印 (可选,添加缩进)
ET.indent(tree, space=" ", level=0)
try:
tree.write("output.xml", encoding="utf-8", xml_declaration=True)
print("Generated output.xml")
# 已生成 output.xml
except IOError as e:
print(f"Error writing XML file: {e}")
# 写入 XML 文件错误