Python XML 处理
Python XML 处理
Section titled “Python XML 处理”XML (Extensible Markup Language,可扩展标记语言) 是一种标记语言,旨在以人类可读和机器可读的格式存储和传输数据。它广泛用于配置文件、数据交换和文档格式。
什么是 XML?
Section titled “什么是 XML?”XML 定义了一套编码文档的规则。它使用标签(类似于 HTML)来定义元素,并使用属性来提供关于这些元素的元数据。与 HTML 不同,XML 不预定义标签;用户根据其数据定义自己的标签。
主要特点:
- 可扩展:您可以定义自己的标签。
- 层次化:数据以树形结构组织。
- 文本化:使用 Unicode 字符。
- 自描述:标签通常描述它们包含的数据。
- 平台/语言无关:一个开放标准。
尽管 JSON 因其简单性以及与 JavaScript 对象的直接映射而在 Web API 中变得更受欢迎,但 XML 在许多企业系统、文档标准(如 DocBook、Office Open XML)和配置方面仍然很重要。
Python XML 解析 API
Section titled “Python XML 解析 API”Python 的标准库提供了几个用于处理 XML 的模块,支持不同的解析策略:
xml.etree.ElementTree: (推荐用于大多数常见任务) 提供了一种简单且符合 Python 风格的方式来解析和创建 XML。它将整个文档加载到内存中的树状结构(类似于 DOM),但提供了一个更轻量级且通常更易于使用的 API。xml.dom.minidom: 实现了 Document Object Model (DOM) API,一个 W3C 标准。它将整个 XML 文档解析到内存中的对象树中,允许完全访问和修改。对于非常大的文件可能占用大量内存。xml.sax: 实现了 Simple API for XML (SAX),一个事件驱动的解析器。它按顺序读取 XML 文件,在解析器遇到特定 XML 结构(如开始标签、结束标签和字符数据)时触发回调函数(处理程序)。它不加载整个文档,因此内存效率高,适用于非常大的文件或流处理,但对于随机访问可能更复杂。
选择 API:
- 对于一般的 XML 解析和创建,首先使用
ElementTree。 - 如果需要严格遵循 DOM 标准或进行复杂的树修改,使用
minidom。 - 对于处理非常大且内存有限的 XML 文件,使用
SAX。
示例 XML 文件(movies.xml)
Section titled “示例 XML 文件(movies.xml)”我们将使用以下 XML 文件作为示例:
<?xml version="1.0" encoding="UTF-8"?><collection shelf="New Arrivals"> <movie title="Enemy Behind"> <type>War, Thriller</type> <format>DVD</format> <year>2003</year> <rating>PG</rating> <stars>10</stars> <description>Talk about a US-Japan war</description> </movie> <movie title="Transformers"> <type>Anime, Science Fiction</type> <format>DVD</format> <year>1989</year> <rating>R</rating> <stars>8</stars> <description>A scientific fiction</description> </movie> <movie title="Trigun"> <type>Anime, Action</type> <format>DVD</nFormat> <episodes>4</episodes> <rating>PG</rating> <stars>10</stars> <description>Vash the Stampede!</description> </movie> <movie title="Ishtar"> <type>Comedy</type> <format>VHS</format> <rating>PG</rating> <stars>2</stars> <description>Viewable boredom</description> </movie></collection>使用 ElementTree 解析 XML(推荐)
Section titled “使用 ElementTree 解析 XML(推荐)”ElementTree 将 XML 文档表示为 Element 对象的树。
#!/usr/bin/env python3
import xml.etree.ElementTree as ET
xml_file = 'movies.xml'
try: tree = ET.parse(xml_file) root = tree.getroot() # Get the root element ('collection') # 获取根元素 ('collection')
print(f"Root element: {root.tag}, Shelf: {root.get('shelf')}")
# Iterate through 'movie' elements under the root # 遍历根元素下的 'movie' 元素 for movie in root.findall('movie'): # findall finds direct children # findall 查找直接子元素 title = movie.get('title') # Get attribute value # 获取属性值 year_element = movie.find('year') # find gets the first matching child # find 获取第一个匹配的子元素 type_element = movie.find('type')
# Element text content is in .text attribute # 元素的文本内容在 .text 属性中 movie_type = type_element.text if type_element is not None else 'N/A' year = year_element.text if year_element is not None else 'N/A'
print(f"\nMovie Title: {title}") print(f" Type: {movie_type}") print(f" Year: {year}")
# Accessing potentially missing elements safely # 安全地访问可能缺失的元素 episodes_element = movie.find('episodes') if episodes_element is not None: print(f" Episodes: {episodes_element.text}")
except ET.ParseError as e: print(f"Error parsing XML: {e}")except FileNotFoundError: print(f"Error: File '{xml_file}' not found.")输出:
Root element: collection, Shelf: New Arrivals
Movie Title: Enemy Behind Type: War, Thriller Year: 2003
Movie Title: Transformers Type: Anime, Science Fiction Year: 1989
Movie Title: Trigun Type: Anime, Action Year: N/A Episodes: 4
Movie Title: Ishtar Type: Comedy Year: N/AElementTree 关键概念:
ET.parse(source): 从文件路径或类文件对象解析 XML。ET.fromstring(text): 从字符串解析 XML。tree.getroot(): 返回解析树的根元素。element.tag: 元素的标签名(字符串)。element.attrib: 元素的属性字典。element.text: 元素内直接的文本内容。element.get(attr_name): 获取属性值。element.find(match): 查找与标签名或路径匹配的第一个直接子元素。element.findall(match): 查找与标签名或路径匹配的所有直接子元素。element.iter(tag): 创建一个迭代器,遍历所有匹配标签的子元素(直接和嵌套)。
使用 DOM(minidom)解析 XML
Section titled “使用 DOM(minidom)解析 XML”minidom 将 XML 解析为标准的 DOM 树结构。
#!/usr/bin/env python3
from xml.dom import minidomimport xml.parsers.expat # Often needed to catch specific parse errors# 通常需要捕捉特定的解析错误
xml_file = 'movies.xml'
try: # Parse the XML file into a DOM tree # 将 XML 文件解析成 DOM 树 dom_tree = minidom.parse(xml_file)
# Get the collection element (root) # 获取 collection 元素(根元素) collection = dom_tree.documentElement # The root <collection> element # 根元素 <collection>
if collection.hasAttribute("shelf"): print(f"Root element: <{collection.tagName}>, Shelf: {collection.getAttribute('shelf')}")
# Get all 'movie' elements # 获取所有 'movie' 元素 movies = collection.getElementsByTagName("movie")
# Print detail of each movie # 打印每部电影的详细信息 for movie in movies: print("\n***** Movie *****") if movie.hasAttribute("title"): print(f"Title: {movie.getAttribute('title')}")
# Helper function to get text from first child element # 获取第一个子元素的文本内容的辅助函数 def get_element_text(parent, tag_name): elements = parent.getElementsByTagName(tag_name) if elements and elements[0].firstChild: # Check if firstChild exists and is a text node # 检查 firstChild 是否存在且是文本节点 if elements[0].firstChild.nodeType == elements[0].TEXT_NODE: return elements[0].firstChild.data.strip() return 'N/A'
print(f"Type: {get_element_text(movie, 'type')}") print(f"Format: {get_element_text(movie, 'format')}") print(f"Rating: {get_element_text(movie, 'rating')}") print(f"Description: {get_element_text(movie, 'description')}")
except xml.parsers.expat.ExpatError as e: print(f"Error parsing XML: {e}")except FileNotFoundError: print(f"Error: File '{xml_file}' not found.")except Exception as e: print(f"An unexpected error occurred: {e}")输出(类似于 ElementTree,演示 DOM 访问):
Root element: <collection>, Shelf: New Arrivals
***** Movie *****Title: Enemy BehindType: War, ThrillerFormat: DVDRating: PGDescription: Talk about a US-Japan war
***** Movie *****Title: TransformersType: Anime, Science FictionFormat: DVDRating: RDescription: A scientific fiction...DOM 关键概念:
minidom.parse(source): 解析 XML。dom.documentElement: 根元素节点。element.tagName: 元素的标签名。element.getAttribute(name): 获取属性值。element.hasAttribute(name): 检查属性是否存在。element.getElementsByTagName(name): 获取所有具有给定标签名的后代元素的列表。node.firstChild、node.childNodes、node.parentNode: 遍历树。node.nodeType: 节点的类型 (ELEMENT_NODE 元素节点, TEXT_NODE 文本节点等)。node.data: TEXT_NODE 的文本内容。
使用 SAX(xml.sax)解析 XML
Section titled “使用 SAX(xml.sax)解析 XML”SAX 解析需要创建一个 ‘Content Handler’ 类,该类定义了当解析器遇到特定 XML 结构(开始/结束标签、字符数据)时要调用的方法。
#!/usr/bin/env python3
import xml.sax
# Define the Content Handler class# 定义内容处理程序类class MovieHandler(xml.sax.ContentHandler): def __init__(self): super().__init__() self.current_data = "" # Stores the tag name being processed # 存储正在处理的标签名 self.type = "" self.format = "" self.year = "" self.rating = "" self.stars = "" self.description = "" self.episodes = "" # Added for Trigun # 为 Trigun 添加 self.current_movie_title = ""
# Called when an element starts # 当元素开始时调用 def startElement(self, tag, attributes): self.current_data = tag if tag == "movie": print("\n***** Movie *****") self.current_movie_title = attributes.get("title", "N/A") print(f"Title: {self.current_movie_title}") # Reset movie-specific fields # 重置电影特定字段 self.type = self.format = self.year = self.rating = "" self.stars = self.description = self.episodes = ""
# Called when an element ends # 当元素结束时调用 def endElement(self, tag): if tag == "type": print(f" Type: {self.type}") elif tag == "format": print(f" Format: {self.format}") elif tag == "year": print(f" Year: {self.year}") elif tag == "rating": print(f" Rating: {self.rating}") elif tag == "stars": print(f" Stars: {self.stars}") elif tag == "description": print(f" Description: {self.description}") elif tag == "episodes": print(f" Episodes: {self.episodes}")
# Reset current tag tracker # 重置当前标签追踪器 self.current_data = ""
# Called with character data between tags # 在标签之间的字符数据时调用 def characters(self, content): # Only store content if inside a relevant tag # 只有在相关标签内时才存储内容 if self.current_data == "type": self.type += content.strip() elif self.current_data == "format": self.format += content.strip() elif self.current_data == "year": self.year += content.strip() elif self.current_data == "rating": self.rating += content.strip() elif self.current_data == "stars": self.stars += content.strip() elif self.current_data == "description": self.description += content.strip() elif self.current_data == "episodes": self.episodes += content.strip()
# --- Main execution ---# --- 主执行流程 ---xml_file = 'movies.xml'
# Create a SAX parser# 创建一个 SAX 解析器parser = xml.sax.make_parser()# Turn off namespaces (optional, simplifies example)# 关闭命名空间 (可选,简化示例)parser.setFeature(xml.sax.handler.feature_namespaces, 0)
# Create an instance of our handler# 创建处理程序实例handler = MovieHandler()# Set the handler for the parser# 设置解析器的处理程序parser.setContentHandler(handler)
print("--- Starting SAX Parse ---")# --- SAX 解析开始 ---try: parser.parse(xml_file) print("\n--- SAX Parse Finished ---") # --- SAX 解析完成 ---except xml.sax.SAXParseException as e: print(f"SAX Parsing Error: {e}") # SAX 解析错误except FileNotFoundError: print(f"Error: File '{xml_file}' not found.") # 错误:文件未找到输出(内容相似,处理流程不同):
--- Starting SAX Parse ---
***** Movie *****Title: Enemy Behind Type: War, Thriller Format: DVD Year: 2003 Rating: PG Stars: 10 Description: Talk about a US-Japan war
***** Movie *****Title: Transformers...SAX 关键概念:
xml.sax.make_parser(): 创建一个 SAX 解析器对象。xml.sax.ContentHandler: 处理解析事件的基类。startElement(tag, attributes): 遇到每个开始标签时调用。endElement(tag): 遇到每个结束标签时调用。characters(content): 在标签之间的字符数据块时调用。parser.setContentHandler(handler): 将您的处理程序与解析器关联。parser.parse(source): 启动解析过程。
XML 生成
Section titled “XML 生成”ElementTree 也方便用于创建 XML 文档。
import xml.etree.ElementTree as ET
# Create root element# 创建根元素root = ET.Element("data")
# Create child elements# 创建子元素item1 = ET.SubElement(root, "item")item1.set("name", "item1_name") # Set attribute# 设置属性item1.text = "This is the first item."
item2 = ET.SubElement(root, "item", attrib={"name": "item2_name", "type": "testing"})item2.text = "Second item content."
# Create a tree and write to file# 创建一个树并写入文件tree = ET.ElementTree(root)
# Pretty print (optional, adds indentation)# 漂亮打印 (可选,添加缩进)ET.indent(tree, space=" ", level=0)
try: tree.write("output.xml", encoding="utf-8", xml_declaration=True) print("Generated output.xml") # 已生成 output.xmlexcept IOError as e: print(f"Error writing XML file: {e}") # 写入 XML 文件错误