Skip to content

R - XML 文件

使用 xml2 在 R 中进行现代 XML 处理

Section titled “使用 xml2 在 R 中进行现代 XML 处理”

XML(可扩展标记语言)是一种多功能格式,用于在 Web 和其他系统中构建和共享数据。与 HTML 类似,它使用标签,但与定义内容呈现方式的 HTML 不同,XML 标签定义了数据本身的含义和结构。这使其成为数据交换和配置文件一个强大的选择。

尽管 XML 包长期以来一直用于此目的,但现代 R 生态系统,特别是 Tidyverse(一组用于数据科学的 R 包,倡导整洁数据原则),更青睐 xml2 包。xml2 包基于强大的 libxml2 C 库构建,提供了一个更直观、高效且管道友好型接口。

xml2 包是 tidyverse 的核心部分,因此您可以通过安装整个 tidyverse 套件来安装它。这还会为您提供强大的工具,例如我们将用于数据操作的 dplyr 和 purrr。

# 最佳实践:安装完整的 tidyverse 以获得一致的工作流程
install.packages("tidyverse")
# 加载必要的库
library(xml2)
library(dplyr) # 用于数据操作,将随 tidyverse 加载
library(purrr) # 用于函数式编程,也包含在 tidyverse 中

输入数据:一个结构良好的 XML 示例

Section titled “输入数据:一个结构良好的 XML 示例”

让我们使用一个干净、结构良好的 XML 文件作为示例。将以下内容保存为 employees.xml。请注意,它只有一个根元素 <records>,并且没有用 <html> 标签包裹,这对于数据文件来说是标准做法。

<?xml version="1.0" encoding="UTF-8"?>
<records>
<employee id="101">
<name>Rick</name>
<salary>623.30</salary>
<start_date>2012-01-01</start_date>
<department>IT</department>
</employee>
<employee id="102">
<name>Dan</name>
<salary>515.20</salary>
<start_date>2013-09-23</start_date>
<department>Operations</department>
</employee>
<employee id="103">
<name>Michelle</name>
<salary>611.00</salary>
<start_date>2014-11-15</start_date>
<department>IT</department>
</employee>
<employee id="104">
<name>Ryan</name>
<salary>729.00</salary>
<start_date>2014-05-11</start_date>
<department>HR</department>
</employee>
<employee id="105">
<name>Gary</name>
<salary>843.25</salary>
<start_date>2015-03-27</start_date>
<department>Finance</department>
</employee>
</records>

read_xml() 函数将 XML 文件解析为 R 对象。然后我们可以使用 xml_children() 等函数检查其内容。

# 将 XML 文件读取为 xml_document 对象
doc <- read_xml("employees.xml")
# 打印文档以查看其结构
print(doc)

执行上述代码后,将产生以下结果:

{xml_document}
<records>
[1] <employee id="101">\n <name>Rick</name>\n <salary>623.3</salary>\n <start_date>2012-01-01</start_date>\n <department>IT</department>\n</employee>
[2] <employee id="102">\n <name>Dan</name>\n <salary>515.2</salary>\n <start_date>2013-09-23</start_date>\n <department>Operations</department>\n</employee>
[3] <employee id="103">\n <name>Michelle</name>\n <salary>611</salary>\n <start_date>2014-11-15</start_date>\n <department>IT</department>\n</employee>
[4] <employee id="104">\n <name>Ryan</name>\n <salary>729</salary>\n <start_date>2014-05-11</start_date>\n <department>HR</department>\n</employee>
[5] <employee id="105">\n <name>Gary</name>\n <salary>843.25</salary>\n <start_date>2015-03-27</start_date>\n <department>Finance</department>\n</employee>

与手动索引不同,查询 XML 的现代标准方法是使用 XPath。xml_find_all() 函数使用 XPath 表达式来选择节点。这种方式极其强大且精确。

# 查找根目录下所有“employee”节点
employee_nodes <- xml_find_all(doc, ".//employee")
# 获取找到的 employee 节点数量
num_employees <- length(employee_nodes)
cat("Number of employees found:", num_employees, "\n")
# 获取第一个 employee 节点
first_employee <- employee_nodes[[1]]
print(first_employee)
# 从第一个 employee 中提取特定信息
# 获取“id”属性
first_id <- xml_attr(first_employee, "id")
cat("First employee ID:", first_id, "\n")
# 在第一个 employee 中查找“name”标签并获取其文本
first_name <- xml_find_first(first_employee, ".//name") |> xml_text()
cat("First employee Name:", first_name, "\n")

执行上述代码后,将产生以下结果:

Number of employees found: 5
{xml_node}
<employee id="101">
[1] <name>Rick</name>
[2] <salary>623.3</salary>
[3] <start_date>2012-01-01</start_date>
[4] <department>IT</department>
First employee ID: 101
First employee Name: Rick

最佳实践:将 XML 转换为整洁数据框

Section titled “最佳实践:将 XML 转换为整洁数据框”

对于数据分析,最常见的有用格式是数据框(或 tibble,一种现代化的数据框)。虽然旧的 XML 包提供了 xmlToDataFrame(),但它在数据类型处理上通常不可靠。现代 tidyverse 方法更具明确性、健壮性和可读性。我们将遍历每个 <employee> 节点,并将其信息提取到单独的行中。

# 一个用于处理单个 <employee> 节点的辅助函数
parse_employee_node <- function(node) {
# 定义一个安全的提取函数以处理缺失的标签
safe_get_text <- function(n, xpath) {
child <- xml_find_first(n, xpath)
if (is.na(child)) NA_character_ else xml_text(child)
}
tibble(
id = xml_attr(node, "id") |> as.integer(),
name = safe_get_text(node, ".//name"),
salary = safe_get_text(node, ".//salary") |> as.numeric(),
start_date = safe_get_text(node, ".//start_date") |> as.Date(),
department = safe_get_text(node, ".//department")
)
}
# 查找所有 employee 节点
employee_nodes <- xml_find_all(doc, ".//employee")
# 使用 purrr::map_dfr 将函数应用于每个节点并按行绑定结果
employees_df <- map_dfr(employee_nodes, parse_employee_node)
print(employees_df)

此方法可让您完全控制解析过程,包括处理缺失值和设置正确的数据类型。执行上述代码后,将生成一个干净、整洁的 tibble:

# 一个 tibble: 5 × 5
id name salary start_date department
<int> <chr> <dbl> <date> <chr>
1 101 Rick 623. 2012-01-01 IT
2 102 Dan 515. 2013-09-23 Operations
3 103 Michelle 611 2014-11-15 IT
4 104 Ryan 729 2014-05-11 HR
5 105 Gary 843. 2015-03-27 Finance

1. 格式错误的 XML: 如果 XML 格式不正确,read_xml() 将抛出错误。请使用在线验证器检查您的文件。

2. 缺失节点/属性: 如果 xml_find_first() 未找到节点,它将返回 NA。如果您尝试对 NA 使用 xml_text() 或 xml_attr(),您将会得到一个错误。我们的 safe_get_text 辅助函数展示了一种优雅地处理这种情况的方法。

3. 命名空间: 实际的 XML 通常使用命名空间(例如,<ns1:tag>)。要查询这些命名空间,您必须使用 xml_ns() 注册命名空间,并在 XPath 查询中使用它。这是一个高级主题,但也是一个常见的障碍。