R - XML 文件
使用 xml2 在 R 中进行现代 XML 处理
Section titled “使用 xml2 在 R 中进行现代 XML 处理”XML(可扩展标记语言)是一种多功能格式,用于在 Web 和其他系统中构建和共享数据。与 HTML 类似,它使用标签,但与定义内容呈现方式的 HTML 不同,XML 标签定义了数据本身的含义和结构。这使其成为数据交换和配置文件一个强大的选择。
尽管 XML 包长期以来一直用于此目的,但现代 R 生态系统,特别是 Tidyverse(一组用于数据科学的 R 包,倡导整洁数据原则),更青睐 xml2 包。xml2 包基于强大的 libxml2 C 库构建,提供了一个更直观、高效且管道友好型接口。
先决条件:设置您的环境
Section titled “先决条件:设置您的环境”xml2 包是 tidyverse 的核心部分,因此您可以通过安装整个 tidyverse 套件来安装它。这还会为您提供强大的工具,例如我们将用于数据操作的 dplyr 和 purrr。
# 最佳实践:安装完整的 tidyverse 以获得一致的工作流程install.packages("tidyverse")
# 加载必要的库library(xml2)library(dplyr) # 用于数据操作,将随 tidyverse 加载library(purrr) # 用于函数式编程,也包含在 tidyverse 中输入数据:一个结构良好的 XML 示例
Section titled “输入数据:一个结构良好的 XML 示例”让我们使用一个干净、结构良好的 XML 文件作为示例。将以下内容保存为 employees.xml。请注意,它只有一个根元素 <records>,并且没有用 <html> 标签包裹,这对于数据文件来说是标准做法。
<?xml version="1.0" encoding="UTF-8"?><records> <employee id="101"> <name>Rick</name> <salary>623.30</salary> <start_date>2012-01-01</start_date> <department>IT</department> </employee> <employee id="102"> <name>Dan</name> <salary>515.20</salary> <start_date>2013-09-23</start_date> <department>Operations</department> </employee> <employee id="103"> <name>Michelle</name> <salary>611.00</salary> <start_date>2014-11-15</start_date> <department>IT</department> </employee> <employee id="104"> <name>Ryan</name> <salary>729.00</salary> <start_date>2014-05-11</start_date> <department>HR</department> </employee> <employee id="105"> <name>Gary</name> <salary>843.25</salary> <start_date>2015-03-27</start_date> <department>Finance</department> </employee></records>读取和检查 XML 文件
Section titled “读取和检查 XML 文件”read_xml() 函数将 XML 文件解析为 R 对象。然后我们可以使用 xml_children() 等函数检查其内容。
# 将 XML 文件读取为 xml_document 对象doc <- read_xml("employees.xml")
# 打印文档以查看其结构print(doc)执行上述代码后,将产生以下结果:
{xml_document}<records>[1] <employee id="101">\n <name>Rick</name>\n <salary>623.3</salary>\n <start_date>2012-01-01</start_date>\n <department>IT</department>\n</employee>[2] <employee id="102">\n <name>Dan</name>\n <salary>515.2</salary>\n <start_date>2013-09-23</start_date>\n <department>Operations</department>\n</employee>[3] <employee id="103">\n <name>Michelle</name>\n <salary>611</salary>\n <start_date>2014-11-15</start_date>\n <department>IT</department>\n</employee>[4] <employee id="104">\n <name>Ryan</name>\n <salary>729</salary>\n <start_date>2014-05-11</start_date>\n <department>HR</department>\n</employee>[5] <employee id="105">\n <name>Gary</name>\n <salary>843.25</salary>\n <start_date>2015-03-27</start_date>\n <department>Finance</department>\n</employee>使用 XPath 导航 XML
Section titled “使用 XPath 导航 XML”与手动索引不同,查询 XML 的现代标准方法是使用 XPath。xml_find_all() 函数使用 XPath 表达式来选择节点。这种方式极其强大且精确。
# 查找根目录下所有“employee”节点employee_nodes <- xml_find_all(doc, ".//employee")
# 获取找到的 employee 节点数量num_employees <- length(employee_nodes)cat("Number of employees found:", num_employees, "\n")
# 获取第一个 employee 节点first_employee <- employee_nodes[[1]]print(first_employee)
# 从第一个 employee 中提取特定信息# 获取“id”属性first_id <- xml_attr(first_employee, "id")cat("First employee ID:", first_id, "\n")
# 在第一个 employee 中查找“name”标签并获取其文本first_name <- xml_find_first(first_employee, ".//name") |> xml_text()cat("First employee Name:", first_name, "\n")执行上述代码后,将产生以下结果:
Number of employees found: 5{xml_node}<employee id="101">[1] <name>Rick</name>[2] <salary>623.3</salary>[3] <start_date>2012-01-01</start_date>[4] <department>IT</department>
First employee ID: 101First employee Name: Rick最佳实践:将 XML 转换为整洁数据框
Section titled “最佳实践:将 XML 转换为整洁数据框”对于数据分析,最常见的有用格式是数据框(或 tibble,一种现代化的数据框)。虽然旧的 XML 包提供了 xmlToDataFrame(),但它在数据类型处理上通常不可靠。现代 tidyverse 方法更具明确性、健壮性和可读性。我们将遍历每个 <employee> 节点,并将其信息提取到单独的行中。
# 一个用于处理单个 <employee> 节点的辅助函数parse_employee_node <- function(node) { # 定义一个安全的提取函数以处理缺失的标签 safe_get_text <- function(n, xpath) { child <- xml_find_first(n, xpath) if (is.na(child)) NA_character_ else xml_text(child) }
tibble( id = xml_attr(node, "id") |> as.integer(), name = safe_get_text(node, ".//name"), salary = safe_get_text(node, ".//salary") |> as.numeric(), start_date = safe_get_text(node, ".//start_date") |> as.Date(), department = safe_get_text(node, ".//department") )}
# 查找所有 employee 节点employee_nodes <- xml_find_all(doc, ".//employee")
# 使用 purrr::map_dfr 将函数应用于每个节点并按行绑定结果employees_df <- map_dfr(employee_nodes, parse_employee_node)
print(employees_df)此方法可让您完全控制解析过程,包括处理缺失值和设置正确的数据类型。执行上述代码后,将生成一个干净、整洁的 tibble:
# 一个 tibble: 5 × 5 id name salary start_date department <int> <chr> <dbl> <date> <chr>1 101 Rick 623. 2012-01-01 IT2 102 Dan 515. 2013-09-23 Operations3 103 Michelle 611 2014-11-15 IT4 104 Ryan 729 2014-05-11 HR5 105 Gary 843. 2015-03-27 Finance常见错误和调试
Section titled “常见错误和调试”1. 格式错误的 XML: 如果 XML 格式不正确,read_xml() 将抛出错误。请使用在线验证器检查您的文件。
2. 缺失节点/属性: 如果 xml_find_first() 未找到节点,它将返回 NA。如果您尝试对 NA 使用 xml_text() 或 xml_attr(),您将会得到一个错误。我们的 safe_get_text 辅助函数展示了一种优雅地处理这种情况的方法。
3. 命名空间: 实际的 XML 通常使用命名空间(例如,<ns1:tag>)。要查询这些命名空间,您必须使用 xml_ns() 注册命名空间,并在 XPath 查询中使用它。这是一个高级主题,但也是一个常见的障碍。