Ruby Ruby/XML, XSLT
Ruby: XML、XPath 和 XSLT
Section titled “Ruby: XML、XPath 和 XSLT”什么是 XML?
Section titled “什么是 XML?”可扩展标记语言(XML,Extensible Markup Language)是一种标记语言,旨在传输数据,而不是显示数据(与 HTML 不同)。它是 W3C 推荐标准和开放标准,使得不同应用程序和系统之间能够交换结构化数据。
XML 对于数据存储、配置文件以及作为网络上的数据交换格式都很有价值。
XML 解析方法
Section titled “XML 解析方法”解析 XML 文档主要有两种方法:
- DOM (Document Object Model 文档对象模型):解析器将整个 XML 文档读入内存并构建一个代表文档的树形结构。这使得导航和修改 XML 变得容易。然而,对于非常大的文档,它可能会占用大量内存。Nokogiri 主要使用 DOM 风格的方法来解析整个文档。
- SAX (Simple API for XML 简单 XML 应用编程接口):这是一种基于事件的方法。解析器按顺序读取文档,并在遇到元素、属性、文本等时触发事件(回调)。对于大型文档来说,由于不需要将整个文档存储在内存中,因此它更省内存。解析过程中通常是只读的。
在 DOM 和 SAX 之间选择取决于任务:DOM 通常更易于随机访问和修改,而 SAX 更适合处理大型文件或关注内存使用的情况。
Ruby 中的 XML 处理:Nokogiri 和 REXML
Section titled “Ruby 中的 XML 处理:Nokogiri 和 REXML”虽然 Ruby 的标准库包含 REXML(一个纯 Ruby 的 XML 处理器),但在 Ruby 社区中,用于处理 XML(和 HTML)的事实标准是 Nokogiri gem。Nokogiri 显着更快、功能更丰富,它利用原生 C 库(如 libxml2)来提升性能。
安装 Nokogiri:
Section titled “安装 Nokogiri:”gem install nokogiriREXML: REXML 是标准库的一部分,因此无需安装。它适用于更简单的任务或不希望依赖外部 gem 的情况。
在本教程中,我们将主要关注 Nokogiri,因为它更常用且功能强大。我们将使用以下示例 XML 文件 (movies.xml):
<!-- movies.xml --><collection shelf="New Arrivals"> <movie title="Enemy Behind"> <type>War, Thriller</type> <format>DVD</format> <year>2003</year> <rating>PG</rating> <stars>8</stars> <description>A gripping war drama.</description> </movie> <movie title="Cosmic Odyssey"> <type>Science Fiction</type> <format>Blu-ray</format> <year>2021</year> <rating>PG-13</rating> <stars>9</stars> <description>An epic journey through space.</description> </movie> <movie title="The Silent City"> <type>Mystery</type> <format>Streaming</format> <episodes>10</episodes> <!-- Note: some movies might have episodes instead of year --> <rating>R</rating> <stars>7</stars> <description>A detective uncovers secrets.</description> </movie></collection>使用 Nokogiri 进行 DOM 风格解析
Section titled “使用 Nokogiri 进行 DOM 风格解析”Nokogiri 可以从文件、字符串或 IO 对象解析 XML。
#!/usr/bin/env ruby# frozen_string_literal: true
require 'nokogiri'
# Assuming movies.xml is in the same directorybegin file = File.open("movies.xml") doc = Nokogiri::XML(file) file.close
# Get the root element and its attributes root = doc.root puts "Root element: #{root.name}, Shelf: #{root['shelf']}"
puts "\n--- Movie Titles ---" doc.xpath("//movie").each do |movie_node| puts "Title: #{movie_node['title']}" end
puts "\n--- Movie Types ---" doc.xpath("//movie/type").each do |type_node| puts "Type: #{type_node.content}" # .text or .content gets the text content end
puts "\n--- Movie Descriptions ---" # Using CSS selectors as an alternative to XPath doc.css("collection movie description").each do |desc_node| puts "Description: #{desc_node.text}" end
# Find a specific movie and its details puts "\n--- Details for 'Cosmic Odyssey' ---" cosmic_odyssey = doc.at_xpath("//movie[@title='Cosmic Odyssey']") # .at_xpath finds the first match if cosmic_odyssey puts "Format: #{cosmic_odyssey.at_xpath('format')&.text}" # &. safe navigation operator puts "Year: #{cosmic_odyssey.at_css('year')&.text}" puts "Rating: #{cosmic_odyssey.css('rating').first&.text}" else puts "'Cosmic Odyssey' not found." end
rescue Errno::ENOENT puts "Error: movies.xml not found."rescue StandardError => e puts "An XML parsing error occurred: #{e.message}"end此示例演示了如何读取 XML,使用 XPath 和 CSS 选择器访问元素和属性,以及遍历节点。
使用 Nokogiri 进行 SAX 风格解析
Section titled “使用 Nokogiri 进行 SAX 风格解析”对于 SAX 解析,您需要创建一个继承自 Nokogiri::XML::SAX::Document 的自定义处理程序类,并覆盖其事件方法。
#!/usr/bin/env ruby# frozen_string_literal: true
require 'nokogiri'
class MovieSAXHandler < Nokogiri::XML::SAX::Document def initialize @current_element = nil @in_movie_title = false end
# 在元素开始时调用 def start_element(name, attrs = []) @current_element = name # 为简单起见,如果需要可以转换为哈希:Hash[attrs] # attrs is an array of [localname, prefix, URI, value] or [localname, value] puts "Start Element: #{name}, Attributes: #{Hash[attrs]}" if name == "movie" title = Hash[attrs]["title"] puts " -> Found movie: #{title}" if title end @in_movie_title = true if name == "title" && @current_element_path&.end_with?("/movie/title") end
# 在元素内部遇到字符数据时调用 def characters(string) # 避免打印过多的空白字符,修剪以提高可读性 text_content = string.strip if text_content.length > 0 puts " Characters: '#{text_content}' (inside #{@current_element})" end end
# 在元素结束时调用 def end_element(name) puts "End Element: #{name}" @current_element = nil @in_movie_title = false if name == "title" end
def error(message) puts "SAX Error: #{message}" end
def warning(message) puts "SAX Warning: #{message}" endend
begin parser = Nokogiri::XML::SAX::Parser.new(MovieSAXHandler.new) file_path = "movies.xml" puts "Starting SAX parsing of #{file_path}...\n" parser.parse(File.open(file_path)) puts "\nSAX parsing finished."rescue Errno::ENOENT puts "Error: movies.xml not found."rescue StandardError => e puts "An error occurred: #{e.message}"endSAX 解析的设置更繁琐,但对于大型 XML 流非常高效。
使用 Nokogiri 的 XPath 和 CSS 选择器
Section titled “使用 Nokogiri 的 XPath 和 CSS 选择器”XPath 是一种用于在 XML 文档中导航(查找信息)的语言。CSS 选择器,通常用于 HTML,也可以与 Nokogiri 一起用于 XML。Nokogiri 在文档和节点对象上提供了 xpath() 和 css() 方法。at_xpath() 和 at_css() 方法用于查找第一个匹配的节点。
#!/usr/bin/env ruby# frozen_string_literal: truerequire 'nokogiri'
xml_data = <<~XML <catalog> <book id="bk101"> <author>Gambardella, Matthew</author> <title>XML Developer's Guide</title> <genre>Computer</genre> <price>44.95</price> </book> <book id="bk102"> <author>Ralls, Kim</author> <title>Midnight Rain</title> <genre>Fantasy</genre> <price>5.95</price> </book> </catalog>XML
doc = Nokogiri::XML(xml_data)
# XPath 示例puts "--- XPath --- "puts "All book titles:"doc.xpath("//book/title").each { |title| puts " - #{title.text}" }
puts "Author of book with id 'bk101': #{doc.at_xpath("//book[@id='bk101']/author")&.text}"
# CSS 选择器示例puts "\n--- CSS Selectors --- "puts "All prices:"doc.css("book price").each { |price| puts " - #{price.text}" }
puts "Genre of book with id 'bk102': #{doc.at_css("book[id='bk102'] genre")&.text}"使用 Nokogiri 进行 XSLT 处理
Section titled “使用 Nokogiri 进行 XSLT 处理”XSLT (Extensible Stylesheet Language Transformations 可扩展样式表语言转换) 是一种用于将 XML 文档转换为其他 XML 文档或 HTML、纯文本等其他格式的语言。Nokogiri 可以应用 XSLT 样式表。
假设您有一个 XSLT 样式表 (movies_to_html.xsl):
<!-- movies_to_html.xsl --><xsl:stylesheet version="1.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform"> <xsl:template match="/"> <html> <body> <h2>Movie Collection</h2> <table border="1"> <tr> <th>Title</th> <th>Type</th> <th>Year/Episodes</th> </tr> <xsl:for-each select="collection/movie"> <tr> <td><xsl:value-of select="@title"/></td> <td><xsl:value-of select="type"/></td> <td> <xsl:choose> <xsl:when test="year"> <xsl:value-of select="year"/> </xsl:when> <xsl:otherwise> <xsl:value-of select="episodes"/> episodes </xsl:otherwise> </xsl:choose> </td> </tr> </xsl:for-each> </table> </body> </html> </xsl:template></xsl:stylesheet>您可以使用 Nokogiri 应用此转换:
#!/usr/bin/env ruby# frozen_string_literal: true
require 'nokogiri'
begin xml_doc = Nokogiri::XML(File.open("movies.xml")) xslt_doc = Nokogiri::XSLT(File.open("movies_to_html.xsl"))
html_output = xslt_doc.transform(xml_doc)
puts "--- Transformed HTML Output ---" puts html_output.to_html # 如果输出是 XML,可以使用 .to_xml
# 保存到 HTML 文件 File.open("movies_report.html", "w") { |f| f.write(html_output.to_html) } puts "\nReport saved to movies_report.html"
rescue Errno::ENOENT => e puts "Error: Missing XML or XSLT file. #{e.message}"rescue StandardError => e puts "An error occurred during XSLT processing: #{e.message}"end进一步阅读:
Section titled “进一步阅读:”- Nokogiri Official Documentation:
https://nokogiri.org/ - REXML Documentation (Ruby Standard Library): 在 Ruby 核心文档中搜索 REXML。
- W3Schools XPath Tutorial:
https://www.w3schools.com/xml/xpath_intro.asp - W3Schools XSLT Tutorial:
https://www.w3schools.com/xml/xsl_intro.asp