Sed 特殊字符
理解 sed 中的特殊字符
Section titled “理解 sed 中的特殊字符”流编辑器 sed 使用特殊字符作为命令或在脚本中具有独特的含义。掌握这些特殊字符可以显著增强你的文本处理能力。本节重点介绍两个关键的特殊字符:= 和 &。
= 命令:显示行号
Section titled “= 命令:显示行号”= 命令用于将当前行的行号打印到 standard output(标准输出)。它通常后跟一个换行符,然后 sed 会继续打印该行本身(除非抑制了默认输出)。
[/pattern/]=[address1[,address2]]=假设我们有一个名为 books.txt 的文件,内容如下:
1) A Storm of Swords, George R. R. Martin, 12162) The Two Towers, J. R. R. Tolkien, 3523) The Alchemist, Paulo Coelho, 1974) The Fellowship of the Ring, J. R. R. Tolkien, 4325) The Pilgrimage, Paulo Coelho, 2886) A Game of Thrones, George R. R. Martin, 864示例 1:打印所有行的行号。
$ sed '=' books.txt输出:
11) A Storm of Swords, George R. R. Martin, 121622) The Two Towers, J. R. R. Tolkien, 35233) The Alchemist, Paulo Coelho, 19744) The Fellowship of the Ring, J. R. R. Tolkien, 43255) The Pilgrimage, Paulo Coelho, 28866) A Game of Thrones, George R. R. Martin, 864解释:对于每一行,sed 首先打印行号,后跟一个换行符,然后由于其默认行为,打印行内容本身。
示例 2:打印指定行范围(例如,前四行)的行号。
$ sed '1,4=' books.txt输出:
11) A Storm of Swords, George R. R. Martin, 121622) The Two Towers, J. R. R. Tolkien, 35233) The Alchemist, Paulo Coelho, 19744) The Fellowship of the Ring, J. R. R. Tolkien, 4325) The Pilgrimage, Paulo Coelho, 2886) A Game of Thrones, George R. R. Martin, 864解释:为第 1 到第 4 行打印了行号。第 5 行和第 6 行没有在其前面打印行号,因为它们超出了 = 命令指定的地址范围。
示例 3:仅打印匹配特定模式的行的行号。
$ sed '/Paulo/=' books.txt输出:
1) A Storm of Swords, George R. R. Martin, 12162) The Two Towers, J. R. R. Tolkien, 35233) The Alchemist, Paulo Coelho, 1974) The Fellowship of the Ring, J. R. R. Tolkien, 43255) The Pilgrimage, Paulo Coelho, 28866) A Game of Thrones, George R. R. Martin, 864解释:当一行包含模式 “Paulo” 时,其行号会在该行本身之前打印。
示例 4:计算总行数(模拟 wc -l)。
$ sed -n '$=' books.txt输出:
6解释:$ 地址表示最后一行。= 命令打印其行号。-n 选项抑制了 sed 的默认打印每行的行为。因此,只打印了最后一行的行号(即总行数)。
& 字符:重用匹配的模式
Section titled “& 字符:重用匹配的模式”与号 & 是一个特殊字符,主要用于替换命令(s/pattern/replacement/)的 replacement(替换部分)。它代表 pattern(模式)所匹配的整个字符串。
示例 1:在匹配的模式前添加文本。
假设我们想在 books.txt 中每行开头的数字 ID 前添加 “Book ID: ”。
$ sed 's/^[0-9]/Book ID: &/' books.txt这里,^[0-9] 匹配行开头的第一个数字。[[:digit:]] 是一个更具可移植性的 POSIX 字符类,表示一个数字,所以 s/^[[:digit:]]/Book ID: &/' 也能工作。
输出:
Book ID: 1) A Storm of Swords, George R. R. Martin, 1216Book ID: 2) The Two Towers, J. R. R. Tolkien, 352Book ID: 3) The Alchemist, Paulo Coelho, 197Book ID: 4) The Fellowship of the Ring, J. R. R. Tolkien, 432Book ID: 5) The Pilgrimage, Paulo Coelho, 288Book ID: 6) A Game of Thrones, George R. R. Martin, 864解释:模式 ^[0-9](或 ^[[:digit:]])匹配开头的数字(例如 ‘1’, ‘2’)。替换字符串中的 & 代表这个匹配到的数字。所以 ‘1’ 变成了 ‘Book ID: 1’。
示例 2:在页码周围添加上下文。
假设 books.txt 中每行的最后一个数字代表页数。我们希望将其格式化为 “Pages: [数字]”。
$ sed 's/[0-9]*$/Pages: &/' books.txt或者,使用 POSIX 字符类:sed 's/[[:digit:]]*$/Pages: &/' books.txt
输出:
1) A Storm of Swords, George R. R. Martin, Pages: 12162) The Two Towers, J. R. Tolkien, Pages: 3523) The Alchemist, Paulo Coelho, Pages: 1974) The Fellowship of the Ring, J. R. R. Tolkien, Pages: 4325) The Pilgrimage, Paulo Coelho, Pages: 2886) A Game of Thrones, George R. R. Martin, Pages: 864解释:模式 [0-9]*$(或 [[:digit:]]*$)匹配行尾的零个或多个数字序列(页码)。然后 & 将这个匹配到的页码插入到替换字符串中。
& 字符对于需要包装匹配文本(例如用 HTML 标签)、在特定单词周围添加引号,或在保留原始匹配内容的同时重新格式化行的部分内容等任务非常宝贵。例如,将所有出现的 ‘Warning:’ 转换为 ‘Warning:‘。
- GNU sed manual: 关于 sed 命令和特殊字符的详细信息。
- Regular-Expressions.info: 一个全面的资源,用于理解正则表达式,这对于有效使用 sed 至关重要。