Skip to content

Sed 模式缓冲区

SED 中的一个核心概念是 模式缓冲区(pattern buffer)。当 SED 处理输入时,它会逐行读取内容(除非被 N 等命令修改)到这个临时存储区域。大多数 SED 命令都作用于模式缓冲区中的文本。在脚本中的所有命令都应用于模式缓冲区中的行之后,SED 通常会将缓冲区的内容打印到标准输出,然后读取下一行,重复这个循环。

让我们创建一个名为 books.txt 的示例文件,后续示例将使用该文件。请确保您的文件内容与此一致,以便获得一致的结果。

文件: `books.txt`
1) A Storm of Swords, George R. R. Martin, 1216
2) The Two Towers, J. R. R. Tolkien, 352
3) The Alchemist, Paulo Coelho, 197
4) The Fellowship of the Ring, J. R. R. Tolkien, 432
5) The Pilgrimage, Paulo Coelho, 288
6) A Game of Thrones, George R. R. Martin, 864

该 p 命令会显式地打印模式缓冲区的当前内容。

[jerry]$ sed 'p' books.txt

该命令产生以下结果:

1) A Storm of Swords, George R. R. Martin, 1216
1) A Storm of Swords, George R. R. Martin, 1216
2) The Two Towers, J. R. R. Tolkien, 352
2) The Two Towers, J. R. R. Tolkien, 352
3) The Alchemist, Paulo Coelho, 197
3) The Alchemist, Paulo Coelho, 197
4) The Fellowship of the Ring, J. R. R. Tolkien, 432
4) The Fellowship of the Ring, J. R. R. Tolkien, 432
5) The Pilgrimage, Paulo Coelho, 288
5) The Pilgrimage, Paulo Coelho, 288
6) A Game of Thrones, George R. R. Martin, 864
6) A Game of Thrones, George R. R. Martin, 864

每行出现两次是因为:

  1. 默认情况下,SED 在执行完针对一行的所有命令后,会打印模式缓冲区的内容。
  2. 我们显式使用了 p 命令,该命令也会打印模式缓冲区的内容。 要抑制默认的打印行为,可以使用 -n(no-print,不打印)选项。
[jerry]$ sed -n 'p' books.txt

现在,每行只打印一次,这是完全由于 p 命令的作用:

1) A Storm of Swords, George R. R. Martin, 1216
2) The Two Towers, J. R. R. Tolkien, 352
3) The Alchemist, Paulo Coelho, 197
4) The Fellowship of the Ring, J. R. R. Tolkien, 432
5) The Pilgrimage, Paulo Coelho, 288
6) A Game of Thrones, George R. R. Martin, 864

可以使用地址来限制 SED 命令只作用于特定的行。地址可以是行号、正则表达式或范围。

打印指定行号的行(例如,第 3 行):

[jerry]$ sed -n '3p' books.txt

输出:

3) The Alchemist, Paulo Coelho, 197

打印一个行的范围(例如,第 2 行到第 5 行):

[jerry]$ sed -n '2,5p' books.txt

输出:

2) The Two Towers, J. R. R. Tolkien, 352
3) The Alchemist, Paulo Coelho, 197
4) The Fellowship of the Ring, J. R. R. Tolkien, 432
5) The Pilgrimage, Paulo Coelho, 288

$ 字符代表输入的最后一行。

打印最后一行:

[jerry]$ sed -n '$p' books.txt

输出:

6) A Game of Thrones, George R. R. Martin, 864

从第 3 行打印到最后一行:

[jerry]$ sed -n '3,$p' books.txt

输出:

3) The Alchemist, Paulo Coelho, 197
4) The Fellowship of the Ring, J. R. R. Tolkien, 432
5) The Pilgrimage, Paulo Coelho, 288
6) A Game of Thrones, George R. R. Martin, 864

GNU SED 提供了更灵活的地址方案。

  1. M,+n:匹配第 M 行及其后的 n 行。因此,它总共作用于 n+1 行。

示例:打印第 2 行及随后的 3 行(总共 4 行:2, 3, 4, 5)。

[jerry]$ sed -n '2,+3p' books.txt

输出:

2) The Two Towers, J. R. R. Tolkien, 352
3) The Alchemist, Paulo Coelho, 197
4) The Fellowship of the Ring, J. R. R. Tolkien, 432
5) The Pilgrimage, Paulo Coelho, 288
  1. M~n:匹配第 M 行,然后是第 M 行之后的每隔 n 行。

示例:只打印奇数行(从第 1 行开始,步长为 2)。

[jerry]$ sed -n '1~2p' books.txt

输出:

1) A Storm of Swords, George R. R. Martin, 1216
3) The Alchemist, Paulo Coelho, 197
5) The Pilgrimage, Paulo Coelho, 288

示例:只打印偶数行(从第 2 行开始,步长为 2)。

[jerry]$ sed -n '2~2p' books.txt

输出:

2) The Two Towers, J. R. R. Tolkien, 352
4) The Fellowship of the Ring, J. R. R. Tolkien, 432
6) A Game of Thrones, George R. R. Martin, 864

理解模式缓冲区和地址对于有效地使用 SED 进行文本处理至关重要。请记住,默认情况下,这些操作不会改变原始文件;它们只影响输出流。要在原地(in-place)修改文件,请使用 -i 选项(例如,sed -i 'command' file'),但使用时要小心,最好先备份文件。