Skip to content

Sed 快速指南

Sed,Stream EDitor 的缩写,是一个强大而灵活的命令行工具,用于解析和转换文本。它由 Lee E. McMahon 于 1973-74 年在贝尔实验室开发,如今已成为几乎所有类 Unix 操作系统(包括 GNU/Linux 和 macOS)上的必备工具。

McMahon 将 Sed 设计为一个通用的、面向行的编辑器,其灵感和特性源自早期的 ed 编辑器。Sed 从诞生之日起就具备的一项关键优势是它支持正则表达式,这使得复杂的模式匹配和文本操作成为可能。Sed 处理来自文件、管道或标准输入流的输入,这使其在各种文本处理任务中都非常灵活。

GNU/Linux 系统上常见的版本是 GNU Sed,它是 GNU 项目的一部分并由其维护。虽然 Sed 的语法对新手来说可能显得晦涩难懂,但掌握它就能用简洁的脚本解决复杂的文本操作问题。这种效率是 Sed 设计的一个显著特点。

Sed 通常用于广泛的任务,包括:

  • 文本替换(查找和替换)。
  • 选择性地打印文件中的行。
  • 原地编辑文件(直接修改文件)。
  • 非交互式批量编辑文本。
  • 在管道中过滤和转换文本。
  • 数据清理和重新格式化(例如,CSV 或日志文件)。

本节介绍如何确保 Sed 可用,以及必要时如何在 GNU/Linux 系统上安装或更新它。

Sed 通常预装在大多数 GNU/Linux 发行版和 macOS 上。您可以通过运行以下命令来检查它是否可用并查看其版本:

[user]$ sed --version

输出可能类似于(版本号和详细信息会有所不同):

sed (GNU sed) 4.8
Copyright (C) 2020 Free Software Foundation, Inc.
License GPLv3+: GNU GPL version 3 or later <https://gnu.org/licenses/gpl.html>.
This is free software: you are free to change and redistribute it.
There is NO WARRANTY, to the extent permitted by law.
Written by Jay Fenlason, Tom Lord, Ken Pizzini,
and Paolo Bonzini.
GNU sed home page: <https://www.gnu.org/software/sed/>.
General help using GNU software: <https://www.gnu.org/gethelp/>.
E-mail bug reports to: <bug-sed@gnu.org>.

如果找不到该命令,或者您需要特定版本(例如,在默认使用 BSD Sed 的 macOS 上安装 GNU Sed),您可能需要安装它。

在基于 Debian 的系统上(如 Ubuntu, Mint):

[user]$ sudo apt update
[user]$ sudo apt install sed

在基于 RPM 的系统上(如 Fedora, CentOS, RHEL):较旧的系统使用 yum,较新的系统使用 dnf。

# For DNF-based systems (e.g., Fedora, RHEL 8+)
[root]# dnf install sed
# For YUM-based systems (e.g., CentOS 7)
[root]# yum install sed

在 macOS 上(获取 GNU Sed,因为默认是 BSD Sed):推荐使用 Homebrew。

[user]$ brew install gnu-sed

通过 Homebrew 安装后,GNU Sed 通常以 gsed 的名称提供。您可能需要将其添加到 PATH 中,或者明确使用 gsed 命令。

对于 Sed 这样的工具,从源代码安装很少是必要的,但对于高级用户或有特定需求的用户来说是一个选择。始终从官方 GNU 镜像下载源代码。

一般步骤:

  1. 从 GNU FTP 服务器(例如,https://ftp.gnu.org/gnu/sed/)下载最新的源代码 tar 包(例如,sed-latest.tar.gz 或 sed-latest.tar.xz)。 [user]$ wget https://ftp.gnu.org/gnu/sed/sed-latest.tar.xz (替换为实际的最新版本链接)
  2. 解压存档文件: [user]$ tar xvf sed-latest.tar.xz
  3. 进入解压后的目录: [user]$ cd sed-*/ (目录名称将与版本匹配)
  4. 配置构建环境: [user]$ ./configure
  5. 编译源代码: [user]$ make
  6. 运行测试(可选但推荐): [user]$ make check
  7. 安装编译好的工具(需要超级用户权限): [user]$ sudo make install

完成这些步骤后,使用 sed --version 验证安装。

理解 Sed 的内部工作流程是有效使用它的关键。Sed 按照循环方式处理每一行输入。涉及的核心组件是模式空间(pattern space)和保持空间(hold space)。

每行的基本工作流程是:

  1. 读取(Read):Sed 从输入流(文件、管道或标准输入)中读取一行。末尾的换行符在内部被移除,然后将其放入模式空间(pattern space),这是一个用于执行编辑操作的内部缓冲区。
  2. 执行(Execute):脚本中指定的所有 Sed 命令按顺序应用于模式空间中的文本。可以使用地址将命令限制在特定行。
  3. 显示(Display):在对模式空间中的当前行执行完所有命令后,Sed 默认会将模式空间的内容打印到输出流,并在其后加上一个换行符。可以使用 -n 选项来抑制这种默认打印。
  4. 重复(Repeat):然后 Sed 清空模式空间(除非特定的命令如 N 修改此行为),并对下一行输入重复此循环,直到输入结束。
  • 模式空间(Pattern Space):一个临时的内存缓冲区,用于存放当前输入行并由 Sed 命令进行操作。对于每一行新的输入,其内容通常会被清空并重新填充。
  • 保持空间(Hold Space):另一个临时的内存缓冲区。与模式空间不同,保持空间在循环之间不会自动清空。数据可以从模式空间复制到保持空间(例如,使用 h 或 H 命令),也可以从保持空间复制回模式空间(例如,使用 g 或 G 命令)。这使得更复杂的多行操作和跨行数据保存成为可能。Sed 命令不能直接操作保持空间;数据必须先移动到模式空间。
  • 输入文件不变:默认情况下,Sed 不会修改原始输入文件。它将输出发送到标准输出。要原地修改文件,GNU Sed 提供了 -i 选项。
  • 全局应用:如果未指定命令的行地址,则该命令将应用于输入中的所有行。
  • 标准输入 (stdin):如果没有提供输入文件,Sed 将从 stdin 读取。

让我们创建一个文件 quote.txt: `There is only one thing that makes a dream impossible to achieve: the fear of failure.

  • Paulo Coelho, The Alchemist`

命令:sed 's/dream/goal/' quote.txt

  • 循环 1:
    1. 读取:Sed 将 “There is only one thing that makes a dream impossible to achieve: the fear of failure.” 读取到模式空间中。
    1. 执行:应用命令 s/dream/goal/。找到 “dream” 并替换为 “goal”。模式空间现在变为:“There is only one thing that makes a goal impossible to achieve: the fear of failure.”
    1. 显示:Sed 将模式空间的内容打印到标准输出。
  • 循环 2:
    1. 读取:Sed 将 ”- Paulo Coelho, The Alchemist” 读取到模式空间中。
    1. 执行:应用命令 s/dream/goal/。未找到 “dream”。模式空间保持不变。
    1. 显示:Sed 将模式空间的内容打印到标准输出。

命令的输出:

There is only one thing that makes a goal impossible to achieve: the fear of failure.
- Paulo Coelho, The Alchemist

Sed - 基本语法(快速指南重申)

Section titled “Sed - 基本语法(快速指南重申)”

如前所述,Sed 的基本调用形式是:

sed [options] 'command(s)' [file(s)]
sed [options] -f scriptfile [file(s)]

关键选项包括 -n(抑制默认输出)、-e(指定脚本)和 -f(指定脚本文件)。GNU Sed 提供 -r(或 -E)用于扩展正则表达式和 -i 用于原地编辑。

使用 books.txt 的示例(请参阅“基本语法和选项”章节获取 books.txt 内容):删除第 1、2 和 5 行。

[user]$ sed -e '1d' -e '2d' -e '5d' books.txt

输出:

3) The Alchemist, Paulo Coelho, 197
4) The Fellowship of the Ring, J. R. R. Tolkien, 432
6) A Game of Thrones, George R. R. Martin, 864

Sed 提供了用于基本循环和条件分支的命令,允许实现更复杂的脚本逻辑,类似于其他语言中的 goto 语句。

标签是 Sed 脚本中的一个标记,您可以跳转到该标记处。使用冒号后跟标签名来定义标签:

:mylabel
:start_loop
:process_data

b 命令会分支(跳转)到指定的 label。如果未给出 label,它会分支到脚本的末尾,有效地以新的输入行开始下一个循环。

语法:[address]b [label]

示例:处理 books.txt(假设标题和作者像本节“管理模式”章节示例中那样交替出现)。合并标题和作者,然后对包含 “Paulo” 的行在其开头添加 ”- ”。

本节的示例文件 books_alternate.txt: A Storm of Swords George R. R. Martin The Two Towers J. R. R. Tolkien The Alchemist Paulo Coelho The Fellowship of the Ring J. R. R. Tolkien The Pilgrimage Paulo Coelho A Game of Thrones George R. R. Martin

[user]$ sed -n 'N; s/\n/, /; /Paulo/!b Print; s/^/- /; :Print; p' books_alternate.txt

解释:

  1. N:将下一行附加到模式空间,用换行符分隔。
  2. s/\n/, /:将嵌入的换行符替换为 ”, “。模式空间现在为:“Title, Author”。
  3. /Paulo/!b Print:如果行不包含(!)“Paulo”,则分支(b)到标签 Print。
  4. s/^/- /:如果未执行分支(即,行包含 “Paulo”),则在开头添加 ”- ”。
  5. :Print:定义标签 Print。
  6. p:打印模式空间。

输出:

A Storm of Swords, George R. R. Martin
The Two Towers, J. R. R. Tolkien
- The Alchemist, Paulo Coelho
The Fellowship of the Ring, J. R. R. Tolkien
- The Pilgrimage, Paulo Coelho
A Game of Thrones, George R. R. Martin

命令可以写在脚本文件的不同行以提高可读性,或者用分号分隔(尽管 N、标签和分支通常在新行上效果最佳)。

t 命令提供条件分支:它仅在上次读取输入行或上次执行 t 命令以来,有 s///(替换)命令成功执行了替换操作时,才会跳转到指定的 label。

语法:[address]t [label]

示例:使用 books_alternate.txt(标题/作者交替出现在不同行),对提及 “Paulo” 的行,在其开头最多添加四个连字符。

[user]$ sed -n 'N; s/\n/, /; :Loop; /Paulo/s/^/-/;/----/!t Loop; p' books_alternate.txt

解释:

  1. N; s/\n/, /:像之前一样合并标题和作者。
  2. :Loop:定义标签 Loop。
  3. /Paulo/s/^/-/:如果行包含 “Paulo”,则在开头添加一个连字符。这是 t 命令将测试的 s/// 命令。
  4. /----/!t Loop:如果行还未以四个连字符开头,并且之前的 s/^/-/ 命令成功执行了替换,则分支(t)回到 Loop 标签。
  5. p:打印模式空间。

输出:

A Storm of Swords, George R. R. Martin
The Two Towers, J. R. R. Tolkien
----The Alchemist, Paulo Coelho
The Fellowship of the Ring, J. R. R. Tolkien
----The Pilgrimage, Paulo Coelho
A Game of Thrones, George R. R. Martin

t 命令对于创建基于替换操作成功与否来终止的循环至关重要。

模式空间是 Sed 对当前行执行操作的地方。保持空间是一个辅助缓冲区,用于在行处理周期之间存储数据。理解如何在这些空间之间移动数据是编写高级 Sed 脚本的关键。

除非另有说明,否则以下示例使用 books.txt 文件(原始格式:ID) Title, Author, Pages)。

p:打印模式空间的全部当前内容。 P(大写):打印模式空间中直到第一个嵌入换行符(\n)的部分。与 N 等多行命令结合使用时非常有用。

示例:只打印第 3 行(使用 -n 抑制默认输出)。

[user]$ sed -n '3p' books.txt

输出:

3) The Alchemist, Paulo Coelho, 197

命令可以限定在特定的行或范围:

  • N:行号 N(例如,3p 打印第 3 行)。
  • N,M:从 N 行到 M 行(例如,2,5p 打印第 2 到 5 行)。
  • $:输入的最后一行(例如,$p 打印最后一行)。
  • N,$:从 N 行到最后一行(例如,3,$p 打印从第 3 行到末尾)。
  • M,+n (GNU 扩展):从 M 行开始,以及接下来的 n 行(例如,2,+2p 打印第 2、3、4 行)。
  • M~n (GNU 扩展):从 M 行开始,每隔 n 行(例如,1~2p 打印奇数行:1, 3, 5…;2~2p 打印偶数行:2, 4, 6…)。
  • /pattern/:匹配正则表达式 pattern 的行。

示例:打印 books.txt 中的奇数行。

[user]$ sed -n '1~2p' books.txt

输出:

1) A Storm of Swords, George R. R. Martin, 1216
3) The Alchemist, Paulo Coelho, 197
5) The Pilgrimage, Paulo Coelho, 288

命令也可以通过模式指定地址:

  • /pattern/command:对匹配 pattern 的行执行 command。
  • /pattern1/,/pattern2/command:对从第一次匹配 pattern1 的行到之后第一次匹配 pattern2 的行(包括这两行)执行 command。
  • /pattern/,Np:从匹配 pattern 的行到行号为 N 的行。

示例:打印所有 “Paulo Coelho” 的书。

[user]$ sed -n '/Paulo Coelho/p' books.txt

输出:

3) The Alchemist, Paulo Coelho, 197
5) The Pilgrimage, Paulo Coelho, 288

示例:打印从第一次出现 “Two Towers” 到 “Pilgrimage” 的行。

[user]$ sed -n '/Two Towers/,/Pilgrimage/p' books.txt

输出:

2) The Two Towers, J. R. R. Tolkien, 352
3) The Alchemist, Paulo Coelho, 197
4) The Fellowship of the Ring, J. R. R. Tolkien, 432
5) The Pilgrimage, Paulo Coelho, 288

以下示例使用 books_alternate.txt 文件(标题在一行,作者在下一行):

  • h:将模式空间复制到保持空间(覆盖保持空间)。
  • H:将模式空间附加到保持空间(添加一个换行符,然后添加模式空间内容)。
  • g:将保持空间复制到模式空间(覆盖模式空间)。
  • G:将保持空间附加到模式空间(添加一个换行符,然后添加保持空间内容)。
  • x:交换模式空间和保持空间的内容。
  • n:打印当前模式空间(如果未被 -n 抑制),然后将下一行读入模式空间。
  • N:将下一行附加到模式空间(用 \n 分隔)。

示例:只打印 books_alternate.txt 中的作者姓名。

[user]$ sed -n 'h; n; p; x; d' books_alternate.txt # More direct: sed -n 'n;p' books_alternate.txt

一个更简单的方法来打印每隔一行(作者):

[user]$ sed -n 'n;p' books_alternate.txt

输出(作者):

George R. R. Martin
J. R. R. Tolkien
Paulo Coelho
J. R. R. Tolkien
Paulo Coelho
George R. R. Martin

示例:颠倒标题和作者的顺序(先打印作者,再打印标题)。

[user]$ sed -n 'h; n; p; g; p' books_alternate.txt

解释:

  1. h:将当前行(标题)复制到保持空间。
  2. n:将下一行(作者)读入模式空间。
  3. p:打印模式空间(作者)。
  4. g:将保持空间(标题)复制回模式空间。
  5. p:打印模式空间(标题)。

输出:

George R. R. Martin
A Storm of Swords
J. R. R. Tolkien
The Two Towers
...

本节介绍 Sed 的核心文本操作命令,除非另有说明,否则使用 books.txt 文件(原始格式:ID) Title, Author, Pages)。

[address1[,address2]]d 删除模式空间的内容,因此该行不会被打印。立即开始一个新的循环。

示例:删除第 4 行。

[user]$ sed '4d' books.txt

输出(缺少第 4 行):

1) A Storm of Swords, George R. R. Martin, 1216
2) The Two Towers, J. R. R. Tolkien, 352
3) The Alchemist, Paulo Coelho, 197
5) The Pilgrimage, Paulo Coelho, 288
6) A Game of Thrones, George R. R. Martin, 864

[address1[,address2]]w filename 将模式空间的内容写入 filename 文件。如果文件不存在则创建,如果存在则覆盖(在 Sed 脚本开始时截断)。

示例:将包含 “Tolkien” 的行写入 tolkien_books.txt 文件。

[user]$ sed -n '/Tolkien/w tolkien_books.txt' books.txt # Use -n if only writing, not printing
[user]$ cat tolkien_books.txt

cat tolkien_books.txt 的输出:

2) The Two Towers, J. R. R. Tolkien, 352
4) The Fellowship of the Ring, J. R. R. Tolkien, 432

[address]a\ 要附加的文本 将“要附加的文本”附加到当前行的后面。如果文本与 a 在同一行,则必须在其前面加上反斜杠 \,或者从下一行开始。

示例:在第 2 行之后附加文本。

[user]$ sed '2a\
---> Inserted Line <---' books.txt

输出:

1) A Storm of Swords, George R. R. Martin, 1216
2) The Two Towers, J. R. R. Tolkien, 352
---> Inserted Line <---
3) The Alchemist, Paulo Coelho, 197
...

[address]i\ 要插入的文本 将“要插入的文本”插入到当前行的前面。文本的语法类似于 a 命令。

示例:在第 1 行之前插入文本。

[user]$ sed '1i\
*** START OF FILE ***' books.txt

输出:

*** START OF FILE ***
1) A Storm of Swords, George R. R. Martin, 1216
2) The Two Towers, J. R. R. Tolkien, 352
...

[address1[,address2]]c\ 新文本 将匹配地址范围的行替换为“新文本”。如果给定了范围,则该范围内的所有行都将被替换为单个实例的“新文本”。

示例:更改第 3 行。

[user]$ sed '3c\
REPLACED: This was line 3.' books.txt

输出:

1) A Storm of Swords, George R. R. Martin, 1216
2) The Two Towers, J. R. R. Tolkien, 352
REPLACED: This was line 3.
4) The Fellowship of the Ring, J. R. R. Tolkien, 432
...

[address1[,address2]]y/源字符集/目标字符集/ 执行字符转换。源字符集 中的每个字符都将被 目标字符集 中对应位置的字符替换。源字符集 和 目标字符集 的长度必须相同。此处不使用正则表达式。

示例:对包含 “Alchemist” 的行,将小写 ‘aeio’ 转换为大写 ‘AEIO’。

[user]$ sed '/Alchemist/y/aeio/AEIO/' books.txt

输出(第 3 行被修改):

1) A Storm of Swords, George R. R. Martin, 1216
2) The Two Towers, J. R. R. Tolkien, 352
3) ThE AlchEmIst, PAUlO COElhO, 197
4) The Fellowship of the Ring, J. R. R. Tolkien, 432
...

[address1[,address2]]l [长度](长度是 GNU 扩展) 以明确无歧义的形式打印模式空间:非打印字符显示为转义序列(例如,制表符显示为 \t),长行会被换行。GNU Sed 允许指定换行 长度。

示例:以 30 的换行长度显示 books.txt。

[user]$ sed -n 'l 30' books.txt

输出(第一行示例):

1) A Storm of Swords, George \
R. R. Martin, 1216$

[地址]q [退出码](退出码是 GNU 扩展) Sed 会立即停止处理并退出。如果给定了 地址,则当到达该行时就会退出。GNU Sed 允许指定 退出码。

示例:打印前 3 行(类似于 head -3)。

[user]$ sed '3q' books.txt

输出:

1) A Storm of Swords, George R. R. Martin, 1216
2) The Two Towers, J. R. R. Tolkien, 352
3) The Alchemist, Paulo Coelho, 197

[地址]r 文件名 读取 filename 的内容,并在处理匹配 address 的行之后(如果未指定地址,则在脚本结束时)将其附加到输出中。如果 address 为 0(GNU Sed),则在开始处读取。

[user]$ echo "--- Additional Content ---" > extra.txt
[user]$ sed '3r extra.txt' books.txt

输出(extra.txt 的内容插入到第 3 行之后):

1) A Storm of Swords, George R. R. Martin, 1216
2) The Two Towers, J. R. R. Tolkien, 352
3) The Alchemist, Paulo Coelho, 197
--- Additional Content ---
4) The Fellowship of the Ring, J. R. R. Tolkien, 432
...

[地址1[,地址2]]e [命令] 执行 命令。命令 的输出被发送到标准输出。如果省略了 命令,则将模式空间的内容作为命令执行。这是一个强大但潜在危险的 GNU Sed 扩展。

示例:在第 1 行之后运行 date 命令。

[user]$ sed '1e date' books.txt

输出(date 命令的输出出现在第 1 行之后,原始第 1 行也会被打印):

1) A Storm of Swords, George R. R. Martin, 1216
Wed Mar 27 10:30:00 EDT 2024 (Actual date will vary)
2) The Two Towers, J. R. R. Tolkien, 352
...

Sed 将几个特殊字符视为命令本身,提供了独特的功能。

[地址1[,地址2]]= 将当前输入行号打印到标准输出,后跟一个换行符。

示例:在每行之前打印行号。

[user]$ sed '=' books.txt

输出:

1
1) A Storm of Swords, George R. R. Martin, 1216
2
2) The Two Towers, J. R. R. Tolkien, 352
...

示例:计算文件中的行数(类似于 wc -l)。

[user]$ sed -n '$=' books.txt

输出(总行数):

6

如果 Sed 命令的第一个字符是 #,则整行被视为注释并被 Sed 忽略(除非它是 #n,如果出现在脚本的第一行,它等同于 -n 选项)。注释对于给 Sed 脚本添加说明非常有用。

在一个脚本文件 comment_example.sed 中的示例:

# This is a comment, it will be ignored.
# Print lines containing 'Tolkien'
/Tolkien/p
[user]$ sed -n -f comment_example.sed books.txt

输出(仅显示 Tolkien 的书):

2) The Two Towers, J. R. R. Tolkien, 352
4) The Fellowship of the Ring, J. R. R. Tolkien, 432

虽然本身不是一个命令,但与号 & 在替换命令 (s/pattern/replacement/) 的替换部分中具有特殊含义。它代表了由 pattern 匹配到的整个文本。

示例:将所有 “George R. R. Martin” 的出现用星号括起来。

[user]$ sed 's/George R. R. Martin/**&**/g' books.txt

输出:

1) A Storm of Swords, **George R. R. Martin**, 1216
2) The Two Towers, J. R. R. Tolkien, 352
...
6) A Game of Thrones, **George R. R. Martin**, 864

Sed - 字符串操作(快速指南重申)

Section titled “Sed - 字符串操作(快速指南重申)”

s(替换)命令是 Sed 进行字符串操作的主要工具。语法:[地址]s/模式/替换文本/[标志]。

常用标志:g(全局)、N(第 N 次出现)、p(打印)、w 文件(写入文件)、i(忽略大小写 - GNU 扩展)。

示例:将 books.txt 中的所有逗号替换为分号。

[user]$ sed 's/,/;/g' books.txt

输出:

1) A Storm of Swords; George R. R. Martin; 1216
2) The Two Towers; J. R. R. Tolkien; 352
...

捕获组(在 BRE 中使用 \(...\),在 ERE 中使用 (...))允许在替换文本中使用 \1、\2 等来重新排序和重用匹配文本的部分。

示例:交换作者和标题(假设格式为 ‘Title, Author’)。

[user]$ echo "The Hobbit, J. R. R. Tolkien" | sed -E 's/([^,]+), *(.+)/\2 - \1/'

输出:

J. R. R. Tolkien - The Hobbit

GNU Sed 在替换文本中提供了 \L、\U、\l、\u、\E 用于大小写转换。

Sed - 正则表达式(快速指南重申)

Section titled “Sed - 正则表达式(快速指南重申)”

Sed 的强大之处源于正则表达式。默认情况下,它使用基本正则表达式 (BRE)。可以通过 -E (POSIX) 或 -r (GNU) 启用扩展正则表达式 (ERE)。

关键元字符(BRE 示例,ERE 通常更简单):

  • ^:行首。
  • $:行尾。
  • .:任意单个字符(换行符除外)。
  • [...]:字符集(例如,[abc])。
  • [^...]:非字符集(排除字符集中的字符)。
  • *:前一个项的零个或多个。
  • \?:零个或一个(BRE:\?,ERE:?)。
  • \+:一个或多个(BRE:\+,ERE:+)。
  • \{N\}:恰好 N 次(BRE:\{N\},ERE:{N})。
  • \{N,\}:至少 N 次。
  • \{N,M\}:介于 N 和 M 次之间。
  • \(...\):用于捕获的分组(BRE:\(...\),ERE:(...))。
  • \|:交替或 (OR)(BRE:在 \(\) 内部使用 \|,ERE:|)。

POSIX 字符类:[[:alnum:]]、[[:alpha:]]、[[:digit:]]、[[:space:]] 等。提供了匹配字符类型的可移植方式。

示例:打印以数字开头的行。

[user]$ sed -n '/^[[:digit:]]/p' books.txt

输出(books.txt 中的所有行,因为它们都以数字开头):

1) A Storm of Swords, George R. R. Martin, 1216
2) The Two Towers, J. R. R. Tolkien, 352
3) The Alchemist, Paulo Coelho, 197
4) The Fellowship of the Ring, J. R. R. Tolkien, 432
5) The Pilgrimage, Paulo Coelho, 288
6) A Game of Thrones, George R. R. Martin, 864

GNU Sed 元字符:\b(单词边界)、\s(空白字符)、\w(单词字符)等,提供了类似于 Perl 的便利功能。

本节通过模拟常见的 Unix 工具和执行实用任务来展示 Sed 的多功能性。这些示例展示了各种 Sed 命令和技术。

默认行为(空脚本打印行):

[user]$ sed '' books.txt

使用带 -n 的 print 命令:

[user]$ sed -n 'p' books.txt

模式 ^$ 匹配空行。

[user]$ echo -e "Line #1\n\n\nLine #2" | sed '/^$/d'

输出:

Line #1
Line #2

移除注释行(例如,C++ 的 // 注释)

Section titled “移除注释行(例如,C++ 的 // 注释)”

示例文件 hello.cpp:

#include <iostream>
// using namespace std; // A commented out line
int main(void) {
// Displays message on stdout.
std::cout << "Hello, World !!!" << std::endl; // Another comment
return 0; /* Not this type of comment */
}
[user]$ sed 's|//.*||g' hello.cpp

输出(带有 // 注释的行被移除,或者只移除注释部分):

#include <iostream>
int main(void) {
std::cout << "Hello, World !!!" << std::endl;
return 0; /* Not this type of comment */
}

注意:这个正则表达式比较基础,可能无法完美处理所有 C++ 注释场景(例如,字符串内部的 //)。对于健壮的解析,专用的工具更好。

在特定行前添加注释(例如,Shell 脚本)

Section titled “在特定行前添加注释(例如,Shell 脚本)”

示例文件 hello.sh: #!/bin/bash pwd hostname uname -a who who -r lsb_release -a

[user]$ sed '3,5s/^/# /' hello.sh

输出(第 3 到 5 行被注释掉了):

#!/bin/bash
pwd
# hostname
# uname -a
# who
who -r
lsb_release -a
[user]$ sed -n '$=' hello.sh

输出(hello.sh 中的行数):

7

示例:head -3 books_alternate.txt

[user]$ sed '3q' books_alternate.txt

输出:

A Storm of Swords
George R. R. Martin
The Two Towers
[user]$ sed -n '$p' books_alternate.txt

输出:

George R. R. Martin

更健壮的 tail -N 需要更复杂的 Sed 脚本(通常涉及保持空间和循环),或者可以使用 tac file | sed Nq | tac 命令组合。

模拟 dos2unix(将 CRLF 转换为 LF)

Section titled “模拟 dos2unix(将 CRLF 转换为 LF)”

DOS/Windows 使用 CR+LF(\r\n)作为行尾符。Unix 使用 LF(\n)。^M 通常代表 CR。

[user]$ echo -e "Line #1\r\nLine #2\r" > test_dos.txt
[user]$ file test_dos.txt

file 的输出:

test_dos.txt: ASCII text, with CRLF line terminators
[user]$ sed 's/\r$//' test_dos.txt > test_unix.txt # More portable
[user]$ file test_unix.txt

file 的输出:

test_unix.txt: ASCII text

s/\r$// 命令会移除行尾($)的 Carriage Return(\r)。

模拟 unix2dos(将 LF 转换为 CRLF)

Section titled “模拟 unix2dos(将 LF 转换为 CRLF)”
[user]$ echo -e "Line #1\nLine #2" > test_unix2.txt
[user]$ sed 's/$/\r/' test_unix2.txt > test_dos2.txt
[user]$ file test_dos2.txt

file 的输出:

test_dos2.txt: ASCII text, with CRLF line terminators

s/$/\r/ 命令会在每一行($)的末尾添加一个 Carriage Return(\r)。

[user]$ sed 's/$/$/' books.txt # Replace end-of-line with literal $

输出(第一行示例):

1) A Storm of Swords, George R. R. Martin, 1216$
[user]$ echo -e "Column1\tColumn2\nValue1\tValue2" > tab_file.txt
[user]$ sed 's/\t/^I/g' tab_file.txt # Note: ^I is a literal tab character, type Ctrl-V then Tab

或者,如果不需要精确地显示 ^I,可以使用 l 命令(默认显示为 \t),或者通过管道传递给 tr 命令:

[user]$ sed -n 'l' tab_file.txt # Shows tabs as \t
[user]$ sed 's/\t/TEMP_TAB/g' tab_file.txt | tr 'TEMP_TAB' '\t' | cat -T # Complex way for literal ^I
[user]$ sed '=' books.txt | sed 'N; s/\n/\t/'

解释:

  • sed '=' books.txt:打印行号,然后是该行。
  • sed 'N; s/\n/\t/':合并每对行(行号和内容行),并将它们之间的换行符替换为制表符。

输出(第一行示例):

1 1) A Storm of Swords, George R. R. Martin, 1216
[user]$ sed -n 'w books_copy.txt' books.txt
[user]$ diff books.txt books_copy.txt && echo "Files are identical"

输出(如果成功):

Files are identical

模拟 expand(将 Tab 转换为空格)

Section titled “模拟 expand(将 Tab 转换为空格)”

此示例将每个制表符转换为 4 个空格。真正的 expand 命令具有更复杂的制表位逻辑。

[user]$ sed 's/\t/ /g' tab_file.txt

输出:

Column1 Column2
Value1 Value2
[user]$ echo -e "Data line 1\nData line 2" | sed 'p; w output.txt'

解释:p 打印到标准输出。w output.txt 写入文件。由于默认打印是开启的,p 会导致内容被打印到标准输出两次。使用 -n 以获得与 tee 完全一致的行为。

[user]$ echo -e "Data line 1\nData line 2" | sed -n 'p; w output.txt'
[user]$ cat output.txt

输出到屏幕和 output.txt 文件中的内容:

Data line 1
Data line 2

将连续的多个空行替换为一个空行。

[user]$ echo -e "Line 1\n\n\nLine 2\n\nLine 3" > blank_lines.txt
[user]$ sed '/^$/{N;/^\n$/D;}' blank_lines.txt

/^$/{N;/^\n$/D;} 的解释:

  • /^$/:如果当前行为空行…
  • N:附加下一行。模式空间可能变为 \n(如果下一行也是空行)。
  • /^\n$/D:如果模式空间仅为 \n(即连续两个空行),则删除直到 \n 的部分,并使用剩余内容重新开始循环。这实际上删除了其中一个空行,并重新评估了新的模式空间。

输出:

Line 1
Line 2
Line 3

模拟 grep pattern(打印匹配的行)

Section titled “模拟 grep pattern(打印匹配的行)”
[user]$ sed -n '/Alchemist/p' books.txt

输出:

3) The Alchemist, Paulo Coelho, 197

模拟 grep -v pattern(打印不匹配的行)

Section titled “模拟 grep -v pattern(打印不匹配的行)”
[user]$ sed -n '/Alchemist/!p' books.txt

输出(除包含 “Alchemist” 的行之外的所有行):

1) A Storm of Swords, George R. R. Martin, 1216
2) The Two Towers, J. R. R. Tolkien, 352
4) The Fellowship of the Ring, J. R. R. Tolkien, 432
5) The Pilgrimage, Paulo Coelho, 288
6) A Game of Thrones, George R. R. Martin, 864
[user]$ echo "ABCDE" | sed 'y/ACE/XZE/'

输出:

XBDZE