Sed 快速指南
Sed - 流编辑器快速指南
Section titled “Sed - 流编辑器快速指南”Sed - 概述
Section titled “Sed - 概述”Sed,Stream EDitor 的缩写,是一个强大而灵活的命令行工具,用于解析和转换文本。它由 Lee E. McMahon 于 1973-74 年在贝尔实验室开发,如今已成为几乎所有类 Unix 操作系统(包括 GNU/Linux 和 macOS)上的必备工具。
McMahon 将 Sed 设计为一个通用的、面向行的编辑器,其灵感和特性源自早期的 ed 编辑器。Sed 从诞生之日起就具备的一项关键优势是它支持正则表达式,这使得复杂的模式匹配和文本操作成为可能。Sed 处理来自文件、管道或标准输入流的输入,这使其在各种文本处理任务中都非常灵活。
GNU/Linux 系统上常见的版本是 GNU Sed,它是 GNU 项目的一部分并由其维护。虽然 Sed 的语法对新手来说可能显得晦涩难懂,但掌握它就能用简洁的脚本解决复杂的文本操作问题。这种效率是 Sed 设计的一个显著特点。
Sed 的典型用途
Section titled “Sed 的典型用途”Sed 通常用于广泛的任务,包括:
- 文本替换(查找和替换)。
- 选择性地打印文件中的行。
- 原地编辑文件(直接修改文件)。
- 非交互式批量编辑文本。
- 在管道中过滤和转换文本。
- 数据清理和重新格式化(例如,CSV 或日志文件)。
Sed - 环境设置
Section titled “Sed - 环境设置”本节介绍如何确保 Sed 可用,以及必要时如何在 GNU/Linux 系统上安装或更新它。
检查 Sed 是否可用
Section titled “检查 Sed 是否可用”Sed 通常预装在大多数 GNU/Linux 发行版和 macOS 上。您可以通过运行以下命令来检查它是否可用并查看其版本:
[user]$ sed --version输出可能类似于(版本号和详细信息会有所不同):
sed (GNU sed) 4.8Copyright (C) 2020 Free Software Foundation, Inc.License GPLv3+: GNU GPL version 3 or later <https://gnu.org/licenses/gpl.html>.This is free software: you are free to change and redistribute it.There is NO WARRANTY, to the extent permitted by law.
Written by Jay Fenlason, Tom Lord, Ken Pizzini,and Paolo Bonzini.GNU sed home page: <https://www.gnu.org/software/sed/>.General help using GNU software: <https://www.gnu.org/gethelp/>.E-mail bug reports to: <bug-sed@gnu.org>.如果找不到该命令,或者您需要特定版本(例如,在默认使用 BSD Sed 的 macOS 上安装 GNU Sed),您可能需要安装它。
使用包管理器安装
Section titled “使用包管理器安装”在基于 Debian 的系统上(如 Ubuntu, Mint):
[user]$ sudo apt update[user]$ sudo apt install sed在基于 RPM 的系统上(如 Fedora, CentOS, RHEL):较旧的系统使用 yum,较新的系统使用 dnf。
# For DNF-based systems (e.g., Fedora, RHEL 8+)[root]# dnf install sed
# For YUM-based systems (e.g., CentOS 7)[root]# yum install sed在 macOS 上(获取 GNU Sed,因为默认是 BSD Sed):推荐使用 Homebrew。
[user]$ brew install gnu-sed通过 Homebrew 安装后,GNU Sed 通常以 gsed 的名称提供。您可能需要将其添加到 PATH 中,或者明确使用 gsed 命令。
从源代码安装(高级)
Section titled “从源代码安装(高级)”对于 Sed 这样的工具,从源代码安装很少是必要的,但对于高级用户或有特定需求的用户来说是一个选择。始终从官方 GNU 镜像下载源代码。
一般步骤:
- 从 GNU FTP 服务器(例如,
https://ftp.gnu.org/gnu/sed/)下载最新的源代码 tar 包(例如,sed-latest.tar.gz或sed-latest.tar.xz)。[user]$ wget https://ftp.gnu.org/gnu/sed/sed-latest.tar.xz(替换为实际的最新版本链接) - 解压存档文件:
[user]$ tar xvf sed-latest.tar.xz - 进入解压后的目录:
[user]$ cd sed-*/(目录名称将与版本匹配) - 配置构建环境:
[user]$ ./configure - 编译源代码:
[user]$ make - 运行测试(可选但推荐):
[user]$ make check - 安装编译好的工具(需要超级用户权限):
[user]$ sudo make install
完成这些步骤后,使用 sed --version 验证安装。
Sed - 工作流程
Section titled “Sed - 工作流程”理解 Sed 的内部工作流程是有效使用它的关键。Sed 按照循环方式处理每一行输入。涉及的核心组件是模式空间(pattern space)和保持空间(hold space)。
每行的基本工作流程是:
- 读取(Read):Sed 从输入流(文件、管道或标准输入)中读取一行。末尾的换行符在内部被移除,然后将其放入模式空间(pattern space),这是一个用于执行编辑操作的内部缓冲区。
- 执行(Execute):脚本中指定的所有 Sed 命令按顺序应用于模式空间中的文本。可以使用地址将命令限制在特定行。
- 显示(Display):在对模式空间中的当前行执行完所有命令后,Sed 默认会将模式空间的内容打印到输出流,并在其后加上一个换行符。可以使用
-n选项来抑制这种默认打印。 - 重复(Repeat):然后 Sed 清空模式空间(除非特定的命令如
N修改此行为),并对下一行输入重复此循环,直到输入结束。
- 模式空间(Pattern Space):一个临时的内存缓冲区,用于存放当前输入行并由 Sed 命令进行操作。对于每一行新的输入,其内容通常会被清空并重新填充。
- 保持空间(Hold Space):另一个临时的内存缓冲区。与模式空间不同,保持空间在循环之间不会自动清空。数据可以从模式空间复制到保持空间(例如,使用
h或H命令),也可以从保持空间复制回模式空间(例如,使用g或G命令)。这使得更复杂的多行操作和跨行数据保存成为可能。Sed 命令不能直接操作保持空间;数据必须先移动到模式空间。 - 输入文件不变:默认情况下,Sed 不会修改原始输入文件。它将输出发送到标准输出。要原地修改文件,GNU Sed 提供了
-i选项。 - 全局应用:如果未指定命令的行地址,则该命令将应用于输入中的所有行。
- 标准输入 (stdin):如果没有提供输入文件,Sed 将从 stdin 读取。
工作流程示例演练
Section titled “工作流程示例演练”让我们创建一个文件 quote.txt:
`There is only one thing that makes a dream impossible to achieve: the fear of failure.
- Paulo Coelho, The Alchemist`
命令:sed 's/dream/goal/' quote.txt
- 循环 1:
-
- 读取:Sed 将 “There is only one thing that makes a dream impossible to achieve: the fear of failure.” 读取到模式空间中。
-
- 执行:应用命令
s/dream/goal/。找到 “dream” 并替换为 “goal”。模式空间现在变为:“There is only one thing that makes a goal impossible to achieve: the fear of failure.”
- 执行:应用命令
-
- 显示:Sed 将模式空间的内容打印到标准输出。
- 循环 2:
-
- 读取:Sed 将 ”- Paulo Coelho, The Alchemist” 读取到模式空间中。
-
- 执行:应用命令
s/dream/goal/。未找到 “dream”。模式空间保持不变。
- 执行:应用命令
-
- 显示:Sed 将模式空间的内容打印到标准输出。
命令的输出:
There is only one thing that makes a goal impossible to achieve: the fear of failure.- Paulo Coelho, The AlchemistSed - 基本语法(快速指南重申)
Section titled “Sed - 基本语法(快速指南重申)”如前所述,Sed 的基本调用形式是:
sed [options] 'command(s)' [file(s)]sed [options] -f scriptfile [file(s)]关键选项包括 -n(抑制默认输出)、-e(指定脚本)和 -f(指定脚本文件)。GNU Sed 提供 -r(或 -E)用于扩展正则表达式和 -i 用于原地编辑。
使用 books.txt 的示例(请参阅“基本语法和选项”章节获取 books.txt 内容):删除第 1、2 和 5 行。
[user]$ sed -e '1d' -e '2d' -e '5d' books.txt输出:
3) The Alchemist, Paulo Coelho, 1974) The Fellowship of the Ring, J. R. R. Tolkien, 4326) A Game of Thrones, George R. R. Martin, 864Sed - 循环与分支
Section titled “Sed - 循环与分支”Sed 提供了用于基本循环和条件分支的命令,允许实现更复杂的脚本逻辑,类似于其他语言中的 goto 语句。
标签(Labels)
Section titled “标签(Labels)”标签是 Sed 脚本中的一个标记,您可以跳转到该标记处。使用冒号后跟标签名来定义标签:
:mylabel:start_loop:process_data分支(b 命令)
Section titled “分支(b 命令)”b 命令会分支(跳转)到指定的 label。如果未给出 label,它会分支到脚本的末尾,有效地以新的输入行开始下一个循环。
语法:[address]b [label]
示例:处理 books.txt(假设标题和作者像本节“管理模式”章节示例中那样交替出现)。合并标题和作者,然后对包含 “Paulo” 的行在其开头添加 ”- ”。
本节的示例文件 books_alternate.txt:
A Storm of Swords
George R. R. Martin
The Two Towers
J. R. R. Tolkien
The Alchemist
Paulo Coelho
The Fellowship of the Ring
J. R. R. Tolkien
The Pilgrimage
Paulo Coelho
A Game of Thrones
George R. R. Martin
[user]$ sed -n 'N; s/\n/, /; /Paulo/!b Print; s/^/- /; :Print; p' books_alternate.txt解释:
N:将下一行附加到模式空间,用换行符分隔。s/\n/, /:将嵌入的换行符替换为 ”, “。模式空间现在为:“Title, Author”。/Paulo/!b Print:如果行不包含(!)“Paulo”,则分支(b)到标签Print。s/^/- /:如果未执行分支(即,行包含 “Paulo”),则在开头添加 ”- ”。:Print:定义标签Print。p:打印模式空间。
输出:
A Storm of Swords, George R. R. MartinThe Two Towers, J. R. R. Tolkien- The Alchemist, Paulo CoelhoThe Fellowship of the Ring, J. R. R. Tolkien- The Pilgrimage, Paulo CoelhoA Game of Thrones, George R. R. Martin命令可以写在脚本文件的不同行以提高可读性,或者用分号分隔(尽管 N、标签和分支通常在新行上效果最佳)。
Sed - 条件分支(t 命令)
Section titled “Sed - 条件分支(t 命令)”t 命令提供条件分支:它仅在上次读取输入行或上次执行 t 命令以来,有 s///(替换)命令成功执行了替换操作时,才会跳转到指定的 label。
语法:[address]t [label]
示例:使用 books_alternate.txt(标题/作者交替出现在不同行),对提及 “Paulo” 的行,在其开头最多添加四个连字符。
[user]$ sed -n 'N; s/\n/, /; :Loop; /Paulo/s/^/-/;/----/!t Loop; p' books_alternate.txt解释:
N; s/\n/, /:像之前一样合并标题和作者。:Loop:定义标签Loop。/Paulo/s/^/-/:如果行包含 “Paulo”,则在开头添加一个连字符。这是t命令将测试的s///命令。/----/!t Loop:如果行还未以四个连字符开头,并且之前的s/^/-/命令成功执行了替换,则分支(t)回到Loop标签。p:打印模式空间。
输出:
A Storm of Swords, George R. R. MartinThe Two Towers, J. R. R. Tolkien----The Alchemist, Paulo CoelhoThe Fellowship of the Ring, J. R. R. Tolkien----The Pilgrimage, Paulo CoelhoA Game of Thrones, George R. R. Martint 命令对于创建基于替换操作成功与否来终止的循环至关重要。
Sed - 使用模式空间和保持空间
Section titled “Sed - 使用模式空间和保持空间”模式空间是 Sed 对当前行执行操作的地方。保持空间是一个辅助缓冲区,用于在行处理周期之间存储数据。理解如何在这些空间之间移动数据是编写高级 Sed 脚本的关键。
除非另有说明,否则以下示例使用 books.txt 文件(原始格式:ID) Title, Author, Pages)。
打印命令:p, P
Section titled “打印命令:p, P”p:打印模式空间的全部当前内容。
P(大写):打印模式空间中直到第一个嵌入换行符(\n)的部分。与 N 等多行命令结合使用时非常有用。
示例:只打印第 3 行(使用 -n 抑制默认输出)。
[user]$ sed -n '3p' books.txt输出:
3) The Alchemist, Paulo Coelho, 197命令的行地址(Line Addressing)
Section titled “命令的行地址(Line Addressing)”命令可以限定在特定的行或范围:
N:行号 N(例如,3p打印第 3 行)。N,M:从 N 行到 M 行(例如,2,5p打印第 2 到 5 行)。$:输入的最后一行(例如,$p打印最后一行)。N,$:从 N 行到最后一行(例如,3,$p打印从第 3 行到末尾)。M,+n(GNU 扩展):从 M 行开始,以及接下来的n行(例如,2,+2p打印第 2、3、4 行)。M~n(GNU 扩展):从 M 行开始,每隔 n 行(例如,1~2p打印奇数行:1, 3, 5…;2~2p打印偶数行:2, 4, 6…)。/pattern/:匹配正则表达式pattern的行。
示例:打印 books.txt 中的奇数行。
[user]$ sed -n '1~2p' books.txt输出:
1) A Storm of Swords, George R. R. Martin, 12163) The Alchemist, Paulo Coelho, 1975) The Pilgrimage, Paulo Coelho, 288模式地址(Pattern Addressing)
Section titled “模式地址(Pattern Addressing)”命令也可以通过模式指定地址:
/pattern/command:对匹配pattern的行执行command。/pattern1/,/pattern2/command:对从第一次匹配pattern1的行到之后第一次匹配pattern2的行(包括这两行)执行command。/pattern/,Np:从匹配pattern的行到行号为N的行。
示例:打印所有 “Paulo Coelho” 的书。
[user]$ sed -n '/Paulo Coelho/p' books.txt输出:
3) The Alchemist, Paulo Coelho, 1975) The Pilgrimage, Paulo Coelho, 288示例:打印从第一次出现 “Two Towers” 到 “Pilgrimage” 的行。
[user]$ sed -n '/Two Towers/,/Pilgrimage/p' books.txt输出:
2) The Two Towers, J. R. R. Tolkien, 3523) The Alchemist, Paulo Coelho, 1974) The Fellowship of the Ring, J. R. R. Tolkien, 4325) The Pilgrimage, Paulo Coelho, 288保持空间和模式空间命令
Section titled “保持空间和模式空间命令”以下示例使用 books_alternate.txt 文件(标题在一行,作者在下一行):
h:将模式空间复制到保持空间(覆盖保持空间)。H:将模式空间附加到保持空间(添加一个换行符,然后添加模式空间内容)。g:将保持空间复制到模式空间(覆盖模式空间)。G:将保持空间附加到模式空间(添加一个换行符,然后添加保持空间内容)。x:交换模式空间和保持空间的内容。n:打印当前模式空间(如果未被-n抑制),然后将下一行读入模式空间。N:将下一行附加到模式空间(用\n分隔)。
示例:只打印 books_alternate.txt 中的作者姓名。
[user]$ sed -n 'h; n; p; x; d' books_alternate.txt # More direct: sed -n 'n;p' books_alternate.txt一个更简单的方法来打印每隔一行(作者):
[user]$ sed -n 'n;p' books_alternate.txt输出(作者):
George R. R. MartinJ. R. R. TolkienPaulo CoelhoJ. R. R. TolkienPaulo CoelhoGeorge R. R. Martin示例:颠倒标题和作者的顺序(先打印作者,再打印标题)。
[user]$ sed -n 'h; n; p; g; p' books_alternate.txt解释:
h:将当前行(标题)复制到保持空间。n:将下一行(作者)读入模式空间。p:打印模式空间(作者)。g:将保持空间(标题)复制回模式空间。p:打印模式空间(标题)。
输出:
George R. R. MartinA Storm of SwordsJ. R. R. TolkienThe Two Towers...Sed - 核心编辑命令
Section titled “Sed - 核心编辑命令”本节介绍 Sed 的核心文本操作命令,除非另有说明,否则使用 books.txt 文件(原始格式:ID) Title, Author, Pages)。
删除行 (d)
Section titled “删除行 (d)”[address1[,address2]]d
删除模式空间的内容,因此该行不会被打印。立即开始一个新的循环。
示例:删除第 4 行。
[user]$ sed '4d' books.txt输出(缺少第 4 行):
1) A Storm of Swords, George R. R. Martin, 12162) The Two Towers, J. R. R. Tolkien, 3523) The Alchemist, Paulo Coelho, 1975) The Pilgrimage, Paulo Coelho, 2886) A Game of Thrones, George R. R. Martin, 864写入文件 (w)
Section titled “写入文件 (w)”[address1[,address2]]w filename
将模式空间的内容写入 filename 文件。如果文件不存在则创建,如果存在则覆盖(在 Sed 脚本开始时截断)。
示例:将包含 “Tolkien” 的行写入 tolkien_books.txt 文件。
[user]$ sed -n '/Tolkien/w tolkien_books.txt' books.txt # Use -n if only writing, not printing[user]$ cat tolkien_books.txtcat tolkien_books.txt 的输出:
2) The Two Towers, J. R. R. Tolkien, 3524) The Fellowship of the Ring, J. R. R. Tolkien, 432附加文本 (a)
Section titled “附加文本 (a)”[address]a\ 要附加的文本
将“要附加的文本”附加到当前行的后面。如果文本与 a 在同一行,则必须在其前面加上反斜杠 \,或者从下一行开始。
示例:在第 2 行之后附加文本。
[user]$ sed '2a\---> Inserted Line <---' books.txt输出:
1) A Storm of Swords, George R. R. Martin, 12162) The Two Towers, J. R. R. Tolkien, 352---> Inserted Line <---3) The Alchemist, Paulo Coelho, 197...插入文本 (i)
Section titled “插入文本 (i)”[address]i\ 要插入的文本
将“要插入的文本”插入到当前行的前面。文本的语法类似于 a 命令。
示例:在第 1 行之前插入文本。
[user]$ sed '1i\*** START OF FILE ***' books.txt输出:
*** START OF FILE ***1) A Storm of Swords, George R. R. Martin, 12162) The Two Towers, J. R. R. Tolkien, 352...更改行 (c)
Section titled “更改行 (c)”[address1[,address2]]c\ 新文本
将匹配地址范围的行替换为“新文本”。如果给定了范围,则该范围内的所有行都将被替换为单个实例的“新文本”。
示例:更改第 3 行。
[user]$ sed '3c\REPLACED: This was line 3.' books.txt输出:
1) A Storm of Swords, George R. R. Martin, 12162) The Two Towers, J. R. R. Tolkien, 352REPLACED: This was line 3.4) The Fellowship of the Ring, J. R. R. Tolkien, 432...转换字符 (y)
Section titled “转换字符 (y)”[address1[,address2]]y/源字符集/目标字符集/
执行字符转换。源字符集 中的每个字符都将被 目标字符集 中对应位置的字符替换。源字符集 和 目标字符集 的长度必须相同。此处不使用正则表达式。
示例:对包含 “Alchemist” 的行,将小写 ‘aeio’ 转换为大写 ‘AEIO’。
[user]$ sed '/Alchemist/y/aeio/AEIO/' books.txt输出(第 3 行被修改):
1) A Storm of Swords, George R. R. Martin, 12162) The Two Towers, J. R. R. Tolkien, 3523) ThE AlchEmIst, PAUlO COElhO, 1974) The Fellowship of the Ring, J. R. R. Tolkien, 432...明确列出 (l)
Section titled “明确列出 (l)”[address1[,address2]]l [长度](长度是 GNU 扩展)
以明确无歧义的形式打印模式空间:非打印字符显示为转义序列(例如,制表符显示为 \t),长行会被换行。GNU Sed 允许指定换行 长度。
示例:以 30 的换行长度显示 books.txt。
[user]$ sed -n 'l 30' books.txt输出(第一行示例):
1) A Storm of Swords, George \R. R. Martin, 1216$退出 (q)
Section titled “退出 (q)”[地址]q [退出码](退出码是 GNU 扩展)
Sed 会立即停止处理并退出。如果给定了 地址,则当到达该行时就会退出。GNU Sed 允许指定 退出码。
示例:打印前 3 行(类似于 head -3)。
[user]$ sed '3q' books.txt输出:
1) A Storm of Swords, George R. R. Martin, 12162) The Two Towers, J. R. R. Tolkien, 3523) The Alchemist, Paulo Coelho, 197读取文件 (r)
Section titled “读取文件 (r)”[地址]r 文件名
读取 filename 的内容,并在处理匹配 address 的行之后(如果未指定地址,则在脚本结束时)将其附加到输出中。如果 address 为 0(GNU Sed),则在开始处读取。
[user]$ echo "--- Additional Content ---" > extra.txt[user]$ sed '3r extra.txt' books.txt输出(extra.txt 的内容插入到第 3 行之后):
1) A Storm of Swords, George R. R. Martin, 12162) The Two Towers, J. R. R. Tolkien, 3523) The Alchemist, Paulo Coelho, 197--- Additional Content ---4) The Fellowship of the Ring, J. R. R. Tolkien, 432...执行外部命令 (e - GNU Sed)
Section titled “执行外部命令 (e - GNU Sed)”[地址1[,地址2]]e [命令]
执行 命令。命令 的输出被发送到标准输出。如果省略了 命令,则将模式空间的内容作为命令执行。这是一个强大但潜在危险的 GNU Sed 扩展。
示例:在第 1 行之后运行 date 命令。
[user]$ sed '1e date' books.txt输出(date 命令的输出出现在第 1 行之后,原始第 1 行也会被打印):
1) A Storm of Swords, George R. R. Martin, 1216Wed Mar 27 10:30:00 EDT 2024 (Actual date will vary)2) The Two Towers, J. R. R. Tolkien, 352...Sed - 作为命令的特殊字符
Section titled “Sed - 作为命令的特殊字符”Sed 将几个特殊字符视为命令本身,提供了独特的功能。
打印行号 (=)
Section titled “打印行号 (=)”[地址1[,地址2]]=
将当前输入行号打印到标准输出,后跟一个换行符。
示例:在每行之前打印行号。
[user]$ sed '=' books.txt输出:
11) A Storm of Swords, George R. R. Martin, 121622) The Two Towers, J. R. R. Tolkien, 352...示例:计算文件中的行数(类似于 wc -l)。
[user]$ sed -n '$=' books.txt输出(总行数):
6注释 (#)
Section titled “注释 (#)”如果 Sed 命令的第一个字符是 #,则整行被视为注释并被 Sed 忽略(除非它是 #n,如果出现在脚本的第一行,它等同于 -n 选项)。注释对于给 Sed 脚本添加说明非常有用。
在一个脚本文件 comment_example.sed 中的示例:
# This is a comment, it will be ignored.# Print lines containing 'Tolkien'/Tolkien/p[user]$ sed -n -f comment_example.sed books.txt输出(仅显示 Tolkien 的书):
2) The Two Towers, J. R. R. Tolkien, 3524) The Fellowship of the Ring, J. R. R. Tolkien, 432替换命令中的与号 (&)
Section titled “替换命令中的与号 (&)”虽然本身不是一个命令,但与号 & 在替换命令 (s/pattern/replacement/) 的替换部分中具有特殊含义。它代表了由 pattern 匹配到的整个文本。
示例:将所有 “George R. R. Martin” 的出现用星号括起来。
[user]$ sed 's/George R. R. Martin/**&**/g' books.txt输出:
1) A Storm of Swords, **George R. R. Martin**, 12162) The Two Towers, J. R. R. Tolkien, 352...6) A Game of Thrones, **George R. R. Martin**, 864Sed - 字符串操作(快速指南重申)
Section titled “Sed - 字符串操作(快速指南重申)”s(替换)命令是 Sed 进行字符串操作的主要工具。语法:[地址]s/模式/替换文本/[标志]。
常用标志:g(全局)、N(第 N 次出现)、p(打印)、w 文件(写入文件)、i(忽略大小写 - GNU 扩展)。
示例:将 books.txt 中的所有逗号替换为分号。
[user]$ sed 's/,/;/g' books.txt输出:
1) A Storm of Swords; George R. R. Martin; 12162) The Two Towers; J. R. R. Tolkien; 352...捕获组(在 BRE 中使用 \(...\),在 ERE 中使用 (...))允许在替换文本中使用 \1、\2 等来重新排序和重用匹配文本的部分。
示例:交换作者和标题(假设格式为 ‘Title, Author’)。
[user]$ echo "The Hobbit, J. R. R. Tolkien" | sed -E 's/([^,]+), *(.+)/\2 - \1/'输出:
J. R. R. Tolkien - The HobbitGNU Sed 在替换文本中提供了 \L、\U、\l、\u、\E 用于大小写转换。
Sed - 正则表达式(快速指南重申)
Section titled “Sed - 正则表达式(快速指南重申)”Sed 的强大之处源于正则表达式。默认情况下,它使用基本正则表达式 (BRE)。可以通过 -E (POSIX) 或 -r (GNU) 启用扩展正则表达式 (ERE)。
关键元字符(BRE 示例,ERE 通常更简单):
^:行首。$:行尾。.:任意单个字符(换行符除外)。[...]:字符集(例如,[abc])。[^...]:非字符集(排除字符集中的字符)。*:前一个项的零个或多个。\?:零个或一个(BRE:\?,ERE:?)。\+:一个或多个(BRE:\+,ERE:+)。\{N\}:恰好 N 次(BRE:\{N\},ERE:{N})。\{N,\}:至少 N 次。\{N,M\}:介于 N 和 M 次之间。\(...\):用于捕获的分组(BRE:\(...\),ERE:(...))。\|:交替或 (OR)(BRE:在\(\)内部使用\|,ERE:|)。
POSIX 字符类:[[:alnum:]]、[[:alpha:]]、[[:digit:]]、[[:space:]] 等。提供了匹配字符类型的可移植方式。
示例:打印以数字开头的行。
[user]$ sed -n '/^[[:digit:]]/p' books.txt输出(books.txt 中的所有行,因为它们都以数字开头):
1) A Storm of Swords, George R. R. Martin, 12162) The Two Towers, J. R. R. Tolkien, 3523) The Alchemist, Paulo Coelho, 1974) The Fellowship of the Ring, J. R. R. Tolkien, 4325) The Pilgrimage, Paulo Coelho, 2886) A Game of Thrones, George R. R. Martin, 864GNU Sed 元字符:\b(单词边界)、\s(空白字符)、\w(单词字符)等,提供了类似于 Perl 的便利功能。
Sed - 有用的技巧集
Section titled “Sed - 有用的技巧集”本节通过模拟常见的 Unix 工具和执行实用任务来展示 Sed 的多功能性。这些示例展示了各种 Sed 命令和技术。
模拟 cat(显示文件内容)
Section titled “模拟 cat(显示文件内容)”默认行为(空脚本打印行):
[user]$ sed '' books.txt使用带 -n 的 print 命令:
[user]$ sed -n 'p' books.txt模式 ^$ 匹配空行。
[user]$ echo -e "Line #1\n\n\nLine #2" | sed '/^$/d'输出:
Line #1Line #2移除注释行(例如,C++ 的 // 注释)
Section titled “移除注释行(例如,C++ 的 // 注释)”示例文件 hello.cpp:
#include <iostream>// using namespace std; // A commented out line
int main(void) { // Displays message on stdout. std::cout << "Hello, World !!!" << std::endl; // Another comment return 0; /* Not this type of comment */}[user]$ sed 's|//.*||g' hello.cpp输出(带有 // 注释的行被移除,或者只移除注释部分):
#include <iostream>
int main(void) {
std::cout << "Hello, World !!!" << std::endl; return 0; /* Not this type of comment */}注意:这个正则表达式比较基础,可能无法完美处理所有 C++ 注释场景(例如,字符串内部的 //)。对于健壮的解析,专用的工具更好。
在特定行前添加注释(例如,Shell 脚本)
Section titled “在特定行前添加注释(例如,Shell 脚本)”示例文件 hello.sh:
#!/bin/bash pwd hostname uname -a who who -r lsb_release -a
[user]$ sed '3,5s/^/# /' hello.sh输出(第 3 到 5 行被注释掉了):
#!/bin/bashpwd# hostname# uname -a# whowho -rlsb_release -a模拟 wc -l(计算行数)
Section titled “模拟 wc -l(计算行数)”[user]$ sed -n '$=' hello.sh输出(hello.sh 中的行数):
7模拟 head -N(打印前 N 行)
Section titled “模拟 head -N(打印前 N 行)”示例:head -3 books_alternate.txt
[user]$ sed '3q' books_alternate.txt输出:
A Storm of SwordsGeorge R. R. MartinThe Two Towers模拟 tail -1(打印最后一行)
Section titled “模拟 tail -1(打印最后一行)”[user]$ sed -n '$p' books_alternate.txt输出:
George R. R. Martin更健壮的 tail -N 需要更复杂的 Sed 脚本(通常涉及保持空间和循环),或者可以使用 tac file | sed Nq | tac 命令组合。
模拟 dos2unix(将 CRLF 转换为 LF)
Section titled “模拟 dos2unix(将 CRLF 转换为 LF)”DOS/Windows 使用 CR+LF(\r\n)作为行尾符。Unix 使用 LF(\n)。^M 通常代表 CR。
[user]$ echo -e "Line #1\r\nLine #2\r" > test_dos.txt[user]$ file test_dos.txtfile 的输出:
test_dos.txt: ASCII text, with CRLF line terminators[user]$ sed 's/\r$//' test_dos.txt > test_unix.txt # More portable[user]$ file test_unix.txtfile 的输出:
test_unix.txt: ASCII texts/\r$// 命令会移除行尾($)的 Carriage Return(\r)。
模拟 unix2dos(将 LF 转换为 CRLF)
Section titled “模拟 unix2dos(将 LF 转换为 CRLF)”[user]$ echo -e "Line #1\nLine #2" > test_unix2.txt[user]$ sed 's/$/\r/' test_unix2.txt > test_dos2.txt[user]$ file test_dos2.txtfile 的输出:
test_dos2.txt: ASCII text, with CRLF line terminatorss/$/\r/ 命令会在每一行($)的末尾添加一个 Carriage Return(\r)。
模拟 cat -E(在行尾显示 $)
Section titled “模拟 cat -E(在行尾显示 $)”[user]$ sed 's/$/$/' books.txt # Replace end-of-line with literal $输出(第一行示例):
1) A Storm of Swords, George R. R. Martin, 1216$模拟 cat -T(将 Tab 显示为 ^I)
Section titled “模拟 cat -T(将 Tab 显示为 ^I)”[user]$ echo -e "Column1\tColumn2\nValue1\tValue2" > tab_file.txt[user]$ sed 's/\t/^I/g' tab_file.txt # Note: ^I is a literal tab character, type Ctrl-V then Tab或者,如果不需要精确地显示 ^I,可以使用 l 命令(默认显示为 \t),或者通过管道传递给 tr 命令:
[user]$ sed -n 'l' tab_file.txt # Shows tabs as \t[user]$ sed 's/\t/TEMP_TAB/g' tab_file.txt | tr 'TEMP_TAB' '\t' | cat -T # Complex way for literal ^I模拟 nl(给行加编号)
Section titled “模拟 nl(给行加编号)”[user]$ sed '=' books.txt | sed 'N; s/\n/\t/'解释:
sed '=' books.txt:打印行号,然后是该行。sed 'N; s/\n/\t/':合并每对行(行号和内容行),并将它们之间的换行符替换为制表符。
输出(第一行示例):
1 1) A Storm of Swords, George R. R. Martin, 1216模拟 cp(复制文件)
Section titled “模拟 cp(复制文件)”[user]$ sed -n 'w books_copy.txt' books.txt[user]$ diff books.txt books_copy.txt && echo "Files are identical"输出(如果成功):
Files are identical模拟 expand(将 Tab 转换为空格)
Section titled “模拟 expand(将 Tab 转换为空格)”此示例将每个制表符转换为 4 个空格。真正的 expand 命令具有更复杂的制表位逻辑。
[user]$ sed 's/\t/ /g' tab_file.txt输出:
Column1 Column2Value1 Value2模拟 tee(输出到屏幕和文件)
Section titled “模拟 tee(输出到屏幕和文件)”[user]$ echo -e "Data line 1\nData line 2" | sed 'p; w output.txt'解释:p 打印到标准输出。w output.txt 写入文件。由于默认打印是开启的,p 会导致内容被打印到标准输出两次。使用 -n 以获得与 tee 完全一致的行为。
[user]$ echo -e "Data line 1\nData line 2" | sed -n 'p; w output.txt'[user]$ cat output.txt输出到屏幕和 output.txt 文件中的内容:
Data line 1Data line 2模拟 cat -s(压缩空行)
Section titled “模拟 cat -s(压缩空行)”将连续的多个空行替换为一个空行。
[user]$ echo -e "Line 1\n\n\nLine 2\n\nLine 3" > blank_lines.txt[user]$ sed '/^$/{N;/^\n$/D;}' blank_lines.txt/^$/{N;/^\n$/D;} 的解释:
/^$/:如果当前行为空行…N:附加下一行。模式空间可能变为\n(如果下一行也是空行)。/^\n$/D:如果模式空间仅为\n(即连续两个空行),则删除直到\n的部分,并使用剩余内容重新开始循环。这实际上删除了其中一个空行,并重新评估了新的模式空间。
输出:
Line 1
Line 2
Line 3模拟 grep pattern(打印匹配的行)
Section titled “模拟 grep pattern(打印匹配的行)”[user]$ sed -n '/Alchemist/p' books.txt输出:
3) The Alchemist, Paulo Coelho, 197模拟 grep -v pattern(打印不匹配的行)
Section titled “模拟 grep -v pattern(打印不匹配的行)”[user]$ sed -n '/Alchemist/!p' books.txt输出(除包含 “Alchemist” 的行之外的所有行):
1) A Storm of Swords, George R. R. Martin, 12162) The Two Towers, J. R. R. Tolkien, 3524) The Fellowship of the Ring, J. R. R. Tolkien, 4325) The Pilgrimage, Paulo Coelho, 2886) A Game of Thrones, George R. R. Martin, 864模拟 tr 'abc' 'xyz'(转换字符)
Section titled “模拟 tr 'abc' 'xyz'(转换字符)”[user]$ echo "ABCDE" | sed 'y/ACE/XZE/'输出:
XBDZE