Skip to content

Sed 字符串

替换命令 s 是 Sed 最强大的功能之一,用于查找和替换文本。其基本语法是:

[address1[,address2]]s/pattern/replacement/[flags]

其中:

  • address1, address2: 可选的行地址或模式,用于限制命令的作用范围。
  • s: 替换命令。
  • /: 分隔符。虽然 / 很常用,但可以使用任何字符(除了反斜杠或换行符),例如 s|pattern|replacement| 或 s#pattern#replacement#。如果 pattern 或 replacement 包含斜杠,这会很有用。
  • pattern: 用于搜索的正则表达式。
  • replacement: 用于替换匹配到的 pattern 的字符串。可以在此处使用特殊字符,例如 & (整个匹配到的 pattern) 和 \1 到 \9 (捕获组)。
  • flags: 改变替换行为的可选修饰符。

常用标志 (Flags):

  • g: 全局 (Global)。替换行上所有匹配到的 pattern,而不是只替换第一个。
  • N (数字): 只替换第 N 次匹配。
  • p: 打印 (Print)。如果发生替换,则打印(修改后的)行。常与 -n 一起使用,只打印更改的行。
  • w file: 写入 (Write)。如果发生替换,则将(修改后的)行写入 file。
  • i 或 I (GNU Sed 扩展): 对 pattern 进行不区分大小写的匹配。

让我们使用一个示例 books.txt 文件:

  1. A Storm of Swords, George R. R. Martin, 1216
  2. The Two Towers, J. R. R. Tolkien, 352
  3. The Alchemist, Paulo Coelho, 197
  4. The Fellowship of the Ring, J. R. R. Tolkien, 432
  5. The Pilgrimage, Paulo Coelho, 288
  6. A Game of Thrones, George R. R. Martin, 864

示例 1: 将第一个逗号替换为管道符。

[user]$ sed 's/,/ | /' books.txt

输出:

1) A Storm of Swords | George R. R. Martin, 1216
2) The Two Towers | J. R. R. Tolkien, 352
3) The Alchemist | Paulo Coelho, 197
4) The Fellowship of the Ring | J. R. R. Tolkien, 432
5) The Pilgrimage | Paulo Coelho, 288
6) A Game of Thrones | George R. R. Martin, 864

注意,只替换了每行的第一个逗号。要替换所有逗号,请使用 g 标志:

[user]$ sed 's/,/ | /g' books.txt

输出:

1) A Storm of Swords | George R. R. Martin | 1216
2) The Two Towers | J. R. R. Tolkien | 352
3) The Alchemist | Paulo Coelho | 197
4) The Fellowship of the Ring | J. R. R. Tolkien | 432
5) The Pilgrimage | Paulo Coelho | 288
6) A Game of Thrones | George R. R. Martin | 864

示例 2: 只替换包含 “The Pilgrimage” 的行中的逗号。

[user]$ sed '/The Pilgrimage/s/,/ | /g' books.txt

输出:

1) A Storm of Swords, George R. R. Martin, 1216
2) The Two Towers, J. R. R. Tolkien, 352
3) The Alchemist, Paulo Coelho, 197
4) The Fellowship of the Ring, J. R. R. Tolkien, 432
5) The Pilgrimage | Paulo Coelho | 288
6) A Game of Thrones, George R. R. Martin, 864

示例 3: 只将第二个逗号替换为管道符。

[user]$ sed 's/,/ | /2' books.txt

输出:

1) A Storm of Swords, George R. R. Martin | 1216
2) The Two Towers, J. R. R. Tolkien | 352
3) The Alchemist, Paulo Coelho | 197
4) The Fellowship of the Ring, J. R. R. Tolkien | 432
5) The Pilgrimage, Paulo Coelho | 288
6) A Game of Thrones, George R. R. Martin | 864

示例 4: 只打印更改的行 (使用 -n 和 p 标志)。

[user]$ sed -n 's/Paulo Coelho/PAULO COELHO/p' books.txt

输出:

3) The Alchemist, PAULO COELHO, 197
5) The Pilgrimage, PAULO COELHO, 288

示例 5: 将更改的行写入文件 (使用 w 标志)。

[user]$ sed 's/Paulo Coelho/PAULO COELHO/w changed_authors.txt' books.txt > /dev/null # suppress normal output
[user]$ cat changed_authors.txt

cat changed_authors.txt 的输出:

3) The Alchemist, PAULO COELHO, 197
5) The Pilgrimage, PAULO COELHO, 288

注意:原始命令 sed -n 's/.../w file' 将不会向标准输出打印任何内容。如果您希望既有默认输出 又 将更改的行写入文件,请使用 sed 's/.../&/w file'。& 指代整个匹配项。上述示例明确地重定向了主要输出。

示例 6: 不区分大小写的替换 (GNU Sed 的 i 标志)。

[user]$ sed -n 's/pAuLo CoElHo/PAULO COELHO/pi' books.txt

输出:

3) The Alchemist, PAULO COELHO, 197
5) The Pilgrimage, PAULO COELHO, 288

示例 7: 使用不同的分隔符来替换路径。

如果您需要将 /usr/local/bin 替换为 /opt/bin,使用 / 作为分隔符需要进行大量的转义:

[user]$ echo "Path is /usr/local/bin" | sed 's/\/usr\/local\/bin/\/opt\/bin/'

输出:

Path is /opt/bin

使用像 |、# 或 ! 这样的替代分隔符会使其更易读:

[user]$ echo "Path is /usr/local/bin" | sed 's|/usr/local/bin|/opt/bin|'

输出:

Path is /opt/bin

Sed 允许您捕获匹配到的 pattern 的一部分,并在 replacement 字符串中重用它们。这是通过使用捕获组来实现的。

在基本正则表达式(BRE,Sed 的默认模式)中,组使用 \( 和 \) 定义。捕获到的文本可以使用 \1 引用第一个组,\2 引用第二个组,依此类推。

如果使用扩展正则表达式(sed -E 或 sed -r),组使用 ( 和 ) 定义,并且仍然通过 \1、\2 等引用。

考虑输入字符串:“Three One Two”。我们想将其重排为 “One Two Three”。

[user]$ echo "Three One Two" | sed 's|\(\w\+\) \(\w\+\) \(\w\+\)|\2 \3 \1|'

解释:

  • \(\w\+\): 这是一个捕获组。
    • \w\+: 匹配一个或多个单词字符(GNU Sed 扩展;[[:alnum:]_]+ 对于“单词字符”的概念更具可移植性)。BRE 中需要转义 \+。
  • 该模式 \(\w\+\) \(\w\+\) \(\w\+\) 捕获三个由空格分隔的单词。
    • \1 捕获 “Three”
    • \2 捕获 “One”
    • \3 捕获 “Two”
  • \2 \3 \1: replacement 字符串重新排列这些捕获组。

输出:

One Two Three

使用 ERE,命令更简洁:

[user]$ echo "Three One Two" | sed -E 's|(\w+) (\w+) (\w+)|\2 \3 \1|'

逗号分隔值的示例:

[user]$ echo "Smith,John,123 Main St" | sed -E 's|([^,]+),([^,]+),(.+)|\2 \1, \3|'

输出(重排为 Firstname Lastname, Address):

John Smith, 123 Main St

在这里,[^,]+ 匹配一个或多个不是逗号的字符。

Replacement 中的字符串大小写修改 (GNU Sed 扩展)

Section titled “Replacement 中的字符串大小写修改 (GNU Sed 扩展)”

GNU Sed 在 replacement 字符串中提供特殊序列,用于改变匹配文本或捕获组的大小写。

  • \L: 将 replacement 的其余部分(或直到 \E)转换为小写。
  • \l: 将下一个字符转换为小写。
  • \U: 将 replacement 的其余部分(或直到 \E)转换为大写。
  • \u: 将下一个字符转换为大写。
  • \E: 停止由 \L 或 \U 开始的大小写转换。

示例 1: 使用 \L

[user]$ sed -n 's/Paulo/PA\LULO COELHO/p' books.txt

输出 (PA 后跟小写的 ulo):

3) The Alchemist, PAulo COELHO, 197
5) The Pilgrimage, PAulo COELHO, 288

示例 2: 使用 \u

[user]$ sed -n 's/Paulo/p\uaul\uo/p' books.txt

输出 (p,然后是大写 A,然后是 ul,然后是大写 O):

3) The Alchemist, pAulO Coelho, 197
5) The Pilgrimage, pAulO Coelho, 288

示例 3: 使用 \U

[user]$ sed -n 's/Paulo/\Upaulo/p' books.txt

输出 (将 ‘paulo’ 转换为 ‘PAULO’):

3) The Alchemist, PAULO Coelho, 197
5) The Pilgrimage, PAULO Coelho, 288

示例 4: 使用 \U 和 \E

[user]$ sed -n 's/Paulo Coelho/Title: \U&\E - Author./p' books.txt # & refers to the whole match 'Paulo Coelho'

输出 (在 replacement 中将 ‘Paulo Coelho’ 转换为大写):

3) The Alchemist, Title: PAULO COELHO - Author., 197
5) The Pilgrimage, Title: PAULO COELHO - Author., 288

实际应用: 将 CSV 转换为格式化字符串。

输入 data.csv: name,age,city Alice,30,New York Bob,24,Los Angeles

[user]$ sed -E '1d; s/([^,]+),([^,]+),(.+)/Name: \U\1\E, Age: \2, City: \u\3/' data.csv

解释:

  • 1d: 删除标题行。
  • s/([^,]+),([^,]+),(.+)/.../: 捕获三个逗号分隔的字段。
    • \1 是姓名,\2 是年龄,\3 是城市。
  • Name: \U\1\E: 打印 “Name: “,后跟大写的姓名。
  • Age: \2: 打印 “Age: “,后跟年龄。
  • City: \u\3: 打印 “City: “,后跟首字母大写的城市名。

输出:

Name: ALICE, Age: 30, City: New York
Name: BOB, Age: 24, City: Los Angeles