Skip to content

Perl 正则表达式

正则表达式(regex)是定义搜索模式的字符序列。Perl 的正则表达式功能是其最强大的特性之一,深受 grep、sed 和 awk 等工具的影响,但已得到显著扩展。它们用于模式匹配、文本操作和数据验证。

应用正则表达式的主要方式是使用绑定操作符 =~(匹配)和 !~(不匹配)。

Perl 提供三种主要的正则表达式操作符:

  • 匹配 (Match): m/模式/修饰符(如果使用斜杠作为定界符,也可以简写为 /模式/修饰符)
  • 替换 (Substitute): s/模式/替换内容/修饰符
  • 转写 (Transliterate): tr/查找列表/替换列表/修饰符(或 y///)

正斜杠 / 是传统的定界符(delimiters),但你几乎可以使用任何非字母数字、非空白字符(例如,m!PATTERN!, s{PATTERN}{REPLACEMENT}, tr[SEARCHLIST][REPLACEMENTLIST])。如果你使用单引号作为定界符(例如,m'PATTERN'),模式中不会发生变量插值(interpolation)。

匹配操作符 m// 将字符串与正则表达式进行测试。在标量上下文(scalar context)中,如果模式匹配则返回真,否则返回假。在列表上下文(list context)中,它返回一个捕获的子字符串列表(来自模式中的括号)。

#!/usr/bin/perl
use strict;
use warnings;
my $text = "Hello Perl World, Perl is powerful.";
if ($text =~ m/Perl/) {
print "'Perl' found in the text.\n";
} else {
print "'Perl' not found.\n";
}
# Using different delimiters
if ($text =~ m!powerful!) {
print "'powerful' found using 'm!!' operator.\n";
}
# Omitting 'm' with / delimiters
if ($text =~ /World/) {
print "'World' found using '//' operator.\n";
}

输出:

'Perl' found in the text.
'powerful' found using 'm!!' operator.
'World' found using '//' operator.

在列表上下文,m// 返回捕获组:

#!/usr/bin/perl
use strict;
use warnings;
my $time_string = "Time: 10:30:45";
my ($hours, $minutes, $seconds) = ($time_string =~ m/(\d+):(\d+):(\d+)/);
if (defined $hours) {
print "Hours: $hours, Minutes: $minutes, Seconds: $seconds\n";
} else {
print "Time pattern not found.\n";
}

输出:

Hours: 10, Minutes: 30, Seconds: 45

修饰符(modifiers)改变匹配操作符的行为:

修饰符描述
i不区分大小写匹配。
m多行模式:^ 和 $ 除了匹配字符串的开始/结束外,还会匹配字符串内部行的开始/结束。
s单行模式:. (点) 匹配任何单个字符,包括换行符 (\n)。如果没有 /s,. 不匹配 \n。
x扩展模式:允许在模式中使用空白和注释 (# ...) 以提高可读性。
g全局匹配:查找所有匹配项。在列表上下文,返回所有匹配项的所有捕获。在 while 循环的标量上下文,迭代查找匹配项。
o只编译一次:即使模式中包含插值变量,也只编译一次正则表达式。由于优化,在现代 Perl 版本中重要性降低,但有时仍有用。
c继续搜索:与 /g 一起使用时,允许在全局匹配在某个位置失败后继续搜索(阻止在失败时重置 pos())。

不再依赖数字捕获组,如 $1、$2,现代 Perl (5.10+) 允许使用 (?<name>PATTERN) 语法进行命名捕获(named captures)。这些捕获可通过 %+ 哈希访问。

#!/usr/bin/perl
use strict;
use warnings;
use v5.10; # Enables named captures and %+ hash
my $data = "Date: 2023-10-26";
if ($data =~ m/Date: (?<year>\d{4})-(?<month>\d{2})-(?<day>\d{2})/) {
print "Year: $ENV{'+'}{'year'}\n"; # Access via %+ hash (older syntax was $^{+}{name})
print "Month: $ENV{'+'}{'month'}\n";
print "Day: $ENV{'+'}{'day'}\n";
} else {
print "Date not found or format incorrect.\n";
}

输出:

Year: 2023
Month: 10
Day: 26

Perl 提供一些特殊变量,这些变量在成功匹配后会被设置:

  • $&(或使用 use English; 后的 $MATCH):正则表达式匹配的整个字符串。
  • $(或使用 use English; 后的 $PREMATCH):匹配部分之前的字符串。
  • $'(或使用 use English; 后的 $POSTMATCH):匹配部分之后的字符串。
  • $1, $2, … :模式中第一个、第二个等括号对捕获的字符串。
  • %+:包含命名捕获的哈希(如果使用)。
  • %-:包含所有已成功匹配组的捕获的哈希,按组名索引。与 (?|...) 分支重置分组一起使用时有用。

注意:在旧版本的 Perl 中,使用 $, $& 或 $' 可能会带来性能损失,因为 Perl 必须为每次匹配填充它们。现代 Perl (5.16+) 已对此进行了显著优化。对于 Perl 5.10+,如果你需要这些,使用 ${^PREMATCH}、${^MATCH}、${^POSTMATCH}(需要 use re 'eval'; 或在 use v5.10 下可用)有时性能更好,或者使用 /p 修饰符 (Perl 5.20+) 来填充 @{-} 和 @{-}。通常,建议使用捕获组 ($1、$2 或命名捕获) 以提高清晰度和性能。

#!/usr/bin/perl
use strict;
use warnings;
my $string = "The quick brown fox jumps over the lazy_dog.";
if ($string =~ /(brown fox)/) {
print "Matched: '$&'\n";
print "Before: '$`'\n";
print "After: '$''\n";
print "Captured (\$1): '$1'\n";
}

输出:

Matched: 'brown fox'
Before: 'The quick '
After: ' jumps over the lazy_dog.'
Captured ($1): 'brown fox'

替换操作符 s/模式/替换内容/修饰符 搜索 模式,并将其替换为 替换内容。

#!/usr/bin/perl
use strict;
use warnings;
my $sentence = "I like apples and apples are tasty.";
$sentence =~ s/apples/bananas/;
# Replaces only the first occurrence
print "1. $sentence\n";
$sentence =~ s/bananas/oranges/g;
# Replaces all occurrences (due to /g)
print "2. $sentence\n";

输出:

1. I like bananas and apples are tasty.
2. I like oranges and oranges are tasty.

包含匹配修饰符中的 i、m、s、x、g、o,以及:

修饰符描述
e将 替换内容 评估为 Perl 代码,并使用其返回值作为替换字符串。使用 ee 进行两次评估。
r非破坏性替换:返回修改后的字符串,不改变原始字符串。这非常有用,因为它避免就地修改变量,并允许更方便的链式操作。

/e 和 /r 的示例:

#!/usr/bin/perl
use strict;
use warnings;
my $text = "Value: 10, Value: 20";
# Using /e to perform calculation
$text =~ s/(\d+)/$1 * 2/eg;
print "Evaluated: $text\n"; # Output: Evaluated: Value: 20, Value: 40
my $original = "hello world";
my $modified = $original =~ s/world/Perl/r;
print "Original: $original\n"; # Output: Original: hello world
print "Modified: $modified\n"; # Output: Modified: hello Perl

tr/查找列表/替换列表/修饰符 将字符从 查找列表 转写到 替换列表 中对应的字符。它不使用正则表达式模式。

#!/usr/bin/perl
use strict;
use warnings;
my $string = "Hello World 123";
$string =~ tr/A-Za-z/a-zA-Z/; # Swap case
print "Case swapped: $string\n";
$string =~ tr/0-9/X/; # Replace all digits with X
print "Digits replaced: $string\n";

输出:

Case swapped: hELLO wORLD 123
Digits replaced: hELLO wORLD XXX
修饰符描述
c补充 查找列表。针对 查找列表 中 不 包含的所有字符。
d删除找到但未被替换的字符(如果 替换列表 短于 查找列表)。
s将重复的替换字符压缩为单个字符。
#!/usr/bin/perl
use strict;
use warnings;
my $data = "Perl---Programming---Rocks!!!";
$data =~ tr/-/ /s; # Squash multiple hyphens to a single space
print "Squashed: $data\n";
my $text = "abc123xyz";
$text =~ tr/0-9//d; # Delete all digits
print "Digits deleted: $text\n";

输出:

Squashed: Perl Programming Rocks!!!
Digits deleted: abcxyz

Perl 的正则表达式语法非常丰富。这里总结了一些常用元素:

模式描述
^匹配字符串的开始处(使用 /m 时也匹配行的开始处)。
$匹配字符串的结束处(使用 /m 时也匹配行的结束处)。
.匹配任何单个字符,除了换行符(除非使用 /s 修饰符)。
[…]字符类:匹配括号内的任何单个字符(例如 [aeiou])。可以使用范围(例如 [a-z0-9])。
[^…]否定字符类:匹配 不 在括号内的任何单个字符。
*匹配前一个表达式 0 次或多次。
+匹配前一个表达式 1 次或多次。
?匹配前一个表达式 0 次或 1 次(使其可选)。
{n}精确匹配 n 次。
{n,}匹配 n 次或更多次。
{n,m}匹配至少 n 次且至多 m 次。
*?, +?, ??, {n,m}?量词的非贪婪(最小)匹配版本。
|选择:匹配其之前或之后的表达式(例如 cat|dog)。
(…)分组:将表达式分组并捕获匹配的文本。通过 $1、$2 等或命名捕获访问。
(?:…)非捕获组:将表达式分组但不捕获。
\w匹配“词”字符(字母数字加 ”_”)。等同于 [A-Za-z0-9_]。取决于语言环境。
\W匹配非词字符。等同于 [^A-Za-z0-9_]。取决于语言环境。
\s匹配空白字符(空格、制表符、换行符、回车符、换页符)。
\S匹配非空白字符。
\d匹配数字。等同于 [0-9]。
\D匹配非数字。等同于 [^0-9]。
\A只匹配字符串的开始处。
\Z只匹配字符串的结束处,或结束处换行符之前。
\z只匹配字符串的绝对结束处。
\G匹配上一个 /g 匹配结束的位置。对于迭代匹配非常有用。
\b匹配词边界(\w 和 \W 之间的位置)。
\B匹配非词边界。
\n, \t, \r, \f分别匹配换行符、制表符、回车符、换页符。
\1, \2…后向引用:匹配被第 N 个分组捕获的相同文本。
\k命名后向引用:匹配被命名分组捕获的文本。

演示锚点和词边界的示例:

#!/usr/bin/perl
use strict;
use warnings;
my $string = "The cat scattered its catapult.";
# Matches 'cat' as a whole word
if ($string =~ /\bcat\b/) {
print "Found whole word 'cat'. Matched: '$&'\n";
}
# Matches 'cat' NOT as a whole word (e.g., part of 'catapult')
if ($string =~ /\Bcat\B/) {
# This would match 'cat' inside 'scattered' if it were 'scatered'.
# For 'catapult', \bcat\B would match 'cat' at the start of 'catapult'.
print "Found 'cat' within a word (e.g., 'scattered' if applicable). Matched: '$&' (if any)\n";
}
# Matches 'cat' at the beginning of 'catapult'
if ($string =~ /\bcat\B/) {
print "Found 'cat' at start of a longer word. Matched: '$&'\n"; # Matches 'cat' in 'catapult'
}

输出会根据匹配情况而异。对于提供的字符串和正则表达式:

Found whole word 'cat'. Matched: 'cat'
Found 'cat' at start of a longer word. Matched: 'cat'

G 断言与 /g 修饰符一起使用,将匹配锚定在上一个匹配结束的位置(pos($string))。这对于逐段解析或标记化字符串非常有用。

#!/usr/bin/perl
use strict;
use warnings;
my $text = "key1=value1;key2=value2;key3=value3";
while ($text =~ m/\G(?<key>\w+)=(?<value>\w+);?/gc) {
print "Key: $ENV{'+'}{'key'}, Value: $ENV{'+'}{'value'}\n";
}
# The /c modifier prevents reset of pos() on final non-match of the loop condition.

输出:

Key: key1, Value: value1
Key: key2, Value: value2
Key: key3, Value: value3

Perl 的正则表达式功能非常强大且灵活。对于复杂的模式,强烈推荐使用 /x 修饰符来添加注释和空白以提高可读性。掌握正则表达式是高效 Perl 编程的关键技能。请查阅 perldoc perlre 获取完整的指南。