Perl 正则表达式
Perl - 正则表达式
Section titled “Perl - 正则表达式”正则表达式(regex)是定义搜索模式的字符序列。Perl 的正则表达式功能是其最强大的特性之一,深受 grep、sed 和 awk 等工具的影响,但已得到显著扩展。它们用于模式匹配、文本操作和数据验证。
应用正则表达式的主要方式是使用绑定操作符 =~(匹配)和 !~(不匹配)。
Perl 提供三种主要的正则表达式操作符:
- 匹配 (Match):
m/模式/修饰符(如果使用斜杠作为定界符,也可以简写为/模式/修饰符) - 替换 (Substitute):
s/模式/替换内容/修饰符 - 转写 (Transliterate):
tr/查找列表/替换列表/修饰符(或y///)
正斜杠 / 是传统的定界符(delimiters),但你几乎可以使用任何非字母数字、非空白字符(例如,m!PATTERN!, s{PATTERN}{REPLACEMENT}, tr[SEARCHLIST][REPLACEMENTLIST])。如果你使用单引号作为定界符(例如,m'PATTERN'),模式中不会发生变量插值(interpolation)。
匹配操作符 (m//)
Section titled “匹配操作符 (m//)”匹配操作符 m// 将字符串与正则表达式进行测试。在标量上下文(scalar context)中,如果模式匹配则返回真,否则返回假。在列表上下文(list context)中,它返回一个捕获的子字符串列表(来自模式中的括号)。
#!/usr/bin/perluse strict;use warnings;
my $text = "Hello Perl World, Perl is powerful.";
if ($text =~ m/Perl/) { print "'Perl' found in the text.\n";} else { print "'Perl' not found.\n";}
# Using different delimitersif ($text =~ m!powerful!) { print "'powerful' found using 'm!!' operator.\n";}
# Omitting 'm' with / delimitersif ($text =~ /World/) { print "'World' found using '//' operator.\n";}输出:
'Perl' found in the text.'powerful' found using 'm!!' operator.'World' found using '//' operator.在列表上下文,m// 返回捕获组:
#!/usr/bin/perluse strict;use warnings;
my $time_string = "Time: 10:30:45";my ($hours, $minutes, $seconds) = ($time_string =~ m/(\d+):(\d+):(\d+)/);
if (defined $hours) { print "Hours: $hours, Minutes: $minutes, Seconds: $seconds\n";} else { print "Time pattern not found.\n";}输出:
Hours: 10, Minutes: 30, Seconds: 45匹配操作符修饰符
Section titled “匹配操作符修饰符”修饰符(modifiers)改变匹配操作符的行为:
| 修饰符 | 描述 |
|---|---|
| i | 不区分大小写匹配。 |
| m | 多行模式:^ 和 $ 除了匹配字符串的开始/结束外,还会匹配字符串内部行的开始/结束。 |
| s | 单行模式:. (点) 匹配任何单个字符,包括换行符 (\n)。如果没有 /s,. 不匹配 \n。 |
| x | 扩展模式:允许在模式中使用空白和注释 (# ...) 以提高可读性。 |
| g | 全局匹配:查找所有匹配项。在列表上下文,返回所有匹配项的所有捕获。在 while 循环的标量上下文,迭代查找匹配项。 |
| o | 只编译一次:即使模式中包含插值变量,也只编译一次正则表达式。由于优化,在现代 Perl 版本中重要性降低,但有时仍有用。 |
| c | 继续搜索:与 /g 一起使用时,允许在全局匹配在某个位置失败后继续搜索(阻止在失败时重置 pos())。 |
命名捕获 (现代 Perl)
Section titled “命名捕获 (现代 Perl)”不再依赖数字捕获组,如 $1、$2,现代 Perl (5.10+) 允许使用 (?<name>PATTERN) 语法进行命名捕获(named captures)。这些捕获可通过 %+ 哈希访问。
#!/usr/bin/perluse strict;use warnings;use v5.10; # Enables named captures and %+ hash
my $data = "Date: 2023-10-26";if ($data =~ m/Date: (?<year>\d{4})-(?<month>\d{2})-(?<day>\d{2})/) { print "Year: $ENV{'+'}{'year'}\n"; # Access via %+ hash (older syntax was $^{+}{name}) print "Month: $ENV{'+'}{'month'}\n"; print "Day: $ENV{'+'}{'day'}\n";} else { print "Date not found or format incorrect.\n";}输出:
Year: 2023Month: 10Day: 26正则表达式变量
Section titled “正则表达式变量”Perl 提供一些特殊变量,这些变量在成功匹配后会被设置:
$&(或使用use English;后的$MATCH):正则表达式匹配的整个字符串。$(或使用use English;后的$PREMATCH):匹配部分之前的字符串。$'(或使用use English;后的$POSTMATCH):匹配部分之后的字符串。$1,$2, … :模式中第一个、第二个等括号对捕获的字符串。%+:包含命名捕获的哈希(如果使用)。%-:包含所有已成功匹配组的捕获的哈希,按组名索引。与(?|...)分支重置分组一起使用时有用。
注意:在旧版本的 Perl 中,使用 $, $& 或 $' 可能会带来性能损失,因为 Perl 必须为每次匹配填充它们。现代 Perl (5.16+) 已对此进行了显著优化。对于 Perl 5.10+,如果你需要这些,使用 ${^PREMATCH}、${^MATCH}、${^POSTMATCH}(需要 use re 'eval'; 或在 use v5.10 下可用)有时性能更好,或者使用 /p 修饰符 (Perl 5.20+) 来填充 @{-} 和 @{-}。通常,建议使用捕获组 ($1、$2 或命名捕获) 以提高清晰度和性能。
#!/usr/bin/perluse strict;use warnings;
my $string = "The quick brown fox jumps over the lazy_dog.";if ($string =~ /(brown fox)/) { print "Matched: '$&'\n"; print "Before: '$`'\n"; print "After: '$''\n"; print "Captured (\$1): '$1'\n";}输出:
Matched: 'brown fox'Before: 'The quick 'After: ' jumps over the lazy_dog.'Captured ($1): 'brown fox'替换操作符 (s///)
Section titled “替换操作符 (s///)”替换操作符 s/模式/替换内容/修饰符 搜索 模式,并将其替换为 替换内容。
#!/usr/bin/perluse strict;use warnings;
my $sentence = "I like apples and apples are tasty.";$sentence =~ s/apples/bananas/;# Replaces only the first occurrenceprint "1. $sentence\n";
$sentence =~ s/bananas/oranges/g;# Replaces all occurrences (due to /g)print "2. $sentence\n";输出:
1. I like bananas and apples are tasty.2. I like oranges and oranges are tasty.替换操作符修饰符
Section titled “替换操作符修饰符”包含匹配修饰符中的 i、m、s、x、g、o,以及:
| 修饰符 | 描述 |
|---|---|
| e | 将 替换内容 评估为 Perl 代码,并使用其返回值作为替换字符串。使用 ee 进行两次评估。 |
| r | 非破坏性替换:返回修改后的字符串,不改变原始字符串。这非常有用,因为它避免就地修改变量,并允许更方便的链式操作。 |
/e 和 /r 的示例:
#!/usr/bin/perluse strict;use warnings;
my $text = "Value: 10, Value: 20";# Using /e to perform calculation$text =~ s/(\d+)/$1 * 2/eg;print "Evaluated: $text\n"; # Output: Evaluated: Value: 20, Value: 40
my $original = "hello world";my $modified = $original =~ s/world/Perl/r;print "Original: $original\n"; # Output: Original: hello worldprint "Modified: $modified\n"; # Output: Modified: hello Perl转写操作符 (tr/// 或 y///)
Section titled “转写操作符 (tr/// 或 y///)”tr/查找列表/替换列表/修饰符 将字符从 查找列表 转写到 替换列表 中对应的字符。它不使用正则表达式模式。
#!/usr/bin/perluse strict;use warnings;
my $string = "Hello World 123";$string =~ tr/A-Za-z/a-zA-Z/; # Swap caseprint "Case swapped: $string\n";
$string =~ tr/0-9/X/; # Replace all digits with Xprint "Digits replaced: $string\n";输出:
Case swapped: hELLO wORLD 123Digits replaced: hELLO wORLD XXX转写操作符修饰符
Section titled “转写操作符修饰符”| 修饰符 | 描述 |
|---|---|
| c | 补充 查找列表。针对 查找列表 中 不 包含的所有字符。 |
| d | 删除找到但未被替换的字符(如果 替换列表 短于 查找列表)。 |
| s | 将重复的替换字符压缩为单个字符。 |
#!/usr/bin/perluse strict;use warnings;
my $data = "Perl---Programming---Rocks!!!";$data =~ tr/-/ /s; # Squash multiple hyphens to a single spaceprint "Squashed: $data\n";
my $text = "abc123xyz";$text =~ tr/0-9//d; # Delete all digitsprint "Digits deleted: $text\n";输出:
Squashed: Perl Programming Rocks!!!Digits deleted: abcxyz核心正则表达式语法
Section titled “核心正则表达式语法”Perl 的正则表达式语法非常丰富。这里总结了一些常用元素:
| 模式 | 描述 |
|---|---|
| ^ | 匹配字符串的开始处(使用 /m 时也匹配行的开始处)。 |
| $ | 匹配字符串的结束处(使用 /m 时也匹配行的结束处)。 |
| . | 匹配任何单个字符,除了换行符(除非使用 /s 修饰符)。 |
| […] | 字符类:匹配括号内的任何单个字符(例如 [aeiou])。可以使用范围(例如 [a-z0-9])。 |
| [^…] | 否定字符类:匹配 不 在括号内的任何单个字符。 |
| * | 匹配前一个表达式 0 次或多次。 |
| + | 匹配前一个表达式 1 次或多次。 |
| ? | 匹配前一个表达式 0 次或 1 次(使其可选)。 |
| {n} | 精确匹配 n 次。 |
| {n,} | 匹配 n 次或更多次。 |
| {n,m} | 匹配至少 n 次且至多 m 次。 |
| *?, +?, ??, {n,m}? | 量词的非贪婪(最小)匹配版本。 |
| | | 选择:匹配其之前或之后的表达式(例如 cat|dog)。 |
| (…) | 分组:将表达式分组并捕获匹配的文本。通过 $1、$2 等或命名捕获访问。 |
| (?:…) | 非捕获组:将表达式分组但不捕获。 |
| \w | 匹配“词”字符(字母数字加 ”_”)。等同于 [A-Za-z0-9_]。取决于语言环境。 |
| \W | 匹配非词字符。等同于 [^A-Za-z0-9_]。取决于语言环境。 |
| \s | 匹配空白字符(空格、制表符、换行符、回车符、换页符)。 |
| \S | 匹配非空白字符。 |
| \d | 匹配数字。等同于 [0-9]。 |
| \D | 匹配非数字。等同于 [^0-9]。 |
| \A | 只匹配字符串的开始处。 |
| \Z | 只匹配字符串的结束处,或结束处换行符之前。 |
| \z | 只匹配字符串的绝对结束处。 |
| \G | 匹配上一个 /g 匹配结束的位置。对于迭代匹配非常有用。 |
| \b | 匹配词边界(\w 和 \W 之间的位置)。 |
| \B | 匹配非词边界。 |
| \n, \t, \r, \f | 分别匹配换行符、制表符、回车符、换页符。 |
| \1, \2… | 后向引用:匹配被第 N 个分组捕获的相同文本。 |
| \k | 命名后向引用:匹配被命名分组捕获的文本。 |
演示锚点和词边界的示例:
#!/usr/bin/perluse strict;use warnings;
my $string = "The cat scattered its catapult.";
# Matches 'cat' as a whole wordif ($string =~ /\bcat\b/) { print "Found whole word 'cat'. Matched: '$&'\n";}
# Matches 'cat' NOT as a whole word (e.g., part of 'catapult')if ($string =~ /\Bcat\B/) { # This would match 'cat' inside 'scattered' if it were 'scatered'. # For 'catapult', \bcat\B would match 'cat' at the start of 'catapult'. print "Found 'cat' within a word (e.g., 'scattered' if applicable). Matched: '$&' (if any)\n";}
# Matches 'cat' at the beginning of 'catapult'if ($string =~ /\bcat\B/) { print "Found 'cat' at start of a longer word. Matched: '$&'\n"; # Matches 'cat' in 'catapult'}输出会根据匹配情况而异。对于提供的字符串和正则表达式:
Found whole word 'cat'. Matched: 'cat'Found 'cat' at start of a longer word. Matched: 'cat'用于连续匹配的 G 断言
Section titled “用于连续匹配的 G 断言”G 断言与 /g 修饰符一起使用,将匹配锚定在上一个匹配结束的位置(pos($string))。这对于逐段解析或标记化字符串非常有用。
#!/usr/bin/perluse strict;use warnings;
my $text = "key1=value1;key2=value2;key3=value3";while ($text =~ m/\G(?<key>\w+)=(?<value>\w+);?/gc) { print "Key: $ENV{'+'}{'key'}, Value: $ENV{'+'}{'value'}\n";}# The /c modifier prevents reset of pos() on final non-match of the loop condition.输出:
Key: key1, Value: value1Key: key2, Value: value2Key: key3, Value: value3Perl 的正则表达式功能非常强大且灵活。对于复杂的模式,强烈推荐使用 /x 修饰符来添加注释和空白以提高可读性。掌握正则表达式是高效 Perl 编程的关键技能。请查阅 perldoc perlre 获取完整的指南。