Python 正则表达式
Python 正则表达式(re 模块)
Section titled “Python 正则表达式(re 模块)”正则表达式(regular expressions,简称 regex 或 RE)是强大的字符序列,用于定义搜索模式。它们用于在字符串中进行模式匹配,从而实现查找特定文本格式(电子邮件、URL)、验证、搜索和替换子字符串等任务。
Python 内置的 re 模块为正则表达式提供了全面的支持,其语法大部分与 Perl 的语法兼容。如果在正则表达式的编译或使用过程中发生错误,re 模块将引发 re.error 异常。
强烈建议为正则表达式模式使用 Python 的原始字符串记法(raw string notation)(r"...")。这可以防止正则表达式模式中的反斜杠(\)被解释为 Python 字符串转义字符。
核心函数:match() vs search()
Section titled “核心函数:match() vs search()”re 模块提供了多种用于模式匹配的函数。其中两个基础函数是 match() 和 search()。
1. re.match(pattern, string, flags=0)
Section titled “1. re.match(pattern, string, flags=0)”尝试仅从 string 的开头匹配 pattern。如果模式匹配字符串的开头,则返回一个匹配对象(match object);否则,返回 None。
参数:
| 参数 | 描述 |
|---|---|
pattern | 正则表达式字符串(使用原始字符串 r'')。 |
string | 要在其中搜索的字符串。 |
flags | 可选修饰符(例如 re.IGNORECASE、re.MULTILINE),使用按位或 (|) 组合。默认为 0。 |
2. re.search(pattern, string, flags=0)
Section titled “2. re.search(pattern, string, flags=0)”扫描整个 string,查找 pattern 产生匹配的第一个位置。如果在字符串中的任何位置找到匹配项,则返回一个匹配对象;否则,返回 None。
参数与 re.match() 相同。
match() 和 search() 成功时都会返回一个 Match 对象。匹配对象具有有用的方法:
| 方法 | 描述 |
|---|---|
group(num=0) | 返回整个匹配的字符串,或特定的捕获组(capturing group)num(基于 1 的索引)。group(0) 或 group() 返回整个匹配项。 |
groups() | 返回一个包含所有捕获的子组的元组(如果没有捕获组,则返回空元组)。 |
start(num=0) | 返回匹配项(或子组 num)的起始索引。 |
end(num=0) | 返回匹配项(或子组 num)的结束索引(不包含该位置)。 |
span(num=0) | 返回匹配项(或子组 num)的元组 (start, end)。 |
示例:match() vs search()
Section titled “示例:match() vs search()”import re
text = "The quick brown fox jumps over the lazy dog."pattern = r"fox"
# match() - Fails because 'fox' is not at the beginningmatch_obj = re.match(pattern, text)if match_obj: print(f"match() found: '{match_obj.group()}' at index {match_obj.start()}")else: print("match() found nothing.")
# search() - Succeeds because 'fox' is found within the stringsearch_obj = re.search(pattern, text)if search_obj: print(f"search() found: '{search_obj.group()}' at index {search_obj.start()}")else: print("search() found nothing.")
# Example with capturing groupsline = "Cats are smarter than dogs"pattern_groups = r'(.*) are (.*?) .*' # Capture groups with ()search_groups = re.search(pattern_groups, line, re.IGNORECASE)
if search_groups: print(f"\nFull match: {search_groups.group(0)}") print(f"Group 1: {search_groups.group(1)}") print(f"Group 2: {search_groups.group(2)}") print(f"All groups: {search_groups.groups()}")else: print("Pattern with groups not found!")输出:
match() found nothing.search() found: 'fox' at index 16
Full match: Cats are smarter than dogsGroup 1: CatsGroup 2: smarterAll groups: ('Cats', 'smarter')其他有用的 re 函数
Section titled “其他有用的 re 函数”re.findall(pattern, string, flags=0)
Section titled “re.findall(pattern, string, flags=0)”在 string 中找到 pattern 的所有非重叠匹配项,并将它们作为字符串列表返回。如果模式包含捕获组,则返回一个元组列表,其中每个元组包含对应匹配项的捕获组。
re.finditer(pattern, string, flags=0)
Section titled “re.finditer(pattern, string, flags=0)”类似于 findall(),但返回一个迭代器,为每个匹配项生成 Match 对象。这对于大量匹配项来说更节省内存。
re.sub(pattern, repl, string, count=0, flags=0)
Section titled “re.sub(pattern, repl, string, count=0, flags=0)”将 string 中 pattern 的出现替换为 repl。repl 可以是一个字符串(允许使用反向引用,如 \1、\g<name>),也可以是一个接受匹配对象并返回替换字符串的函数。如果指定了 count 且非零,则只替换前 count 个匹配项。返回修改后的字符串。
示例:findall() 和 sub()
Section titled “示例:findall() 和 sub()”import re
text = "Contact us at info@example.com or support@example.org for help."
# Find all email addressesemail_pattern = r'[\w.-]+@[\w.-]+\.\w+'emails = re.findall(email_pattern, text)print(f"Emails found: {emails}")
# Replace domain namesphone = "Call 2004-959-559 for support # This is a Phone Number"
# Delete commentsnum_no_comment = re.sub(r'#.*$', '', phone).strip()print(f"Phone without comment: '{num_no_comment}'")
# Remove non-digitsnum_digits_only = re.sub(r'\D', '', phone)print(f"Phone digits only: '{num_digits_only}'")输出:
Emails found: ['info@example.com', 'support@example.org']Phone without comment: 'Call 2004-959-559 for support'Phone digits only: '2004959559'正则表达式标志
Section titled “正则表达式标志”标志(Flags)修改模式的解释方式:
| 标志 | 别名 | 描述 |
|---|---|---|
re.IGNORECASE | re.I | 执行不区分大小写的匹配。 |
re.MULTILINE | re.M | 使 ^ 匹配每行的开头(换行符之后)和 $ 匹配每行的结尾(换行符之前),此外还匹配整个字符串的开头/结尾。 |
re.DOTALL | re.S | 使点 (.) 特殊字符匹配包括换行符在内的任何字符;如果没有此标志,. 匹配除换行符以外的任何字符。 |
re.VERBOSE | re.X | 允许您编写更具可读性的正则表达式模式,忽略空白字符(转义或在字符集内除外),并将 # 视作注释标记。 |
re.ASCII | re.A | 使 \w, \W, \b, \B, \s, \S, \d, \D 执行 ASCII 字符集的匹配,而不是完整的 Unicode 匹配(Python 3 中的默认行为)。 |
re.UNICODE | re.U | 在 Python 3 中已弃用,因为 Unicode 匹配是默认行为。包含此项是为了向后兼容性。 |
常见正则表达式元字符和语法
Section titled “常见正则表达式元字符和语法”正则表达式模式由字面字符和特殊元字符(metacharacters)构成:
| 模式 | 描述 |
|---|---|
. | 匹配除换行符以外的任何字符(除非使用 re.DOTALL)。 |
^ | 匹配字符串的开头(如果使用 re.MULTILINE,则匹配行的开头)。 |
$ | 匹配字符串的结尾(如果使用 re.MULTILINE,则匹配行的结尾)。 |
* | 匹配前一个表达式的 0 次或多次重复(贪婪)。 |
+ | 匹配前一个表达式的 1 次或多次重复(贪婪)。 |
? | 匹配前一个表达式的 0 次或 1 次重复(贪婪)。 |
*?, +?, ?? | *, +, ? 的非贪婪(lazy)版本。匹配尽可能少的字符。 |
{m} | 精确匹配前一个表达式的 m 次重复。 |
{m,n} | 匹配前一个表达式的 m 到 n 次重复(包含 m 和 n,贪婪)。 |
{m,n}? | {m,n} 的非贪婪版本。 |
[...] | 字符集:匹配方括号内的任何单个字符。[aeiou] 匹配任何元音字母。可以使用范围,如 [a-z] 或 [0-9]。 |
[^...] | 否定字符集:匹配不在方括号内的任何单个字符。 |
\ | 转义字符:转义特殊字符(例如 \., \*, \\)或标记特殊序列。 |
\d | 匹配任何 Unicode 十进制数字(如果使用 re.ASCII,则等同于 [0-9])。 |
\D | 匹配任何非十进制数字的字符。 |
\s | 匹配任何 Unicode 空白字符(空格、制表符、换行符等)。 |
\S | 匹配任何非空白字符的字符。 |
\w | 匹配 Unicode 词字符(字母、数字加下划线)。 |
\W | 匹配任何非词字符的字符。 |
\b | 匹配单词开头或结尾的空字符串(词边界)。 |
\B | 匹配非词边界的空字符串。 |
(...) | 捕获组:匹配括号内的表达式并捕获匹配的文本。可以使用 group(n) 检索。 |
(?:...) | 非捕获组:分组表达式但不捕获匹配项。 |
| | 或:匹配 | 前面或后面的表达式。cat|dog 匹配 ‘cat’ 或 ‘dog’。 |
(?P<name>...) | 命名捕获组:捕获匹配项并为其指定一个名称。通过 group('name') 访问。 |
(?=...) | 正向先行断言(Positive Lookahead Assertion):如果子模式在当前位置后面匹配,则匹配成功,但不消耗任何字符。 |
(?!...) | 负向先行断言(Negative Lookahead Assertion):如果子模式在当前位置后面不匹配,则匹配成功。 |
(?<=...) | 正向后行断言(Positive Lookbehind Assertion):如果子模式在当前位置前面匹配,则匹配成功,但不消耗任何字符。 |
(?<!...) | 负向后行断言(Negative Lookbehind Assertion):如果子模式在当前位置前面不匹配,则匹配成功。 |
有关更多详细信息和高级模式,请查阅官方 Python re 模块文档。