C# - 正则表达式
C# - 正则表达式
Section titled “C# - 正则表达式”正则表达式(或 regex)是定义搜索模式的字符序列。在 C# 中,System.Text.RegularExpressions 命名空间提供了一个强大的引擎,用于字符串的模式匹配、验证、替换和拆分。
核心 Regex 概念与语法
Section titled “核心 Regex 概念与语法”Regex 模式由字面量(如 a、b、1)和具有特殊含义的元字符构建而成。
| 类别 | 构造 | 描述 | 示例 |
|---|---|---|---|
| 字符类 | . | 匹配除换行符外的任意单个字符。 | h.t 匹配 ‘hat’, ‘hot’, ‘h&t’ |
\d | 匹配任意数字 (0-9)。\D 匹配非数字字符。 | \d{3} 匹配 ‘123’ | |
\w | 匹配任意单词字符 (a-z, A-Z, 0-9, _)。\W 匹配非单词字符。 | \w+ 匹配 ‘hello_123’ | |
\s | 匹配任意空白字符。\S 匹配非空白字符。 | hello\sworld | |
[...] | 匹配集合中的任意单个字符。 | [aeiou] 匹配任意元音 | |
| 锚点 | ^ | 匹配字符串的开头。 | ^Hello |
$ | 匹配字符串的末尾。 | world$ | |
\b | 匹配单词边界。 | \bcat\b 匹配 ‘cat’, 而非 ‘category’ | |
| 量词 | * | 零次或多次出现。 | a* 匹配 ”, ‘a’, ‘aa’ |
+ | 一次或多次出现。 | a+ 匹配 ‘a’, ‘aa’ | |
? | 零次或一次出现。 | colou?r 匹配 ‘color’ 和 ‘colour’ | |
{n} | 恰好 n 次出现。 | \d{4} 匹配一个 4 位数字 |
注意: 在 C# 字符串中,反斜杠 \ 是一个转义字符。要为你的正则表达式模式编写一个字面量反斜杠,你必须将其转义("\\d")或使用逐字字符串字面量(@"\d")。强烈建议使用逐字字符串语法以提高可读性。
Regex 类
Section titled “Regex 类”Regex 类的静态方法方便用于一次性使用,而如果你多次使用相同的模式,创建 Regex 实例则更高效。
Regex.IsMatch(input, pattern):如果模式在输入中找到匹配项,则返回true。Regex.Match(input, pattern):返回找到的第一个匹配项的Match对象。Regex.Matches(input, pattern):返回所有找到的匹配项的MatchCollection。Regex.Replace(input, pattern, replacement):将所有匹配项替换为替换字符串。Regex.Split(input, pattern):在每个匹配项处拆分输入字符串。
示例:验证电话号码
Section titled “示例:验证电话号码”让我们检查一个字符串是否符合简单的北美电话号码格式,例如 XXX-XXX-XXXX。
using System.Text.RegularExpressions;
string[] phoneNumbers = { "123-456-7890", "123 456 7890", "999-999-999" };string pattern = @"^\d{3}-\d{3}-\d{4}$";
foreach (var number in phoneNumbers){ bool isValid = Regex.IsMatch(number, pattern); Console.WriteLine($"'{number}' is a valid format: {isValid}");}
// 输出:// '{number}' 是有效格式:True// '{number}' 是有效格式:False// '{number}' 是有效格式:False示例:从文本中提取所有标签
Section titled “示例:从文本中提取所有标签”这里,我们将实例化一个 Regex 对象来查找所有以 # 开头的单词。
using System.Text.RegularExpressions;
string text = "Check out this #cool #feature in #CSharp! It's #awesome.";string pattern = @"#\w+";
Regex hashtagRegex = new Regex(pattern);MatchCollection matches = hashtagRegex.Matches(text);
Console.WriteLine($"Found {matches.Count} hashtags:");foreach (Match match in matches){ Console.WriteLine(match.Value);}
// 输出:// 找到 3 个标签:// #cool// #feature// #CSharp// #awesome示例:替换多个空格
Section titled “示例:替换多个空格”Replace 方法对于清理用户输入很有用。
using System.Text.RegularExpressions;
string messyInput = "This string has too much space.";string pattern = @"\s+";string replacement = " ";
string cleanOutput = Regex.Replace(messyInput, pattern, replacement);
Console.WriteLine($"Original: '{messyInput}'");Console.WriteLine($"Cleaned: '{cleanOutput}'");
// 输出:// 原始: 'This string has too much space.'// 清理后: 'This string has too much space.'现代 C# 与高性能 Regex
Section titled “现代 C# 与高性能 Regex”对于性能敏感型应用程序,.NET 7 引入了 Regex 源生成器。通过使用 [GeneratedRegex] 特性装饰分部方法,编译器会在构建时预处理你的模式,生成高度优化的代码。这避免了运行时解析模式的开销,是现代 C# 推荐的方法。
示例:使用 Regex 源生成器
Section titled “示例:使用 Regex 源生成器”要使用此功能,你的包含类必须是 partial。
using System.Text.RegularExpressions;
public partial class MyValidators{ // 源生成器在编译时创建此方法的实现。 [GeneratedRegex(@"^[\w-\.]+@([\w-]+\.)+[\w-]{2,4}$", RegexOptions.IgnoreCase)] public static partial Regex EmailRegex();}
// --- 用法 ---string email = "test@example.com";if (MyValidators.EmailRegex().IsMatch(email)){ Console.WriteLine("Email is valid.");} else { Console.WriteLine("Email is invalid.");}
// 注意:提供的电子邮件正则表达式仅用于演示。// 完美的电子邮件验证是出了名的复杂。使用源生成器,你可以获得编译并缓存的 Regex 实例的性能,同时保持静态方法调用的简洁语法。