Skip to content

C# - 正则表达式

正则表达式(或 regex)是定义搜索模式的字符序列。在 C# 中,System.Text.RegularExpressions 命名空间提供了一个强大的引擎,用于字符串的模式匹配、验证、替换和拆分。

Regex 模式由字面量(如 a、b、1)和具有特殊含义的元字符构建而成。

类别构造描述示例
字符类.匹配除换行符外的任意单个字符。h.t 匹配 ‘hat’, ‘hot’, ‘h&t’
\d匹配任意数字 (0-9)。\D 匹配非数字字符。\d{3} 匹配 ‘123’
\w匹配任意单词字符 (a-z, A-Z, 0-9, _)。\W 匹配非单词字符。\w+ 匹配 ‘hello_123’
\s匹配任意空白字符。\S 匹配非空白字符。hello\sworld
[...]匹配集合中的任意单个字符。[aeiou] 匹配任意元音
锚点^匹配字符串的开头。^Hello
$匹配字符串的末尾。world$
\b匹配单词边界。\bcat\b 匹配 ‘cat’, 而非 ‘category’
量词*零次或多次出现。a* 匹配 ”, ‘a’, ‘aa’
+一次或多次出现。a+ 匹配 ‘a’, ‘aa’
?零次或一次出现。colou?r 匹配 ‘color’ 和 ‘colour’
{n}恰好 n 次出现。\d{4} 匹配一个 4 位数字

注意: 在 C# 字符串中,反斜杠 \ 是一个转义字符。要为你的正则表达式模式编写一个字面量反斜杠,你必须将其转义("\\d")或使用逐字字符串字面量(@"\d")。强烈建议使用逐字字符串语法以提高可读性。

Regex 类的静态方法方便用于一次性使用,而如果你多次使用相同的模式,创建 Regex 实例则更高效。

  • Regex.IsMatch(input, pattern):如果模式在输入中找到匹配项,则返回 true。
  • Regex.Match(input, pattern):返回找到的第一个匹配项的 Match 对象。
  • Regex.Matches(input, pattern):返回所有找到的匹配项的 MatchCollection。
  • Regex.Replace(input, pattern, replacement):将所有匹配项替换为替换字符串。
  • Regex.Split(input, pattern):在每个匹配项处拆分输入字符串。

让我们检查一个字符串是否符合简单的北美电话号码格式,例如 XXX-XXX-XXXX。

using System.Text.RegularExpressions;
string[] phoneNumbers = { "123-456-7890", "123 456 7890", "999-999-999" };
string pattern = @"^\d{3}-\d{3}-\d{4}$";
foreach (var number in phoneNumbers)
{
bool isValid = Regex.IsMatch(number, pattern);
Console.WriteLine($"'{number}' is a valid format: {isValid}");
}
// 输出:
// '{number}' 是有效格式:True
// '{number}' 是有效格式:False
// '{number}' 是有效格式:False

这里,我们将实例化一个 Regex 对象来查找所有以 # 开头的单词。

using System.Text.RegularExpressions;
string text = "Check out this #cool #feature in #CSharp! It's #awesome.";
string pattern = @"#\w+";
Regex hashtagRegex = new Regex(pattern);
MatchCollection matches = hashtagRegex.Matches(text);
Console.WriteLine($"Found {matches.Count} hashtags:");
foreach (Match match in matches)
{
Console.WriteLine(match.Value);
}
// 输出:
// 找到 3 个标签:
// #cool
// #feature
// #CSharp
// #awesome

Replace 方法对于清理用户输入很有用。

using System.Text.RegularExpressions;
string messyInput = "This string has too much space.";
string pattern = @"\s+";
string replacement = " ";
string cleanOutput = Regex.Replace(messyInput, pattern, replacement);
Console.WriteLine($"Original: '{messyInput}'");
Console.WriteLine($"Cleaned: '{cleanOutput}'");
// 输出:
// 原始: 'This string has too much space.'
// 清理后: 'This string has too much space.'

对于性能敏感型应用程序,.NET 7 引入了 Regex 源生成器。通过使用 [GeneratedRegex] 特性装饰分部方法,编译器会在构建时预处理你的模式,生成高度优化的代码。这避免了运行时解析模式的开销,是现代 C# 推荐的方法。

要使用此功能,你的包含类必须是 partial。

using System.Text.RegularExpressions;
public partial class MyValidators
{
// 源生成器在编译时创建此方法的实现。
[GeneratedRegex(@"^[\w-\.]+@([\w-]+\.)+[\w-]{2,4}$", RegexOptions.IgnoreCase)]
public static partial Regex EmailRegex();
}
// --- 用法 ---
string email = "test@example.com";
if (MyValidators.EmailRegex().IsMatch(email))
{
Console.WriteLine("Email is valid.");
} else {
Console.WriteLine("Email is invalid.");
}
// 注意:提供的电子邮件正则表达式仅用于演示。
// 完美的电子邮件验证是出了名的复杂。

使用源生成器,你可以获得编译并缓存的 Regex 实例的性能,同时保持静态方法调用的简洁语法。