处理文本数据
Pandas - 处理文本数据
Section titled “Pandas - 处理文本数据”文本数据(字符串)无处不在,Pandas 提供了强大且高效的工具来处理存储在 Series 或 Index 对象中的字符串数据。这些工具通过 .str 访问器来访问。
使用 .str 访问器的一个主要优点是,它的方法会自动处理缺失值(NaN),对于涉及缺失值的任何操作,通常会返回 NaN,从而避免了错误。
大多数 .str 方法模仿标准的 Python 字符串方法(如 .lower()、.split()、.replace() 等),但它们是向量化的,这意味着它们可以高效地对整个 Series 进行操作,无需显式的 Python 循环。
让我们使用一个示例 Series 来探索一些常见的字符串操作:
import pandas as pdimport numpy as np
s = pd.Series([' Tom ', 'William Rick', 'JOHN', ' Alber@t', np.nan, '12345', 'Steve Smith ', 'Python & Pandas'])
print("Original Series:")print(s)原始 Series 输出:
0 Tom1 William Rick2 JOHN3 Alber@t4 NaN5 123456 Steve Smith7 Python & Pandasdtype: object通过 .str 访问器访问的常用字符串方法
Section titled “通过 .str 访问器访问的常用字符串方法”| 方法 | 说明 |
|---|---|
.str.lower() | 将字符串转换为小写。 |
.str.upper() | 将字符串转换为大写。 |
.str.len() | 计算每个字符串的长度。 |
.str.strip() | 移除开头和结尾的空白字符(包括换行符)。 |
.str.lstrip() | 移除开头的空白字符。 |
.str.rstrip() | 移除结尾的空白字符。 |
.str.split(pat=None, n=-1, expand=False) | 根据指定的分隔符/模式 (pat) 拆分每个字符串。默认返回一个列表 Series,如果 expand=True 则返回一个 DataFrame。 |
.str.cat(sep='') | 使用给定分隔符连接 Series/Index 中的字符串。 |
.str.contains(pat, case=True, regex=True) | 检查每个字符串中是否包含某个模式 (pat)。返回一个布尔 Series。 |
.str.startswith(pat) | 检查每个字符串是否以指定的模式 (pat) 开头。返回一个布尔 Series。 |
.str.endswith(pat) | 检查每个字符串是否以指定的模式 (pat) 结尾。返回一个布尔 Series。 |
.str.replace(pat, repl, n=-1, regex=True) | 在每个字符串中替换 pat 的出现项为 repl。 |
.str.repeat(repeats) | 将每个字符串重复 repeats 次。 |
.str.count(pat) | 计算 pat 在每个字符串中出现的次数。 |
.str.find(sub) | 返回子字符串 sub 在每个字符串中的最低索引;如果未找到则返回 -1。 |
.str.findall(pat) | 在每个字符串中查找模式 pat(可以是正则表达式)的所有出现项;返回一个列表 Series。 |
.str.get(i) | 提取每个字符串在位置 i 的元素(如果可索引)。 |
.str.slice(start=None, stop=None, step=None) | 从 Series/Index 中的每个元素切片子字符串。 |
.str.islower(),.str.isupper(),.str.isnumeric() 等 | 检查每个字符串的字符属性(例如,全部小写,全部大写,全部数字)。返回一个布尔 Series。 |
pd.get_dummies(s) 或 s.str.get_dummies(sep='|') | 将分类字符串数据转换为哑变量/指示变量(独热编码,One-Hot Encoding)。可以先按 sep 拆分字符串。 |
让我们将其中一些方法应用于示例 Series s。
# Case Conversionprint("\n--- Case Conversion ---")print("Lower Case:\n", s.str.lower())print("\nUpper Case:\n", s.str.upper())
# Lengthprint("\n--- Length ---")print("String Lengths:\n", s.str.len())
# Stripping Whitespaceprint("\n--- Stripping Whitespace ---")print("Stripped (both sides):\n", s.str.strip())print("\nLeft Stripped:\n", s.str.lstrip())print("\nRight Stripped:\n", s.str.rstrip())
# Splittingprint("\n--- Splitting ---")# Split by space, return listsprint("Split by space (into lists):\n", s.str.split(' '))# Split by space, expand into DataFrame columnsprint("\nSplit by space (expand=True):\n", s.str.split(' ', expand=True))
# Checking Contentprint("\n--- Checking Content ---")print("Contains 'Rick'?\n", s.str.contains('Rick'))print("\nStarts with 'P'?\n", s.str.startswith('P'))print("\nEnds with ' '?\n", s.str.endswith(' '))
# Replacingprint("\n--- Replacing ---")print("Replace '@' with '#':\n", s.str.replace('@', '#'))print("\nReplace ' ' with '_':\n", s.str.replace(' ', '_')) # Note: affects internal spaces too
# Find / Countprint("\n--- Find / Count ---")print("Find first 'e':\n", s.str.find('e'))print("\nCount occurrences of 'n':\n", s.str.count('n'))
# Character Type Checksprint("\n--- Character Type Checks ---")print("Is Lowercase?\n", s.str.islower())print("\nIs Uppercase?\n", s.str.isupper())print("\nIs Numeric?\n", s.str.isnumeric())
# Get Dummies (Example on a simpler series)s_cat = pd.Series(['A|B', 'B', 'A|C', 'A'])print("\n--- Get Dummies --- ")print("Categorical Series:\n", s_cat)print("\nDummy Variables (split by '|'):\n", s_cat.str.get_dummies(sep='|'))示例输出片段:
--- Case Conversion ---Lower Case: 0 tom1 william rick2 john3 alber@t4 NaN5 123456 steve smith7 python & pandasdtype: object
Upper Case: 0 TOM1 WILLIAM RICK2 JOHN3 ALBER@T4 NaN5 123456 STEVE SMITH7 PYTHON & PANDASdtype: object
--- Length ---String Lengths: 0 9.01 12.02 4.03 8.04 NaN5 5.06 12.07 15.0dtype: float64
--- Stripping Whitespace ---Stripped (both sides): 0 Tom1 William Rick2 JOHN3 Alber@t4 NaN5 123456 Steve Smith7 Python & Pandasdtype: object...--- Splitting ---Split by space (into lists): 0 [, , , Tom, , , ]1 [William, Rick]2 [JOHN]3 [, Alber@t]4 NaN5 [12345]6 [Steve, Smith, ]7 [Python, &, Pandas]dtype: object
Split by space (expand=True): 0 1 20 Tom1 William Rick None2 JOHN None None3 Alber@t None4 None None None5 12345 None None6 Steve Smith7 Python & Pandas...--- Checking Content ---Contains 'Rick'? 0 False1 True2 False3 False4 NaN5 False6 False7 Falsedtype: object...--- Replacing ---Replace '@' with '#': 0 Tom1 William Rick2 JOHN3 Alber#t4 NaN5 123456 Steve Smith7 Python & Pandasdtype: object...--- Find / Count ---Find first 'e': 0 -1.01 3.02 -1.03 4.04 NaN5 -1.06 3.07 -1.0dtype: float64
Count occurrences of 'n': 0 0.01 1.02 1.03 0.04 NaN5 0.06 0.07 2.0dtype: float64...--- Get Dummies ---Categorical Series: 0 A|B1 B2 A|C3 Adtype: object
Dummy Variables (split by '|'): A B C0 1 1 01 0 1 02 1 0 13 1 0 0这些由 .str 访问器提供的向量化字符串方法对于在 Pandas 中高效地清洗、准备和分析文本数据至关重要。