Skip to content

处理文本数据

文本数据(字符串)无处不在,Pandas 提供了强大且高效的工具来处理存储在 Series 或 Index 对象中的字符串数据。这些工具通过 .str 访问器来访问。

使用 .str 访问器的一个主要优点是,它的方法会自动处理缺失值(NaN),对于涉及缺失值的任何操作,通常会返回 NaN,从而避免了错误。

大多数 .str 方法模仿标准的 Python 字符串方法(如 .lower()、.split()、.replace() 等),但它们是向量化的,这意味着它们可以高效地对整个 Series 进行操作,无需显式的 Python 循环。

让我们使用一个示例 Series 来探索一些常见的字符串操作:

import pandas as pd
import numpy as np
s = pd.Series([' Tom ', 'William Rick', 'JOHN', ' Alber@t',
np.nan, '12345', 'Steve Smith ',
'Python & Pandas'])
print("Original Series:")
print(s)

原始 Series 输出:

0 Tom
1 William Rick
2 JOHN
3 Alber@t
4 NaN
5 12345
6 Steve Smith
7 Python & Pandas
dtype: object

通过 .str 访问器访问的常用字符串方法

Section titled “通过 .str 访问器访问的常用字符串方法”
方法说明
.str.lower()将字符串转换为小写。
.str.upper()将字符串转换为大写。
.str.len()计算每个字符串的长度。
.str.strip()移除开头和结尾的空白字符(包括换行符)。
.str.lstrip()移除开头的空白字符。
.str.rstrip()移除结尾的空白字符。
.str.split(pat=None, n=-1, expand=False)根据指定的分隔符/模式 (pat) 拆分每个字符串。默认返回一个列表 Series,如果 expand=True 则返回一个 DataFrame。
.str.cat(sep='')使用给定分隔符连接 Series/Index 中的字符串。
.str.contains(pat, case=True, regex=True)检查每个字符串中是否包含某个模式 (pat)。返回一个布尔 Series。
.str.startswith(pat)检查每个字符串是否以指定的模式 (pat) 开头。返回一个布尔 Series。
.str.endswith(pat)检查每个字符串是否以指定的模式 (pat) 结尾。返回一个布尔 Series。
.str.replace(pat, repl, n=-1, regex=True)在每个字符串中替换 pat 的出现项为 repl。
.str.repeat(repeats)将每个字符串重复 repeats 次。
.str.count(pat)计算 pat 在每个字符串中出现的次数。
.str.find(sub)返回子字符串 sub 在每个字符串中的最低索引;如果未找到则返回 -1。
.str.findall(pat)在每个字符串中查找模式 pat(可以是正则表达式)的所有出现项;返回一个列表 Series。
.str.get(i)提取每个字符串在位置 i 的元素(如果可索引)。
.str.slice(start=None, stop=None, step=None)从 Series/Index 中的每个元素切片子字符串。
.str.islower(),.str.isupper(),.str.isnumeric() 等检查每个字符串的字符属性(例如,全部小写,全部大写,全部数字)。返回一个布尔 Series。
pd.get_dummies(s) 或 s.str.get_dummies(sep='|')将分类字符串数据转换为哑变量/指示变量(独热编码,One-Hot Encoding)。可以先按 sep 拆分字符串。

让我们将其中一些方法应用于示例 Series s。

# Case Conversion
print("\n--- Case Conversion ---")
print("Lower Case:\n", s.str.lower())
print("\nUpper Case:\n", s.str.upper())
# Length
print("\n--- Length ---")
print("String Lengths:\n", s.str.len())
# Stripping Whitespace
print("\n--- Stripping Whitespace ---")
print("Stripped (both sides):\n", s.str.strip())
print("\nLeft Stripped:\n", s.str.lstrip())
print("\nRight Stripped:\n", s.str.rstrip())
# Splitting
print("\n--- Splitting ---")
# Split by space, return lists
print("Split by space (into lists):\n", s.str.split(' '))
# Split by space, expand into DataFrame columns
print("\nSplit by space (expand=True):\n", s.str.split(' ', expand=True))
# Checking Content
print("\n--- Checking Content ---")
print("Contains 'Rick'?\n", s.str.contains('Rick'))
print("\nStarts with 'P'?\n", s.str.startswith('P'))
print("\nEnds with ' '?\n", s.str.endswith(' '))
# Replacing
print("\n--- Replacing ---")
print("Replace '@' with '#':\n", s.str.replace('@', '#'))
print("\nReplace ' ' with '_':\n", s.str.replace(' ', '_')) # Note: affects internal spaces too
# Find / Count
print("\n--- Find / Count ---")
print("Find first 'e':\n", s.str.find('e'))
print("\nCount occurrences of 'n':\n", s.str.count('n'))
# Character Type Checks
print("\n--- Character Type Checks ---")
print("Is Lowercase?\n", s.str.islower())
print("\nIs Uppercase?\n", s.str.isupper())
print("\nIs Numeric?\n", s.str.isnumeric())
# Get Dummies (Example on a simpler series)
s_cat = pd.Series(['A|B', 'B', 'A|C', 'A'])
print("\n--- Get Dummies --- ")
print("Categorical Series:\n", s_cat)
print("\nDummy Variables (split by '|'):\n", s_cat.str.get_dummies(sep='|'))

示例输出片段:

--- Case Conversion ---
Lower Case:
0 tom
1 william rick
2 john
3 alber@t
4 NaN
5 12345
6 steve smith
7 python & pandas
dtype: object
Upper Case:
0 TOM
1 WILLIAM RICK
2 JOHN
3 ALBER@T
4 NaN
5 12345
6 STEVE SMITH
7 PYTHON & PANDAS
dtype: object
--- Length ---
String Lengths:
0 9.0
1 12.0
2 4.0
3 8.0
4 NaN
5 5.0
6 12.0
7 15.0
dtype: float64
--- Stripping Whitespace ---
Stripped (both sides):
0 Tom
1 William Rick
2 JOHN
3 Alber@t
4 NaN
5 12345
6 Steve Smith
7 Python & Pandas
dtype: object
...
--- Splitting ---
Split by space (into lists):
0 [, , , Tom, , , ]
1 [William, Rick]
2 [JOHN]
3 [, Alber@t]
4 NaN
5 [12345]
6 [Steve, Smith, ]
7 [Python, &, Pandas]
dtype: object
Split by space (expand=True):
0 1 2
0 Tom
1 William Rick None
2 JOHN None None
3 Alber@t None
4 None None None
5 12345 None None
6 Steve Smith
7 Python & Pandas
...
--- Checking Content ---
Contains 'Rick'?
0 False
1 True
2 False
3 False
4 NaN
5 False
6 False
7 False
dtype: object
...
--- Replacing ---
Replace '@' with '#':
0 Tom
1 William Rick
2 JOHN
3 Alber#t
4 NaN
5 12345
6 Steve Smith
7 Python & Pandas
dtype: object
...
--- Find / Count ---
Find first 'e':
0 -1.0
1 3.0
2 -1.0
3 4.0
4 NaN
5 -1.0
6 3.0
7 -1.0
dtype: float64
Count occurrences of 'n':
0 0.0
1 1.0
2 1.0
3 0.0
4 NaN
5 0.0
6 0.0
7 2.0
dtype: float64
...
--- Get Dummies ---
Categorical Series:
0 A|B
1 B
2 A|C
3 A
dtype: object
Dummy Variables (split by '|'):
A B C
0 1 1 0
1 0 1 0
2 1 0 1
3 1 0 0

这些由 .str 访问器提供的向量化字符串方法对于在 Pandas 中高效地清洗、准备和分析文本数据至关重要。