Python Pandas - 注意事项与陷阱
Pandas - 注意事项和常见陷阱
Section titled “Pandas - 注意事项和常见陷阱”虽然 Pandas 功能强大,但存在一些常见的陷阱(‘gotchas’)和注意事项(‘caveats’),用户(特别是初学者)应该了解这些,以避免意外的结果或错误。
Series/DataFrame 的真值(Truth Value)具有歧义性
Section titled “Series/DataFrame 的真值(Truth Value)具有歧义性”尝试在布尔上下文(例如 if 语句或 Python 的 and、or、not 运算符)中直接使用 Pandas 的 Series 或 DataFrame 会引发 ValueError。这是因为 Pandas 遵循 NumPy 的约定:对于一个多元素对象,不清楚它应该评估为 True(例如,因为它不为空,或者包含至少一个 True)还是 False(例如,因为它包含至少一个 False),因此真值是模糊的。
考虑以下示例:
import pandas as pd
s = pd.Series([False, True, False])
# 这将引发错误:# if s:# print('This will not print')
try: if s: passexcept ValueError as e: print(f"Caught expected error: {e}")输出:
Caught expected error: The truth value of a Series is ambiguous. Use a.empty, a.bool(), a.item(), a.any() or a.all().错误消息提示了明确检查你预期条件的正确方法:
s.empty: 检查Series/DataFrame是否为空(如果包含零个元素则为True)。s.bool(): 仅适用于Series/DataFrame包含恰好一个元素的情况;返回该元素的布尔值。否则引发ValueError。s.item(): 类似于.bool(),提取单个元素。s.any(): 如果Series/DataFrame中 至少一个 元素为真(评估为 True),则返回True。s.all(): 如果Series/DataFrame中 所有 元素都为真,则返回True。
import pandas as pd
s = pd.Series([False, True, False])
if s.any(): print("Series 中至少包含一个 True 值。")
if not s.all(): # 检查是否并非所有元素都为 True print("Series 不包含所有 True 值。")
s_single = pd.Series([True])if s_single.bool(): # .bool() 适用于单元素 Series print("单元素 Series 为 True。")输出:
Series 中至少包含一个 True 值。Series 不包含所有 True 值。单元素 Series 为 True。使用位运算符(Bitwise Boolean Operators) (&, |, ~)
Section titled “使用位运算符(Bitwise Boolean Operators) (&, |, ~)”当对 Series/DataFrame 进行逐元素布尔条件组合时,请使用位运算符 & (AND)、| (OR) 和 ~ (NOT)。标准的 Python 逻辑运算符 and、or、not 作用于 整个 对象的模糊真值,从而导致前面讨论的 ValueError。
比较运算符如 ==, !=, >, < 会返回布尔型的 Series,这通常正是你进行筛选(filtering)时所需要的。
import pandas as pd
s = pd.Series(range(5))print("原始 Series:")print(s)
# 使用位运算符和括号(控制优先级)进行正确筛选filtered = s[(s > 1) & (s < 4)] # 选择大于 1 且小于 4 的元素print("\n筛选后的 Series (s > 1) & (s < 4):")print(filtered)
# 使用 == 的示例print("\n布尔型 Series (s == 4):")print(s == 4)输出:
原始 Series:0 01 12 23 34 4dtype: int64
筛选后的 Series (s > 1) & (s < 4):2 23 3dtype: int64
布尔型 Series (s == 4):0 False1 False2 False3 False4 Truedtype: bool用于成员检查的 isin() 操作
Section titled “用于成员检查的 isin() 操作”isin() 方法提供了一种高效的方式来检查 Series 的每个元素是否包含在给定序列(list、set 等)中。它返回一个布尔型的 Series。
import pandas as pd
s = pd.Series(list('abcde'))allowed_values = ['a', 'c', 'e', 'g']
mask = s.isin(allowed_values)print(f"原始 Series:\n{s}")print(f"\n允许的值:{allowed_values}")print(f"\n布尔掩码 (isin):")print(mask)
print(f"\n使用 isin 筛选后的 Series:")print(s[mask])输出:
原始 Series:0 a1 b2 c3 d4 edtype: object
允许的值: ['a', 'c', 'e', 'g']
布尔掩码 (isin):0 True1 False2 True3 False4 Truedtype: bool
使用 isin 筛选后的 Series:0 a2 c4 edtype: object索引陷阱:.loc、.iloc 和整型索引
Section titled “索引陷阱:.loc、.iloc 和整型索引”对 DataFrame 进行索引时,尤其是那些使用基于整数的索引的 DataFrame,经常会出现一个常见的困惑点。Pandas 之前有一个索引器 .ix,它试图智能地处理基于标签和基于位置的索引,但这通常会产生歧义,现在已被 弃用并移除。
现代 Pandas 依赖于两个明确的索引器:
.loc[]: 基于标签 的索引。根据索引标签和列名选择数据。.iloc[]: 基于整数位置 的索引。根据整数位置(类似于 Python 列表切片)选择数据,从 0 开始。
潜在的陷阱在你拥有整型索引时出现。考虑以下示例:
import pandas as pd
# 带有整型索引标签的 DataFrame (0, 1, 2...)df = pd.DataFrame({'A': [10, 20, 30], 'B': [40, 50, 60]})print("带有整型索引的 DataFrame:")print(df)
# 将 .loc 与整型标签一起使用print("\n使用 .loc[0]:(选择标签为 0 的行)")print(df.loc[0])
# 将 .iloc 与整数位置一起使用print("\n使用 .iloc[0]:(选择位置为 0 的行)")print(df.iloc[0])输出:
带有整型索引的 DataFrame: A B0 10 401 20 502 30 60
使用 .loc[0]:(选择标签为 0 的行)A 10B 40Name: 0, dtype: int64
使用 .iloc[0]:(选择位置为 0 的行)A 10B 40Name: 0, dtype: int64在这种情况下,当标签恰好与位置匹配(0, 1, 2)时,.loc[0] 和 .iloc[0] 返回相同的结果。然而,情况并非总是如此!
现在,考虑一个 DataFrame,其整型索引标签 不 是顺序的或不从 0 开始:
import pandas as pd
# 带有非顺序整型索引标签的 DataFramedf_nonseq = pd.DataFrame({'A': [10, 20, 30], 'B': [40, 50, 60]}, index=[5, 2, 8])print("\n带有非顺序整型索引的 DataFrame:")print(df_nonseq)
# 将 .loc 与整型标签一起使用print("\n使用 .loc[2]:(选择标签为 2 的行)")print(df_nonseq.loc[2])
# 将 .iloc 与整数位置一起使用print("\n使用 .iloc[2]:(选择位置为 2 的行 - 第三行)")print(df_nonseq.iloc[2])
# 尝试使用不存在的标签调用 .loc 会引发 KeyError# print(df_nonseq.loc[0]) # 引发 KeyError
# 尝试使用越界位置调用 .iloc 会引发 IndexError# print(df_nonseq.iloc[3]) # 引发 IndexError输出:
带有非顺序整型索引的 DataFrame: A B5 10 402 20 508 30 60
使用 .loc[2]:(选择标签为 2 的行)A 20B 50Name: 2, dtype: int64
使用 .iloc[2]:(选择位置为 2 的行 - 第三行)A 30B 60Name: 8, dtype: int64关键要点: 始终使用 .loc[] 进行基于标签的选择,使用 .iloc[] 进行基于位置的选择,以避免歧义,尤其是在处理整型索引时。避免使用已弃用的 .ix[] 索引器。
SettingWithCopyWarning
Section titled “SettingWithCopyWarning”另一个常见的陷阱是 SettingWithCopyWarning。当你尝试修改一个可能是另一个对象的 视图(view) 而非 副本(copy) 的 DataFrame 或 Series 时,就会出现此警告。执行修改可能会无意中改变原始对象,或者可能没有任何效果,导致不可预测的行为。
这通常发生在链式索引(chained indexing)期间(例如 df[col1][row_indexer] = value)。Pandas 无法总是保证 df[col1] 返回的是视图还是副本。
可能触发此警告的示例:
import pandas as pd
df = pd.DataFrame({'A': [1, 2, 3], 'B': [4, 5, 6]})
# 潜在的 SettingWithCopyWarning 场景subset = df[df['A'] > 1]try: # 尝试修改 subset 可能会发出警告,因为 'subset' 可能是视图或副本 subset['B'] = 99except pd.errors.SettingWithCopyWarning: print("捕获到 SettingWithCopyWarning(行为取决于 Pandas 版本/上下文)")
print("\n原始 DataFrame(可能未改变):")print(df)print("\n子集 DataFrame(修改可能反映或不反映):")print(subset)为了避免这种歧义并确保正确应用修改,请使用 .loc[] 根据标签/条件设置值:
import pandas as pd
df = pd.DataFrame({'A': [1, 2, 3], 'B': [4, 5, 6]})
# 使用 .loc 设置值的推荐方法df.loc[df['A'] > 1, 'B'] = 99 # 直接在 df 上修改 'A' > 1 的行的 'B' 列
print("\n使用 .loc 修改后的 DataFrame:")print(df)输出:
使用 .loc 修改后的 DataFrame: A B0 1 41 2 992 3 99同时使用 .loc 进行行/列选择和值赋值是避免 SettingWithCopyWarning 并确保预期修改的最安全方法。