Python Pandas - 迭代
现代 Pandas - 迭代技巧
Section titled “现代 Pandas - 迭代技巧”直接迭代 Pandas 对象时,行为会因对象类型而异。迭代 Series 会产生其值,类似于 NumPy 数组。迭代 DataFrame,遵循类似字典的约定,会产生其列标签(keys)。
总结来说,基本迭代(for item in object)会产生:
- Series: 值 (values)
- DataFrame: 列标签 (column labels)
迭代 DataFrame(列标签)
Section titled “迭代 DataFrame(列标签)”直接迭代 DataFrame 会提供列名。这对于快速访问列标签很有用,但无法访问数据本身。
import pandas as pdimport numpy as np
# Use a modern random number generator if possible, though np.random is still commonrng = np.random.default_rng(42) # 设置种子以保证结果可复现
N = 5 # N 设置小一些以便输出更清晰df = pd.DataFrame({ 'A': pd.date_range(start='2023-01-01', periods=N, freq='D'), 'x': np.linspace(0, stop=N - 1, num=N), 'y': rng.random(N), 'C': rng.choice(['Low', 'Medium', 'High'], N).tolist(), 'D': rng.normal(100, 10, size=N).tolist()})
print('DataFrame:\n', df)print('\nIterating over DataFrame (yields column labels):')for col_label in df: print(col_label)示例输出:
DataFrame: A x y C D0 2023-01-01 0.0 0.773956 Low 105.6505471 2023-01-02 1.0 0.438878 High 94.4857182 2023-01-03 2.0 0.858598 Low 104.1124783 2023-01-04 3.0 0.697368 Medium 101.2107744 2023-01-05 4.0 0.094177 Medium 104.057573
Iterating over DataFrame (yields column labels):AxyCD虽然直接迭代很简单,但通常不如矢量化操作 (vectorized operations) 高效。如果确实需要进行逐行操作 (row-wise operations),Pandas 提供了专门的方法:
.items(): 迭代 DataFrame 的列,生成(label, Series)对。.iterrows(): 迭代 DataFrame 的行,生成(index, Series)对。**注意:**通常效率低下,并可能改变数据类型 (dtypes)。.itertuples(): 迭代 DataFrame 的行,生成 namedtuple。通常比iterrows更高效,并且更好地保留数据类型。
最佳实践: 尽可能避免迭代。使用矢量化操作(直接对 Series/DataFrames 应用函数)以获得更好的性能。如果迭代不可避免,优先选择 itertuples() 而不是 iterrows()。
迭代列:.items()
Section titled “迭代列:.items()”.items() 方法迭代 DataFrame 的列,生成每个列的标签以及作为 Series 对象的列数据。当需要单独对每个列 Series 执行操作时,此方法很有用。
import pandas as pdimport numpy as np
rng = np.random.default_rng(42)df = pd.DataFrame(rng.standard_normal((4, 3)), columns=['col1', 'col2', 'col3'])
print('DataFrame:\n', df)print('\nIterating with .items():')for label, content_series in df.items(): print(f'\nColumn Label: {label}') print(content_series)示例输出:
DataFrame: col1 col2 col30 0.496714 -0.138264 0.6476891 1.523030 -0.234153 -0.2341372 1.579213 0.767435 -0.4694743 0.542560 -0.463418 -0.465730
Iterating with .items():
Column Label: col10 0.4967141 1.5230302 1.5792133 0.542560Name: col1, dtype: float64
Column Label: col20 -0.1382641 -0.2341532 0.7674353 -0.463418Name: col2, dtype: float64
Column Label: col30 0.6476891 -0.2341372 -0.4694743 -0.465730Name: col3, dtype: float64请注意每个列都作为 (label, Series) 元组 (tuple) 生成。
迭代行:.iterrows()(谨慎使用)
Section titled “迭代行:.iterrows()(谨慎使用)”.iterrows() 方法将每一行作为包含行索引 (index) 和代表该行数据的 Series 的元组 (tuple) 生成。警告: 对于大型 DataFrame 而言,此方法非常慢,并且由于每行都作为单个 Series 返回,可能意外地转换行数据类型(例如,将整数转换为浮点数)。通常应避免使用此方法,优先选择 vectorized operations 或 itertuples()。
import pandas as pdimport numpy as np
rng = np.random.default_rng(43)df = pd.DataFrame(rng.standard_normal((3, 3)), columns=['col1', 'col2', 'col3'])
print('DataFrame:\n', df)print('\nIterating with .iterrows():')for index, row_series in df.iterrows(): print(f'\nIndex: {index}') print(row_series)示例输出:
DataFrame: col1 col2 col30 -0.535952 -0.875490 -0.2690011 -0.087176 0.125819 0.7471782 -1.255743 -0.977989 -1.034384
Iterating with .iterrows():
Index: 0col1 -0.535952col2 -0.875490col3 -0.269001Name: 0, dtype: float64
Index: 1col1 -0.087176col2 0.125819col3 0.747178Name: 1, dtype: float64
Index: 2col1 -1.255743col2 -0.977989col3 -1.034384Name: 2, dtype: float64重要提示: 因为 iterrows() 将每一行作为 Series 返回,如果行中有混合数据类型,它不会保留原有的数据类型。例如,一个整数列在返回的 Series 中可能会被转换为浮点数 (float)。这是在需要行迭代时优先选择 itertuples() 的主要原因之一。
高效迭代行:.itertuples()
Section titled “高效迭代行:.itertuples()”.itertuples() 是在需要行迭代时推荐的方法。它将每一行作为轻量级的 namedtuple 生成。第一个元素是行的索引 (index),后续元素是行的值。此方法显著快于 iterrows(),并且通常会保留数据类型。
import pandas as pdimport numpy as np
rng = np.random.default_rng(44)df = pd.DataFrame(rng.standard_normal((3, 3)), columns=['col1', 'col2', 'col3'])
print('DataFrame:\n', df)print('\nIterating with .itertuples():')for row_tuple in df.itertuples(): # index=True 是默认值 print(row_tuple)
print('\nIterating with .itertuples(index=False, name="MyRow"):')for row_tuple in df.itertuples(index=False, name='MyRow'): # 自定义名称,不包含索引 print(row_tuple)示例输出:
DataFrame: col1 col2 col30 -0.846179 0.058019 -0.6162891 0.277575 -0.447319 -1.0491082 0.223398 -0.204669 -0.358949
Iterating with .itertuples():Pandas(Index=0, col1=-0.8461786296694368, col2=0.05801868623661116, col3=-0.6162891895287465)Pandas(Index=1, col1=0.277575258121937, col2=-0.4473188014249641, col3=-1.0491081632783335)Pandas(Index=2, col1=0.22339849625659426, col2=-0.20466887033086502, col3=-0.3589488659138892)
Iterating with .itertuples(index=False, name='MyRow'):MyRow(col1=-0.8461786296694368, col2=0.05801868623661116, col3=-0.6162891895287465)MyRow(col1=0.277575258121937, col2=-0.4473188014249641, col3=-1.0491081632783335)MyRow(col1=0.22339849625659426, col2=-0.20466887033086502, col3=-0.3589488659138892)可以通过属性名(例如 row_tuple.col1)或索引(例如 row_tuple[1],如果 index=True)访问元素。index=False 选项会从 tuple 中排除索引,而 name 参数允许自定义 namedtuple 的类型名称。
重要提示: 永远不要在迭代 DataFrame 的循环内部修改该 DataFrame(使用 .iterrows() 或 .itertuples())。这些方法通常返回的是副本(或视图,行为可能有所不同),而不是原始数据的直接引用。在循环内部对 row_series 或 row_tuple 所做的修改很可能不会改变原始 DataFrame。
import pandas as pdimport numpy as np
rng = np.random.default_rng(45)df = pd.DataFrame(rng.standard_normal((3, 3)), columns=['col1', 'col2', 'col3'])
print('Original DataFrame:\n', df)
for index, row_series in df.iterrows(): # 尝试修改 - 这很可能只会影响副本 row_series['col1'] = 1000
print('\nDataFrame after attempting modification during iterrows():')print(df) # 观察:原始 df 没有反映出任何变化
# 正确的修改方式需要重新赋值或使用 .loc 等其他方法# 示例(不使用迭代):df.loc[df['col1'] < 0, 'col1'] = 0示例输出:
Original DataFrame: col1 col2 col30 -0.815408 0.797837 1.2328031 -0.228229 0.209149 -1.1019922 -0.375991 -1.043386 -0.826451
DataFrame after attempting modification during iterrows(): col1 col2 col30 -0.815408 0.797837 1.2328031 -0.228229 0.209149 -1.1019922 -0.375991 -1.043386 -0.826451如果需要根据条件修改数据,请使用 .loc 或 .iloc 进行矢量化赋值,或谨慎使用 .apply() 方法。