Skip to content

Python Pandas - 迭代

直接迭代 Pandas 对象时,行为会因对象类型而异。迭代 Series 会产生其值,类似于 NumPy 数组。迭代 DataFrame,遵循类似字典的约定,会产生其列标签(keys)。

总结来说,基本迭代(for item in object)会产生:

  • Series: 值 (values)
  • DataFrame: 列标签 (column labels)

直接迭代 DataFrame 会提供列名。这对于快速访问列标签很有用,但无法访问数据本身。

import pandas as pd
import numpy as np
# Use a modern random number generator if possible, though np.random is still common
rng = np.random.default_rng(42) # 设置种子以保证结果可复现
N = 5 # N 设置小一些以便输出更清晰
df = pd.DataFrame({
'A': pd.date_range(start='2023-01-01', periods=N, freq='D'),
'x': np.linspace(0, stop=N - 1, num=N),
'y': rng.random(N),
'C': rng.choice(['Low', 'Medium', 'High'], N).tolist(),
'D': rng.normal(100, 10, size=N).tolist()
})
print('DataFrame:\n', df)
print('\nIterating over DataFrame (yields column labels):')
for col_label in df:
print(col_label)

示例输出:

DataFrame:
A x y C D
0 2023-01-01 0.0 0.773956 Low 105.650547
1 2023-01-02 1.0 0.438878 High 94.485718
2 2023-01-03 2.0 0.858598 Low 104.112478
3 2023-01-04 3.0 0.697368 Medium 101.210774
4 2023-01-05 4.0 0.094177 Medium 104.057573
Iterating over DataFrame (yields column labels):
A
x
y
C
D

虽然直接迭代很简单,但通常不如矢量化操作 (vectorized operations) 高效。如果确实需要进行逐行操作 (row-wise operations),Pandas 提供了专门的方法:

  • .items(): 迭代 DataFrame 的列,生成 (label, Series) 对。
  • .iterrows(): 迭代 DataFrame 的行,生成 (index, Series) 对。**注意:**通常效率低下,并可能改变数据类型 (dtypes)。
  • .itertuples(): 迭代 DataFrame 的行,生成 namedtuple。通常比 iterrows 更高效,并且更好地保留数据类型。

最佳实践: 尽可能避免迭代。使用矢量化操作(直接对 Series/DataFrames 应用函数)以获得更好的性能。如果迭代不可避免,优先选择 itertuples() 而不是 iterrows()。

.items() 方法迭代 DataFrame 的列,生成每个列的标签以及作为 Series 对象的列数据。当需要单独对每个列 Series 执行操作时,此方法很有用。

import pandas as pd
import numpy as np
rng = np.random.default_rng(42)
df = pd.DataFrame(rng.standard_normal((4, 3)), columns=['col1', 'col2', 'col3'])
print('DataFrame:\n', df)
print('\nIterating with .items():')
for label, content_series in df.items():
print(f'\nColumn Label: {label}')
print(content_series)

示例输出:

DataFrame:
col1 col2 col3
0 0.496714 -0.138264 0.647689
1 1.523030 -0.234153 -0.234137
2 1.579213 0.767435 -0.469474
3 0.542560 -0.463418 -0.465730
Iterating with .items():
Column Label: col1
0 0.496714
1 1.523030
2 1.579213
3 0.542560
Name: col1, dtype: float64
Column Label: col2
0 -0.138264
1 -0.234153
2 0.767435
3 -0.463418
Name: col2, dtype: float64
Column Label: col3
0 0.647689
1 -0.234137
2 -0.469474
3 -0.465730
Name: col3, dtype: float64

请注意每个列都作为 (label, Series) 元组 (tuple) 生成。

.iterrows() 方法将每一行作为包含行索引 (index) 和代表该行数据的 Series 的元组 (tuple) 生成。警告: 对于大型 DataFrame 而言,此方法非常慢,并且由于每行都作为单个 Series 返回,可能意外地转换行数据类型(例如,将整数转换为浮点数)。通常应避免使用此方法,优先选择 vectorized operations 或 itertuples()。

import pandas as pd
import numpy as np
rng = np.random.default_rng(43)
df = pd.DataFrame(rng.standard_normal((3, 3)), columns=['col1', 'col2', 'col3'])
print('DataFrame:\n', df)
print('\nIterating with .iterrows():')
for index, row_series in df.iterrows():
print(f'\nIndex: {index}')
print(row_series)

示例输出:

DataFrame:
col1 col2 col3
0 -0.535952 -0.875490 -0.269001
1 -0.087176 0.125819 0.747178
2 -1.255743 -0.977989 -1.034384
Iterating with .iterrows():
Index: 0
col1 -0.535952
col2 -0.875490
col3 -0.269001
Name: 0, dtype: float64
Index: 1
col1 -0.087176
col2 0.125819
col3 0.747178
Name: 1, dtype: float64
Index: 2
col1 -1.255743
col2 -0.977989
col3 -1.034384
Name: 2, dtype: float64

重要提示: 因为 iterrows() 将每一行作为 Series 返回,如果行中有混合数据类型,它不会保留原有的数据类型。例如,一个整数列在返回的 Series 中可能会被转换为浮点数 (float)。这是在需要行迭代时优先选择 itertuples() 的主要原因之一。

.itertuples() 是在需要行迭代时推荐的方法。它将每一行作为轻量级的 namedtuple 生成。第一个元素是行的索引 (index),后续元素是行的值。此方法显著快于 iterrows(),并且通常会保留数据类型。

import pandas as pd
import numpy as np
rng = np.random.default_rng(44)
df = pd.DataFrame(rng.standard_normal((3, 3)), columns=['col1', 'col2', 'col3'])
print('DataFrame:\n', df)
print('\nIterating with .itertuples():')
for row_tuple in df.itertuples(): # index=True 是默认值
print(row_tuple)
print('\nIterating with .itertuples(index=False, name="MyRow"):')
for row_tuple in df.itertuples(index=False, name='MyRow'): # 自定义名称,不包含索引
print(row_tuple)

示例输出:

DataFrame:
col1 col2 col3
0 -0.846179 0.058019 -0.616289
1 0.277575 -0.447319 -1.049108
2 0.223398 -0.204669 -0.358949
Iterating with .itertuples():
Pandas(Index=0, col1=-0.8461786296694368, col2=0.05801868623661116, col3=-0.6162891895287465)
Pandas(Index=1, col1=0.277575258121937, col2=-0.4473188014249641, col3=-1.0491081632783335)
Pandas(Index=2, col1=0.22339849625659426, col2=-0.20466887033086502, col3=-0.3589488659138892)
Iterating with .itertuples(index=False, name='MyRow'):
MyRow(col1=-0.8461786296694368, col2=0.05801868623661116, col3=-0.6162891895287465)
MyRow(col1=0.277575258121937, col2=-0.4473188014249641, col3=-1.0491081632783335)
MyRow(col1=0.22339849625659426, col2=-0.20466887033086502, col3=-0.3589488659138892)

可以通过属性名(例如 row_tuple.col1)或索引(例如 row_tuple[1],如果 index=True)访问元素。index=False 选项会从 tuple 中排除索引,而 name 参数允许自定义 namedtuple 的类型名称。

重要提示: 永远不要在迭代 DataFrame 的循环内部修改该 DataFrame(使用 .iterrows() 或 .itertuples())。这些方法通常返回的是副本(或视图,行为可能有所不同),而不是原始数据的直接引用。在循环内部对 row_series 或 row_tuple 所做的修改很可能不会改变原始 DataFrame。

import pandas as pd
import numpy as np
rng = np.random.default_rng(45)
df = pd.DataFrame(rng.standard_normal((3, 3)), columns=['col1', 'col2', 'col3'])
print('Original DataFrame:\n', df)
for index, row_series in df.iterrows():
# 尝试修改 - 这很可能只会影响副本
row_series['col1'] = 1000
print('\nDataFrame after attempting modification during iterrows():')
print(df) # 观察:原始 df 没有反映出任何变化
# 正确的修改方式需要重新赋值或使用 .loc 等其他方法
# 示例(不使用迭代):df.loc[df['col1'] < 0, 'col1'] = 0

示例输出:

Original DataFrame:
col1 col2 col3
0 -0.815408 0.797837 1.232803
1 -0.228229 0.209149 -1.101992
2 -0.375991 -1.043386 -0.826451
DataFrame after attempting modification during iterrows():
col1 col2 col3
0 -0.815408 0.797837 1.232803
1 -0.228229 0.209149 -1.101992
2 -0.375991 -1.043386 -0.826451

如果需要根据条件修改数据,请使用 .loc 或 .iloc 进行矢量化赋值,或谨慎使用 .apply() 方法。