Skip to content

函数应用

Pandas 允许你将自定义函数或来自其他库的函数应用于其对象(Series、DataFrame)。理解如何有效地应用函数是充分利用 Pandas 进行复杂数据处理的关键。函数应用有三种主要方法:

  • DataFrame 级别: pipe() - 用于链接预期接收 DataFrame 或 Series 作为输入的函数。
  • 行/列级别: apply() - 用于沿着一个轴(行或列)应用函数。
  • 元素级别: applymap() (DataFrame) / map() (Series) - 用于将一个 Python 函数应用于每个单独的元素。

性能说明: 向量化操作(直接在 Series/DataFrames 上使用内置的 Pandas/NumPy 函数)几乎总是比使用这些方法应用自定义 Python 函数更快。当没有向量化解决方案或解决方案过于复杂时,才使用 apply、applymap 或 map。

.pipe() 方法允许你以更具可读性的方式链接自定义函数,特别是那些预期将 DataFrame 或 Series 作为第一个参数的函数。它将调用该方法的 DataFrame(或 Series)作为第一个参数传递给提供的函数,同时传递任何其他指定的参数。

考虑这样一个场景:你有一个函数,它处理一个 DataFrame,也许是添加一个常量值。

示例: 通过 pipe() 应用 add_value 函数

Section titled “示例: 通过 pipe() 应用 add_value 函数”

我们来定义一个简单的函数 add_value,它将一个指定数字添加到 DataFrame 的所有元素中。

import pandas as pd
import numpy as np
def add_value(input_df, value_to_add):
"""为 DataFrame 的所有元素添加一个常量值。"""
return input_df + value_to_add
rng = np.random.default_rng(42)
df = pd.DataFrame(rng.standard_normal((5, 3)), columns=['col1', 'col2', 'col3'])
print("原始 DataFrame:")
print(df)
# 使用 pipe 应用函数
df_added = df.pipe(add_value, value_to_add=2)
print("\n使用 df.pipe(add_value, 2) 后的 DataFrame:")
print(df_added)
# 不使用 pipe(对于复杂链式调用可读性较差):
# df_added_alt = add_value(df, 2)
# print("\n使用 add_value(df, 2) 后的 DataFrame:")
# print(df_added_alt)

输出示例:

Original DataFrame:
col1 col2 col3
0 0.496714 -0.138264 0.647689
1 1.523030 -0.234153 -0.234137
2 1.579213 0.767435 -0.469474
3 0.542560 -0.463418 -0.465730
4 -0.465729 1.104638 -0.122903
DataFrame after df.pipe(add_value, 2):
col1 col2 col3
0 2.496714 1.861736 2.647689
1 3.523030 1.765847 1.765863
2 3.579213 2.767435 1.530526
3 2.542560 1.536582 1.534270
4 1.534271 3.104638 1.877097

pipe 在你需要按顺序应用多个自定义处理步骤时特别有用。

.apply() 方法沿着 DataFrame 的一个轴应用函数。传递给 apply 的函数接收一个 Series(表示一行或一列)作为输入。

  • axis=0 (默认): 将函数应用于每列(函数接收每列 Series)。
  • axis=1: 将函数应用于每行(函数接收每行 Series)。

注意: 使用 apply(..., axis=1) 可能会很慢,因为它通常在内部涉及逐行迭代。如果可能,总是优先使用向量化操作。

import pandas as pd
import numpy as np
rng = np.random.default_rng(43)
df = pd.DataFrame(rng.standard_normal((5, 3)), columns=['col1', 'col2', 'col3'])
print("DataFrame:")
print(df)
# 计算每列的均值
# 函数 np.mean 接收每列 Series
column_means = df.apply(np.mean, axis=0)
print("\n列均值 (df.apply(np.mean, axis=0)):")
print(column_means)
# 这等价于更直接且通常更推荐的方法:
# column_means_direct = df.mean(axis=0)
# print("\n列均值 (df.mean(axis=0)):")
# print(column_means_direct)

输出示例:

DataFrame:
col1 col2 col3
0 -0.535952 -0.875490 -0.269001
1 0.125819 0.747178 -1.255743
2 -1.034384 -0.977989 -0.375991
3 0.417561 -0.166471 0.674515
4 -0.563337 -0.508814 -0.494112
Column Means (df.apply(np.mean, axis=0)):
col1 -0.318059
col2 -0.356317
col3 -0.344066
dtype: float64
# 使用示例 1 中的 df
# 计算每行的均值
# 函数 np.mean 接收每行 Series
row_means = df.apply(np.mean, axis=1)
print("\n行均值 (df.apply(np.mean, axis=1)):")
print(row_means)
# 再次,等价于更直接的方法:
# row_means_direct = df.mean(axis=1)
# print("\n行均值 (df.mean(axis=1)):")
# print(row_means_direct)

输出示例:

Row Means (df.apply(np.mean, axis=1)):
0 -0.560148
1 -0.127582
2 -0.796121
3 -0.021465
4 -0.522088
dtype: float64

示例 3: 使用 Lambda 计算每行的范围 (最大值 - 最小值)

Section titled “示例 3: 使用 Lambda 计算每行的范围 (最大值 - 最小值)”

你可以使用 lambda 函数进行简洁的自定义操作。

# 使用示例 1 中的 df
# 计算每行的范围 (max - min)
row_range = df.apply(lambda s: s.max() - s.min(), axis=1)
print("\n行范围 (df.apply(lambda s: s.max() - s.min(), axis=1)):")
print(row_range)

输出示例:

Row Range (df.apply(lambda s: s.max() - s.min(), axis=1)):
0 0.606489
1 2.002921
2 0.658393
3 1.854339
4 0.069225
dtype: float64

元素级别应用: .applymap() (DataFrame) 和 .map() (Series)

Section titled “元素级别应用: .applymap() (DataFrame) 和 .map() (Series)”

当你需要将一个作用于单个值的 Python 函数应用于 DataFrame 或 Series 的每个元素时,分别使用 applymap() 或 map()。

  • .applymap(func): 将 func 应用于 DataFrame 的每个元素。
  • .map(func): 将 func 应用于 Series 的每个元素。

性能: 这些方法在 Python 层面逐个元素迭代,可能比向量化操作甚至 apply 慢得多。只有当你需要将纯 Python 函数应用于每个元素且向量化不可行时才使用它们。

我们来将 Series 中的数字格式化为货币字符串。

import pandas as pd
import numpy as np
rng = np.random.default_rng(44)
s = pd.Series(rng.random(5) * 100)
print("原始 Series:")
print(s)
# 使用 map 应用格式化函数
formatted_series = s.map(lambda x: f'${x:,.2f}') # 格式化为 $xxx.xx
print("\n格式化后的 Series (s.map(lambda x: f'${x:,.2f}')):")
print(formatted_series)

输出示例:

Original Series:
0 67.841675
1 24.097500
2 94.916909
3 23.539393
4 26.238116
dtype: float64
Formatted Series (s.map(lambda x: f'${x:,.2f}')):
0 $67.84
1 $24.10
2 $94.92
3 $23.54
4 $26.24
dtype: object

示例 2: 在 DataFrame 上使用 .applymap()

Section titled “示例 2: 在 DataFrame 上使用 .applymap()”

我们来将相同的货币格式化函数应用于 DataFrame 的所有元素。

# 使用示例 1 中的 df
print("原始 DataFrame:")
print(df)
# 使用 applymap 应用格式化函数
formatted_df = df.applymap(lambda x: f'{x:.2f}') # 格式化为保留两位小数
print("\n格式化后的 DataFrame (df.applymap(lambda x: f'{x:.2f}')):")
print(formatted_df)

输出示例:

Original DataFrame:
col1 col2 col3
0 -0.535952 -0.875490 -0.269001
1 0.125819 0.747178 -1.255743
2 -1.034384 -0.977989 -0.375991
3 0.417561 -0.166471 0.674515
4 -0.563337 -0.508814 -0.494112
Formatted DataFrame (df.applymap(lambda x: f'{x:.2f}')):
col1 col2 col3
0 -0.54 -0.88 -0.27
1 0.13 0.75 -1.26
2 -1.03 -0.98 -0.38
3 0.42 -0.17 0.67
4 -0.56 -0.51 -0.49