函数应用
Pandas - 应用自定义函数
Section titled “Pandas - 应用自定义函数”Pandas 允许你将自定义函数或来自其他库的函数应用于其对象(Series、DataFrame)。理解如何有效地应用函数是充分利用 Pandas 进行复杂数据处理的关键。函数应用有三种主要方法:
- DataFrame 级别:
pipe()- 用于链接预期接收 DataFrame 或 Series 作为输入的函数。 - 行/列级别:
apply()- 用于沿着一个轴(行或列)应用函数。 - 元素级别:
applymap()(DataFrame) /map()(Series) - 用于将一个 Python 函数应用于每个单独的元素。
性能说明: 向量化操作(直接在 Series/DataFrames 上使用内置的 Pandas/NumPy 函数)几乎总是比使用这些方法应用自定义 Python 函数更快。当没有向量化解决方案或解决方案过于复杂时,才使用 apply、applymap 或 map。
DataFrame 级别应用: .pipe()
Section titled “DataFrame 级别应用: .pipe()”.pipe() 方法允许你以更具可读性的方式链接自定义函数,特别是那些预期将 DataFrame 或 Series 作为第一个参数的函数。它将调用该方法的 DataFrame(或 Series)作为第一个参数传递给提供的函数,同时传递任何其他指定的参数。
考虑这样一个场景:你有一个函数,它处理一个 DataFrame,也许是添加一个常量值。
示例: 通过 pipe() 应用 add_value 函数
Section titled “示例: 通过 pipe() 应用 add_value 函数”我们来定义一个简单的函数 add_value,它将一个指定数字添加到 DataFrame 的所有元素中。
import pandas as pdimport numpy as np
def add_value(input_df, value_to_add): """为 DataFrame 的所有元素添加一个常量值。""" return input_df + value_to_add
rng = np.random.default_rng(42)df = pd.DataFrame(rng.standard_normal((5, 3)), columns=['col1', 'col2', 'col3'])
print("原始 DataFrame:")print(df)
# 使用 pipe 应用函数df_added = df.pipe(add_value, value_to_add=2)
print("\n使用 df.pipe(add_value, 2) 后的 DataFrame:")print(df_added)
# 不使用 pipe(对于复杂链式调用可读性较差):# df_added_alt = add_value(df, 2)# print("\n使用 add_value(df, 2) 后的 DataFrame:")# print(df_added_alt)输出示例:
Original DataFrame: col1 col2 col30 0.496714 -0.138264 0.6476891 1.523030 -0.234153 -0.2341372 1.579213 0.767435 -0.4694743 0.542560 -0.463418 -0.4657304 -0.465729 1.104638 -0.122903
DataFrame after df.pipe(add_value, 2): col1 col2 col30 2.496714 1.861736 2.6476891 3.523030 1.765847 1.7658632 3.579213 2.767435 1.5305263 2.542560 1.536582 1.5342704 1.534271 3.104638 1.877097pipe 在你需要按顺序应用多个自定义处理步骤时特别有用。
行或列级别应用: .apply()
Section titled “行或列级别应用: .apply()”.apply() 方法沿着 DataFrame 的一个轴应用函数。传递给 apply 的函数接收一个 Series(表示一行或一列)作为输入。
axis=0(默认): 将函数应用于每列(函数接收每列 Series)。axis=1: 将函数应用于每行(函数接收每行 Series)。
注意: 使用 apply(..., axis=1) 可能会很慢,因为它通常在内部涉及逐行迭代。如果可能,总是优先使用向量化操作。
示例 1: 列的均值 (axis=0)
Section titled “示例 1: 列的均值 (axis=0)”import pandas as pdimport numpy as np
rng = np.random.default_rng(43)df = pd.DataFrame(rng.standard_normal((5, 3)), columns=['col1', 'col2', 'col3'])
print("DataFrame:")print(df)
# 计算每列的均值# 函数 np.mean 接收每列 Seriescolumn_means = df.apply(np.mean, axis=0)print("\n列均值 (df.apply(np.mean, axis=0)):")print(column_means)
# 这等价于更直接且通常更推荐的方法:# column_means_direct = df.mean(axis=0)# print("\n列均值 (df.mean(axis=0)):")# print(column_means_direct)输出示例:
DataFrame: col1 col2 col30 -0.535952 -0.875490 -0.2690011 0.125819 0.747178 -1.2557432 -1.034384 -0.977989 -0.3759913 0.417561 -0.166471 0.6745154 -0.563337 -0.508814 -0.494112
Column Means (df.apply(np.mean, axis=0)):col1 -0.318059col2 -0.356317col3 -0.344066dtype: float64示例 2: 行的均值 (axis=1)
Section titled “示例 2: 行的均值 (axis=1)”# 使用示例 1 中的 df
# 计算每行的均值# 函数 np.mean 接收每行 Seriesrow_means = df.apply(np.mean, axis=1)print("\n行均值 (df.apply(np.mean, axis=1)):")print(row_means)
# 再次,等价于更直接的方法:# row_means_direct = df.mean(axis=1)# print("\n行均值 (df.mean(axis=1)):")# print(row_means_direct)输出示例:
Row Means (df.apply(np.mean, axis=1)):0 -0.5601481 -0.1275822 -0.7961213 -0.0214654 -0.522088dtype: float64示例 3: 使用 Lambda 计算每行的范围 (最大值 - 最小值)
Section titled “示例 3: 使用 Lambda 计算每行的范围 (最大值 - 最小值)”你可以使用 lambda 函数进行简洁的自定义操作。
# 使用示例 1 中的 df
# 计算每行的范围 (max - min)row_range = df.apply(lambda s: s.max() - s.min(), axis=1)print("\n行范围 (df.apply(lambda s: s.max() - s.min(), axis=1)):")print(row_range)输出示例:
Row Range (df.apply(lambda s: s.max() - s.min(), axis=1)):0 0.6064891 2.0029212 0.6583933 1.8543394 0.069225dtype: float64元素级别应用: .applymap() (DataFrame) 和 .map() (Series)
Section titled “元素级别应用: .applymap() (DataFrame) 和 .map() (Series)”当你需要将一个作用于单个值的 Python 函数应用于 DataFrame 或 Series 的每个元素时,分别使用 applymap() 或 map()。
.applymap(func): 将func应用于 DataFrame 的每个元素。.map(func): 将func应用于 Series 的每个元素。
性能: 这些方法在 Python 层面逐个元素迭代,可能比向量化操作甚至 apply 慢得多。只有当你需要将纯 Python 函数应用于每个元素且向量化不可行时才使用它们。
示例 1: 在 Series 上使用 .map()
Section titled “示例 1: 在 Series 上使用 .map()”我们来将 Series 中的数字格式化为货币字符串。
import pandas as pdimport numpy as np
rng = np.random.default_rng(44)s = pd.Series(rng.random(5) * 100)
print("原始 Series:")print(s)
# 使用 map 应用格式化函数formatted_series = s.map(lambda x: f'${x:,.2f}') # 格式化为 $xxx.xxprint("\n格式化后的 Series (s.map(lambda x: f'${x:,.2f}')):")print(formatted_series)输出示例:
Original Series:0 67.8416751 24.0975002 94.9169093 23.5393934 26.238116dtype: float64
Formatted Series (s.map(lambda x: f'${x:,.2f}')):0 $67.841 $24.102 $94.923 $23.544 $26.24dtype: object示例 2: 在 DataFrame 上使用 .applymap()
Section titled “示例 2: 在 DataFrame 上使用 .applymap()”我们来将相同的货币格式化函数应用于 DataFrame 的所有元素。
# 使用示例 1 中的 df
print("原始 DataFrame:")print(df)
# 使用 applymap 应用格式化函数formatted_df = df.applymap(lambda x: f'{x:.2f}') # 格式化为保留两位小数print("\n格式化后的 DataFrame (df.applymap(lambda x: f'{x:.2f}')):")print(formatted_df)输出示例:
Original DataFrame: col1 col2 col30 -0.535952 -0.875490 -0.2690011 0.125819 0.747178 -1.2557432 -1.034384 -0.977989 -0.3759913 0.417561 -0.166471 0.6745154 -0.563337 -0.508814 -0.494112
Formatted DataFrame (df.applymap(lambda x: f'{x:.2f}')): col1 col2 col30 -0.54 -0.88 -0.271 0.13 0.75 -1.262 -1.03 -0.98 -0.383 0.42 -0.17 0.674 -0.56 -0.51 -0.49