统计函数
Pandas - 统计函数
Section titled “Pandas - 统计函数”Pandas 对象(Series, DataFrame)内置了广泛的统计方法,这些方法对于数据分析和理解数据行为至关重要。这些方法默认通常会排除缺失数据(missing data) (NaN)。
百分比变化率 (.pct_change())
Section titled “百分比变化率 (.pct_change())”pct_change() 函数计算当前元素与先前元素之间的百分比变化率(默认情况下,是紧邻的前一个元素)。这对于分析时间序列数据或顺序测量非常有用。
import pandas as pdimport numpy as np
# Example with a Seriess = pd.Series([1, 2, 3, 4, 5, 4, 5, 6])print("Original Series:")print(s)print("\nPercentage Change in Series:")print(s.pct_change())
# Example with a DataFramedf = pd.DataFrame(np.random.randn(5, 2), columns=['A', 'B'])print("\nOriginal DataFrame:")print(df)print("\nPercentage Change in DataFrame (column-wise default):")print(df.pct_change()) # Default periods=1输出 (DataFrame 中的随机值可能不同):
Original Series:0 11 1.0000002 0.5000003 0.3333334 0.2500005 -0.2000006 0.2500007 0.200000dtype: float64
Percentage Change in Series:0 NaN # (NaN - 1) / NaN1 1.000000 # (2-1)/12 0.500000 # (3-2)/23 0.333333 # (4-3)/34 0.250000 # (5-4)/45 -0.200000 # (4-5)/56 0.250000 # (5-4)/47 0.200000 # (6-5)/5dtype: float64
Original DataFrame: A B0 0.123456 -0.9876541 -0.321098 1.0987652 0.543210 -0.1098763 -1.876543 0.7654324 0.112233 -0.445566
Percentage Change in DataFrame (column-wise default): A B0 NaN NaN1 -3.601135 -2.1123882 -2.691861 -1.1000003 -4.454545 -7.9660224 -1.059806 -1.581996第一个值始终是 NaN,因为没有先前的元素可以进行比较。您可以使用 periods 参数更改比较周期(例如,periods=2 与前两步的元素进行比较),并使用 axis=1 计算按行的变化率。
协方差 (.cov())
Section titled “协方差 (.cov())”协方差(Covariance)衡量两个随机变量(或 Pandas 中的两个 Series)的联合变异性。正协方差表示变量倾向于一起增加或减少,而负协方差表示它们朝相反的方向移动。
.cov() 方法可以应用于一个 Series 以计算其与另一个 Series 的协方差,也可以直接应用于一个 DataFrame 以计算所有列之间的成对协方差(pairwise covariance)。
两个 Series 之间的协方差
Section titled “两个 Series 之间的协方差”import pandas as pdimport numpy as np
s1 = pd.Series(np.random.randn(10))s2 = pd.Series(np.random.randn(10))
print(f"Covariance between s1 and s2: {s1.cov(s2)}")输出 (一个浮点数值,会变化):
Covariance between s1 and s2: -0.12978405324DataFrame 的协方差矩阵
Section titled “DataFrame 的协方差矩阵”import pandas as pdimport numpy as np
frame = pd.DataFrame(np.random.randn(10, 5), columns=['a', 'b', 'c', 'd', 'e'])
# Covariance between specific columns 'a' and 'b'print(f"Covariance between column 'a' and 'b': {frame['a'].cov(frame['b'])}")
# Compute the full covariance matrix for the DataFrameprint("\nCovariance matrix for the DataFrame:")print(frame.cov())输出 (值会变化):
Covariance between column 'a' and 'b': -0.5831292115274144
Covariance matrix for the DataFrame: a b c d ea 1.780628 -0.583129 -0.185575 0.003679 -0.136558b -0.583129 1.297011 0.136530 -0.523719 0.251064c -0.185575 0.136530 0.915227 -0.053881 -0.058926d 0.003679 -0.523719 -0.053881 1.521426 -0.487694e -0.136558 0.251064 -0.058926 -0.487694 0.960761结果 DataFrame 显示了每对列之间的协方差。对角线元素表示每列的方差 (cov(X, X) = var(X))。注意,单独计算的列 ‘a’ 和 ‘b’ 之间的协方差值与矩阵中对应的条目相匹配。
相关系数 (.corr())
Section titled “相关系数 (.corr())”相关系数(Correlation)是协方差的标准化版本,它指示两个变量之间线性关系的强度和方向。相关系数值范围从 -1(完全负线性关系)到 +1(完全正线性关系),0 表示没有线性关系。
Pandas 提供了 .corr() 方法,其工作方式类似于 .cov()。它可以计算两个 Series 之间的相关系数,或计算 DataFrame 的成对相关矩阵(pairwise correlation matrix)。有几种可用的相关方法:
'pearson'(默认):标准相关系数(Standard correlation coefficient)。'kendall':Kendall Tau 相关系数(Kendall Tau correlation coefficient)(基于排名)。'spearman':Spearman 秩相关系数(Spearman rank correlation coefficient)(基于排名)。
import pandas as pdimport numpy as np
frame = pd.DataFrame(np.random.randn(10, 5), columns=['a', 'b', 'c', 'd', 'e'])
# Correlation between specific columns 'a' and 'b'print(f"Pearson Correlation between 'a' and 'b': {frame['a'].corr(frame['b'])}")
# Compute the full Pearson correlation matrix for the DataFrameprint("\nPearson Correlation matrix:")print(frame.corr(method='pearson')) # 'pearson' is default
# Compute the Spearman rank correlation matrixprint("\nSpearman Correlation matrix:")print(frame.corr(method='spearman'))输出 (值会变化):
Pearson Correlation between 'a' and 'b': -0.383712785514
Pearson Correlation matrix: a b c d ea 1.000000 -0.383713 -0.145368 0.002235 -0.104405b -0.383713 1.000000 0.125311 -0.372821 0.224908c -0.145368 0.125311 1.000000 -0.045661 -0.062840d 0.002235 -0.372821 -0.045661 1.000000 -0.403380e -0.104405 0.224908 -0.062840 -0.403380 1.000000
Spearman Correlation matrix: a b c d ea 1.000000 -0.393939 -0.139394 0.078788 -0.163636b -0.393939 1.000000 0.163636 -0.333333 0.163636c -0.139394 0.163636 1.000000 -0.066667 -0.103030d 0.078788 -0.333333 -0.066667 1.000000 -0.406061e -0.163636 0.163636 -0.103030 -0.406061 1.000000与 .cov() 一样,非数值列会自动从相关计算中排除。对角线元素始终为 1(变量与其自身的相关性)。
数据排名 (.rank())
Section titled “数据排名 (.rank())”排名(Ranking)根据元素的大小或在 DataFrame 中沿某个轴的大小,为 Series 或 DataFrame 中的每个元素分配一个排名(例如,第 1 名、第 2 名、第 3 名)。它对于理解相对顺序(relative order)或作为非参数统计(non-parametric statistics)的基础很有用。
.rank() 方法使用 method 参数指定的不同策略处理并列(ties)(相等值):
'average'(默认):为并列元素分配平均排名。'min':为并列组中的所有元素分配最小排名。'max':为并列组中的所有元素分配最大排名。'first':根据元素在数据中出现的顺序顺序分配排名。'dense':类似于 ‘min’,但排名在组之间仅增加 1(无间隔)。
默认情况下,排名以升序分配(较小的值获得较低的排名)。使用 ascending=False 进行降序排名。
import pandas as pdimport numpy as np
# Create a Series with a ties = pd.Series(np.random.randn(5), index=list('abcde'))s['d'] = s['b'] # Introduce a tie between 'b' and 'd'
print("Original Series with a tie:")print(s)
# Calculate ranks (default: average method, ascending)print("\nRanks (method='average', ascending=True):")print(s.rank())
# Calculate ranks using 'first' methodprint("\nRanks (method='first', ascending=True):")print(s.rank(method='first'))
# Calculate ranks in descending orderprint("\nRanks (method='average', ascending=False):")print(s.rank(ascending=False))输出 (随机值会变化,影响排名,但并列处理逻辑保持不变):
Original Series with a tie:a -0.500000b 1.200000 # Tied valuec 0.300000d 1.200000 # Tied valuee 1.800000dtype: float64
Ranks (method='average', ascending=True):a 1.0 # Smallest valueb 3.5 # Average of ranks 3 and 4c 2.0d 3.5 # Average of ranks 3 and 4e 5.0 # Largest valuedtype: float64
Ranks (method='first', ascending=True):a 1.0b 3.0 # Gets rank 3 firstc 2.0d 4.0 # Gets rank 4 nexte 5.0dtype: float64
Ranks (method='average', ascending=False):a 5.0 # Largest rank (smallest value)b 2.5 # Average of ranks 2 and 3c 4.0d 2.5 # Average of ranks 2 and 3e 1.0 # Smallest rank (largest value)dtype: float64这些统计函数是 Pandas 中更复杂数据分析工作流程的基础构建块(building blocks)。