Skip to content

统计函数

Pandas 对象(Series, DataFrame)内置了广泛的统计方法,这些方法对于数据分析和理解数据行为至关重要。这些方法默认通常会排除缺失数据(missing data) (NaN)。

pct_change() 函数计算当前元素与先前元素之间的百分比变化率(默认情况下,是紧邻的前一个元素)。这对于分析时间序列数据或顺序测量非常有用。

import pandas as pd
import numpy as np
# Example with a Series
s = pd.Series([1, 2, 3, 4, 5, 4, 5, 6])
print("Original Series:")
print(s)
print("\nPercentage Change in Series:")
print(s.pct_change())
# Example with a DataFrame
df = pd.DataFrame(np.random.randn(5, 2), columns=['A', 'B'])
print("\nOriginal DataFrame:")
print(df)
print("\nPercentage Change in DataFrame (column-wise default):")
print(df.pct_change()) # Default periods=1

输出 (DataFrame 中的随机值可能不同):

Original Series:
0 1
1 1.000000
2 0.500000
3 0.333333
4 0.250000
5 -0.200000
6 0.250000
7 0.200000
dtype: float64
Percentage Change in Series:
0 NaN # (NaN - 1) / NaN
1 1.000000 # (2-1)/1
2 0.500000 # (3-2)/2
3 0.333333 # (4-3)/3
4 0.250000 # (5-4)/4
5 -0.200000 # (4-5)/5
6 0.250000 # (5-4)/4
7 0.200000 # (6-5)/5
dtype: float64
Original DataFrame:
A B
0 0.123456 -0.987654
1 -0.321098 1.098765
2 0.543210 -0.109876
3 -1.876543 0.765432
4 0.112233 -0.445566
Percentage Change in DataFrame (column-wise default):
A B
0 NaN NaN
1 -3.601135 -2.112388
2 -2.691861 -1.100000
3 -4.454545 -7.966022
4 -1.059806 -1.581996

第一个值始终是 NaN,因为没有先前的元素可以进行比较。您可以使用 periods 参数更改比较周期(例如,periods=2 与前两步的元素进行比较),并使用 axis=1 计算按行的变化率。

协方差(Covariance)衡量两个随机变量(或 Pandas 中的两个 Series)的联合变异性。正协方差表示变量倾向于一起增加或减少,而负协方差表示它们朝相反的方向移动。

.cov() 方法可以应用于一个 Series 以计算其与另一个 Series 的协方差,也可以直接应用于一个 DataFrame 以计算所有列之间的成对协方差(pairwise covariance)。

import pandas as pd
import numpy as np
s1 = pd.Series(np.random.randn(10))
s2 = pd.Series(np.random.randn(10))
print(f"Covariance between s1 and s2: {s1.cov(s2)}")

输出 (一个浮点数值,会变化):

Covariance between s1 and s2: -0.12978405324
import pandas as pd
import numpy as np
frame = pd.DataFrame(np.random.randn(10, 5), columns=['a', 'b', 'c', 'd', 'e'])
# Covariance between specific columns 'a' and 'b'
print(f"Covariance between column 'a' and 'b': {frame['a'].cov(frame['b'])}")
# Compute the full covariance matrix for the DataFrame
print("\nCovariance matrix for the DataFrame:")
print(frame.cov())

输出 (值会变化):

Covariance between column 'a' and 'b': -0.5831292115274144
Covariance matrix for the DataFrame:
a b c d e
a 1.780628 -0.583129 -0.185575 0.003679 -0.136558
b -0.583129 1.297011 0.136530 -0.523719 0.251064
c -0.185575 0.136530 0.915227 -0.053881 -0.058926
d 0.003679 -0.523719 -0.053881 1.521426 -0.487694
e -0.136558 0.251064 -0.058926 -0.487694 0.960761

结果 DataFrame 显示了每对列之间的协方差。对角线元素表示每列的方差 (cov(X, X) = var(X))。注意,单独计算的列 ‘a’ 和 ‘b’ 之间的协方差值与矩阵中对应的条目相匹配。

相关系数(Correlation)是协方差的标准化版本,它指示两个变量之间线性关系的强度和方向。相关系数值范围从 -1(完全负线性关系)到 +1(完全正线性关系),0 表示没有线性关系。

Pandas 提供了 .corr() 方法,其工作方式类似于 .cov()。它可以计算两个 Series 之间的相关系数,或计算 DataFrame 的成对相关矩阵(pairwise correlation matrix)。有几种可用的相关方法:

  • 'pearson' (默认):标准相关系数(Standard correlation coefficient)。
  • 'kendall':Kendall Tau 相关系数(Kendall Tau correlation coefficient)(基于排名)。
  • 'spearman':Spearman 秩相关系数(Spearman rank correlation coefficient)(基于排名)。
import pandas as pd
import numpy as np
frame = pd.DataFrame(np.random.randn(10, 5), columns=['a', 'b', 'c', 'd', 'e'])
# Correlation between specific columns 'a' and 'b'
print(f"Pearson Correlation between 'a' and 'b': {frame['a'].corr(frame['b'])}")
# Compute the full Pearson correlation matrix for the DataFrame
print("\nPearson Correlation matrix:")
print(frame.corr(method='pearson')) # 'pearson' is default
# Compute the Spearman rank correlation matrix
print("\nSpearman Correlation matrix:")
print(frame.corr(method='spearman'))

输出 (值会变化):

Pearson Correlation between 'a' and 'b': -0.383712785514
Pearson Correlation matrix:
a b c d e
a 1.000000 -0.383713 -0.145368 0.002235 -0.104405
b -0.383713 1.000000 0.125311 -0.372821 0.224908
c -0.145368 0.125311 1.000000 -0.045661 -0.062840
d 0.002235 -0.372821 -0.045661 1.000000 -0.403380
e -0.104405 0.224908 -0.062840 -0.403380 1.000000
Spearman Correlation matrix:
a b c d e
a 1.000000 -0.393939 -0.139394 0.078788 -0.163636
b -0.393939 1.000000 0.163636 -0.333333 0.163636
c -0.139394 0.163636 1.000000 -0.066667 -0.103030
d 0.078788 -0.333333 -0.066667 1.000000 -0.406061
e -0.163636 0.163636 -0.103030 -0.406061 1.000000

与 .cov() 一样,非数值列会自动从相关计算中排除。对角线元素始终为 1(变量与其自身的相关性)。

排名(Ranking)根据元素的大小或在 DataFrame 中沿某个轴的大小,为 Series 或 DataFrame 中的每个元素分配一个排名(例如,第 1 名、第 2 名、第 3 名)。它对于理解相对顺序(relative order)或作为非参数统计(non-parametric statistics)的基础很有用。

.rank() 方法使用 method 参数指定的不同策略处理并列(ties)(相等值):

  • 'average' (默认):为并列元素分配平均排名。
  • 'min':为并列组中的所有元素分配最小排名。
  • 'max':为并列组中的所有元素分配最大排名。
  • 'first':根据元素在数据中出现的顺序顺序分配排名。
  • 'dense':类似于 ‘min’,但排名在组之间仅增加 1(无间隔)。

默认情况下,排名以升序分配(较小的值获得较低的排名)。使用 ascending=False 进行降序排名。

import pandas as pd
import numpy as np
# Create a Series with a tie
s = pd.Series(np.random.randn(5), index=list('abcde'))
s['d'] = s['b'] # Introduce a tie between 'b' and 'd'
print("Original Series with a tie:")
print(s)
# Calculate ranks (default: average method, ascending)
print("\nRanks (method='average', ascending=True):")
print(s.rank())
# Calculate ranks using 'first' method
print("\nRanks (method='first', ascending=True):")
print(s.rank(method='first'))
# Calculate ranks in descending order
print("\nRanks (method='average', ascending=False):")
print(s.rank(ascending=False))

输出 (随机值会变化,影响排名,但并列处理逻辑保持不变):

Original Series with a tie:
a -0.500000
b 1.200000 # Tied value
c 0.300000
d 1.200000 # Tied value
e 1.800000
dtype: float64
Ranks (method='average', ascending=True):
a 1.0 # Smallest value
b 3.5 # Average of ranks 3 and 4
c 2.0
d 3.5 # Average of ranks 3 and 4
e 5.0 # Largest value
dtype: float64
Ranks (method='first', ascending=True):
a 1.0
b 3.0 # Gets rank 3 first
c 2.0
d 4.0 # Gets rank 4 next
e 5.0
dtype: float64
Ranks (method='average', ascending=False):
a 5.0 # Largest rank (smallest value)
b 2.5 # Average of ranks 2 and 3
c 4.0
d 2.5 # Average of ranks 2 and 3
e 1.0 # Smallest rank (largest value)
dtype: float64

这些统计函数是 Pandas 中更复杂数据分析工作流程的基础构建块(building blocks)。