Skip to content

使用统计学理解数据

在机器学习中,数据是基础。模型的质量很大程度上取决于数据的质量和理解程度。在将数据输入任何算法之前,了解其特征至关重要。

数据可视化提供图形洞察(在另一章节介绍),而描述性统计(Descriptive statistics)提供数据属性的定量摘要。这种初步分析有助于识别潜在问题,如异常值(Outliers)、缺失值(Missing values)或偏态分布(Skewed distributions),从而指导数据预处理(Preprocessing)步骤。

本章重点介绍如何使用 Python,特别是 Pandas 库,对数据集执行基本的统计分析。

第一步总是查看实际数据值。这能让你对数据集有一个定性感官认识,并能立即发现诸如格式不一致或意外值等问题。

Pandas 中的 head() 和 tail() 方法非常适合此目的。我们加载 Pima Indian Diabetes 数据集(一个常用于二分类 Binary classification 的数据集)并查看前几行。

注意:确保已安装 ‘pandas’ 和 ‘scikit-learn’ 库(pip install pandas scikit-learn)。我们将使用 scikit-learn 的实用工具轻松获取数据集,避免硬编码文件路径。

import pandas as pd
from sklearn.datasets import fetch_openml
# Fetch the Pima Indians Diabetes dataset from OpenML
# Note: fetch_openml might require an internet connection the first time
pima = fetch_openml(name='diabetes', version=1, as_frame=True, parser='pandas')
df = pima.frame
# Define more descriptive column names (based on dataset description)
df.columns = ['num_pregnancies', 'glucose_concentration', 'blood_pressure', 'skin_thickness', 'insulin', 'bmi', 'diabetes_pedigree', 'age', 'class']
# Display the first 10 rows
print("First 10 rows of the dataset:")
print(df.head(10))
# Display the last 5 rows
print("\nLast 5 rows of the dataset:")
print(df.tail(5))
First 10 rows of the dataset:
num_pregnancies glucose_concentration blood_pressure skin_thickness insulin bmi diabetes_pedigree age class
0 6.0 148.0 72.0 35.0 0.0 33.6 0.627 50 1
1 1.0 85.0 66.0 29.0 0.0 26.6 0.351 31 0
2 8.0 183.0 64.0 0.0 0.0 23.3 0.672 32 1
3 1.0 89.0 66.0 23.0 94.0 28.1 0.167 21 0
4 0.0 137.0 40.0 35.0 168.0 43.1 2.288 33 1
5 5.0 116.0 74.0 0.0 0.0 25.6 0.201 30 0
6 3.0 78.0 50.0 32.0 88.0 31.0 0.248 26 1
7 10.0 115.0 0.0 0.0 0.0 35.3 0.134 29 0
8 2.0 197.0 70.0 45.0 543.0 30.5 0.158 53 1
9 8.0 125.0 96.0 0.0 0.0 0.0 0.232 54 1
Last 5 rows of the dataset:
num_pregnancies glucose_concentration blood_pressure skin_thickness insulin bmi diabetes_pedigree age class
763 10.0 101.0 76.0 48.0 180.0 32.9 0.171 63 0
764 2.0 122.0 70.0 27.0 0.0 36.8 0.340 27 0
765 5.0 121.0 72.0 23.0 112.0 26.2 0.245 30 0
766 1.0 126.0 60.0 0.0 0.0 30.1 0.349 47 1
767 1.0 93.0 70.0 31.0 0.0 30.4 0.315 23 0

观察前几行可以发现数据类型(多数为数值型)并大致了解数值范围。请注意,像 ‘blood_pressure’、‘skin_thickness’、‘insulin’ 和 ‘bmi’ 等一些列存在零值,这可能表示这些生理测量值的缺失数据,需要进一步调查或填充(Imputation)。

了解数据集的大小(行数/样本数和列数/特征数)对于规划至关重要:

  • 特征过多可能导致“维度灾难”(Curse of dimensionality),需要进行特征选择或特征工程。
  • 样本过少可能限制模型的泛化能力。
  • 大型数据集需要更多计算资源,可能需要不同的算法方法(例如,增量学习 Incremental learning、分布式计算 Distributed computing)。

Pandas DataFrame 的 shape 属性提供了这些信息。

import pandas as pd
from sklearn.datasets import load_iris
# Load the Iris dataset
iris = load_iris(as_frame=True)
df_iris = iris.frame
# Add the target column for completeness (optional for shape)
df_iris['target'] = iris.target
print(f"Dataset dimensions (rows, columns): {df_iris.shape}")
Dataset dimensions (rows, columns): (150, 5)

输出表明 Iris 数据集有 150 个样本和 5 列(4 个特征 + 1 个目标变量)。

机器学习算法通常需要特定的数据类型。例如,大多数算法期望数值输入。理解数据类型(dtypes)有助于识别可能需要转换的列(例如,将分类字符串转换为数值表示,如 One-Hot 编码 One-hot encoding 或 Label 编码 Label encoding)。

Pandas 在加载时会自动推断数据类型,但使用 dtypes 属性或 info() 方法(它也显示非空计数)进行验证是良好的实践。

import pandas as pd
from sklearn.datasets import load_iris
# Load the Iris dataset
iris = load_iris(as_frame=True)
df_iris = iris.frame
df_iris['target'] = iris.target
print("Data types of each column:")
print(df_iris.dtypes)
print("\n--- More Detailed Info ---")
df_iris.info()
Data types of each column:
sepal length (cm) float64
sepal width (cm) float64
petal length (cm) float64
petal width (cm) float64
target int64
dtype: object
--- More Detailed Info ---
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 150 entries, 0 to 149
Data columns (total 5 columns):
# Column Non-Null Count Dtype
--- ------ -------------- -----
0 sepal length (cm) 150 non-null float64
1 sepal width (cm) 150 non-null float64
2 petal length (cm) 150 non-null float64
3 petal width (cm) 150 non-null float64
4 target 150 non-null int64
dtypes: float64(4), int64(1)
memory usage: 6.0 KB

输出确认所有特征都是 float64,目标变量是 int64,适用于大多数数值算法。info() 输出也确认在这个特定的数据集中没有缺失值。

除了维度和类型,对每个数值属性的中心趋势(Central tendency)、离散度(Dispersion)和分布形状的摘要也至关重要。Pandas 中的 describe() 方法提供了关键统计量:

  • Count:非空观测值数量。
  • Mean:均值(平均值)。
  • Std:标准差(Standard Deviation),衡量围绕均值的离散度。
  • Min:最小值。
  • 25%:第一四分位数(Q1)。
  • 50%:中位数(Median,Q2)。
  • 75%:第三四分位数(Q3)。
  • Max:最大值。

这个摘要有助于识别不同属性的尺度(可能需要进行特征缩放 Scaling/Normalization)、异常值(中位数和均值之间的巨大差异,或极端的最小值/最大值)的存在,以及基本分布。

import pandas as pd
from sklearn.datasets import fetch_openml
# Fetch the Pima Indians Diabetes dataset from OpenML
pima = fetch_openml(name='diabetes', version=1, as_frame=True, parser='pandas')
df = pima.frame
df.columns = ['num_pregnancies', 'glucose_concentration', 'blood_pressure', 'skin_thickness', 'insulin', 'bmi', 'diabetes_pedigree', 'age', 'class']
# Set display options for better readability
pd.set_option('display.width', 100)
pd.set_option('display.precision', 3)
print("Dataset dimensions:")
print(df.shape)
print("\nStatistical Summary:")
print(df.describe())
Dataset dimensions:
(768, 9)
Statistical Summary:
num_pregnancies glucose_concentration blood_pressure skin_thickness insulin bmi diabetes_pedigree age class
count 768.000 768.000 768.000 768.000 768.000 768.000 768.000 768.000 768.000
mean 3.845 120.895 69.105 20.536 79.799 31.993 0.472 33.241 0.349
std 3.370 31.973 19.356 15.952 115.244 7.884 0.331 11.760 0.477
min 0.000 0.000 0.000 0.000 0.000 0.000 0.078 21.000 0.000
25% 1.000 99.000 62.000 0.000 0.000 27.300 0.244 24.000 0.000
50% 3.000 117.000 72.000 23.000 30.500 32.000 0.372 29.000 0.000
75% 6.000 140.250 80.000 32.000 127.250 36.600 0.626 41.000 1.000
max 17.000 199.000 122.000 99.000 846.000 67.100 2.420 81.000 1.000

摘要确认有 768 个样本。再次观察到 ‘glucose_concentration’、‘blood_pressure’、‘skin_thickness’、‘insulin’ 和 ‘bmi’ 等列的最小值为 0。这再次证实了将零值编码为缺失值的猜测,需要谨慎处理(例如,数据填充 Imputation)。

在分类任务中,理解目标变量(类别标签)的分布至关重要。不平衡类别(Imbalanced classes,指数据集中不同类别的样本数量差异显著的情况)可能使模型偏向于多数类别,可能需要特定的技术,如重采样(Resampling,过采样 Oversampling/欠采样 Undersampling)或使用不同的评估指标(例如,F1 分数 F1-score、AUC)。

我们可以使用 Pandas 的 groupby() 方法,然后结合 size() 来计算每个类别的出现次数。

import pandas as pd
from sklearn.datasets import fetch_openml
# Fetch the Pima Indians Diabetes dataset
pima = fetch_openml(name='diabetes', version=1, as_frame=True, parser='pandas')
df = pima.frame
df.columns = ['num_pregnancies', 'glucose_concentration', 'blood_pressure', 'skin_thickness', 'insulin', 'bmi', 'diabetes_pedigree', 'age', 'class']
# Convert class column to integer if it's not already (fetch_openml might load it as categorical)
df['class'] = pd.to_numeric(df['class'])
# Calculate class distribution
class_counts = df.groupby('class').size()
print("Class Distribution:")
print(class_counts)
Class Distribution:
class
0 500
1 268
dtype: int64

输出显示类别 0(无糖尿病)有 500 个样本,类别 1(有糖尿病)有 268 个样本。虽然不算极端不平衡,但差异很明显,在建模时可能需要考虑。

相关性(Correlation)衡量数值属性对之间的线性关系。Pearson 相关系数(Pearson correlation coefficient)是常用的指标:

  • 值接近 +1:强正线性相关。
  • 值接近 -1:强负线性相关。
  • 值接近 0:弱或无线性相关。

识别高度相关的特征很重要,因为:

  1. 冗余(Redundancy): 高度相关的特征可能提供冗余信息。
  2. 模型稳定性(Model Stability): 某些算法(如线性回归 Linear Regression)在面对高度相关的预测变量(多重共线性 Multicollinearity)时可能变得不稳定。 可能需要进行特征选择或降维技术(如 PCA)。

Pandas 的 corr() 方法计算成对相关矩阵。

import pandas as pd
from sklearn.datasets import fetch_openml
# Fetch the Pima Indians Diabetes dataset
pima = fetch_openml(name='diabetes', version=1, as_frame=True, parser='pandas')
df = pima.frame
df.columns = ['num_pregnancies', 'glucose_concentration', 'blood_pressure', 'skin_thickness', 'insulin', 'bmi', 'diabetes_pedigree', 'age', 'class']
df['class'] = pd.to_numeric(df['class'])
# Set display options
pd.set_option('display.width', 100)
pd.set_option('display.precision', 3)
# Calculate correlation matrix
correlations = df.corr(method='pearson')
print("Correlation Matrix:")
print(correlations)
Correlation Matrix:
num_pregnancies glucose_concentration blood_pressure skin_thickness insulin bmi diabetes_pedigree age class
num_pregnancies 1.000 0.129 0.141 -0.082 -0.074 0.018 -0.034 0.544 0.222
glucose_concentration 1.000 1.000 0.153 0.057 0.331 0.221 0.137 0.264 0.467
blood_pressure 0.141 0.153 1.000 0.207 0.089 0.282 0.041 0.240 0.065
skin_thickness -0.082 0.057 0.207 1.000 0.437 0.393 0.184 -0.114 0.075
insulin -0.074 0.331 0.089 0.437 1.000 0.198 0.185 -0.042 0.131
bmi 0.018 0.221 0.282 0.393 0.198 1.000 0.141 0.036 0.293
diabetes_pedigree -0.034 0.137 0.041 0.184 0.185 0.141 1.000 0.034 0.174
age 0.544 0.264 0.240 -0.114 -0.042 0.036 0.034 1.000 0.238
class 0.222 0.467 0.065 0.075 0.131 0.293 0.174 0.238 1.000

该矩阵显示了所有属性对之间的相关性。例如,‘age’ 和 ‘num_pregnancies’ 具有中等程度的正相关性(0.544)。‘glucose_concentration’ 与目标变量 ‘class’ 的相关性最高(0.467),在所示特征中。

偏度(Skewness)衡量概率分布的不对称性。许多机器学习算法(特别是线性模型和假设正态性的算法)在特征呈对称(类似高斯分布)分布时表现更好。

  • 正偏态(Positive Skew,右偏):尾部在右侧更长。均值 > 中位数。
  • 负偏态(Negative Skew,左偏):尾部在左侧更长。均值 < 中位数。
  • 零偏态(Zero Skew):对称分布(如完美的正态分布 Gaussian)。

显著的偏度可能需要在预处理过程中进行数据转换(例如,对数 Log、平方根 Square root、Box-Cox 转换)。Pandas 的 skew() 方法计算每个数值列的偏度。

import pandas as pd
from sklearn.datasets import fetch_openml
# Fetch the Pima Indians Diabetes dataset
pima = fetch_openml(name='diabetes', version=1, as_frame=True, parser='pandas')
df = pima.frame
df.columns = ['num_pregnancies', 'glucose_concentration', 'blood_pressure', 'skin_thickness', 'insulin', 'bmi', 'diabetes_pedigree', 'age', 'class']
df['class'] = pd.to_numeric(df['class'])
# Calculate skewness
skewness = df.skew()
print("Skewness of each attribute:")
print(skewness)
Skewness of each attribute:
num_pregnancies 0.902
glucose_concentration 0.174
blood_pressure -1.844
skin_thickness 0.109
insulin 2.272
bmi -0.429
diabetes_pedigree 1.920
age 1.130
class 0.635
dtype: float64

接近 0 的值表示低偏度(例如,‘glucose_concentration’、‘skin_thickness’)。与 0 显著不同的值表示偏度。例如,‘insulin’ (2.272) 和 ‘diabetes_pedigree’ (1.920) 显示显著的正偏态,而 ‘blood_pressure’ (-1.844) 显示显著的负偏态(可能受到零值的影响)。这些属性可能从转换中受益。